system
The system automatically transcribes and summarizes meeting audio, compares it with past records, and notifies users of delays, enhancing meeting efficiency and project management by addressing inefficiencies in manual processes.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-01
- Publication Date
- 2026-04-13
AI Technical Summary
Existing meeting recording and management systems require manual conversion of audio to text and summarization, which is time-consuming and inefficient, and lack real-time progress evaluation capabilities, making it difficult to detect delays and discrepancies promptly.
A system that records meeting audio in real time, transcribes it into text using speech recognition, summarizes the content with natural language processing, compares it with past meeting minutes, and visualizes and notifies users of delays or discrepancies.
Enables efficient and accurate real-time generation of meeting summaries and progress evaluations, allowing for immediate user notification of significant delays or discrepancies, thereby improving meeting efficiency and project management.
Smart Images

Figure 2026063784000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
[0006] "Means for recording meeting audio" refers to an apparatus or method for capturing and recording audio during a meeting as a digital signal using a microphone or other audio input device.
[0007] "Means for converting recorded audio data into text data" refers to an apparatus or method for converting captured audio data into natural language text format using speech recognition technology.
[0008] "Means for summarizing converted text data" refers to a device or method that uses natural language processing technology to extract important points from converted text data and express them in a shortened form.
[0009] "Means for comparing with past meeting minutes data" refers to a device or method for comparing generated summary data with past meeting minutes data stored in a database to identify differences and similarities.
[0010] "Means for detecting delays or discrepancies in progress" refers to a device or method for identifying parts that are not progressing as planned or that deviate from the plan, based on summarized data compared with past meeting minutes data.
[0011] "Means for visualizing comparison results" refers to a device or method that displays detected progress delays or discrepancies in a visual format such as graphs or charts, making them easily recognizable by users.
[0012] "Means of notifying the user of the results" refers to a device or method that communicates information about detected significant progress delays or discrepancies to the user as an alert or notification message.
[0013] A "speech recognition engine" is software or algorithms that analyze audio signals and convert that audio into text format. [Brief explanation of the drawing]
[0014] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12]It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when the emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when the emotion engine is combined.
Mode for Carrying Out the Invention
[0015] Hereinafter, an example of an embodiment of the system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0016] First, the language used in the following description will be explained.
[0017] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), APU (Accelerated Processing Unit), and the like.
[0018] In the following embodiments, the numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0019] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.
[0020] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0021] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0022] [First Embodiment]
[0023] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0024] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0025] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0026] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0027] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0028] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0029] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0030] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0031] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0032] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0033] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0034] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0035] The present invention provides a system that records meeting audio in real time, automatically transcribes and summarizes the content, and further evaluates the progress by comparing it with past meeting minutes data. The embodiments for carrying out the present invention are described in detail below.
[0036] Audio recording and conversion
[0037] Terminal:
[0038] When the meeting begins, audio data is captured via the microphone on the terminal. The captured audio data is sent to the server in real time.
[0039] server:
[0040] The server inputs the received audio data into a speech recognition engine and converts it into text data. A highly accurate speech recognition engine (e.g., speech recognition software or a cloud API) is used.
[0041] Summary generation using natural language processing
[0042] server:
[0043] Once the text data is generated, the server analyzes it using natural language processing (NLP) algorithms. The NLP algorithms extract key points and keywords, generating a summary of the text data. This summarization process provides long meeting minutes in a concise, summarized form.
[0044] Comparison with past meeting minutes and progress evaluation
[0045] server:
[0046] The generated summary data is compared with past meeting minutes data stored in the database. The comparison algorithm detects delays in progress and discrepancies with the plan. Specifically, the current status is compared with milestones and plans in past meeting minutes.
[0047] Visualization and notification
[0048] server:
[0049] The detected results are reflected in a visually displayed dashboard, allowing users to see the progress at a glance. For example, delays in progress can be displayed in graph or chart format, making it easy for users to understand.
[0050] Terminal:
[0051] Results sent from the server are updated in real time on the device. If significant delays or discrepancies are detected, the device will notify the user and display the details.
[0052] Specific example
[0053] For example, suppose a weekly project meeting is discussing the progress of new feature A. If, during the meeting, someone says, "Development of new feature A is on track and we've entered the testing phase":
[0054] 1. Terminal: Captures audio and sends it to the server.
[0055] 2. Server: Converts the audio data into text and generates the text, "Development of new feature A is on track and has entered the testing phase."
[0056] 3. Server: Summarize using an NLP algorithm and conclude with "New Feature A: Proceeding to Test Phase".
[0057] 4. Server: Compare the current meeting minutes with those stored in the database to assess whether progress is on schedule.
[0058] 5. Server: Visualize the progress on a dashboard and notify users.
[0059] This process allows users to immediately check the progress after the meeting ends and take prompt action. The system of this invention enables companies and project teams to achieve efficient and effective meeting management and project management.
[0060] The following describes the processing flow.
[0061] Step 1:
[0062] Terminal: As soon as the meeting starts, the microphone on the terminal captures audio data. The captured audio data is sent to the server in real time.
[0063] Step 2:
[0064] Server: Passes the received audio data to the speech recognition engine, which converts it into text data. The speech recognition engine analyzes the audio signal and transcribes it into sentences such as, "The next milestone is next Monday."
[0065] Step 3:
[0066] Server: Inputs the converted text data into a natural language processing (NLP) algorithm to generate a summary. For example, it summarizes the sentence "The next milestone is next Monday" as "Milestone: Monday".
[0067] Step 4:
[0068] Server: Compares the generated summary data with past meeting minutes data stored in the database. Uses a comparison algorithm to detect delays and discrepancies in progress.
[0069] Step 5:
[0070] Server: Analyzes comparison results to identify significant delays and discrepancies. The identified results are displayed on a dashboard in graph and chart format using visualization tools.
[0071] Step 6:
[0072] Terminal: Receives visualized data sent from the server and displays it on the user's screen in real time. If there are significant delays or discrepancies in progress, a notification will pop up.
[0073] Step 7:
[0074] User: Check the latest meeting minutes and progress displayed on the terminal. The user will consider and take appropriate action regarding any delays or discrepancies in progress.
[0075] To explain with a concrete example, it would be as follows:
[0076] In the weekly project progress meeting:
[0077] 1. Device: Captures meeting audio in real time.
[0078] 2. Server: Converts audio data into text data.
[0079] 3. Server: Generates summaries using NLP algorithms.
[0080] 4. Server: Compares with past meeting minutes data and detects discrepancies.
[0081] 5. Server: Visualize the results on a dashboard.
[0082] 6. Terminal: Displays results and important notifications on the user's terminal.
[0083] 7. User: Review the results and respond to the project's progress.
[0084] In this way, the system enables automatic summarization of meeting content, progress monitoring, and important notifications at each step.
[0085] (Example 1)
[0086] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0087] Traditional meeting recording systems required manually converting audio data to text and summarizing its contents, which was time-consuming and labor-intensive. Furthermore, it was difficult to evaluate progress by comparing current meeting minutes with past ones. This not only hindered prompt responses after meetings but also made it difficult to detect project delays and discrepancies early on.
[0088] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0089] In this invention, the server includes means for converting received audio data into text data using a speech recognition engine, means for summarizing the text data using a natural language processing algorithm, and means for comparing the summarized text data with past meeting minutes data. This makes it possible to automatically generate accurate meeting records in real time and to quickly perform summarization and progress evaluation.
[0090] 1. "Means for capturing meeting audio" refers to a hardware or software system for collecting audio during a meeting in digital format.
[0091] 2. "Means for transmitting captured audio data in real time" refers to communication means for transmitting collected audio data to a server without delay.
[0092] 3. "Means for converting received audio data into text data using a speech recognition engine" refers to algorithms or software for analyzing audio data and converting it into corresponding text data.
[0093] 4. "Means for summarizing text data using natural language processing algorithms" refers to methods for analyzing generated text data, extracting important information, and summarizing it concisely.
[0094] 5. "Means for comparing summarized text data with past meeting minutes data" refers to a method for comparing newly generated summarized text with past meeting minutes to evaluate progress and detect discrepancies.
[0095] 6. "Means for detecting delays or discrepancies in progress based on comparison results" refers to a mechanism that analyzes the comparison results of text data and automatically identifies differences from the plan or delays in progress.
[0096] 7. "Means for visualizing comparison results" refers to display methods in the form of graphs or charts that allow users to intuitively understand the analyzed comparison results.
[0097] 8. "Means for notifying users of visualized results" refers to a notification system that informs users in real time of important progress and any detected discrepancies.
[0098] 9. "Means for converting audio data to text in real time through parallel processing" refers to a method of efficiently utilizing computing resources to quickly convert audio data to text by simultaneously executing multiple audio analysis tasks.
[0099] 10. "Means of using generative AI models in natural language processing algorithms" refers to advanced algorithms that utilize generative AI models to analyze and summarize text data.
[0100] This invention provides a system that records meeting audio in real time, automatically transcribes and summarizes the content, and evaluates progress by comparing it with past meeting minutes data. The following describes how to implement this system in detail.
[0101] Audio recording and conversion
[0102] Terminal:
[0103] When the meeting begins, audio data is captured via the microphone on the terminal. The captured audio data is sent to the server in real time. The audio data is saved to a buffer at regular intervals, and when the buffer size reaches a certain level, it is sent back to the server.
[0104] server:
[0105] The server inputs the received audio data into a speech recognition engine (e.g., Google® Cloud Speech-to-Text API) and converts it into text data. The text conversion speed is optimized by dividing the audio data into small chunks and executing API requests in parallel.
[0106] Summary generation using natural language processing
[0107] server:
[0108] Once text data is generated, the server analyzes it using a natural language processing (NLP) algorithm (e.g., OpenAI's generative AI model). The NLP algorithm extracts key points and keywords, generating a summary of the text data. The following is an example of a prompt: "Use an NLP algorithm to summarize the following text: Text: 'Development of new feature A is on track and has entered the testing phase. The development team is progressing as planned and will report the test results at the next meeting.'"
[0109] Comparison with past meeting minutes and progress evaluation
[0110] server:
[0111] The generated summary data is compared with historical meeting minutes data stored in the database. A comparison algorithm detects delays in progress and discrepancies with the plan. Specifically, historical data is retrieved using SQL queries, and a Python script detects matching keywords and progress differences with the summary data.
[0112] Visualization and notification
[0113] server:
[0114] The detected results are reflected in a visually displayed dashboard, allowing users to check the progress at a glance. Progress delays are displayed in graph and chart format using D3.js, among other tools.
[0115] Terminal:
[0116] Results sent from the server are updated in real time on the device. If significant delays or discrepancies are detected, the device notifies the user and displays the details. A notification pop-up appears in the browser, and clicking it jumps to a detailed screen.
[0117] Specific example
[0118] For example, suppose a weekly project meeting is discussing the progress of new feature A. If, during the meeting, someone says, "Development of new feature A is on track and we've entered the testing phase":
[0119] 1. Terminal: Captures audio and sends it to the server.
[0120] 2. Server: The server inputs the received audio data into the speech recognition engine and generates the text, "Development of new feature A is on track and has entered the testing phase."
[0121] 3. Server: Use an NLP algorithm to summarize and conclude, "New Feature A: Proceeding to Test Phase."
[0122] 4. Server: Compare the current meeting minutes with those stored in the database to assess whether Feature A is progressing as planned. For example, "Previous meeting: Design phase scheduled to be completed," "Current meeting: Proceeding to the testing phase." Assess that it is progressing as planned.
[0123] 5. Server: Display "New feature A is progressing as planned" on the dashboard and notify the user.
[0124] This process allows users to immediately check the progress after the meeting ends and take prompt action.
[0125] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0126] Step 1: Capture and send audio data
[0127] Terminal:
[0128] Input: Capture conference audio via the microphone on the device.
[0129] Data processing: Audio data is saved to a buffer at regular intervals.
[0130] Output: Once the buffer size reaches a certain level, it is sent to the server in real time.
[0131] Specific operation: The terminal captures audio data, and when it reaches a certain buffer size, it sends the data to the server. A low-latency communication protocol is used for this process.
[0132] Step 2: Receiving audio data and converting it to text.
[0133] server:
[0134] Input: Audio data sent from the device.
[0135] Data processing: Divide audio data into small chunks.
[0136] Data processing: Convert audio data into text data using a speech recognition engine (e.g., a cloud-based API).
[0137] Output: Converted real-time text data.
[0138] Specific operation: The server subdivides the received audio data and passes it to the speech recognition engine. Parallel processing is used to perform efficient text conversion.
[0139] Step 3: Text-based summary generation
[0140] server:
[0141] Input: Text data.
[0142] Data processing: Analyze text data using natural language processing algorithms.
[0143] Data processing: Extract and summarize key points using a generative AI model (e.g., OpenAI GPT-3®).
[0144] Output: Summarized text data.
[0145] Specific operation: The server inputs a prompt message into the generation AI model and generates a summary of the text data. Example of a prompt message: "Use an NLP algorithm to summarize the following text. Text: 'Development of new feature A is on track and has entered the testing phase. The development team is progressing as planned and will report the test results at the next meeting.'"
[0146] Step 4: Comparison of summary data and past meeting minutes
[0147] server:
[0148] Input: Summarized text data.
[0149] Data processing: Retrieve past meeting minutes data from the database.
[0150] Data processing: Use a comparison algorithm to match summary data with past meeting minutes.
[0151] Output: Progress evaluation and comparison results of discrepancies.
[0152] Specific operation: The server retrieves past meeting minutes using SQL queries, and a Python script detects similarities and differences between the two sets of data.
[0153] Step 5: Visualizing progress
[0154] server:
[0155] Input: Results of progress evaluations and discrepancies comparisons.
[0156] Data processing: Formatting data for visualization.
[0157] Data processing: Visualization in graph and chart formats (e.g., D3.js).
[0158] Output: The visualization results displayed on the dashboard.
[0159] Specific operation: The server sends visualization data in JSON format to the dashboard API and updates it in real time.
[0160] Step 6: Notifications and detailed display
[0161] Terminal:
[0162] Input: Visualization results sent from the server.
[0163] Data processing: Formatting the data into a form that is easy for users to understand.
[0164] Output: Real-time notifications and detailed display on the device.
[0165] Specific operation: The device displays a notification pop-up, and clicking it jumps to a details screen. The user can easily check the progress.
[0166] These steps enable the system to transcribe, summarize, assess progress, visualize meeting audio in real time, and notify users.
[0167] (Application Example 1)
[0168] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0169] Modern factory operations demand complex project management and rapid response. However, managing meeting audio and progress is often done manually, which is inefficient and makes it difficult to immediately detect delays or discrepancies. This leads to delays in appropriate responses and decreased productivity. Therefore, there is a need for a system that automatically records meeting audio, transcribes it into text, summarizes it, and evaluates progress in real time.
[0170] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0171] In this invention, the server includes means for recording meeting audio, means for converting the recorded audio data into text data, means for summarizing the converted text data, means for comparing it with past meeting minutes data, means for detecting delays or discrepancies in progress, means for visualizing the comparison results, means for notifying the user of the results, means for a robot placed in the factory to record and process the meeting audio, and means for displaying the converted and summarized text data on a dashboard in real time. This enables the automatic recording of audio content of meetings held in the factory, improving the efficiency of progress management and allowing for quicker responses.
[0172] "Means of recording meeting audio" refers to a system that saves the audio from a meeting as digital data using a microphone or similar device.
[0173] "Means for converting recorded audio data into text data" refers to the process of converting recorded audio into text using a speech recognition engine.
[0174] "Methods for summarizing converted text data" refers to the process of using natural language processing techniques to extract only the important points from long texts and make them concise.
[0175] "Methods for comparing with past meeting minutes data" refer to a system that compares the content of the previous meeting stored in a database with the content of the current meeting to check for differences and progress.
[0176] "Methods for detecting delays or discrepancies in progress" refer to a system that compares past meeting minutes data with current data to identify differences between the plan and the actual progress.
[0177] "Methods for visualizing comparison results" refer to the process of displaying data differences and progress in visual formats such as graphs and charts.
[0178] "Means of notifying users of results" refers to a system that informs users of any delays or discrepancies in progress.
[0179] "Means for recording and processing meeting audio using robots deployed in a factory" refers to a process in which robots installed within a factory capture meeting audio and process that audio data in real time.
[0180] "A means of displaying real-time converted and summarized text data on a dashboard" refers to a mechanism that displays real-time converted and summarized text data on a dashboard so that users can understand it at a glance.
[0181] This invention is a system for automating and streamlining the progress management of meetings held in a factory. It records meeting audio, transcribes and summarizes the content, and evaluates the progress by comparing it with past meeting minutes data. Furthermore, this system has a function to visualize the progress in real time and notify the user.
[0182] Audio recording and conversion
[0183] Hardware: Robots deployed in the factory are equipped with high-sensitivity microphones to capture meeting audio in real time. This allows the system to begin collecting audio data as soon as the meeting starts.
[0184] Software: The collected audio data is sent to a server and converted into text data by a speech recognition engine (e.g., Google Web Speech API). This process is done in real time, so speech during a meeting is instantly transcribed into text.
[0185] Summary generation using natural language processing
[0186] Server: Once text data is generated, the server uses natural language processing (NLP) algorithms (e.g., the Hugging Face generative AI model) to summarize the text data. The summarized text contains key points, allowing users to efficiently grasp the information.
[0187] Comparison with past meeting minutes and progress evaluation
[0188] Server: The generated summary data is compared with past meeting minutes data stored in the database. The comparison algorithm detects any delays or discrepancies in progress. This ensures a clear understanding of the current progress against milestones and plans decided in past meetings.
[0189] Visualization and notification
[0190] Server: Detected comparison results are reflected in a visually displayed dashboard. Progress is shown in graph and chart format, allowing users to see the current progress and any issues at a glance. Furthermore, if significant delays or discrepancies are detected, the server will notify the user.
[0191] Specific example: For instance, when reporting on the progress of new feature A in a project meeting at a factory, it would look like this:
[0192] 1. The robot captures the meeting audio and sends it to the server.
[0193] 2. The server converts the audio data into text, specifically generating the text "New feature A has progressed from the development stage to the testing stage."
[0194] 3. Use an NLP algorithm to summarize the text, concluding it as "New Feature A: Testing Phase".
[0195] 4. Compare the current server status with past meeting minutes stored in the database to assess whether it is progressing according to plan.
[0196] 5. Progress is visualized on a dashboard, allowing users to check it in real time. If progress is behind schedule, users can immediately consider countermeasures.
[0197] A possible example of a specific prompt message would be something like, "We have received a progress report on Project X. You need to start on a new task, Task Y."
[0198] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0199] Step 1:
[0200] Terminal: A robot positioned in the factory captures meeting audio and transmits it to a server in real time. The input data is audio data, and the output is audio data transferred to the server. Specifically, the robot's microphone captures the audio and transmits that audio data to the server via the network.
[0201] Step 2:
[0202] Server: Inputs received audio data into a speech recognition engine and converts it into text data. Input data is audio data, and output data is text data. Specifically, the server calls a speech recognition engine (e.g., Google Web Speech API) and converts the audio data into text format.
[0203] Step 3:
[0204] Server: Summarizes the generated text data using a natural language processing (NLP) algorithm. The input data is the transformed text data, and the output data is the summarized text data. Specifically, the server uses the Hugging Face generative AI model to extract important information from long texts and generate summaries.
[0205] Step 4:
[0206] Server: Compares summarized text data with past meeting minutes data stored in the database. The input data consists of summarized text data and past meeting minutes data, and the output data is the comparison result. Specifically, the server retrieves meeting minutes data from the past database, compares it with the current summarized data, and calculates the difference.
[0207] Step 5:
[0208] Server: Detects delays and discrepancies in progress based on comparison results. Input data is the comparison results, and output data is information about delays and discrepancies in progress. Specifically, the server uses an algorithm to evaluate the progress status and identify the gap with the plan.
[0209] Step 6:
[0210] Server: Reflects detected results in a dashboard for visual display. Input data is information about delays and discrepancies in progress, and output data is visualized graphs and charts. Specifically, the server uses the matplotlib library to graphically display the progress.
[0211] Step 7:
[0212] Server: Sends generated visualization data to the terminal and notifies the user. Input data is visualized graphs and charts, and output data is progress displayed on the user interface. Specifically, the server pushes the generated visualization data to the terminal in real time and notifies the user of important difference information.
[0213] Step 8:
[0214] User: Check the progress on the dashboard on the device and decide on countermeasures as needed. The input data is the displayed progress, and the output is the decision on the appropriate countermeasure. Specifically, the user accesses the dashboard on the device and plans the next action based on the visualized progress.
[0215] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0216] This invention combines a system that records meeting audio in real time, automatically transcribes and summarizes its contents, and evaluates progress by comparing it with past meeting minutes data, with an emotion engine that recognizes the user's emotions from the audio data. The embodiments for carrying out this invention will be described in detail below.
[0217] Audio recording and conversion
[0218] Terminal:
[0219] When the meeting begins, audio data is captured via the microphone on the terminal. The captured audio data is sent to the server in real time.
[0220] server:
[0221] The server inputs the received audio data into the speech recognition engine, which converts it into text data. The speech recognition engine analyzes the audio signal and converts the meeting content into text format.
[0222] Summary generation using natural language processing
[0223] server:
[0224] Once the text data is generated, the server analyzes it using natural language processing (NLP) algorithms. The NLP algorithms extract key points and keywords, generating a summary of the text data. This summarization process provides a concise summary of the meeting's content.
[0225] User emotion recognition by an emotion engine
[0226] server:
[0227] Furthermore, the server inputs the voice data into the emotion engine to recognize the user's emotions. The emotion engine analyzes the user's emotions (e.g., joy, anger, sadness, surprise, etc.) from the voice signal and adds them to the text data.
[0228] Comparison with past meeting minutes and progress evaluation
[0229] server:
[0230] The generated summary data and sentiment data are compared with historical meeting minutes data stored in the database. A comparison algorithm detects delays and discrepancies with the plan. Based on the sentiment data, the importance of the delays and discrepancies is also evaluated.
[0231] Visualization and notification
[0232] server:
[0233] The detected results and sentiment data are reflected in a visually displayed dashboard. This allows users to see the progress and its emotional impact at a glance. For example, progress delays are displayed in graphs and charts, and sentiment data is indicated by colors and icons.
[0234] Terminal:
[0235] Results sent from the server are updated in real time on the device. If significant progress delays or discrepancies are detected, a notification will pop up along with sentiment data.
[0236] Specific example
[0237] For example, suppose a weekly project meeting is discussing the progress of new feature A. If, during the meeting, it is stated, "Development of new feature A is on track and has entered the testing phase":
[0238] 1. Terminal: Captures audio and sends it to the server.
[0239] 2. Server: Converts audio data into text data.
[0240] 3. Server: Generates summaries using NLP algorithms.
[0241] 4. Server: The emotion engine analyzes the emotion of "joy" and assigns it to the summary. For example, it might be assigned as "New Feature A: Progressing to Test Phase (Joy)".
[0242] 5. Server: Compare the current meeting minutes with those stored in the database to assess whether progress is on schedule. Evaluate the importance of any delays, including sentiment data.
[0243] 6. Server: Visualize the results on a dashboard and notify the user.
[0244] 7. Device: Displayed on the user's device, allowing them to check progress and sentiment data.
[0245] 8. User: Based on the results, take appropriate action regarding the project's progress.
[0246] By combining this with an emotion engine, it becomes possible to not only automatically summarize meeting content but also to evaluate progress while considering the user's emotions, resulting in more accurate project management. This improves the project success rate and enhances team engagement.
[0247] The following describes the processing flow.
[0248] Step 1:
[0249] Terminal: As soon as the meeting starts, the microphone on the terminal captures audio data. The captured audio data is sent to the server in real time.
[0250] Step 2:
[0251] Server: Inputs received audio data into the speech recognition engine and converts it into text data. For example, it recognizes the audio data "The next milestone is next Monday" and converts it into the corresponding text data.
[0252] Step 3:
[0253] Server: Inputs the converted text data into a natural language processing (NLP) algorithm to generate a summary. For example, it summarizes the sentence "The next milestone is next Monday" as "Milestone: Monday".
[0254] Step 4:
[0255] Server: The generated text data is input into the emotion engine, which analyzes the tone and tempo of the voice to recognize the user's emotions. For example, "The next milestone is next Monday" is tagged with the emotion "joy."
[0256] Step 5:
[0257] Server: Compares summarized text data and sentiment data with past meeting minutes data stored in the database to detect delays or discrepancies in progress. For example, it might detect if a milestone previously set for "this Friday" has been postponed to "next Monday."
[0258] Step 6:
[0259] Server: Analyzes comparison results and sentiment data to identify delays and discrepancies in progress and assess their importance. For example, delays in progress are displayed along with the emotion of "joy," indicating positive progress.
[0260] Step 7:
[0261] Server: Visualizes detected results and sentiment data and displays them on a dashboard. For example, areas with delays are shown in red, and sentiment data is indicated by sentiment icons (e.g., a smile).
[0262] Step 8:
[0263] Terminal: Receives visualization data sent from the server and displays it on the user's screen in real time. If significant progress delays or discrepancies are detected, it notifies the user and displays the details along with sentiment data.
[0264] Step 9:
[0265] User: Review the latest meeting minutes, progress, and sentiment data displayed on the device. The user considers and takes appropriate action as needed.
[0266] For example, consider a situation where the progress of new feature A is being discussed in a weekly project meeting:
[0267] 1. Terminal: Captures meeting audio in real time and sends it to the server.
[0268] 2. Server: Converts audio data into text data.
[0269] 3. Server: Generates summaries using NLP algorithms.
[0270] 4. Server: The emotion engine analyzes the emotion of "joy" and assigns it to the summary. This will result in "New Feature A: Proceeding to Test Phase (Joy)".
[0271] 5. Server: Compare the current meeting minutes with those stored in the database to assess whether progress is on schedule. Consider sentiment data to evaluate the importance of any delays.
[0272] 6. Server: Visualize the results on a dashboard and notify the user.
[0273] 7. Device: The results are displayed on the user's device, allowing them to check the progress and sentiment data.
[0274] 8. User: Respond to project progress based on the results.
[0275] This integrates automated summarization of meeting content, progress evaluation, and consideration of sentiment data, enabling effective project management.
[0276] (Example 2)
[0277] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0278] Traditional meeting minutes systems simply transcribe meeting audio into text, failing to consider the meeting's progress or emotional fluctuations. Furthermore, they lacked automated summarization and comparison with past minutes, requiring users to perform these tasks manually, which was extremely time-consuming. Additionally, the lack of technology to analyze user emotions and reflect them in meeting content limited their ability to improve team engagement and project success rates.
[0279] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0280] In this invention, the server includes means for recording meeting audio, means for converting the recorded audio data into text data, means for summarizing the converted text data, means for comparing the summarized text data with past recorded data, means for detecting delays or discrepancies in progress based on the comparison results, means for visualizing the comparison results, means for notifying the user of the results, means for analyzing the user's emotions from the audio data, and means for adding the analyzed emotion data to the text data. This enables a comprehensive analysis including detailed progress evaluation of meetings and fluctuations in emotions, thereby improving the efficiency of project management and strengthening team engagement.
[0281] "Means for recording meeting audio" refers to devices or programs that capture audio during a meeting in real time and save it as digital data.
[0282] "Means for converting recorded audio data into text data" refers to speech recognition technologies and software that analyze audio signals and recognize and convert them into text information.
[0283] "Means for summarizing converted text data" refers to natural language processing techniques that analyze text data to extract important information and shorten the original text to create a summary.
[0284] The means for "comparing the summarized text data with past record data" is an algorithm for collating the generated summary data with existing meeting record data and performing progress and match / mismatch analysis.
[0285] The means for "detecting progress delays and discrepancies based on the comparison results" is an analysis system for identifying the gap between the expected progress situation and the reality from the comparison results and detecting delays and problems.
[0286] The means for "visualizing the comparison results" is a program or tool for displaying the comparison results and progress information in a visual form such as graphs or charts in a form that is easy for users to understand.
[0287] The means for "notifying the user of the results" is a program or device with a notification function for notifying the user of the analysis results and important information in real time.
[0288] The means for "analyzing the user's emotions from voice data" is a technology for judging the user's emotional state from the intonation, tempo, volume, etc. of the voice and classifying it into emotion categories (e.g., joy, anger, sadness, surprise, etc.).
[0289] The means for "attaching the analyzed emotion data to the text data" is a system for attaching the analyzed emotion information as part of the corresponding text data to明示 the emotional state.
[0290] The present invention combines a technology for recognizing the user's emotions from voice data in addition to a system that records the voice of a meeting in real time, automatically converts its content into text and summarizes it, and further evaluates the progress situation by comparing it with past minutes data. The embodiments for implementing the present invention will be described in detail.
[0291] Voice recording and conversion
[0292] Terminal:
[0293] When a meeting begins, audio data is captured through the terminal's microphone. The captured audio data is sent to the server in real time. The hardware used is either the terminal's built-in microphone or an external microphone.
[0294] server:
[0295] The server inputs the received audio data into a speech recognition engine (e.g., Google Cloud Speech-to-Text) and converts it into text data. It analyzes the audio signal and converts the meeting content into text format.
[0296] Summary generation using natural language processing
[0297] server:
[0298] Once the text data is generated, the server analyzes it using natural language processing (NLP) algorithms (e.g., Google Cloud Natural Language API). The NLP algorithms extract key points and keywords, generating a summary of the text data. This summarization process provides a concise summary of the meeting content.
[0299] User emotion recognition by an emotion engine
[0300] server:
[0301] Furthermore, the server inputs the audio data into an emotion analysis engine (e.g., Amazon Comprehend) to recognize the user's emotions. The emotion analysis engine analyzes the user's emotions (e.g., joy, anger, sadness, surprise, etc.) from the audio signal and adds them to the text data.
[0302] Comparison with past meeting minutes and progress evaluation
[0303] server:
[0304] The generated summary data and sentiment data are compared with the past minutes data stored in a database (e.g., Amazon RDS). Through a comparison algorithm, any delays in progress or discrepancies with the plan are detected. Based on the sentiment data, the importance of the progress delays or discrepancies is also evaluated.
[0305] Visualization and Notification
[0306] Server:
[0307] The detected results and sentiment data are reflected on a dashboard that is visually displayed. This allows users to quickly check the progress status and its emotional impact at a glance. For example, progress delays are shown in the form of graphs or charts, and the sentiment data is indicated by colors or icons.
[0308] Terminal:
[0309] The results sent from the server are updated in real-time on the terminal. When important progress delays or discrepancies are detected, notifications are popped up along with the sentiment data.
[0310] Specific Example
[0311] For example, assume that during a weekly project meeting, discussions are held regarding the progress of new feature A. If someone says during the meeting, "The development of new feature A is on track and has entered the testing phase":
[0312] 1. Terminal: Capture the voice and send it to the server.
[0313] 2. Server: Convert the voice data into text data using Google Cloud Speech-to-Text.
[0314] 3. Server: Generate a summary using the NLP algorithm with Google Cloud Natural Language API.
[0315] 4. Server: Use Amazon Comprehend to analyze the emotion of "joy" and assign it to the summary. For example, assign it as "New Feature A: Progressing to the testing phase (joy)".
[0316] 5. Server: Compare the current meeting minutes with those stored in Amazon RDS to assess whether progress is on schedule. Evaluate the importance of any delays, including sentiment data.
[0317] 6. Server: Visually display the results on a dashboard and notify the user.
[0318] 7. Device: Displayed on the user's device, allowing them to check progress and sentiment data.
[0319] 8. User: Based on the results, take appropriate action regarding the project's progress.
[0320] Example of a prompt
[0321] "Recognize and summarize in text the emotions you would feel when reporting on the progress of new feature A at this week's project meeting."
[0322] Through these concrete actions, the overall system becomes easier to understand, leading to more effective project management and progress evaluation.
[0323] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0324] Step 1: Capture and send audio data
[0325] Terminal: When the meeting begins, audio data is captured through the terminal's microphone. The input is the meeting audio, which is captured as audio data. The captured audio data is sent to the server in real time. Specifically, the terminal's microphone converts the audio signal into digital data and sends it to the server using the HTTPS protocol. The output is the audio data sent to the server.
[0326] Step 2: Converting audio data to text
[0327] Server: The server inputs the received audio data into the Google Cloud Speech-to-Text speech recognition engine. The input is the received audio data, which the server parses and converts into text data. Specifically, the server sends the audio data to the speech recognition engine via an API call and receives the text data extracted from the audio signal. The output is the generated text data.
[0328] Step 3: Summary Generation
[0329] Server: The server inputs the generated text data into the Google Cloud Natural Language API. The input is the generated text data, and the server uses NLP algorithms to extract important points and keywords and summarize the text data. Specifically, the server applies an algorithm that analyzes the text data, extracts the important parts, and generates a summary. The output is the summarized text data.
[0330] Step 4: Emotion Recognition
[0331] Server: The server inputs the summarized text data and the original audio data into Amazon Comprehend's sentiment analysis engine. The input consists of summarized text data and the original audio data, and the sentiment analysis engine analyzes the user's emotions from the audio signal. Specifically, the sentiment analysis engine analyzes the audio data, identifies emotional states (e.g., joy, anger, sadness, surprise, etc.), and adds that information to the summarized text. The output is the summarized text with the sentiment data added.
[0332] Step 5: Comparison with past data and progress evaluation
[0333] Server: The server compares the generated sentiment-data-enhanced summary text with historical meeting minutes data stored in Amazon RDS. The input is the sentiment-data-enhanced summary text and historical meeting minutes data, and the server compares this data to detect delays in progress and discrepancies with the plan. Specifically, the server retrieves historical data and applies a comparison algorithm to detect progress and discrepancies. The output is the evaluation result of progress and discrepancies.
[0334] Step 6: Visualize and notify results
[0335] Server: The server reflects the detected progress and discrepancy evaluation results, as well as sentiment data, in a dashboard for visual display. Inputs are progress and discrepancy evaluation results and sentiment data, which are then visualized on the dashboard in graph and chart format. Specifically, the server generates evaluation results as graphs and charts and displays them to the user via a web interface. Output is the visualized information on the dashboard.
[0336] Terminal: The terminal updates the results sent from the server in real time. The input consists of progress and discrepancy evaluation results and sentiment data sent from the server, which are displayed in real time as notification pop-ups. Specifically, the terminal's notification system receives new data and displays a pop-up notification to the user. The output is the pop-up notification.
[0337] Through these steps, the system can perform detailed progress evaluations and sentiment analysis of meeting content. This allows users to understand project progress in real time and take necessary actions quickly.
[0338] (Application Example 2)
[0339] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0340] Current factory automation systems struggle to record communication between robots and workers, and to track work progress. Furthermore, progress management doesn't consider worker emotions, potentially leading to decreased efficiency and lower worker morale. To address these issues, a system is needed that analyzes and evaluates progress and worker emotions in real time, providing appropriate feedback.
[0341] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0342] In this invention, the server includes means for recording instructions between robots and workers in a factory, means for recognizing user emotions based on voice data, means for incorporating the emotion data into progress evaluation, and means for displaying the voice data and emotion data on a dashboard. This enables real-time progress management and emotion evaluation of robots and workers in a factory.
[0343] "Methods for recording meeting audio" refers to technologies that collect audio information during a meeting using microphones or other audio capture devices and store it as digital data.
[0344] "Means of converting recorded audio data into text data" refers to the process of converting audio data into text format data using conversion technologies such as a speech recognition engine.
[0345] "Methods for summarizing converted text data" refer to techniques that utilize natural language processing (NLP) algorithms to extract key points from large amounts of text data and condense them into short summaries.
[0346] "Means for comparing the summarized text data with past meeting minutes data" refers to a technique that compares newly generated summarized data with previously stored data to identify progress and changes.
[0347] "Means for detecting delays or discrepancies in progress based on the aforementioned comparison results" refers to a mechanism that analyzes the comparison results and evaluates whether the work or project is progressing according to plan.
[0348] "Means for visualizing the comparison results" refers to technologies that display the comparison results in a visual format, such as graphs or charts, so that users can easily understand them.
[0349] "Means for notifying the user of the results" refers to a notification function that communicates detected progress delays, discrepancies, and other important information to the user in real time.
[0350] "A means of recognizing user emotions based on voice data" refers to an emotion recognition engine that analyzes and recognizes human emotions (joy, anger, sadness, surprise, etc.) from voice.
[0351] "Means for incorporating the aforementioned emotional data into progress evaluations" refers to a system that incorporates users' emotional data into work and project progress evaluations to provide a more comprehensive assessment.
[0352] "A means of recording instructions between robots and workers in a factory and evaluating progress and emotions" refers to a technology that records instructions exchanged between robots and workers in a factory and uses that data to evaluate the progress of the work and the emotions of the workers.
[0353] "Means for displaying audio data and sentiment data on a dashboard" refers to technology that provides a dashboard function for visually displaying collected audio data and sentiment data.
[0354] The embodiments for carrying out this invention will be described in detail. A system for recording instructions between a robot and a worker in a factory and evaluating progress and worker emotions consists of the following steps.
[0355] First, the robots in the factory record voice instructions from workers via microphones. The recorded voice data is sent to a server in real time. The server uses a speech recognition engine (e.g., Google Speech Recognition API) to convert this voice data into text data.
[0356] The converted text data is analyzed using natural language processing (NLP) algorithms to extract key points and generate short summaries. These summaries are compared to past work records to assess progress. This helps determine whether the work is on schedule or behind schedule. Simultaneously, the server inputs audio data into an emotion recognition engine to analyze the worker's emotions. This allows the worker's emotions (e.g., joy, anger, sadness, surprise) to be added to the text data.
[0357] This data is stored in a database and used to evaluate progress and discrepancies by comparing it with past work records. The evaluation results are visually displayed on a visualization dashboard, allowing users to grasp the progress and sentiment status at a glance. For example, delays in progress are displayed in graph or chart format, and sentiment data is indicated by colors and icons. If significant delays or discrepancies are detected, a notification pops up on the user's device along with the sentiment data.
[0358] Hardware and software to be used
[0359] The hardware used will be a microphone (which can be built into a smartphone or smart glasses, etc.). The software will include Python, Google Speech Recognition API, NLTK, scikit-learn, a dedicated emotion recognition engine, and a web application for displaying the dashboard (e.g., Dash, Flask).
[0360] Specific example
[0361] For example, consider a scenario in a factory where a robot operates under the instructions of a worker. When the worker says, "I'm ready to proceed to the next work stage," the following process is executed:
[0362] Voice recording: The robot captures the worker's voice and sends it to the server.
[0363] Text Conversion: The Google Speech Recognition API is used to convert the audio data into text data. The resulting text will read, "Ready to proceed to the next work stage."
[0364] Summary generation: Generates summaries using natural language processing algorithms.
[0365] Emotion Analysis: An emotion recognition engine analyzes the emotion of "joy" from the voice.
[0366] Progress evaluation: Evaluate whether progress is "good" by comparing it with past work record data.
[0367] Dashboard display: The generated summary, sentiment data, and progress assessment are displayed on a dashboard for visual confirmation.
[0368] Example of a prompt
[0369] A factory robot receives the instruction, "Ready to proceed to the next work stage." Translate this instruction into text, summarize it, and evaluate the robot's work progress and emotional state.
[0370] Thus, the present invention's system efficiently records and manages communication between robots and workers in a factory, and simultaneously evaluates work progress and emotions, thereby realizing an efficient and harmonious work environment.
[0371] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0372] Step 1:
[0373] Audio recording
[0374] The terminal uses microphones within the factory to record voice instructions from workers in real time. The terminal then transmits the recorded voice data to a server.
[0375] Input: Voice instructions from the worker
[0376] Output: Audio data
[0377] Step 2:
[0378] Text conversion
[0379] The server inputs the received audio data into a speech recognition engine (e.g., Google Speech Recognition API) and converts the audio data into text data.
[0380] Input: Audio data
[0381] Output: Text data
[0382] Step 3:
[0383] Summary generation
[0384] The server analyzes the converted text data using natural language processing (NLP) algorithms, extracts key points, and generates a short summary.
[0385] Input: Text data
[0386] Output: Summary data
[0387] Step 4:
[0388] sentiment analysis
[0389] The server inputs the voice data into an emotion recognition engine and recognizes the worker's emotions (e.g., joy, anger, sadness, surprise, etc.). The recognized emotion data is then added to the text data.
[0390] Input: Audio data
[0391] Output: Sentiment data
[0392] Step 5:
[0393] Progress evaluation
[0394] The server compares the generated summary data and sentiment data with past work record data stored in the database to assess delays and discrepancies. It also evaluates the importance of delays and discrepancies based on the sentiment data.
[0395] Input: Summary data, sentiment data, past work record data
[0396] Output: Progress evaluation data
[0397] Step 6:
[0398] visualization
[0399] The server converts progress evaluation data and sentiment data into a visual format (e.g., graphs and charts) for display on the dashboard.
[0400] Input: Progress evaluation data, sentiment data
[0401] Output: Visualization data for dashboard
[0402] Step 7:
[0403] notification
[0404] If the server detects any significant delays or discrepancies in the project, it will notify the user's terminal in real time.
[0405] Input: Progress evaluation data, sentiment data
[0406] Output: User notification
[0407] These steps enable the system to record the collaboration between robots and workers in the factory, and to monitor and evaluate progress and emotions in real time.
[0408] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0409] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0410] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0411] [Second Embodiment]
[0412] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0413] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0414] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0415] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0416] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0417] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0418] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0419] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0420] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0421] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0422] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0423] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0424] The present invention provides a system that records meeting audio in real time, automatically transcribes and summarizes the content, and further evaluates the progress by comparing it with past meeting minutes data. The embodiments for carrying out the present invention are described in detail below.
[0425] Audio recording and conversion
[0426] Terminal:
[0427] When the meeting begins, audio data is captured via the microphone on the terminal. The captured audio data is sent to the server in real time.
[0428] server:
[0429] The server inputs the received audio data into a speech recognition engine and converts it into text data. A highly accurate speech recognition engine (e.g., speech recognition software or a cloud API) is used.
[0430] Summary generation using natural language processing
[0431] server:
[0432] Once the text data is generated, the server analyzes it using natural language processing (NLP) algorithms. The NLP algorithms extract key points and keywords, generating a summary of the text data. This summarization process provides long meeting minutes in a concise, summarized form.
[0433] Comparison with past meeting minutes and progress evaluation
[0434] server:
[0435] The generated summary data is compared with past meeting minutes data stored in the database. The comparison algorithm detects delays in progress and discrepancies with the plan. Specifically, the current status is compared with milestones and plans in past meeting minutes.
[0436] Visualization and notification
[0437] server:
[0438] The detected results are reflected in a visually displayed dashboard, allowing users to see the progress at a glance. For example, delays in progress can be displayed in graph or chart format, making it easy for users to understand.
[0439] Terminal:
[0440] Results sent from the server are updated in real time on the terminal. If significant delays or discrepancies are detected, the terminal will notify the user and display the details.
[0441] Specific example
[0442] For example, suppose a weekly project meeting is discussing the progress of new feature A. If, during the meeting, someone says, "Development of new feature A is on track and we've entered the testing phase":
[0443] 1. Terminal: Captures audio and sends it to the server.
[0444] 2. Server: Converts the audio data into text and generates the text, "Development of new feature A is on track and has entered the testing phase."
[0445] 3. Server: Use an NLP algorithm to summarize and conclude, "New Feature A: Proceeding to Test Phase."
[0446] 4. Server: Compare the current meeting minutes with those stored in the database to assess whether progress is on schedule.
[0447] 5. Server: Visualize the progress on a dashboard and notify users.
[0448] This process allows users to immediately check the progress after the meeting ends and take prompt action. The system of this invention enables companies and project teams to achieve efficient and effective meeting management and project management.
[0449] The following describes the processing flow.
[0450] Step 1:
[0451] Terminal: As soon as the meeting starts, the microphone on the terminal captures audio data. The captured audio data is sent to the server in real time.
[0452] Step 2:
[0453] Server: Passes the received audio data to the speech recognition engine, which converts it into text data. The speech recognition engine analyzes the audio signal and transcribes it into sentences such as, "The next milestone is next Monday."
[0454] Step 3:
[0455] Server: Inputs the converted text data into a natural language processing (NLP) algorithm to generate a summary. For example, it summarizes the sentence "The next milestone is next Monday" as "Milestone: Monday".
[0456] Step 4:
[0457] Server: Compares the generated summary data with past meeting minutes data stored in the database. Uses a comparison algorithm to detect delays and discrepancies in progress.
[0458] Step 5:
[0459] Server: Analyzes comparison results to identify significant delays and discrepancies. The identified results are displayed on a dashboard in graph and chart format using visualization tools.
[0460] Step 6:
[0461] Terminal: Receives visualized data sent from the server and displays it on the user's screen in real time. If there are significant delays or discrepancies in progress, a notification will pop up.
[0462] Step 7:
[0463] User: Check the latest meeting minutes and progress displayed on the terminal. The user will consider and take appropriate action regarding any delays or discrepancies in progress.
[0464] To explain with a concrete example, it would be as follows:
[0465] In the weekly project progress meeting:
[0466] 1. Device: Captures meeting audio in real time.
[0467] 2. Server: Converts audio data into text data.
[0468] 3. Server: Generates summaries using NLP algorithms.
[0469] 4. Server: Compares with past meeting minutes data and detects discrepancies.
[0470] 5. Server: Visualize the results on a dashboard.
[0471] 6. Terminal: Displays results and important notifications on the user's terminal.
[0472] 7. User: Review the results and respond to the project's progress.
[0473] In this way, the system enables automatic summarization of meeting content, progress monitoring, and important notifications at each step.
[0474] (Example 1)
[0475] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0476] Traditional meeting recording systems required manually converting audio data to text and summarizing its contents, which was time-consuming and labor-intensive. Furthermore, it was difficult to evaluate progress by comparing current meeting minutes with past ones. This not only hindered prompt responses after meetings but also made it difficult to detect project delays and discrepancies early on.
[0477] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0478] In this invention, the server includes means for converting received audio data into text data using a speech recognition engine, means for summarizing the text data using a natural language processing algorithm, and means for comparing the summarized text data with past meeting minutes data. This makes it possible to automatically generate accurate meeting records in real time and to quickly perform summarization and progress evaluation.
[0479] 1. "Means for capturing meeting audio" refers to a hardware or software system for collecting audio during a meeting in digital format.
[0480] 2. "Means for transmitting captured audio data in real time" refers to communication means for transmitting collected audio data to a server without delay.
[0481] 3. "Means for converting received audio data into text data using a speech recognition engine" refers to algorithms or software for analyzing audio data and converting it into corresponding text data.
[0482] 4. "Means for summarizing text data using natural language processing algorithms" refers to methods for analyzing generated text data, extracting important information, and summarizing it concisely.
[0483] 5. "Means for comparing summarized text data with past meeting minutes data" refers to methods for comparing newly generated summarized text with past meeting minutes to evaluate progress and detect discrepancies.
[0484] 6. "Means for detecting delays or discrepancies in progress based on comparison results" refers to a mechanism that analyzes the comparison results of text data and automatically identifies differences from the plan or delays in progress.
[0485] 7. "Means for visualizing comparison results" refers to display methods in the form of graphs or charts that allow users to intuitively understand the analyzed comparison results.
[0486] 8. "Means for notifying users of visualized results" refers to a notification system that informs users in real time of important progress and detected discrepancies.
[0487] 9. "Means for converting audio data to text in real time using parallel processing" refers to a method that efficiently utilizes computing resources to quickly convert audio data to text by simultaneously executing multiple audio analysis tasks.
[0488] 10. "Means of using generative AI models in natural language processing algorithms" refers to advanced algorithms that utilize generative AI models to analyze and summarize text data.
[0489] This invention provides a system that records meeting audio in real time, automatically transcribes and summarizes the content, and evaluates progress by comparing it with past meeting minutes data. The following describes how to implement this system in detail.
[0490] Audio recording and conversion
[0491] Terminal:
[0492] When the meeting begins, audio data is captured via the microphone on the terminal. The captured audio data is sent to the server in real time. The audio data is saved to a buffer at regular intervals, and when the buffer size reaches a certain level, it is sent back to the server.
[0493] server:
[0494] The server inputs the received audio data into a speech recognition engine (e.g., Google Cloud Speech-to-Text API) and converts it into text data. The text conversion speed is optimized by dividing the audio data into small chunks and executing API requests in parallel.
[0495] Summary generation using natural language processing
[0496] server:
[0497] Once text data is generated, the server analyzes it using a natural language processing (NLP) algorithm (e.g., OpenAI's generative AI model). The NLP algorithm extracts key points and keywords, generating a summary of the text data. The following is an example of a prompt: "Use an NLP algorithm to summarize the following text: Text: 'Development of new feature A is on track and has entered the testing phase. The development team is progressing as planned and will report the test results at the next meeting.'"
[0498] Comparison with past meeting minutes and progress evaluation
[0499] server:
[0500] The generated summary data is compared with historical meeting minutes data stored in the database. A comparison algorithm detects delays in progress and discrepancies with the plan. Specifically, historical data is retrieved using SQL queries, and a Python script detects matching keywords and progress differences with the summary data.
[0501] Visualization and notification
[0502] server:
[0503] The detected results are reflected in a visually displayed dashboard, allowing users to check the progress at a glance. Progress delays are displayed in graph and chart format using D3.js, etc.
[0504] Terminal:
[0505] Results sent from the server are updated in real time on the device. If significant delays or discrepancies are detected, the device notifies the user and displays the details. A notification pop-up appears in the browser, and clicking it jumps to a detailed screen.
[0506] Specific example
[0507] For example, suppose a weekly project meeting is discussing the progress of new feature A. If, during the meeting, someone says, "Development of new feature A is on track and we've entered the testing phase":
[0508] 1. Terminal: Captures audio and sends it to the server.
[0509] 2. Server: The server inputs the received audio data into the speech recognition engine and generates the text, "Development of new feature A is on track and has entered the testing phase."
[0510] 3. Server: Use an NLP algorithm to summarize and conclude, "New Feature A: Proceeding to Test Phase."
[0511] 4. Server: Compare the current meeting minutes with those stored in the database to assess whether Feature A is progressing as planned. For example, "Previous meeting: Design phase scheduled to be completed," "Current meeting: Proceeding to the testing phase." Assess that it is progressing as planned.
[0512] 5. Server: Display "New feature A is progressing as planned" on the dashboard and notify the user.
[0513] This process allows users to immediately check the progress after the meeting ends and take prompt action.
[0514] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0515] Step 1: Capture and send audio data
[0516] Terminal:
[0517] Input: Capture conference audio via the microphone on the device.
[0518] Data processing: Audio data is saved to a buffer at regular intervals.
[0519] Output: Once the buffer size reaches a certain level, it is sent to the server in real time.
[0520] Specific operation: The terminal captures audio data, and when it reaches a certain buffer size, it sends the data to the server. A low-latency communication protocol is used for this process.
[0521] Step 2: Receiving audio data and converting it to text.
[0522] server:
[0523] Input: Audio data sent from the device.
[0524] Data processing: Divide audio data into small chunks.
[0525] Data processing: Convert audio data into text data using a speech recognition engine (e.g., a cloud-based API).
[0526] Output: Converted real-time text data.
[0527] Specific operation: The server subdivides the received audio data and passes it to the speech recognition engine. Parallel processing is used to perform efficient text conversion.
[0528] Step 3: Text-based summary generation
[0529] server:
[0530] Input: Text data.
[0531] Data processing: Analyze text data using natural language processing algorithms.
[0532] Data processing: Extract and summarize key points using a generative AI model (e.g., OpenAI GPT-3).
[0533] Output: Summarized text data.
[0534] Specific operation: The server inputs a prompt message into the generation AI model and generates a summary of the text data. Example of a prompt message: "Use an NLP algorithm to summarize the following text. Text: 'Development of new feature A is on track and has entered the testing phase. The development team is progressing as planned and will report the test results at the next meeting.'"
[0535] Step 4: Comparison of summary data and past meeting minutes
[0536] server:
[0537] Input: Summarized text data.
[0538] Data processing: Retrieve past meeting minutes data from the database.
[0539] Data processing: Use a comparison algorithm to match summary data with past meeting minutes.
[0540] Output: Progress evaluation and comparison results of discrepancies.
[0541] Specific operation: The server retrieves past meeting minutes using SQL queries, and a Python script detects similarities and differences between the two sets of data.
[0542] Step 5: Visualizing progress
[0543] server:
[0544] Input: Results of progress evaluations and discrepancies comparisons.
[0545] Data processing: Formatting data for visualization.
[0546] Data processing: Visualization in graph and chart formats (e.g., D3.js).
[0547] Output: The visualization results displayed on the dashboard.
[0548] Specific operation: The server sends visualization data in JSON format to the dashboard API and updates it in real time.
[0549] Step 6: Notifications and detailed display
[0550] Terminal:
[0551] Input: Visualization results sent from the server.
[0552] Data processing: Formatting the data into a form that is easy for users to understand.
[0553] Output: Real-time notifications and detailed display on the device.
[0554] Specific operation: The device displays a notification pop-up, and clicking it jumps to a details screen. The user can easily check the progress.
[0555] These steps enable the system to transcribe, summarize, assess progress, visualize meeting audio in real time, and notify users.
[0556] (Application Example 1)
[0557] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0558] Modern factory operations demand complex project management and rapid response. However, managing meeting audio and progress is often done manually, which is inefficient and makes it difficult to immediately detect delays or discrepancies. This leads to delays in appropriate responses and decreased productivity. Therefore, there is a need for a system that automatically records meeting audio, transcribes it into text, summarizes it, and evaluates progress in real time.
[0559] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0560] In this invention, the server includes means for recording meeting audio, means for converting the recorded audio data into text data, means for summarizing the converted text data, means for comparing it with past meeting minutes data, means for detecting delays or discrepancies in progress, means for visualizing the comparison results, means for notifying the user of the results, means for a robot placed in the factory to record and process the meeting audio, and means for displaying the converted and summarized text data on a dashboard in real time. This enables the automatic recording of audio content of meetings held in the factory, improving the efficiency of progress management and allowing for quicker responses.
[0561] "Means of recording meeting audio" refers to a system that saves the audio from a meeting as digital data using a microphone or similar device.
[0562] "Means of converting recorded audio data into text data" refers to the process of converting recorded audio into text using a speech recognition engine.
[0563] "Methods for summarizing converted text data" refers to the process of using natural language processing techniques to extract only the important points from long texts and make them concise.
[0564] "Methods for comparing with past meeting minutes data" refer to a system that compares the content of the previous meeting stored in a database with the content of the current meeting to check for differences and progress.
[0565] "Methods for detecting delays or discrepancies in progress" refer to a system that compares past meeting minutes data with current data to identify differences between the plan and the actual progress.
[0566] "Methods for visualizing comparison results" refer to the process of displaying data differences and progress in visual formats such as graphs and charts.
[0567] "Means of notifying users of results" refers to a system that informs users of any delays or discrepancies in progress.
[0568] "Means for recording and processing meeting audio using robots deployed in a factory" refers to a process in which robots installed within a factory capture meeting audio and process that audio data in real time.
[0569] "A means of displaying real-time converted and summarized text data on a dashboard" refers to a mechanism that displays real-time converted and summarized text data on a dashboard so that users can understand it at a glance.
[0570] This invention is a system for automating and streamlining the progress management of meetings held in a factory. It records meeting audio, transcribes and summarizes the content, and evaluates the progress by comparing it with past meeting minutes data. Furthermore, this system has a function to visualize the progress in real time and notify the user.
[0571] Audio recording and conversion
[0572] Hardware: Robots deployed in the factory are equipped with high-sensitivity microphones to capture meeting audio in real time. This allows the system to begin collecting audio data as soon as the meeting starts.
[0573] Software: The collected audio data is sent to a server and converted into text data by a speech recognition engine (e.g., Google Web Speech API). This process is done in real time, so speech during a meeting is instantly transcribed into text.
[0574] Summary generation using natural language processing
[0575] Server: Once text data is generated, the server uses natural language processing (NLP) algorithms (e.g., the Hugging Face generative AI model) to summarize the text data. The summarized text contains key points, allowing users to efficiently grasp the information.
[0576] Comparison with past meeting minutes and progress evaluation
[0577] Server: The generated summary data is compared with past meeting minutes data stored in the database. The comparison algorithm detects any delays or discrepancies in progress. This ensures a clear understanding of the current progress against milestones and plans decided in past meetings.
[0578] Visualization and notification
[0579] Server: Detected comparison results are reflected in a visually displayed dashboard. Progress is shown in graph and chart format, allowing users to see the current progress and any issues at a glance. Furthermore, if significant delays or discrepancies are detected, the server will notify the user.
[0580] Specific example: For instance, when reporting on the progress of new feature A in a project meeting at a factory, it would look like this:
[0581] 1. The robot captures the meeting audio and sends it to the server.
[0582] 2. The server converts the audio data into text, specifically generating the text "New feature A has progressed from the development stage to the testing stage."
[0583] 3. Use an NLP algorithm to summarize the text, concluding it as "New Feature A: Testing Phase".
[0584] 4. Compare the current server status with past meeting minutes stored in the database to assess whether it is progressing according to plan.
[0585] 5. Progress is visualized on a dashboard, allowing users to check it in real time. If progress is behind schedule, users can immediately consider countermeasures.
[0586] A possible example of a specific prompt message would be something like, "We have received a progress report on Project X. You need to start on a new task, Task Y."
[0587] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0588] Step 1:
[0589] Terminal: A robot positioned in the factory captures meeting audio and transmits it to a server in real time. The input data is audio data, and the output is audio data transferred to the server. Specifically, the robot's microphone captures the audio and transmits that audio data to the server via the network.
[0590] Step 2:
[0591] Server: Inputs received audio data into a speech recognition engine and converts it into text data. Input data is audio data, and output data is text data. Specifically, the server calls a speech recognition engine (e.g., Google Web Speech API) and converts the audio data into text format.
[0592] Step 3:
[0593] Server: Summarizes the generated text data using a natural language processing (NLP) algorithm. The input data is the transformed text data, and the output data is the summarized text data. Specifically, the server uses the Hugging Face generative AI model to extract important information from long texts and generate summaries.
[0594] Step 4:
[0595] Server: Compares summarized text data with past meeting minutes data stored in the database. The input data consists of summarized text data and past meeting minutes data, and the output data is the comparison result. Specifically, the server retrieves meeting minutes data from the past database, compares it with the current summarized data, and calculates the difference.
[0596] Step 5:
[0597] Server: Detects delays and discrepancies in progress based on comparison results. Input data is the comparison results, and output data is information about delays and discrepancies in progress. Specifically, the server uses an algorithm to evaluate the progress status and identify the gap with the plan.
[0598] Step 6:
[0599] Server: Reflects detected results on a dashboard for visual display. Input data is information about delays and discrepancies in progress, and output data is visualized graphs and charts. Specifically, the server uses the matplotlib library to graphically display the progress.
[0600] Step 7:
[0601] Server: Sends generated visualization data to the terminal and notifies the user. Input data is visualized graphs and charts, and output data is progress displayed on the user interface. Specifically, the server pushes the generated visualization data to the terminal in real time and notifies the user of important difference information.
[0602] Step 8:
[0603] User: Check the progress on the dashboard on the device and decide on countermeasures as needed. The input data is the displayed progress, and the output is the decision on the appropriate countermeasure. Specifically, the user accesses the dashboard on the device and plans the next action based on the visualized progress.
[0604] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0605] This invention combines a system that records meeting audio in real time, automatically transcribes and summarizes its contents, and evaluates progress by comparing it with past meeting minutes data, with an emotion engine that recognizes the user's emotions from the audio data. The embodiments for carrying out this invention will be described in detail below.
[0606] Audio recording and conversion
[0607] Terminal:
[0608] When the meeting begins, audio data is captured via the microphone on the terminal. The captured audio data is sent to the server in real time.
[0609] server:
[0610] The server inputs the received audio data into the speech recognition engine, which converts it into text data. The speech recognition engine analyzes the audio signal and converts the meeting content into text format.
[0611] Summary generation using natural language processing
[0612] server:
[0613] Once the text data is generated, the server analyzes it using natural language processing (NLP) algorithms. The NLP algorithms extract key points and keywords, generating a summary of the text data. This summarization process provides a concise summary of the meeting's content.
[0614] User emotion recognition by an emotion engine
[0615] server:
[0616] Furthermore, the server inputs the voice data into the emotion engine to recognize the user's emotions. The emotion engine analyzes the user's emotions (e.g., joy, anger, sadness, surprise, etc.) from the voice signal and adds them to the text data.
[0617] Comparison with past meeting minutes and progress evaluation
[0618] server:
[0619] The generated summary data and sentiment data are compared with historical meeting minutes data stored in the database. A comparison algorithm detects delays and discrepancies with the plan. Based on the sentiment data, the importance of the delays and discrepancies is also evaluated.
[0620] Visualization and notification
[0621] server:
[0622] The detected results and sentiment data are reflected in a visually displayed dashboard. This allows users to see the progress and its emotional impact at a glance. For example, progress delays are displayed in graphs and charts, and sentiment data is indicated by colors and icons.
[0623] Terminal:
[0624] Results sent from the server are updated in real time on the device. If significant delays or discrepancies are detected, a notification will pop up along with sentiment data.
[0625] Specific example
[0626] For example, suppose a weekly project meeting is held to discuss the progress of new feature A. If, during the meeting, it is stated, "Development of new feature A is on track and has entered the testing phase":
[0627] 1. Terminal: Captures audio and sends it to the server.
[0628] 2. Server: Converts audio data into text data.
[0629] 3. Server: Generates summaries using NLP algorithms.
[0630] 4. Server: The emotion engine analyzes the emotion of "joy" and assigns it to the summary. For example, it might be assigned as "New Feature A: Progressing to Test Phase (Joy)".
[0631] 5. Server: Compare the current meeting minutes with those stored in the database to assess whether progress is on schedule. Evaluate the importance of any delays, including sentiment data.
[0632] 6. Server: Visualize the results on a dashboard and notify the user.
[0633] 7. Device: Displayed on the user's device, allowing them to check progress and sentiment data.
[0634] 8. User: Based on the results, take appropriate action regarding the project's progress.
[0635] By combining this with an emotion engine, it becomes possible to not only automatically summarize meeting content but also to evaluate progress while considering the user's emotions, resulting in more accurate project management. This improves the project success rate and enhances team engagement.
[0636] The following describes the processing flow.
[0637] Step 1:
[0638] Terminal: As soon as the meeting starts, the microphone on the terminal captures audio data. The captured audio data is sent to the server in real time.
[0639] Step 2:
[0640] Server: Inputs received audio data into the speech recognition engine and converts it into text data. For example, it recognizes the audio data "The next milestone is next Monday" and converts it into the corresponding text data.
[0641] Step 3:
[0642] Server: Inputs the converted text data into a natural language processing (NLP) algorithm to generate a summary. For example, it summarizes the sentence "The next milestone is next Monday" as "Milestone: Monday".
[0643] Step 4:
[0644] Server: The generated text data is input into the emotion engine, which analyzes the tone and tempo of the voice to recognize the user's emotions. For example, "The next milestone is next Monday" is tagged with the emotion "joy."
[0645] Step 5:
[0646] Server: Compares summarized text data and sentiment data with past meeting minutes data stored in the database to detect delays or discrepancies in progress. For example, it might detect if a milestone previously set for "this Friday" has been postponed to "next Monday."
[0647] Step 6:
[0648] Server: Analyzes comparison results and sentiment data to identify delays and discrepancies in progress and assess their importance. For example, delays in progress are displayed along with the emotion of "joy," indicating positive progress.
[0649] Step 7:
[0650] Server: Visualizes detected results and sentiment data and displays them on a dashboard. For example, areas with delays are shown in red, and sentiment data is indicated by sentiment icons (e.g., a smile).
[0651] Step 8:
[0652] Terminal: Receives visualization data sent from the server and displays it on the user's screen in real time. If significant progress delays or discrepancies are detected, it notifies the user and displays the details along with sentiment data.
[0653] Step 9:
[0654] User: Review the latest meeting minutes, progress, and sentiment data displayed on the device. The user considers and takes appropriate action as needed.
[0655] As a concrete example, consider a situation where the progress of new feature A is being discussed in a weekly project meeting:
[0656] 1. Terminal: Captures meeting audio in real time and sends it to the server.
[0657] 2. Server: Converts audio data into text data.
[0658] 3. Server: Generates summaries using NLP algorithms.
[0659] 4. Server: The emotion engine analyzes the emotion of "joy" and assigns it to the summary. This will result in "New Feature A: Proceeding to Test Phase (Joy)".
[0660] 5. Server: Compare the current meeting minutes with those stored in the database to assess whether progress is on schedule. Consider sentiment data to evaluate the importance of any delays.
[0661] 6. Server: Visualize the results on a dashboard and notify the user.
[0662] 7. Device: The results are displayed on the user's device, allowing them to check the progress and sentiment data.
[0663] 8. User: Respond to project progress based on the results.
[0664] This integrates automated summarization of meeting content, progress evaluation, and consideration of sentiment data, enabling effective project management.
[0665] (Example 2)
[0666] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0667] Traditional meeting minutes systems simply transcribe meeting audio into text, failing to consider the meeting's progress or emotional fluctuations. Furthermore, they lacked automated summarization and comparison with past minutes, requiring users to perform these tasks manually, which was extremely time-consuming. Additionally, the lack of technology to analyze user emotions and reflect them in meeting content limited their ability to improve team engagement and project success rates.
[0668] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0669] In this invention, the server includes means for recording meeting audio, means for converting the recorded audio data into text data, means for summarizing the converted text data, means for comparing the summarized text data with past recorded data, means for detecting delays or discrepancies in progress based on the comparison results, means for visualizing the comparison results, means for notifying the user of the results, means for analyzing the user's emotions from the audio data, and means for adding the analyzed emotion data to the text data. This enables a comprehensive analysis including detailed progress evaluation of meetings and fluctuations in emotions, thereby improving the efficiency of project management and strengthening team engagement.
[0670] "Means for recording meeting audio" refers to devices or programs that capture audio during a meeting in real time and save it as digital data.
[0671] "Means for converting recorded audio data into text data" refers to speech recognition technologies and software that analyze audio signals and recognize and convert them into text information.
[0672] "Means for summarizing converted text data" refers to natural language processing techniques that analyze text data to extract important information and shorten the original text to create a summary.
[0673] "Means for comparing the summarized text data with past record data" refers to an algorithm for comparing the generated summary data with existing meeting record data and performing analysis of progress, agreements, and disagreements.
[0674] "Means for detecting delays or discrepancies in progress based on the comparison results" refers to an analytical system that identifies the gap between the expected progress and the actual progress based on the comparison results, and detects delays or problems.
[0675] "Means for visualizing the comparison results" refers to programs or tools for displaying comparison results and progress information in a visual format such as graphs or charts that are easy for users to understand.
[0676] "Means for notifying the user of the results" refers to a program or device that has a notification function to inform the user of analysis results and important information in real time.
[0677] "Methods for analyzing user emotions from voice data" refer to technologies that determine a user's emotional state from factors such as voice intonation, tempo, and volume, and classify it into emotional categories (e.g., joy, anger, sadness, surprise, etc.).
[0678] "Means for attaching analyzed emotional data to text data" refers to a system that attaches analyzed emotional information as part of the corresponding text data, thereby clearly indicating the emotional state.
[0679] This invention combines a system that records meeting audio in real time, automatically transcribes and summarizes its contents, and evaluates progress by comparing it with past meeting minutes data, with technology that recognizes user emotions from the audio data. The embodiments for carrying out this invention will be described in detail below.
[0680] Audio recording and conversion
[0681] Terminal:
[0682] When a meeting begins, audio data is captured through the terminal's microphone. The captured audio data is sent to the server in real time. The hardware used is either the terminal's built-in microphone or an external microphone.
[0683] server:
[0684] The server inputs the received audio data into a speech recognition engine (e.g., Google Cloud Speech-to-Text) and converts it into text data. It analyzes the audio signal and converts the meeting content into text format.
[0685] Summary generation using natural language processing
[0686] server:
[0687] Once the text data is generated, the server analyzes it using natural language processing (NLP) algorithms (e.g., Google Cloud Natural Language API). The NLP algorithms extract key points and keywords, generating a summary of the text data. This summarization process provides a concise summary of the meeting content.
[0688] User emotion recognition by an emotion engine
[0689] server:
[0690] Furthermore, the server inputs the audio data into an emotion analysis engine (e.g., Amazon Comprehend) to recognize the user's emotions. The emotion analysis engine analyzes the user's emotions (e.g., joy, anger, sadness, surprise, etc.) from the audio signal and adds them to the text data.
[0691] Comparison with past meeting minutes and progress evaluation
[0692] server:
[0693] The generated summary and sentiment data are compared with historical meeting minutes data stored in a database (e.g., Amazon RDS). A comparison algorithm detects delays and discrepancies with the plan. Based on the sentiment data, the importance of the delays and discrepancies is also assessed.
[0694] Visualization and notification
[0695] server:
[0696] The detected results and sentiment data are reflected in a visually displayed dashboard. This allows users to see the progress and its emotional impact at a glance. For example, progress delays are displayed in graphs and charts, and sentiment data is indicated by colors and icons.
[0697] Terminal:
[0698] Results sent from the server are updated in real time on the device. If significant delays or discrepancies are detected, a notification will pop up along with sentiment data.
[0699] Specific example
[0700] For example, suppose a weekly project meeting is held to discuss the progress of new feature A. If, during the meeting, it is stated, "Development of new feature A is on track and has entered the testing phase":
[0701] 1. Terminal: Captures audio and sends it to the server.
[0702] 2. Server: Converts audio data into text data using Google Cloud Speech-to-Text.
[0703] 3. Server: Use the Google Cloud Natural Language API to generate summaries using NLP algorithms.
[0704] 4. Server: Use Amazon Comprehend to analyze the emotion of "joy" and assign it to the summary. For example, assign it as "New Feature A: Progressing to the testing phase (joy)".
[0705] 5. Server: Compare the current meeting minutes with those stored in Amazon RDS to assess whether progress is on schedule. Evaluate the importance of any delays, including sentiment data.
[0706] 6. Server: Visually display the results on a dashboard and notify the user.
[0707] 7. Device: Displayed on the user's device, allowing them to check progress and sentiment data.
[0708] 8. User: Based on the results, take appropriate action regarding the project's progress.
[0709] Example of a prompt
[0710] "Recognize and summarize in text the emotions you would feel when reporting on the progress of new feature A at this week's project meeting."
[0711] Through these concrete actions, the overall system becomes easier to understand, leading to more effective project management and progress evaluation.
[0712] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0713] Step 1: Capture and send audio data
[0714] Terminal: When the meeting begins, audio data is captured through the terminal's microphone. The input is the meeting audio, which is captured as audio data. The captured audio data is sent to the server in real time. Specifically, the terminal's microphone converts the audio signal into digital data and sends it to the server using the HTTPS protocol. The output is the audio data sent to the server.
[0715] Step 2: Converting audio data to text
[0716] Server: The server inputs the received audio data into the Google Cloud Speech-to-Text speech recognition engine. The input is the received audio data, which the server parses and converts into text data. Specifically, the server sends the audio data to the speech recognition engine via an API call and receives the text data extracted from the audio signal. The output is the generated text data.
[0717] Step 3: Summary Generation
[0718] Server: The server inputs the generated text data into the Google Cloud Natural Language API. The input is the generated text data, and the server uses NLP algorithms to extract important points and keywords and summarize the text data. Specifically, the server applies an algorithm that analyzes the text data, extracts the important parts, and generates a summary. The output is the summarized text data.
[0719] Step 4: Emotion Recognition
[0720] Server: The server inputs the summarized text data and the original audio data into Amazon Comprehend's sentiment analysis engine. The input consists of summarized text data and the original audio data, and the sentiment analysis engine analyzes the user's emotions from the audio signal. Specifically, the sentiment analysis engine analyzes the audio data, identifies emotional states (e.g., joy, anger, sadness, surprise, etc.), and adds that information to the summarized text. The output is the summarized text with the sentiment data added.
[0721] Step 5: Comparison with past data and progress evaluation
[0722] Server: The server compares the generated sentiment-data-enhanced summary text with historical meeting minutes data stored in Amazon RDS. The input is the sentiment-data-enhanced summary text and historical meeting minutes data, and the server compares this data to detect delays in progress and discrepancies with the plan. Specifically, the server retrieves historical data and applies a comparison algorithm to detect progress and discrepancies. The output is the evaluation result of progress and discrepancies.
[0723] Step 6: Visualize and notify results
[0724] Server: The server reflects the detected progress and discrepancy evaluation results, as well as sentiment data, in a dashboard for visual display. Inputs are progress and discrepancy evaluation results and sentiment data, which are then visualized on the dashboard in graph and chart format. Specifically, the server generates evaluation results as graphs and charts and displays them to the user via a web interface. Output is the visualized information on the dashboard.
[0725] Terminal: The terminal updates the results sent from the server in real time. The input consists of progress and discrepancy evaluation results and sentiment data sent from the server, which are displayed in real time as notification pop-ups. Specifically, the terminal's notification system receives new data and displays a pop-up notification to the user. The output is the pop-up notification.
[0726] Through these steps, the system can perform detailed progress evaluations and sentiment analyses of meeting content. This allows users to understand project progress in real time and take necessary actions quickly.
[0727] (Application Example 2)
[0728] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0729] Current factory automation systems struggle to record communication between robots and workers, and to track work progress. Furthermore, progress management doesn't consider worker emotions, potentially leading to decreased efficiency and lower worker morale. To address these issues, a system is needed that analyzes and evaluates progress and worker emotions in real time, providing appropriate feedback.
[0730] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0731] In this invention, the server includes means for recording instructions between robots and workers in a factory, means for recognizing user emotions based on voice data, means for incorporating the emotion data into progress evaluation, and means for displaying the voice data and emotion data on a dashboard. This enables real-time progress management and emotion evaluation of robots and workers in a factory.
[0732] "Methods for recording meeting audio" refers to technologies that collect audio information during a meeting using microphones or other audio capture devices and store it as digital data.
[0733] "Means of converting recorded audio data into text data" refers to the process of converting audio data into text format data using conversion technologies such as a speech recognition engine.
[0734] "Methods for summarizing converted text data" refer to techniques that utilize natural language processing (NLP) algorithms to extract key points from large amounts of text data and condense them into short summaries.
[0735] "Means for comparing the summarized text data with past meeting minutes data" refers to a technique that compares newly generated summarized data with previously stored data to identify progress and changes.
[0736] "Means for detecting delays or discrepancies in progress based on the aforementioned comparison results" refers to a mechanism that analyzes the comparison results and evaluates whether the work or project is progressing according to plan.
[0737] "Means for visualizing the comparison results" refers to technologies that display the comparison results in a visual format, such as graphs or charts, so that users can easily understand them.
[0738] "Means for notifying the user of the results" refers to a notification function that communicates detected progress delays, discrepancies, and other important information to the user in real time.
[0739] "A means of recognizing user emotions based on voice data" refers to an emotion recognition engine that analyzes and recognizes human emotions (joy, anger, sadness, surprise, etc.) from voice.
[0740] "Means for incorporating the aforementioned emotional data into progress evaluations" refers to a system that incorporates users' emotional data into work and project progress evaluations to provide a more comprehensive assessment.
[0741] "A means of recording instructions between robots and workers in a factory and evaluating progress and emotions" refers to a technology that records instructions exchanged between robots and workers in a factory and uses that data to evaluate the progress of the work and the emotions of the workers.
[0742] "Means for displaying audio data and sentiment data on a dashboard" refers to technology that provides a dashboard function for visually displaying collected audio data and sentiment data.
[0743] The embodiments for carrying out this invention will be described in detail. A system for recording instructions between a robot and a worker in a factory and evaluating progress and worker emotions consists of the following steps.
[0744] First, the robots in the factory record voice instructions from workers via microphones. The recorded voice data is sent to a server in real time. The server uses a speech recognition engine (e.g., Google Speech Recognition API) to convert this voice data into text data.
[0745] The converted text data is analyzed using natural language processing (NLP) algorithms to extract key points and generate short summaries. These summaries are compared to past work records to assess progress. This helps determine whether the work is on schedule or behind schedule. Simultaneously, the server inputs audio data into an emotion recognition engine to analyze the worker's emotions. This allows the worker's emotions (e.g., joy, anger, sadness, surprise) to be added to the text data.
[0746] This data is stored in a database and used to evaluate progress and discrepancies by comparing it with past work records. The evaluation results are visually displayed on a visualization dashboard, allowing users to grasp the progress and sentiment status at a glance. For example, delays in progress are displayed in graph or chart format, and sentiment data is indicated by colors and icons. If significant delays or discrepancies are detected, a notification pops up on the user's device along with the sentiment data.
[0747] Hardware and software to be used
[0748] The hardware used will be a microphone (which can be built into a smartphone or smart glasses, etc.). The software will include Python, Google Speech Recognition API, NLTK, scikit-learn, a dedicated emotion recognition engine, and a web application for displaying the dashboard (e.g., Dash, Flask).
[0749] Specific example
[0750] For example, consider a scenario in a factory where a robot operates under the instructions of a worker. When the worker says, "I'm ready to proceed to the next work stage," the following process is executed:
[0751] Voice recording: The robot captures the worker's voice and sends it to the server.
[0752] Text Conversion: The Google Speech Recognition API is used to convert the audio data into text data. The resulting text will read, "Ready to proceed to the next work stage."
[0753] Summary generation: Generates summaries using natural language processing algorithms.
[0754] Emotion Analysis: An emotion recognition engine analyzes the emotion of "joy" from the voice.
[0755] Progress evaluation: Evaluate whether progress is "good" by comparing it with past work record data.
[0756] Dashboard display: The generated summary, sentiment data, and progress assessment are displayed on a dashboard for visual confirmation.
[0757] Example of a prompt
[0758] A factory robot receives the instruction, "Ready to proceed to the next work stage." Translate this instruction into text, summarize it, and evaluate the robot's work progress and emotional state.
[0759] Thus, the present invention's system efficiently records and manages communication between robots and workers in a factory, and simultaneously evaluates work progress and emotions, thereby realizing an efficient and harmonious work environment.
[0760] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0761] Step 1:
[0762] Audio recording
[0763] The terminal uses microphones within the factory to record voice instructions from workers in real time. The terminal then transmits the recorded voice data to a server.
[0764] Input: Voice instructions from the worker
[0765] Output: Audio data
[0766] Step 2:
[0767] Text conversion
[0768] The server inputs the received audio data into a speech recognition engine (e.g., Google Speech Recognition API) and converts the audio data into text data.
[0769] Input: Audio data
[0770] Output: Text data
[0771] Step 3:
[0772] Summary generation
[0773] The server analyzes the converted text data using natural language processing (NLP) algorithms, extracts key points, and generates a short summary.
[0774] Input: Text data
[0775] Output: Summary data
[0776] Step 4:
[0777] sentiment analysis
[0778] The server inputs the voice data into an emotion recognition engine and recognizes the worker's emotions (e.g., joy, anger, sadness, surprise, etc.). The recognized emotion data is then added to the text data.
[0779] Input: Audio data
[0780] Output: Sentiment data
[0781] Step 5:
[0782] Progress evaluation
[0783] The server compares the generated summary data and sentiment data with past work record data stored in the database to assess delays and discrepancies. It also evaluates the importance of delays and discrepancies based on the sentiment data.
[0784] Input: Summary data, sentiment data, past work record data
[0785] Output: Progress evaluation data
[0786] Step 6:
[0787] visualization
[0788] The server converts progress evaluation data and sentiment data into a visual format (e.g., graphs and charts) for display on the dashboard.
[0789] Input: Progress evaluation data, sentiment data
[0790] Output: Visualization data for dashboard
[0791] Step 7:
[0792] notification
[0793] If the server detects any significant delays or discrepancies in the project, it will notify the user's terminal in real time.
[0794] Input: Progress evaluation data, sentiment data
[0795] Output: User notification
[0796] These steps enable the system to record the collaboration between robots and workers in the factory, and to monitor and evaluate progress and emotions in real time.
[0797] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0798] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet Search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0799] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0800] [Third Embodiment]
[0801] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0802] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0803] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0804] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0805] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0806] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0807] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0808] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0809] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0810] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0811] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0812] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0813] The present invention provides a system that records meeting audio in real time, automatically transcribes and summarizes the content, and further evaluates the progress by comparing it with past meeting minutes data. The embodiments for carrying out the present invention are described in detail below.
[0814] Audio recording and conversion
[0815] Terminal:
[0816] When the meeting begins, audio data is captured via the microphone on the terminal. The captured audio data is sent to the server in real time.
[0817] server:
[0818] The server inputs the received audio data into a speech recognition engine and converts it into text data. A highly accurate speech recognition engine (e.g., speech recognition software or a cloud API) is used.
[0819] Summary generation using natural language processing
[0820] server:
[0821] Once the text data is generated, the server analyzes it using natural language processing (NLP) algorithms. The NLP algorithms extract key points and keywords, generating a summary of the text data. This summarization process provides long meeting minutes in a concise, summarized form.
[0822] Comparison with past meeting minutes and progress evaluation
[0823] server:
[0824] The generated summary data is compared with past meeting minutes data stored in the database. The comparison algorithm detects delays in progress and discrepancies with the plan. Specifically, the current status is compared with milestones and plans in past meeting minutes.
[0825] Visualization and notification
[0826] server:
[0827] The detected results are reflected in a visually displayed dashboard, allowing users to see the progress at a glance. For example, delays in progress can be displayed in graph or chart format, making it easy for users to understand.
[0828] Terminal:
[0829] Results sent from the server are updated in real time on the terminal. If significant delays or discrepancies are detected, the terminal will notify the user and display the details.
[0830] Specific example
[0831] For example, suppose a weekly project meeting is discussing the progress of new feature A. If, during the meeting, someone says, "Development of new feature A is on track and we've entered the testing phase":
[0832] 1. Terminal: Captures audio and sends it to the server.
[0833] 2. Server: Converts the audio data into text and generates the text, "Development of new feature A is on track and has entered the testing phase."
[0834] 3. Server: Use an NLP algorithm to summarize and conclude, "New Feature A: Proceeding to Test Phase."
[0835] 4. Server: Compare the current meeting minutes with those stored in the database to assess whether progress is on schedule.
[0836] 5. Server: Visualize the progress on a dashboard and notify users.
[0837] This process allows users to immediately check the progress after the meeting ends and take prompt action. The system of this invention enables companies and project teams to achieve efficient and effective meeting management and project management.
[0838] The following describes the processing flow.
[0839] Step 1:
[0840] Terminal: As soon as the meeting starts, the microphone on the terminal captures audio data. The captured audio data is sent to the server in real time.
[0841] Step 2:
[0842] Server: Passes the received audio data to the speech recognition engine, which converts it into text data. The speech recognition engine analyzes the audio signal and transcribes it into sentences such as, "The next milestone is next Monday."
[0843] Step 3:
[0844] Server: Inputs the converted text data into a natural language processing (NLP) algorithm to generate a summary. For example, it summarizes the sentence "The next milestone is next Monday" as "Milestone: Monday".
[0845] Step 4:
[0846] Server: Compares the generated summary data with past meeting minutes data stored in the database. Uses a comparison algorithm to detect delays and discrepancies in progress.
[0847] Step 5:
[0848] Server: Analyzes comparison results to identify significant delays and discrepancies. The identified results are displayed on a dashboard in graph and chart format using visualization tools.
[0849] Step 6:
[0850] Terminal: Receives visualized data sent from the server and displays it on the user's screen in real time. If there are significant delays or discrepancies in progress, a notification will pop up.
[0851] Step 7:
[0852] User: Check the latest meeting minutes and progress displayed on the terminal. The user will consider and take appropriate action regarding any delays or discrepancies in progress.
[0853] To explain with a concrete example, it would be as follows:
[0854] In the weekly project progress meeting:
[0855] 1. Device: Captures meeting audio in real time.
[0856] 2. Server: Converts audio data into text data.
[0857] 3. Server: Generates summaries using NLP algorithms.
[0858] 4. Server: Compares with past meeting minutes data and detects discrepancies.
[0859] 5. Server: Visualize the results on a dashboard.
[0860] 6. Terminal: Displays results and important notifications on the user's terminal.
[0861] 7. User: Review the results and respond to the project's progress.
[0862] In this way, the system enables automatic summarization of meeting content, progress monitoring, and important notifications at each step.
[0863] (Example 1)
[0864] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0865] Traditional meeting recording systems required manually converting audio data to text and summarizing its contents, which was time-consuming and labor-intensive. Furthermore, it was difficult to evaluate progress by comparing current meeting minutes with past ones. This not only hindered prompt responses after meetings but also made it difficult to detect project delays and discrepancies early on.
[0866] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0867] In this invention, the server includes means for converting received audio data into text data using a speech recognition engine, means for summarizing the text data using a natural language processing algorithm, and means for comparing the summarized text data with past meeting minutes data. This makes it possible to automatically generate accurate meeting records in real time and to quickly perform summarization and progress evaluation.
[0868] 1. "Means for capturing meeting audio" refers to a hardware or software system for collecting audio during a meeting in digital format.
[0869] 2. "Means for transmitting captured audio data in real time" refers to communication means for transmitting collected audio data to a server without delay.
[0870] 3. "Means for converting received audio data into text data using a speech recognition engine" refers to algorithms or software for analyzing audio data and converting it into corresponding text data.
[0871] 4. "Means for summarizing text data using natural language processing algorithms" refers to methods for analyzing generated text data, extracting important information, and summarizing it concisely.
[0872] 5. "Means for comparing summarized text data with past meeting minutes data" refers to methods for comparing newly generated summarized text with past meeting minutes to evaluate progress and detect discrepancies.
[0873] 6. "Means for detecting delays or discrepancies in progress based on comparison results" refers to a mechanism that analyzes the comparison results of text data and automatically identifies differences from the plan or delays in progress.
[0874] 7. "Means for visualizing comparison results" refers to display methods in the form of graphs or charts that allow users to intuitively understand the analyzed comparison results.
[0875] 8. "Means for notifying users of visualized results" refers to a notification system that informs users in real time of important progress and detected discrepancies.
[0876] 9. "Means for converting audio data to text in real time using parallel processing" refers to a method that efficiently utilizes computing resources to quickly convert audio data to text by simultaneously executing multiple audio analysis tasks.
[0877] 10. "Means of using generative AI models in natural language processing algorithms" refers to advanced algorithms that utilize generative AI models to analyze and summarize text data.
[0878] This invention provides a system that records meeting audio in real time, automatically transcribes and summarizes the content, and evaluates progress by comparing it with past meeting minutes data. The following describes how to implement this system in detail.
[0879] Audio recording and conversion
[0880] Terminal:
[0881] When the meeting begins, audio data is captured via the microphone on the terminal. The captured audio data is sent to the server in real time. The audio data is saved to a buffer at regular intervals, and when the buffer size reaches a certain level, it is sent back to the server.
[0882] server:
[0883] The server inputs the received audio data into a speech recognition engine (e.g., Google Cloud Speech-to-Text API) and converts it into text data. The text conversion speed is optimized by dividing the audio data into small chunks and executing API requests in parallel.
[0884] Summary generation using natural language processing
[0885] server:
[0886] Once text data is generated, the server analyzes it using a natural language processing (NLP) algorithm (e.g., OpenAI's generative AI model). The NLP algorithm extracts key points and keywords, generating a summary of the text data. The following is an example of a prompt: "Use an NLP algorithm to summarize the following text: Text: 'Development of new feature A is on track and has entered the testing phase. The development team is progressing as planned and will report the test results at the next meeting.'"
[0887] Comparison with past meeting minutes and progress evaluation
[0888] server:
[0889] The generated summary data is compared with historical meeting minutes data stored in the database. A comparison algorithm detects delays in progress and discrepancies with the plan. Specifically, historical data is retrieved using SQL queries, and a Python script detects matching keywords and progress differences with the summary data.
[0890] Visualization and notification
[0891] server:
[0892] The detected results are reflected in a visually displayed dashboard, allowing users to check the progress at a glance. Progress delays are displayed in graph and chart format using D3.js, etc.
[0893] Terminal:
[0894] Results sent from the server are updated in real time on the device. If significant delays or discrepancies are detected, the device notifies the user and displays the details. A notification pop-up appears in the browser, and clicking it jumps to a detailed screen.
[0895] Specific example
[0896] For example, suppose a weekly project meeting is discussing the progress of new feature A. If, during the meeting, someone says, "Development of new feature A is on track and we've entered the testing phase":
[0897] 1. Terminal: Captures audio and sends it to the server.
[0898] 2. Server: The server inputs the received audio data into the speech recognition engine and generates the text, "Development of new feature A is on track and has entered the testing phase."
[0899] 3. Server: Use an NLP algorithm to summarize and conclude, "New Feature A: Proceeding to Test Phase."
[0900] 4. Server: Compare the current meeting minutes with those stored in the database to assess whether Feature A is progressing as planned. For example, "Previous meeting: Design phase scheduled to be completed," "Current meeting: Proceeding to the testing phase." Assess that it is progressing as planned.
[0901] 5. Server: Display "New feature A is progressing as planned" on the dashboard and notify the user.
[0902] This process allows users to immediately check the progress after the meeting ends and take prompt action.
[0903] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0904] Step 1: Capture and send audio data
[0905] Terminal:
[0906] Input: Capture conference audio via the microphone on the device.
[0907] Data processing: Audio data is saved to a buffer at regular intervals.
[0908] Output: Once the buffer size reaches a certain level, it is sent to the server in real time.
[0909] Specific operation: The terminal captures audio data, and when it reaches a certain buffer size, it sends the data to the server. A low-latency communication protocol is used for this process.
[0910] Step 2: Receiving audio data and converting it to text.
[0911] server:
[0912] Input: Audio data sent from the device.
[0913] Data processing: Divide audio data into small chunks.
[0914] Data processing: Convert audio data into text data using a speech recognition engine (e.g., a cloud-based API).
[0915] Output: Converted real-time text data.
[0916] Specific operation: The server subdivides the received audio data and passes it to the speech recognition engine. Parallel processing is used to perform efficient text conversion.
[0917] Step 3: Text-based summary generation
[0918] server:
[0919] Input: Text data.
[0920] Data processing: Analyze text data using natural language processing algorithms.
[0921] Data processing: Extract and summarize key points using a generative AI model (e.g., OpenAI GPT-3).
[0922] Output: Summarized text data.
[0923] Specific operation: The server inputs a prompt message into the generation AI model and generates a summary of the text data. Example of a prompt message: "Use an NLP algorithm to summarize the following text. Text: 'Development of new feature A is on track and has entered the testing phase. The development team is progressing as planned and will report the test results at the next meeting.'"
[0924] Step 4: Comparison of summary data and past meeting minutes
[0925] server:
[0926] Input: Summarized text data.
[0927] Data processing: Retrieve past meeting minutes data from the database.
[0928] Data processing: Use a comparison algorithm to match summary data with past meeting minutes.
[0929] Output: Progress evaluation and comparison results of discrepancies.
[0930] Specific operation: The server retrieves past meeting minutes using SQL queries, and a Python script detects similarities and differences between the two sets of data.
[0931] Step 5: Visualizing progress
[0932] server:
[0933] Input: Results of progress evaluations and discrepancies comparisons.
[0934] Data processing: Formatting data for visualization.
[0935] Data processing: Visualization in graph and chart formats (e.g., D3.js).
[0936] Output: The visualization results displayed on the dashboard.
[0937] Specific operation: The server sends visualization data in JSON format to the dashboard API and updates it in real time.
[0938] Step 6: Notifications and detailed display
[0939] Terminal:
[0940] Input: Visualization results sent from the server.
[0941] Data processing: Formatting the data into a form that is easy for users to understand.
[0942] Output: Real-time notifications and detailed display on the device.
[0943] Specific operation: The device displays a notification pop-up, and clicking it jumps to a details screen. The user can easily check the progress.
[0944] These steps enable the system to transcribe, summarize, assess progress, visualize meeting audio in real time, and notify users.
[0945] (Application Example 1)
[0946] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0947] Modern factory operations demand complex project management and rapid response. However, managing meeting audio and progress is often done manually, which is inefficient and makes it difficult to immediately detect delays or discrepancies. This leads to delays in appropriate responses and decreased productivity. Therefore, there is a need for a system that automatically records meeting audio, transcribes it into text, summarizes it, and evaluates progress in real time.
[0948] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0949] In this invention, the server includes means for recording meeting audio, means for converting the recorded audio data into text data, means for summarizing the converted text data, means for comparing it with past meeting minutes data, means for detecting delays or discrepancies in progress, means for visualizing the comparison results, means for notifying the user of the results, means for a robot placed in the factory to record and process the meeting audio, and means for displaying the converted and summarized text data on a dashboard in real time. This enables the automatic recording of audio content of meetings held in the factory, improving the efficiency of progress management and allowing for quicker responses.
[0950] "Means of recording meeting audio" refers to a system that saves the audio from a meeting as digital data using a microphone or similar device.
[0951] "Means of converting recorded audio data into text data" refers to the process of converting recorded audio into text using a speech recognition engine.
[0952] "Methods for summarizing converted text data" refers to the process of using natural language processing techniques to extract only the important points from long texts and make them concise.
[0953] "Methods for comparing with past meeting minutes data" refer to a system that compares the content of the previous meeting stored in a database with the content of the current meeting to check for differences and progress.
[0954] "Methods for detecting delays or discrepancies in progress" refer to a system that compares past meeting minutes data with current data to identify differences between the plan and the actual progress.
[0955] "Methods for visualizing comparison results" refer to the process of displaying data differences and progress in visual formats such as graphs and charts.
[0956] "Means of notifying users of results" refers to a system that informs users of any delays or discrepancies in progress.
[0957] "Means for recording and processing meeting audio using robots deployed in a factory" refers to a process in which robots installed within a factory capture meeting audio and process that audio data in real time.
[0958] "A means of displaying real-time converted and summarized text data on a dashboard" refers to a mechanism that displays real-time converted and summarized text data on a dashboard so that users can understand it at a glance.
[0959] This invention is a system for automating and streamlining the progress management of meetings held in a factory. It records meeting audio, transcribes and summarizes the content, and evaluates the progress by comparing it with past meeting minutes data. Furthermore, this system has a function to visualize the progress in real time and notify the user.
[0960] Audio recording and conversion
[0961] Hardware: Robots deployed in the factory are equipped with high-sensitivity microphones to capture meeting audio in real time. This allows the system to begin collecting audio data as soon as the meeting starts.
[0962] Software: The collected audio data is sent to a server and converted into text data by a speech recognition engine (e.g., Google Web Speech API). This process is done in real time, so speech during a meeting is instantly transcribed into text.
[0963] Summary generation using natural language processing
[0964] Server: Once text data is generated, the server uses natural language processing (NLP) algorithms (e.g., the Hugging Face generative AI model) to summarize the text data. The summarized text contains key points, allowing users to efficiently grasp the information.
[0965] Comparison with past meeting minutes and progress evaluation
[0966] Server: The generated summary data is compared with past meeting minutes data stored in the database. The comparison algorithm detects any delays or discrepancies in progress. This ensures a clear understanding of the current progress against milestones and plans decided in past meetings.
[0967] Visualization and notification
[0968] Server: Detected comparison results are reflected in a visually displayed dashboard. Progress is shown in graph and chart format, allowing users to see the current progress and any issues at a glance. Furthermore, if significant delays or discrepancies are detected, the server will notify the user.
[0969] Specific example: For instance, when reporting on the progress of new feature A in a project meeting at a factory, it would look like this:
[0970] 1. The robot captures the meeting audio and sends it to the server.
[0971] 2. The server converts the audio data into text, specifically generating the text "New feature A has progressed from the development stage to the testing stage."
[0972] 3. Use an NLP algorithm to summarize the text, concluding it as "New Feature A: Testing Phase".
[0973] 4. Compare the current server status with past meeting minutes stored in the database to assess whether it is progressing according to plan.
[0974] 5. Progress is visualized on a dashboard, allowing users to check it in real time. If progress is behind schedule, users can immediately consider countermeasures.
[0975] A possible example of a specific prompt message would be something like, "We have received a progress report on Project X. You need to start on a new task, Task Y."
[0976] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0977] Step 1:
[0978] Terminal: A robot positioned in the factory captures meeting audio and transmits it to a server in real time. The input data is audio data, and the output is audio data transferred to the server. Specifically, the robot's microphone captures the audio and transmits that audio data to the server via the network.
[0979] Step 2:
[0980] Server: Inputs received audio data into a speech recognition engine and converts it into text data. Input data is audio data, and output data is text data. Specifically, the server calls a speech recognition engine (e.g., Google Web Speech API) and converts the audio data into text format.
[0981] Step 3:
[0982] Server: Summarizes the generated text data using a natural language processing (NLP) algorithm. The input data is the transformed text data, and the output data is the summarized text data. Specifically, the server uses the Hugging Face generative AI model to extract important information from long texts and generate summaries.
[0983] Step 4:
[0984] Server: Compares summarized text data with past meeting minutes data stored in the database. The input data consists of summarized text data and past meeting minutes data, and the output data is the comparison result. Specifically, the server retrieves meeting minutes data from the past database, compares it with the current summarized data, and calculates the difference.
[0985] Step 5:
[0986] Server: Detects delays and discrepancies in progress based on comparison results. Input data is the comparison results, and output data is information about delays and discrepancies in progress. Specifically, the server uses an algorithm to evaluate the progress status and identify the gap with the plan.
[0987] Step 6:
[0988] Server: Reflects detected results on a dashboard for visual display. Input data is information about delays and discrepancies in progress, and output data is visualized graphs and charts. Specifically, the server uses the matplotlib library to graphically display the progress.
[0989] Step 7:
[0990] Server: Sends generated visualization data to the terminal and notifies the user. Input data is visualized graphs and charts, and output data is progress displayed on the user interface. Specifically, the server pushes the generated visualization data to the terminal in real time and notifies the user of important difference information.
[0991] Step 8:
[0992] User: Check the progress on the dashboard on the device and decide on countermeasures as needed. The input data is the displayed progress, and the output is the decision on the appropriate countermeasure. Specifically, the user accesses the dashboard on the device and plans the next action based on the visualized progress.
[0993] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0994] This invention combines a system that records meeting audio in real time, automatically transcribes and summarizes its contents, and evaluates progress by comparing it with past meeting minutes data, with an emotion engine that recognizes the user's emotions from the audio data. The embodiments for carrying out this invention will be described in detail below.
[0995] Audio recording and conversion
[0996] Terminal:
[0997] When the meeting begins, audio data is captured via the microphone on the terminal. The captured audio data is sent to the server in real time.
[0998] server:
[0999] The server inputs the received audio data into the speech recognition engine, which converts it into text data. The speech recognition engine analyzes the audio signal and converts the meeting content into text format.
[1000] Summary generation using natural language processing
[1001] server:
[1002] Once the text data is generated, the server analyzes it using natural language processing (NLP) algorithms. The NLP algorithms extract key points and keywords, generating a summary of the text data. This summarization process provides a concise summary of the meeting's content.
[1003] User emotion recognition by an emotion engine
[1004] server:
[1005] Furthermore, the server inputs the voice data into the emotion engine to recognize the user's emotions. The emotion engine analyzes the user's emotions (e.g., joy, anger, sadness, surprise, etc.) from the voice signal and adds them to the text data.
[1006] Comparison with past meeting minutes and progress evaluation
[1007] server:
[1008] The generated summary data and sentiment data are compared with historical meeting minutes data stored in the database. A comparison algorithm detects delays and discrepancies with the plan. Based on the sentiment data, the importance of the delays and discrepancies is also evaluated.
[1009] Visualization and notification
[1010] server:
[1011] The detected results and sentiment data are reflected in a visually displayed dashboard. This allows users to see the progress and its emotional impact at a glance. For example, progress delays are displayed in graphs and charts, and sentiment data is indicated by colors and icons.
[1012] Terminal:
[1013] Results sent from the server are updated in real time on the device. If significant delays or discrepancies are detected, a notification will pop up along with sentiment data.
[1014] Specific example
[1015] For example, suppose a weekly project meeting is held to discuss the progress of new feature A. If, during the meeting, it is stated, "Development of new feature A is on track and has entered the testing phase":
[1016] 1. Terminal: Captures audio and sends it to the server.
[1017] 2. Server: Converts audio data into text data.
[1018] 3. Server: Generates summaries using NLP algorithms.
[1019] 4. Server: The emotion engine analyzes the emotion of "joy" and assigns it to the summary. For example, it might be assigned as "New Feature A: Progressing to Test Phase (Joy)".
[1020] 5. Server: Compare the current meeting minutes with those stored in the database to assess whether progress is on schedule. Evaluate the importance of any delays, including sentiment data.
[1021] 6. Server: Visualize the results on a dashboard and notify the user.
[1022] 7. Device: Displayed on the user's device, allowing them to check progress and sentiment data.
[1023] 8. User: Based on the results, take appropriate action regarding the project's progress.
[1024] By combining this with an emotion engine, it becomes possible to not only automatically summarize meeting content but also to evaluate progress while considering the user's emotions, resulting in more accurate project management. This improves the project success rate and enhances team engagement.
[1025] The following describes the processing flow.
[1026] Step 1:
[1027] Terminal: As soon as the meeting starts, the microphone on the terminal captures audio data. The captured audio data is sent to the server in real time.
[1028] Step 2:
[1029] Server: Inputs received audio data into the speech recognition engine and converts it into text data. For example, it recognizes the audio data "The next milestone is next Monday" and converts it into the corresponding text data.
[1030] Step 3:
[1031] Server: Inputs the converted text data into a natural language processing (NLP) algorithm to generate a summary. For example, it summarizes the sentence "The next milestone is next Monday" as "Milestone: Monday".
[1032] Step 4:
[1033] Server: The generated text data is input into the emotion engine, which analyzes the tone and tempo of the voice to recognize the user's emotions. For example, "The next milestone is next Monday" is tagged with the emotion "joy."
[1034] Step 5:
[1035] Server: Compares summarized text data and sentiment data with past meeting minutes data stored in the database to detect delays or discrepancies in progress. For example, it might detect if a milestone previously set for "this Friday" has been postponed to "next Monday."
[1036] Step 6:
[1037] Server: Analyzes comparison results and sentiment data to identify delays and discrepancies in progress and assess their importance. For example, delays in progress are displayed along with the emotion of "joy," indicating positive progress.
[1038] Step 7:
[1039] Server: Visualizes detected results and sentiment data and displays them on a dashboard. For example, areas with delays are shown in red, and sentiment data is indicated by sentiment icons (e.g., a smile).
[1040] Step 8:
[1041] Terminal: Receives visualization data sent from the server and displays it on the user's screen in real time. If significant progress delays or discrepancies are detected, it notifies the user and displays the details along with sentiment data.
[1042] Step 9:
[1043] User: Review the latest meeting minutes, progress, and sentiment data displayed on the device. The user considers and takes appropriate action as needed.
[1044] As a concrete example, consider a situation where the progress of new feature A is being discussed in a weekly project meeting:
[1045] 1. Terminal: Captures meeting audio in real time and sends it to the server.
[1046] 2. Server: Converts audio data into text data.
[1047] 3. Server: Generates summaries using NLP algorithms.
[1048] 4. Server: The emotion engine analyzes the emotion of "joy" and assigns it to the summary. This will result in "New Feature A: Proceeding to Test Phase (Joy)".
[1049] 5. Server: Compare the current meeting minutes with those stored in the database to assess whether progress is on schedule. Consider sentiment data to evaluate the importance of any delays.
[1050] 6. Server: Visualize the results on a dashboard and notify the user.
[1051] 7. Device: The results are displayed on the user's device, allowing them to check the progress and sentiment data.
[1052] 8. User: Respond to project progress based on the results.
[1053] This integrates automated summarization of meeting content, progress evaluation, and consideration of sentiment data, enabling effective project management.
[1054] (Example 2)
[1055] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[1056] Traditional meeting minutes systems simply transcribe meeting audio into text, failing to consider the meeting's progress or emotional fluctuations. Furthermore, they lacked automated summarization and comparison with past minutes, requiring users to perform these tasks manually, which was extremely time-consuming. Additionally, the lack of technology to analyze user emotions and reflect them in meeting content limited their ability to improve team engagement and project success rates.
[1057] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[1058] In this invention, the server includes means for recording meeting audio, means for converting the recorded audio data into text data, means for summarizing the converted text data, means for comparing the summarized text data with past recorded data, means for detecting delays or discrepancies in progress based on the comparison results, means for visualizing the comparison results, means for notifying the user of the results, means for analyzing the user's emotions from the audio data, and means for adding the analyzed emotion data to the text data. This enables a comprehensive analysis including detailed progress evaluation of meetings and fluctuations in emotions, thereby improving the efficiency of project management and strengthening team engagement.
[1059] "Means for recording meeting audio" refers to devices or programs that capture audio during a meeting in real time and save it as digital data.
[1060] "Means for converting recorded audio data into text data" refers to speech recognition technologies and software that analyze audio signals and recognize and convert them into text information.
[1061] "Means for summarizing converted text data" refers to natural language processing techniques that analyze text data to extract important information and shorten the original text to create a summary.
[1062] "Means for comparing the summarized text data with past record data" refers to an algorithm for comparing the generated summary data with existing meeting record data and performing analysis of progress, agreements, and disagreements.
[1063] "Means for detecting delays or discrepancies in progress based on the comparison results" refers to an analytical system that identifies the gap between the expected progress and the actual progress based on the comparison results, and detects delays or problems.
[1064] "Means for visualizing the comparison results" refers to programs or tools for displaying comparison results and progress information in a visual format such as graphs or charts that are easy for users to understand.
[1065] "Means for notifying the user of the results" refers to a program or device that has a notification function to inform the user of analysis results and important information in real time.
[1066] "Methods for analyzing user emotions from voice data" refer to technologies that determine a user's emotional state from factors such as voice intonation, tempo, and volume, and classify it into emotional categories (e.g., joy, anger, sadness, surprise, etc.).
[1067] "Means for attaching analyzed emotional data to text data" refers to a system that attaches analyzed emotional information as part of the corresponding text data, thereby clearly indicating the emotional state.
[1068] This invention combines a system that records meeting audio in real time, automatically transcribes and summarizes its contents, and evaluates progress by comparing it with past meeting minutes data, with technology that recognizes user emotions from the audio data. The embodiments for carrying out this invention will be described in detail below.
[1069] Audio recording and conversion
[1070] Terminal:
[1071] When a meeting begins, audio data is captured through the terminal's microphone. The captured audio data is sent to the server in real time. The hardware used is either the terminal's built-in microphone or an external microphone.
[1072] server:
[1073] The server inputs the received audio data into a speech recognition engine (e.g., Google Cloud Speech-to-Text) and converts it into text data. It analyzes the audio signal and converts the meeting content into text format.
[1074] Summary generation using natural language processing
[1075] server:
[1076] Once the text data is generated, the server analyzes it using natural language processing (NLP) algorithms (e.g., Google Cloud Natural Language API). The NLP algorithms extract key points and keywords, generating a summary of the text data. This summarization process provides a concise summary of the meeting content.
[1077] User emotion recognition by an emotion engine
[1078] server:
[1079] Furthermore, the server inputs the audio data into an emotion analysis engine (e.g., Amazon Comprehend) to recognize the user's emotions. The emotion analysis engine analyzes the user's emotions (e.g., joy, anger, sadness, surprise, etc.) from the audio signal and adds them to the text data.
[1080] Comparison with past meeting minutes and progress evaluation
[1081] server:
[1082] The generated summary and sentiment data are compared with historical meeting minutes data stored in a database (e.g., Amazon RDS). A comparison algorithm detects delays and discrepancies with the plan. Based on the sentiment data, the importance of the delays and discrepancies is also assessed.
[1083] Visualization and notification
[1084] server:
[1085] The detected results and sentiment data are reflected in a visually displayed dashboard. This allows users to see the progress and its emotional impact at a glance. For example, progress delays are displayed in graphs and charts, and sentiment data is indicated by colors and icons.
[1086] Terminal:
[1087] Results sent from the server are updated in real time on the device. If significant delays or discrepancies are detected, a notification will pop up along with sentiment data.
[1088] Specific example
[1089] For example, suppose a weekly project meeting is held to discuss the progress of new feature A. If, during the meeting, it is stated, "Development of new feature A is on track and has entered the testing phase":
[1090] 1. Terminal: Captures audio and sends it to the server.
[1091] 2. Server: Converts audio data into text data using Google Cloud Speech-to-Text.
[1092] 3. Server: Use the Google Cloud Natural Language API to generate summaries using NLP algorithms.
[1093] 4. Server: Use Amazon Comprehend to analyze the emotion of "joy" and assign it to the summary. For example, assign it as "New Feature A: Progressing to the testing phase (joy)".
[1094] 5. Server: Compare the current meeting minutes with those stored in Amazon RDS to assess whether progress is on schedule. Evaluate the importance of any delays, including sentiment data.
[1095] 6. Server: Visually display the results on a dashboard and notify the user.
[1096] 7. Device: Displayed on the user's device, allowing them to check progress and sentiment data.
[1097] 8. User: Based on the results, take appropriate action regarding the project's progress.
[1098] Example of a prompt
[1099] "Recognize and summarize in text the emotions you would feel when reporting on the progress of new feature A at this week's project meeting."
[1100] Through these concrete actions, the overall system becomes easier to understand, leading to more effective project management and progress evaluation.
[1101] The flow of the specific processing in Example 2 will be explained using Figure 13.
[1102] Step 1: Capture and send audio data
[1103] Terminal: When the meeting begins, audio data is captured through the terminal's microphone. The input is the meeting audio, which is captured as audio data. The captured audio data is sent to the server in real time. Specifically, the terminal's microphone converts the audio signal into digital data and sends it to the server using the HTTPS protocol. The output is the audio data sent to the server.
[1104] Step 2: Converting audio data to text
[1105] Server: The server inputs the received audio data into the Google Cloud Speech-to-Text speech recognition engine. The input is the received audio data, which the server parses and converts into text data. Specifically, the server sends the audio data to the speech recognition engine via an API call and receives the text data extracted from the audio signal. The output is the generated text data.
[1106] Step 3: Summary Generation
[1107] Server: The server inputs the generated text data into the Google Cloud Natural Language API. The input is the generated text data, and the server uses NLP algorithms to extract important points and keywords and summarize the text data. Specifically, the server applies an algorithm that analyzes the text data, extracts the important parts, and generates a summary. The output is the summarized text data.
[1108] Step 4: Emotion Recognition
[1109] Server: The server inputs the summarized text data and the original audio data into Amazon Comprehend's sentiment analysis engine. The input consists of summarized text data and the original audio data, and the sentiment analysis engine analyzes the user's emotions from the audio signal. Specifically, the sentiment analysis engine analyzes the audio data, identifies emotional states (e.g., joy, anger, sadness, surprise, etc.), and adds that information to the summarized text. The output is the summarized text with the sentiment data added.
[1110] Step 5: Comparison with past data and progress evaluation
[1111] Server: The server compares the generated sentiment-data-enhanced summary text with historical meeting minutes data stored in Amazon RDS. The input is the sentiment-data-enhanced summary text and historical meeting minutes data, and the server compares this data to detect delays in progress and discrepancies with the plan. Specifically, the server retrieves historical data and applies a comparison algorithm to detect progress and discrepancies. The output is the evaluation result of progress and discrepancies.
[1112] Step 6: Visualize and notify results
[1113] Server: The server reflects the detected progress and discrepancy evaluation results, as well as sentiment data, in a dashboard for visual display. Inputs are progress and discrepancy evaluation results and sentiment data, which are then visualized on the dashboard in graph and chart format. Specifically, the server generates evaluation results as graphs and charts and displays them to the user via a web interface. Output is the visualized information on the dashboard.
[1114] Terminal: The terminal updates the results sent from the server in real time. The input consists of progress and discrepancy evaluation results and sentiment data sent from the server, which are displayed in real time as notification pop-ups. Specifically, the terminal's notification system receives new data and displays a pop-up notification to the user. The output is the pop-up notification.
[1115] Through these steps, the system can perform detailed progress evaluations and sentiment analyses of meeting content. This allows users to understand project progress in real time and take necessary actions quickly.
[1116] (Application Example 2)
[1117] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[1118] Current factory automation systems struggle to record communication between robots and workers, and to track work progress. Furthermore, progress management doesn't consider worker emotions, potentially leading to decreased efficiency and lower worker morale. To address these issues, a system is needed that analyzes and evaluates progress and worker emotions in real time, providing appropriate feedback.
[1119] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[1120] In this invention, the server includes means for recording instructions between robots and workers in a factory, means for recognizing user emotions based on voice data, means for incorporating the emotion data into progress evaluation, and means for displaying the voice data and emotion data on a dashboard. This enables real-time progress management and emotion evaluation of robots and workers in a factory.
[1121] "Methods for recording meeting audio" refers to technologies that collect audio information during a meeting using microphones or other audio capture devices and store it as digital data.
[1122] "Means of converting recorded audio data into text data" refers to the process of converting audio data into text format data using conversion technologies such as a speech recognition engine.
[1123] "Methods for summarizing converted text data" refer to techniques that utilize natural language processing (NLP) algorithms to extract key points from large amounts of text data and condense them into short summaries.
[1124] "Means for comparing the summarized text data with past meeting minutes data" refers to a technique that compares newly generated summarized data with previously stored data to identify progress and changes.
[1125] "Means for detecting delays or discrepancies in progress based on the aforementioned comparison results" refers to a mechanism that analyzes the comparison results and evaluates whether the work or project is progressing according to plan.
[1126] "Means for visualizing the comparison results" refers to technologies that display the comparison results in a visual format, such as graphs or charts, so that users can easily understand them.
[1127] "Means for notifying the user of the results" refers to a notification function that communicates detected progress delays, discrepancies, and other important information to the user in real time.
[1128] "A means of recognizing user emotions based on voice data" refers to an emotion recognition engine that analyzes and recognizes human emotions (joy, anger, sadness, surprise, etc.) from voice.
[1129] "Means for incorporating the aforementioned emotional data into progress evaluations" refers to a system that incorporates users' emotional data into work and project progress evaluations to provide a more comprehensive assessment.
[1130] "A means of recording instructions between robots and workers in a factory and evaluating progress and emotions" refers to a technology that records instructions exchanged between robots and workers in a factory and uses that data to evaluate the progress of the work and the emotions of the workers.
[1131] "Means for displaying audio data and sentiment data on a dashboard" refers to technology that provides a dashboard function for visually displaying collected audio data and sentiment data.
[1132] The embodiments for carrying out this invention will be described in detail. A system for recording instructions between a robot and a worker in a factory and evaluating progress and worker emotions consists of the following steps.
[1133] First, the robots in the factory record voice instructions from workers via microphones. The recorded voice data is sent to a server in real time. The server uses a speech recognition engine (e.g., Google Speech Recognition API) to convert this voice data into text data.
[1134] The converted text data is analyzed using natural language processing (NLP) algorithms to extract key points and generate short summaries. These summaries are compared to past work records to assess progress. This helps determine whether the work is on schedule or behind schedule. Simultaneously, the server inputs audio data into an emotion recognition engine to analyze the worker's emotions. This allows the worker's emotions (e.g., joy, anger, sadness, surprise) to be added to the text data.
[1135] This data is stored in a database and used to evaluate progress and discrepancies by comparing it with past work records. The evaluation results are visually displayed on a visualization dashboard, allowing users to grasp the progress and sentiment status at a glance. For example, delays in progress are displayed in graph or chart format, and sentiment data is indicated by colors and icons. If significant delays or discrepancies are detected, a notification pops up on the user's device along with the sentiment data.
[1136] Hardware and software to be used
[1137] The hardware used will be a microphone (which can be built into a smartphone or smart glasses, etc.). The software will include Python, Google Speech Recognition API, NLTK, scikit-learn, a dedicated emotion recognition engine, and a web application for displaying the dashboard (e.g., Dash, Flask).
[1138] Specific example
[1139] For example, consider a scenario in a factory where a robot operates under the instructions of a worker. When the worker says, "I'm ready to proceed to the next work stage," the following process is executed:
[1140] Voice recording: The robot captures the worker's voice and sends it to the server.
[1141] Text Conversion: The Google Speech Recognition API is used to convert the audio data into text data. The resulting text will read, "Ready to proceed to the next work stage."
[1142] Summary generation: Generates summaries using natural language processing algorithms.
[1143] Emotion Analysis: An emotion recognition engine analyzes the emotion of "joy" from the voice.
[1144] Progress evaluation: Evaluate whether progress is "good" by comparing it with past work record data.
[1145] Dashboard display: The generated summary, sentiment data, and progress assessment are displayed on a dashboard for visual confirmation.
[1146] Example of a prompt
[1147] A factory robot receives the instruction, "Ready to proceed to the next work stage." Translate this instruction into text, summarize it, and evaluate the robot's work progress and emotional state.
[1148] Thus, the present invention's system efficiently records and manages communication between robots and workers in a factory, and simultaneously evaluates work progress and emotions, thereby realizing an efficient and harmonious work environment.
[1149] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[1150] Step 1:
[1151] Audio recording
[1152] The terminal uses microphones within the factory to record voice instructions from workers in real time. The terminal then transmits the recorded voice data to a server.
[1153] Input: Voice instructions from the worker
[1154] Output: Audio data
[1155] Step 2:
[1156] Text conversion
[1157] The server inputs the received audio data into a speech recognition engine (e.g., Google Speech Recognition API) and converts the audio data into text data.
[1158] Input: Audio data
[1159] Output: Text data
[1160] Step 3:
[1161] Summary generation
[1162] The server analyzes the converted text data using natural language processing (NLP) algorithms, extracts key points, and generates a short summary.
[1163] Input: Text data
[1164] Output: Summary data
[1165] Step 4:
[1166] sentiment analysis
[1167] The server inputs the voice data into an emotion recognition engine and recognizes the worker's emotions (e.g., joy, anger, sadness, surprise, etc.). The recognized emotion data is then added to the text data.
[1168] Input: Audio data
[1169] Output: Sentiment data
[1170] Step 5:
[1171] Progress evaluation
[1172] The server compares the generated summary data and sentiment data with past work record data stored in the database to assess delays and discrepancies. It also evaluates the importance of delays and discrepancies based on the sentiment data.
[1173] Input: Summary data, sentiment data, past work record data
[1174] Output: Progress evaluation data
[1175] Step 6:
[1176] visualization
[1177] The server converts progress evaluation data and sentiment data into a visual format (e.g., graphs and charts) for display on the dashboard.
[1178] Input: Progress evaluation data, sentiment data
[1179] Output: Visualization data for dashboard
[1180] Step 7:
[1181] notification
[1182] If the server detects any significant delays or discrepancies in the project, it will notify the user's terminal in real time.
[1183] Input: Progress evaluation data, sentiment data
[1184] Output: User notification
[1185] These steps enable the system to record the collaboration between robots and workers in the factory, and to monitor and evaluate progress and emotions in real time.
[1186] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[1187] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet Search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1188] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[1189] [Fourth Embodiment]
[1190] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[1191] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1192] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1193] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[1194] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[1195] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[1196] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[1197] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[1198] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[1199] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1200] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1201] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[1202] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1203] The present invention provides a system that records meeting audio in real time, automatically transcribes and summarizes the content, and further evaluates the progress by comparing it with past meeting minutes data. The embodiments for carrying out the present invention are described in detail below.
[1204] Audio recording and conversion
[1205] Terminal:
[1206] When the meeting begins, audio data is captured via the microphone on the terminal. The captured audio data is sent to the server in real time.
[1207] server:
[1208] The server inputs the received audio data into a speech recognition engine and converts it into text data. A highly accurate speech recognition engine (e.g., speech recognition software or a cloud API) is used.
[1209] Summary generation using natural language processing
[1210] server:
[1211] Once the text data is generated, the server analyzes it using natural language processing (NLP) algorithms. The NLP algorithms extract key points and keywords, generating a summary of the text data. This summarization process provides long meeting minutes in a concise, summarized form.
[1212] Comparison with past meeting minutes and progress evaluation
[1213] server:
[1214] The generated summary data is compared with past meeting minutes data stored in the database. The comparison algorithm detects delays in progress and discrepancies with the plan. Specifically, the current status is compared with milestones and plans in past meeting minutes.
[1215] Visualization and notification
[1216] server:
[1217] The detected results are reflected in a visually displayed dashboard, allowing users to see the progress at a glance. For example, delays in progress can be displayed in graph or chart format, making it easy for users to understand.
[1218] Terminal:
[1219] Results sent from the server are updated in real time on the terminal. If significant delays or discrepancies are detected, the terminal will notify the user and display the details.
[1220] Specific example
[1221] For example, suppose a weekly project meeting is discussing the progress of new feature A. If, during the meeting, someone says, "Development of new feature A is on track and we've entered the testing phase":
[1222] 1. Terminal: Captures audio and sends it to the server.
[1223] 2. Server: Converts the audio data into text and generates the text, "Development of new feature A is on track and has entered the testing phase."
[1224] 3. Server: Use an NLP algorithm to summarize and conclude, "New Feature A: Proceeding to Test Phase."
[1225] 4. Server: Compare the current meeting minutes with those stored in the database to assess whether progress is on schedule.
[1226] 5. Server: Visualize the progress on a dashboard and notify users.
[1227] This process allows users to immediately check the progress after the meeting ends and take prompt action. The system of this invention enables companies and project teams to achieve efficient and effective meeting management and project management.
[1228] The following describes the processing flow.
[1229] Step 1:
[1230] Terminal: As soon as the meeting starts, the microphone on the terminal captures audio data. The captured audio data is sent to the server in real time.
[1231] Step 2:
[1232] Server: Passes the received audio data to the speech recognition engine, which converts it into text data. The speech recognition engine analyzes the audio signal and transcribes it into sentences such as, "The next milestone is next Monday."
[1233] Step 3:
[1234] Server: Inputs the converted text data into a natural language processing (NLP) algorithm to generate a summary. For example, it summarizes the sentence "The next milestone is next Monday" as "Milestone: Monday".
[1235] Step 4:
[1236] Server: Compares the generated summary data with past meeting minutes data stored in the database. Uses a comparison algorithm to detect delays and discrepancies in progress.
[1237] Step 5:
[1238] Server: Analyzes comparison results to identify significant delays and discrepancies. The identified results are displayed on a dashboard in graph and chart format using visualization tools.
[1239] Step 6:
[1240] Terminal: Receives visualized data sent from the server and displays it on the user's screen in real time. If there are significant delays or discrepancies in progress, a notification will pop up.
[1241] Step 7:
[1242] User: Check the latest meeting minutes and progress displayed on the terminal. The user will consider and take appropriate action regarding any delays or discrepancies in progress.
[1243] To explain with a concrete example, it would be as follows:
[1244] In the weekly project progress meeting:
[1245] 1. Device: Captures meeting audio in real time.
[1246] 2. Server: Converts audio data into text data.
[1247] 3. Server: Generates summaries using NLP algorithms.
[1248] 4. Server: Compares with past meeting minutes data and detects discrepancies.
[1249] 5. Server: Visualize the results on a dashboard.
[1250] 6. Terminal: Displays results and important notifications on the user's terminal.
[1251] 7. User: Review the results and respond to the project's progress.
[1252] In this way, the system enables automatic summarization of meeting content, progress monitoring, and important notifications at each step.
[1253] (Example 1)
[1254] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1255] Traditional meeting recording systems required manually converting audio data to text and summarizing its contents, which was time-consuming and labor-intensive. Furthermore, it was difficult to evaluate progress by comparing current meeting minutes with past ones. This not only hindered prompt responses after meetings but also made it difficult to detect project delays and discrepancies early on.
[1256] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[1257] In this invention, the server includes means for converting received audio data into text data using a speech recognition engine, means for summarizing the text data using a natural language processing algorithm, and means for comparing the summarized text data with past meeting minutes data. This makes it possible to automatically generate accurate meeting records in real time and to quickly perform summarization and progress evaluation.
[1258] 1. "Means for capturing meeting audio" refers to a hardware or software system for collecting audio during a meeting in digital format.
[1259] 2. "Means for transmitting captured audio data in real time" refers to communication means for transmitting collected audio data to a server without delay.
[1260] 3. "Means for converting received audio data into text data using a speech recognition engine" refers to algorithms or software for analyzing audio data and converting it into corresponding text data.
[1261] 4. "Means for summarizing text data using natural language processing algorithms" refers to methods for analyzing generated text data, extracting important information, and summarizing it concisely.
[1262] 5. "Means for comparing summarized text data with past meeting minutes data" refers to methods for comparing newly generated summarized text with past meeting minutes to evaluate progress and detect discrepancies.
[1263] 6. "Means for detecting delays or discrepancies in progress based on comparison results" refers to a mechanism that analyzes the comparison results of text data and automatically identifies differences from the plan or delays in progress.
[1264] 7. "Means for visualizing comparison results" refers to display methods in the form of graphs or charts that allow users to intuitively understand the analyzed comparison results.
[1265] 8. "Means for notifying users of visualized results" refers to a notification system that informs users in real time of important progress and detected discrepancies.
[1266] 9. "Means for converting audio data to text in real time using parallel processing" refers to a method that efficiently utilizes computing resources to quickly convert audio data to text by simultaneously executing multiple audio analysis tasks.
[1267] 10. "Means of using generative AI models in natural language processing algorithms" refers to advanced algorithms that utilize generative AI models to analyze and summarize text data.
[1268] This invention provides a system that records meeting audio in real time, automatically transcribes and summarizes the content, and evaluates progress by comparing it with past meeting minutes data. The following describes how to implement this system in detail.
[1269] Audio recording and conversion
[1270] Terminal:
[1271] When the meeting begins, audio data is captured via the microphone on the terminal. The captured audio data is sent to the server in real time. The audio data is saved to a buffer at regular intervals, and when the buffer size reaches a certain level, it is sent back to the server.
[1272] server:
[1273] The server inputs the received audio data into a speech recognition engine (e.g., Google Cloud Speech-to-Text API) and converts it into text data. The text conversion speed is optimized by dividing the audio data into small chunks and executing API requests in parallel.
[1274] Summary generation using natural language processing
[1275] server:
[1276] Once text data is generated, the server analyzes it using a natural language processing (NLP) algorithm (e.g., OpenAI's generative AI model). The NLP algorithm extracts key points and keywords, generating a summary of the text data. The following is an example of a prompt: "Use an NLP algorithm to summarize the following text: Text: 'Development of new feature A is on track and has entered the testing phase. The development team is progressing as planned and will report the test results at the next meeting.'"
[1277] Comparison with past meeting minutes and progress evaluation
[1278] server:
[1279] The generated summary data is compared with historical meeting minutes data stored in the database. A comparison algorithm detects delays in progress and discrepancies with the plan. Specifically, historical data is retrieved using SQL queries, and a Python script detects matching keywords and progress differences with the summary data.
[1280] Visualization and notification
[1281] server:
[1282] The detected results are reflected in a visually displayed dashboard, allowing users to check the progress at a glance. Progress delays are displayed in graph and chart format using D3.js, etc.
[1283] Terminal:
[1284] Results sent from the server are updated in real time on the device. If significant delays or discrepancies are detected, the device notifies the user and displays the details. A notification pop-up appears in the browser, and clicking it jumps to a detailed screen.
[1285] Specific example
[1286] For example, suppose a weekly project meeting is discussing the progress of new feature A. If, during the meeting, someone says, "Development of new feature A is on track and we've entered the testing phase":
[1287] 1. Terminal: Captures audio and sends it to the server.
[1288] 2. Server: The server inputs the received audio data into the speech recognition engine and generates the text, "Development of new feature A is on track and has entered the testing phase."
[1289] 3. Server: Use an NLP algorithm to summarize and conclude, "New Feature A: Proceeding to Test Phase."
[1290] 4. Server: Compare the current meeting minutes with those stored in the database to assess whether Feature A is progressing as planned. For example, "Previous meeting: Design phase scheduled to be completed," "Current meeting: Proceeding to the testing phase." Assess that it is progressing as planned.
[1291] 5. Server: Display "New feature A is progressing as planned" on the dashboard and notify the user.
[1292] This process allows users to immediately check the progress after the meeting ends and take prompt action.
[1293] The flow of the specific processing in Example 1 will be explained using Figure 11.
[1294] Step 1: Capture and send audio data
[1295] Terminal:
[1296] Input: Capture conference audio via the microphone on the device.
[1297] Data processing: Audio data is saved to a buffer at regular intervals.
[1298] Output: Once the buffer size reaches a certain level, it is sent to the server in real time.
[1299] Specific operation: The terminal captures audio data, and when it reaches a certain buffer size, it sends the data to the server. A low-latency communication protocol is used for this process.
[1300] Step 2: Receiving audio data and converting it to text.
[1301] server:
[1302] Input: Audio data sent from the device.
[1303] Data processing: Divide audio data into small chunks.
[1304] Data processing: Convert audio data into text data using a speech recognition engine (e.g., a cloud-based API).
[1305] Output: Converted real-time text data.
[1306] Specific operation: The server subdivides the received audio data and passes it to the speech recognition engine. Parallel processing is used to perform efficient text conversion.
[1307] Step 3: Text-based summary generation
[1308] server:
[1309] Input: Text data.
[1310] Data processing: Analyze text data using natural language processing algorithms.
[1311] Data processing: Extract and summarize key points using a generative AI model (e.g., OpenAI GPT-3).
[1312] Output: Summarized text data.
[1313] Specific operation: The server inputs a prompt message into the generation AI model and generates a summary of the text data. Example of a prompt message: "Use an NLP algorithm to summarize the following text. Text: 'Development of new feature A is on track and has entered the testing phase. The development team is progressing as planned and will report the test results at the next meeting.'"
[1314] Step 4: Comparison of summary data and past meeting minutes
[1315] server:
[1316] Input: Summarized text data.
[1317] Data processing: Retrieve past meeting minutes data from the database.
[1318] Data processing: Use a comparison algorithm to match summary data with past meeting minutes.
[1319] Output: Progress evaluation and comparison results of discrepancies.
[1320] Specific operation: The server retrieves past meeting minutes using SQL queries, and a Python script detects similarities and differences between the two sets of data.
[1321] Step 5: Visualizing progress
[1322] server:
[1323] Input: Results of progress evaluations and discrepancies comparisons.
[1324] Data processing: Formatting data for visualization.
[1325] Data processing: Visualization in graph and chart formats (e.g., D3.js).
[1326] Output: The visualization results displayed on the dashboard.
[1327] Specific operation: The server sends visualization data in JSON format to the dashboard API and updates it in real time.
[1328] Step 6: Notifications and detailed display
[1329] Terminal:
[1330] Input: Visualization results sent from the server.
[1331] Data processing: Formatting the data into a form that is easy for users to understand.
[1332] Output: Real-time notifications and detailed display on the device.
[1333] Specific operation: The device displays a notification pop-up, and clicking it jumps to a details screen. The user can easily check the progress.
[1334] These steps enable the system to transcribe, summarize, assess progress, visualize meeting audio in real time, and notify users.
[1335] (Application Example 1)
[1336] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1337] Modern factory operations demand complex project management and rapid response. However, managing meeting audio and progress is often done manually, which is inefficient and makes it difficult to immediately detect delays or discrepancies. This leads to delays in appropriate responses and decreased productivity. Therefore, there is a need for a system that automatically records meeting audio, transcribes it into text, summarizes it, and evaluates progress in real time.
[1338] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[1339] In this invention, the server includes means for recording meeting audio, means for converting the recorded audio data into text data, means for summarizing the converted text data, means for comparing it with past meeting minutes data, means for detecting delays or discrepancies in progress, means for visualizing the comparison results, means for notifying the user of the results, means for a robot placed in the factory to record and process the meeting audio, and means for displaying the converted and summarized text data on a dashboard in real time. This enables the automatic recording of audio content of meetings held in the factory, improving the efficiency of progress management and allowing for quicker responses.
[1340] "Means of recording meeting audio" refers to a system that saves the audio from a meeting as digital data using a microphone or similar device.
[1341] "Means of converting recorded audio data into text data" refers to the process of converting recorded audio into text using a speech recognition engine.
[1342] "Methods for summarizing converted text data" refers to the process of using natural language processing techniques to extract only the important points from long texts and make them concise.
[1343] "Methods for comparing with past meeting minutes data" refer to a system that compares the content of the previous meeting stored in a database with the content of the current meeting to check for differences and progress.
[1344] "Methods for detecting delays or discrepancies in progress" refer to a system that compares past meeting minutes data with current data to identify differences between the plan and the actual progress.
[1345] "Methods for visualizing comparison results" refer to the process of displaying data differences and progress in visual formats such as graphs and charts.
[1346] "Means of notifying users of results" refers to a system that informs users of any delays or discrepancies in progress.
[1347] "Means for recording and processing meeting audio using robots deployed in a factory" refers to a process in which robots installed within a factory capture meeting audio and process that audio data in real time.
[1348] "A means of displaying real-time converted and summarized text data on a dashboard" refers to a mechanism that displays real-time converted and summarized text data on a dashboard so that users can understand it at a glance.
[1349] This invention is a system for automating and streamlining the progress management of meetings held in a factory. It records meeting audio, transcribes and summarizes the content, and evaluates the progress by comparing it with past meeting minutes data. Furthermore, this system has a function to visualize the progress in real time and notify the user.
[1350] Audio recording and conversion
[1351] Hardware: Robots deployed in the factory are equipped with high-sensitivity microphones to capture meeting audio in real time. This allows the system to begin collecting audio data as soon as the meeting starts.
[1352] Software: The collected audio data is sent to a server and converted into text data by a speech recognition engine (e.g., Google Web Speech API). This process is done in real time, so speech during a meeting is instantly transcribed into text.
[1353] Summary generation using natural language processing
[1354] Server: Once text data is generated, the server uses natural language processing (NLP) algorithms (e.g., the Hugging Face generative AI model) to summarize the text data. The summarized text contains key points, allowing users to efficiently grasp the information.
[1355] Comparison with past meeting minutes and progress evaluation
[1356] Server: The generated summary data is compared with past meeting minutes data stored in the database. The comparison algorithm detects any delays or discrepancies in progress. This ensures a clear understanding of the current progress against milestones and plans decided in past meetings.
[1357] Visualization and notification
[1358] Server: Detected comparison results are reflected in a visually displayed dashboard. Progress is shown in graph and chart format, allowing users to see the current progress and any issues at a glance. Furthermore, if significant delays or discrepancies are detected, the server will notify the user.
[1359] Specific example: For instance, when reporting on the progress of new feature A in a project meeting at a factory, it would look like this:
[1360] 1. The robot captures the meeting audio and sends it to the server.
[1361] 2. The server converts the audio data into text, specifically generating the text "New feature A has progressed from the development stage to the testing stage."
[1362] 3. Use an NLP algorithm to summarize the text, concluding it as "New Feature A: Testing Phase".
[1363] 4. Compare the current server status with past meeting minutes stored in the database to assess whether it is progressing according to plan.
[1364] 5. Progress is visualized on a dashboard, allowing users to check it in real time. If progress is behind schedule, users can immediately consider countermeasures.
[1365] A possible example of a specific prompt message would be something like, "We have received a progress report on Project X. You need to start on a new task, Task Y."
[1366] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[1367] Step 1:
[1368] Terminal: A robot positioned in the factory captures meeting audio and transmits it to a server in real time. The input data is audio data, and the output is audio data transferred to the server. Specifically, the robot's microphone captures the audio and transmits that audio data to the server via the network.
[1369] Step 2:
[1370] Server: Inputs received audio data into a speech recognition engine and converts it into text data. Input data is audio data, and output data is text data. Specifically, the server calls a speech recognition engine (e.g., Google Web Speech API) and converts the audio data into text format.
[1371] Step 3:
[1372] Server: Summarizes the generated text data using a natural language processing (NLP) algorithm. The input data is the transformed text data, and the output data is the summarized text data. Specifically, the server uses the Hugging Face generative AI model to extract important information from long texts and generate summaries.
[1373] Step 4:
[1374] Server: Compares summarized text data with past meeting minutes data stored in the database. The input data consists of summarized text data and past meeting minutes data, and the output data is the comparison result. Specifically, the server retrieves meeting minutes data from the past database, compares it with the current summarized data, and calculates the difference.
[1375] Step 5:
[1376] Server: Detects delays and discrepancies in progress based on comparison results. Input data is the comparison results, and output data is information about delays and discrepancies in progress. Specifically, the server uses an algorithm to evaluate the progress status and identify the gap with the plan.
[1377] Step 6:
[1378] Server: Reflects detected results on a dashboard for visual display. Input data is information about delays and discrepancies in progress, and output data is visualized graphs and charts. Specifically, the server uses the matplotlib library to graphically display the progress.
[1379] Step 7:
[1380] Server: Sends generated visualization data to the terminal and notifies the user. Input data is visualized graphs and charts, and output data is progress displayed on the user interface. Specifically, the server pushes the generated visualization data to the terminal in real time and notifies the user of important difference information.
[1381] Step 8:
[1382] User: Check the progress on the dashboard on the device and decide on countermeasures as needed. The input data is the displayed progress, and the output is the decision on the appropriate countermeasure. Specifically, the user accesses the dashboard on the device and plans the next action based on the visualized progress.
[1383] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[1384] This invention combines a system that records meeting audio in real time, automatically transcribes and summarizes its contents, and evaluates progress by comparing it with past meeting minutes data, with an emotion engine that recognizes the user's emotions from the audio data. The embodiments for carrying out this invention will be described in detail below.
[1385] Audio recording and conversion
[1386] Terminal:
[1387] When the meeting begins, audio data is captured via the microphone on the terminal. The captured audio data is sent to the server in real time.
[1388] server:
[1389] The server inputs the received audio data into the speech recognition engine, which converts it into text data. The speech recognition engine analyzes the audio signal and converts the meeting content into text format.
[1390] Summary generation using natural language processing
[1391] server:
[1392] Once the text data is generated, the server analyzes it using natural language processing (NLP) algorithms. The NLP algorithms extract key points and keywords, generating a summary of the text data. This summarization process provides a concise summary of the meeting's content.
[1393] User emotion recognition by an emotion engine
[1394] server:
[1395] Furthermore, the server inputs the voice data into the emotion engine to recognize the user's emotions. The emotion engine analyzes the user's emotions (e.g., joy, anger, sadness, surprise, etc.) from the voice signal and adds them to the text data.
[1396] Comparison with past meeting minutes and progress evaluation
[1397] server:
[1398] The generated summary data and sentiment data are compared with historical meeting minutes data stored in the database. A comparison algorithm detects delays and discrepancies with the plan. Based on the sentiment data, the importance of the delays and discrepancies is also evaluated.
[1399] Visualization and notification
[1400] server:
[1401] The detected results and sentiment data are reflected in a visually displayed dashboard. This allows users to see the progress and its emotional impact at a glance. For example, progress delays are displayed in graphs and charts, and sentiment data is indicated by colors and icons.
[1402] Terminal:
[1403] Results sent from the server are updated in real time on the device. If significant delays or discrepancies are detected, a notification will pop up along with sentiment data.
[1404] Specific example
[1405] For example, suppose a weekly project meeting is held to discuss the progress of new feature A. If, during the meeting, it is stated, "Development of new feature A is on track and has entered the testing phase":
[1406] 1. Terminal: Captures audio and sends it to the server.
[1407] 2. Server: Converts audio data into text data.
[1408] 3. Server: Generates summaries using NLP algorithms.
[1409] 4. Server: The emotion engine analyzes the emotion of "joy" and assigns it to the summary. For example, it might be assigned as "New Feature A: Progressing to Test Phase (Joy)".
[1410] 5. Server: Compare the current meeting minutes with those stored in the database to assess whether progress is on schedule. Evaluate the importance of any delays, including sentiment data.
[1411] 6. Server: Visualize the results on a dashboard and notify the user.
[1412] 7. Device: Displayed on the user's device, allowing them to check progress and sentiment data.
[1413] 8. User: Based on the results, take appropriate action regarding the project's progress.
[1414] By combining this with an emotion engine, it becomes possible to not only automatically summarize meeting content but also to evaluate progress while considering the user's emotions, resulting in more accurate project management. This improves the project success rate and enhances team engagement.
[1415] The following describes the processing flow.
[1416] Step 1:
[1417] Terminal: As soon as the meeting starts, the microphone on the terminal captures audio data. The captured audio data is sent to the server in real time.
[1418] Step 2:
[1419] Server: Inputs received audio data into the speech recognition engine and converts it into text data. For example, it recognizes the audio data "The next milestone is next Monday" and converts it into the corresponding text data.
[1420] Step 3:
[1421] Server: Inputs the converted text data into a natural language processing (NLP) algorithm to generate a summary. For example, it summarizes the sentence "The next milestone is next Monday" as "Milestone: Monday".
[1422] Step 4:
[1423] Server: The generated text data is input into the emotion engine, which analyzes the tone and tempo of the voice to recognize the user's emotions. For example, "The next milestone is next Monday" is tagged with the emotion "joy."
[1424] Step 5:
[1425] Server: Compares summarized text data and sentiment data with past meeting minutes data stored in the database to detect delays or discrepancies in progress. For example, it might detect if a milestone previously set for "this Friday" has been postponed to "next Monday."
[1426] Step 6:
[1427] Server: Analyzes comparison results and sentiment data to identify delays and discrepancies in progress and assess their importance. For example, delays in progress are displayed along with the emotion of "joy," indicating positive progress.
[1428] Step 7:
[1429] Server: Visualizes detected results and sentiment data and displays them on a dashboard. For example, areas with delays are shown in red, and sentiment data is indicated by sentiment icons (e.g., a smile).
[1430] Step 8:
[1431] Terminal: Receives visualization data sent from the server and displays it on the user's screen in real time. If significant progress delays or discrepancies are detected, it notifies the user and displays the details along with sentiment data.
[1432] Step 9:
[1433] User: Review the latest meeting minutes, progress, and sentiment data displayed on the device. The user considers and takes appropriate action as needed.
[1434] As a concrete example, consider a situation where the progress of new feature A is being discussed in a weekly project meeting:
[1435] 1. Terminal: Captures meeting audio in real time and sends it to the server.
[1436] 2. Server: Converts audio data into text data.
[1437] 3. Server: Generates summaries using NLP algorithms.
[1438] 4. Server: The emotion engine analyzes the emotion of "joy" and assigns it to the summary. This will result in "New Feature A: Proceeding to Test Phase (Joy)".
[1439] 5. Server: Compare the current meeting minutes with those stored in the database to assess whether progress is on schedule. Consider sentiment data to evaluate the importance of any delays.
[1440] 6. Server: Visualize the results on a dashboard and notify the user.
[1441] 7. Device: The results are displayed on the user's device, allowing them to check the progress and sentiment data.
[1442] 8. User: Respond to project progress based on the results.
[1443] This integrates automated summarization of meeting content, progress evaluation, and consideration of sentiment data, enabling effective project management.
[1444] (Example 2)
[1445] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1446] Traditional meeting minutes systems simply transcribe meeting audio into text, failing to consider the meeting's progress or emotional fluctuations. Furthermore, they lacked automated summarization and comparison with past minutes, requiring users to perform these tasks manually, which was extremely time-consuming. Additionally, the lack of technology to analyze user emotions and reflect them in meeting content limited their ability to improve team engagement and project success rates.
[1447] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[1448] In this invention, the server includes means for recording meeting audio, means for converting the recorded audio data into text data, means for summarizing the converted text data, means for comparing the summarized text data with past recorded data, means for detecting delays or discrepancies in progress based on the comparison results, means for visualizing the comparison results, means for notifying the user of the results, means for analyzing the user's emotions from the audio data, and means for adding the analyzed emotion data to the text data. This enables a comprehensive analysis including detailed progress evaluation of meetings and fluctuations in emotions, thereby improving the efficiency of project management and strengthening team engagement.
[1449] "Means for recording meeting audio" refers to devices or programs that capture audio during a meeting in real time and save it as digital data.
[1450] "Means for converting recorded audio data into text data" refers to speech recognition technologies and software that analyze audio signals and recognize and convert them into text information.
[1451] "Means for summarizing converted text data" refers to natural language processing techniques that analyze text data to extract important information and shorten the original text to create a summary.
[1452] "Means for comparing the summarized text data with past record data" refers to an algorithm for comparing the generated summary data with existing meeting record data and performing analysis of progress, agreements, and disagreements.
[1453] "Means for detecting delays or discrepancies in progress based on the comparison results" refers to an analytical system that identifies the gap between the expected progress and the actual progress based on the comparison results, and detects delays or problems.
[1454] "Means for visualizing the comparison results" refers to programs or tools for displaying comparison results and progress information in a visual format such as graphs or charts that are easy for users to understand.
[1455] "Means for notifying the user of the results" refers to a program or device that has a notification function to inform the user of analysis results and important information in real time.
[1456] "Methods for analyzing user emotions from voice data" refer to technologies that determine a user's emotional state from factors such as voice intonation, tempo, and volume, and classify it into emotional categories (e.g., joy, anger, sadness, surprise, etc.).
[1457] "Means for attaching analyzed emotional data to text data" refers to a system that attaches analyzed emotional information as part of the corresponding text data, thereby clearly indicating the emotional state.
[1458] This invention combines a system that records meeting audio in real time, automatically transcribes and summarizes its contents, and evaluates progress by comparing it with past meeting minutes data, with technology that recognizes user emotions from the audio data. The embodiments for carrying out this invention will be described in detail below.
[1459] Audio recording and conversion
[1460] Terminal:
[1461] When a meeting begins, audio data is captured through the terminal's microphone. The captured audio data is sent to the server in real time. The hardware used is either the terminal's built-in microphone or an external microphone.
[1462] server:
[1463] The server inputs the received audio data into a speech recognition engine (e.g., Google Cloud Speech-to-Text) and converts it into text data. It analyzes the audio signal and converts the meeting content into text format.
[1464] Summary generation using natural language processing
[1465] server:
[1466] Once the text data is generated, the server analyzes it using natural language processing (NLP) algorithms (e.g., Google Cloud Natural Language API). The NLP algorithms extract key points and keywords, generating a summary of the text data. This summarization process provides a concise summary of the meeting content.
[1467] User emotion recognition by an emotion engine
[1468] server:
[1469] Furthermore, the server inputs the audio data into an emotion analysis engine (e.g., Amazon Comprehend) to recognize the user's emotions. The emotion analysis engine analyzes the user's emotions (e.g., joy, anger, sadness, surprise, etc.) from the audio signal and adds them to the text data.
[1470] Comparison with past meeting minutes and progress evaluation
[1471] server:
[1472] The generated summary and sentiment data are compared with historical meeting minutes data stored in a database (e.g., Amazon RDS). A comparison algorithm detects delays and discrepancies with the plan. Based on the sentiment data, the importance of the delays and discrepancies is also assessed.
[1473] Visualization and notification
[1474] server:
[1475] The detected results and sentiment data are reflected in a visually displayed dashboard. This allows users to see the progress and its emotional impact at a glance. For example, progress delays are displayed in graphs and charts, and sentiment data is indicated by colors and icons.
[1476] Terminal:
[1477] Results sent from the server are updated in real time on the device. If significant delays or discrepancies are detected, a notification will pop up along with sentiment data.
[1478] Specific example
[1479] For example, suppose a weekly project meeting is held to discuss the progress of new feature A. If, during the meeting, it is stated, "Development of new feature A is on track and has entered the testing phase":
[1480] 1. Terminal: Captures audio and sends it to the server.
[1481] 2. Server: Converts audio data into text data using Google Cloud Speech-to-Text.
[1482] 3. Server: Use the Google Cloud Natural Language API to generate summaries using NLP algorithms.
[1483] 4. Server: Use Amazon Comprehend to analyze the emotion of "joy" and assign it to the summary. For example, assign it as "New Feature A: Progressing to the testing phase (joy)".
[1484] 5. Server: Compare the current meeting minutes with those stored in Amazon RDS to assess whether progress is on schedule. Evaluate the importance of any delays, including sentiment data.
[1485] 6. Server: Visually display the results on a dashboard and notify the user.
[1486] 7. Device: Displayed on the user's device, allowing them to check progress and sentiment data.
[1487] 8. User: Based on the results, take appropriate action regarding the project's progress.
[1488] Example of a prompt
[1489] "Recognize and summarize in text the emotions you would feel when reporting on the progress of new feature A at this week's project meeting."
[1490] Through these concrete actions, the overall system becomes easier to understand, leading to more effective project management and progress evaluation.
[1491] The flow of the specific processing in Example 2 will be explained using Figure 13.
[1492] Step 1: Capture and send audio data
[1493] Terminal: When the meeting begins, audio data is captured through the terminal's microphone. The input is the meeting audio, which is captured as audio data. The captured audio data is sent to the server in real time. Specifically, the terminal's microphone converts the audio signal into digital data and sends it to the server using the HTTPS protocol. The output is the audio data sent to the server.
[1494] Step 2: Converting audio data to text
[1495] Server: The server inputs the received audio data into the Google Cloud Speech-to-Text speech recognition engine. The input is the received audio data, which the server parses and converts into text data. Specifically, the server sends the audio data to the speech recognition engine via an API call and receives the text data extracted from the audio signal. The output is the generated text data.
[1496] Step 3: Summary Generation
[1497] Server: The server inputs the generated text data into the Google Cloud Natural Language API. The input is the generated text data, and the server uses NLP algorithms to extract important points and keywords and summarize the text data. Specifically, the server applies an algorithm that analyzes the text data, extracts the important parts, and generates a summary. The output is the summarized text data.
[1498] Step 4: Emotion Recognition
[1499] Server: The server inputs the summarized text data and the original audio data into Amazon Comprehend's sentiment analysis engine. The input consists of summarized text data and the original audio data, and the sentiment analysis engine analyzes the user's emotions from the audio signal. Specifically, the sentiment analysis engine analyzes the audio data, identifies emotional states (e.g., joy, anger, sadness, surprise, etc.), and adds that information to the summarized text. The output is the summarized text with the sentiment data added.
[1500] Step 5: Comparison with past data and progress evaluation
[1501] Server: The server compares the generated sentiment-data-enhanced summary text with historical meeting minutes data stored in Amazon RDS. The input is the sentiment-data-enhanced summary text and historical meeting minutes data, and the server compares this data to detect delays in progress and discrepancies with the plan. Specifically, the server retrieves historical data and applies a comparison algorithm to detect progress and discrepancies. The output is the evaluation result of progress and discrepancies.
[1502] Step 6: Visualize and notify results
[1503] Server: The server reflects the detected progress and discrepancy evaluation results, as well as sentiment data, in a dashboard for visual display. Inputs are progress and discrepancy evaluation results and sentiment data, which are then visualized on the dashboard in graph and chart format. Specifically, the server generates evaluation results as graphs and charts and displays them to the user via a web interface. Output is the visualized information on the dashboard.
[1504] Terminal: The terminal updates the results sent from the server in real time. The input consists of progress and discrepancy evaluation results and sentiment data sent from the server, which are displayed in real time as notification pop-ups. Specifically, the terminal's notification system receives new data and displays a pop-up notification to the user. The output is the pop-up notification.
[1505] Through these steps, the system can perform detailed progress evaluations and sentiment analyses of meeting content. This allows users to understand project progress in real time and take necessary actions quickly.
[1506] (Application Example 2)
[1507] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1508] Current factory automation systems struggle to record communication between robots and workers, and to track work progress. Furthermore, progress management doesn't consider worker emotions, potentially leading to decreased efficiency and lower worker morale. To address these issues, a system is needed that analyzes and evaluates progress and worker emotions in real time, providing appropriate feedback.
[1509] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[1510] In this invention, the server includes means for recording instructions between robots and workers in a factory, means for recognizing user emotions based on voice data, means for incorporating the emotion data into progress evaluation, and means for displaying the voice data and emotion data on a dashboard. This enables real-time progress management and emotion evaluation of robots and workers in a factory.
[1511] "Methods for recording meeting audio" refers to technologies that collect audio information during a meeting using microphones or other audio capture devices and store it as digital data.
[1512] "Means of converting recorded audio data into text data" refers to the process of converting audio data into text format data using conversion technologies such as a speech recognition engine.
[1513] "Methods for summarizing converted text data" refer to techniques that utilize natural language processing (NLP) algorithms to extract key points from large amounts of text data and condense them into short summaries.
[1514] "Means for comparing the summarized text data with past meeting minutes data" refers to a technique that compares newly generated summarized data with previously stored data to identify progress and changes.
[1515] "Means for detecting delays or discrepancies in progress based on the aforementioned comparison results" refers to a mechanism that analyzes the comparison results and evaluates whether the work or project is progressing according to plan.
[1516] "Means for visualizing the comparison results" refers to technologies that display the comparison results in a visual format, such as graphs or charts, so that users can easily understand them.
[1517] "Means for notifying the user of the results" refers to a notification function that communicates detected progress delays, discrepancies, and other important information to the user in real time.
[1518] "A means of recognizing user emotions based on voice data" refers to an emotion recognition engine that analyzes and recognizes human emotions (joy, anger, sadness, surprise, etc.) from voice.
[1519] "Means for incorporating the aforementioned emotional data into progress evaluations" refers to a system that incorporates users' emotional data into work and project progress evaluations to provide a more comprehensive assessment.
[1520] "A means of recording instructions between robots and workers in a factory and evaluating progress and emotions" refers to a technology that records instructions exchanged between robots and workers in a factory and uses that data to evaluate the progress of the work and the emotions of the workers.
[1521] "Means for displaying audio data and sentiment data on a dashboard" refers to technology that provides a dashboard function for visually displaying collected audio data and sentiment data.
[1522] The embodiments for carrying out this invention will be described in detail. A system for recording instructions between a robot and a worker in a factory and evaluating progress and worker emotions consists of the following steps.
[1523] First, the robots in the factory record voice instructions from workers via microphones. The recorded voice data is sent to a server in real time. The server uses a speech recognition engine (e.g., Google Speech Recognition API) to convert this voice data into text data.
[1524] The converted text data is analyzed using natural language processing (NLP) algorithms to extract key points and generate short summaries. These summaries are compared to past work records to assess progress. This helps determine whether the work is on schedule or behind schedule. Simultaneously, the server inputs audio data into an emotion recognition engine to analyze the worker's emotions. This allows the worker's emotions (e.g., joy, anger, sadness, surprise) to be added to the text data.
[1525] This data is stored in a database and used to evaluate progress and discrepancies by comparing it with past work records. The evaluation results are visually displayed on a visualization dashboard, allowing users to grasp the progress and sentiment status at a glance. For example, delays in progress are displayed in graph or chart format, and sentiment data is indicated by colors and icons. If significant delays or discrepancies are detected, a notification pops up on the user's device along with the sentiment data.
[1526] Hardware and software to be used
[1527] The hardware used will be a microphone (which can be built into a smartphone or smart glasses, etc.). The software will include Python, Google Speech Recognition API, NLTK, scikit-learn, a dedicated emotion recognition engine, and a web application for displaying the dashboard (e.g., Dash, Flask).
[1528] Specific example
[1529] For example, consider a scenario in a factory where a robot operates under the instructions of a worker. When the worker says, "I'm ready to proceed to the next work stage," the following process is executed:
[1530] Voice recording: The robot captures the worker's voice and sends it to the server.
[1531] Text Conversion: The Google Speech Recognition API is used to convert the audio data into text data. The resulting text will read, "Ready to proceed to the next work stage."
[1532] Summary generation: Generates summaries using natural language processing algorithms.
[1533] Emotion Analysis: An emotion recognition engine analyzes the emotion of "joy" from the voice.
[1534] Progress evaluation: Evaluate whether progress is "good" by comparing it with past work record data.
[1535] Dashboard display: The generated summary, sentiment data, and progress assessment are displayed on a dashboard for visual confirmation.
[1536] Example of a prompt
[1537] A factory robot receives the instruction, "Ready to proceed to the next work stage." Translate this instruction into text, summarize it, and evaluate the robot's work progress and emotional state.
[1538] Thus, the present invention's system efficiently records and manages communication between robots and workers in a factory, and simultaneously evaluates work progress and emotions, thereby realizing an efficient and harmonious work environment.
[1539] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[1540] Step 1:
[1541] Audio recording
[1542] The terminal uses microphones within the factory to record voice instructions from workers in real time. The terminal then transmits the recorded voice data to a server.
[1543] Input: Voice instructions from the worker
[1544] Output: Audio data
[1545] Step 2:
[1546] Text conversion
[1547] The server inputs the received audio data into a speech recognition engine (e.g., Google Speech Recognition API) and converts the audio data into text data.
[1548] Input: Audio data
[1549] Output: Text data
[1550] Step 3:
[1551] Summary generation
[1552] The server analyzes the converted text data using natural language processing (NLP) algorithms, extracts key points, and generates a short summary.
[1553] Input: Text data
[1554] Output: Summary data
[1555] Step 4:
[1556] sentiment analysis
[1557] The server inputs the voice data into an emotion recognition engine and recognizes the worker's emotions (e.g., joy, anger, sadness, surprise, etc.). The recognized emotion data is then added to the text data.
[1558] Input: Audio data
[1559] Output: Sentiment data
[1560] Step 5:
[1561] Progress evaluation
[1562] The server compares the generated summary data and sentiment data with past work record data stored in the database to assess delays and discrepancies. It also evaluates the importance of delays and discrepancies based on the sentiment data.
[1563] Input: Summary data, sentiment data, past work record data
[1564] Output: Progress evaluation data
[1565] Step 6:
[1566] visualization
[1567] The server converts progress evaluation data and sentiment data into a visual format (e.g., graphs and charts) for display on the dashboard.
[1568] Input: Progress evaluation data, sentiment data
[1569] Output: Visualization data for dashboard
[1570] Step 7:
[1571] notification
[1572] If the server detects any significant delays or discrepancies in the project, it will notify the user's terminal in real time.
[1573] Input: Progress evaluation data, sentiment data
[1574] Output: User notification
[1575] These steps enable the system to record the collaboration between robots and workers in the factory, and to monitor and evaluate progress and emotions in real time.
[1576] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[1577] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet Search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1578] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[1579] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1580] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[1581] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[1582] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[1583] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[1584] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[1585] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[1586] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[1587] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[1588] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[1589] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1590] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[1591] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[1592] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[1593] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[1594] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[1595] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[1596] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[1597] The following is further disclosed regarding the embodiments described above.
[1598] (Claim 1)
[1599] Means of recording meeting audio,
[1600] A means of converting recorded audio data into text data,
[1601] A means of summarizing the converted text data,
[1602] A means for comparing the summarized text data with past meeting minutes data,
[1603] A means for detecting delays or discrepancies in progress based on the aforementioned comparison results,
[1604] Means for visualizing the comparison results,
[1605] Means for notifying the user of the results,
[1606] A system that includes this.
[1607] (Claim 2)
[1608] The system according to claim 1, further comprising means for processing the aforementioned audio data in real time.
[1609] (Claim 3)
[1610] The system according to claim 1, further comprising means for using a speech recognition engine to convert the aforementioned speech data into text.
[1611] "Example 1"
[1612] (Claim 1)
[1613] A means of capturing meeting audio,
[1614] A means of transmitting captured audio data in real time,
[1615] A means for converting received audio data into text data using a speech recognition engine,
[1616] A means of summarizing text data using a natural language processing algorithm,
[1617] A means of comparing summarized text data with past meeting minutes data,
[1618] A means for detecting delays or discrepancies in progress based on comparison results,
[1619] A means of visualizing the comparison results,
[1620] A means of notifying the user of the visualized results,
[1621] A system that includes this.
[1622] (Claim 2)
[1623] The system according to claim 1, further comprising means for converting audio data into text in real time by parallel processing.
[1624] (Claim 3)
[1625] The system according to claim 1, comprising means of using a generative AI model in a natural language processing algorithm.
[1626] "Application Example 1"
[1627] (Claim 1)
[1628] Means of recording meeting audio,
[1629] A means of converting recorded audio data into text data,
[1630] A means of summarizing the converted text data,
[1631] A means for comparing the summarized text data with past meeting minutes data,
[1632] A means for detecting delays or discrepancies in progress based on the aforementioned comparison results,
[1633] Means for visualizing the comparison results,
[1634] Means for notifying the user of the results,
[1635] A means for robots deployed in a factory to record and process meeting audio,
[1636] A means of displaying the converted and summarized text data on a dashboard in real time,
[1637] A system that includes this.
[1638] (Claim 2)
[1639] The system according to claim 1, further comprising means for processing the aforementioned audio data in real time.
[1640] (Claim 3)
[1641] The system according to claim 1, further comprising means for using a speech recognition engine to convert the aforementioned speech data into text.
[1642] "Example 2 of combining an emotion engine"
[1643] (Claim 1)
[1644] Means of recording meeting audio,
[1645] A means of converting recorded audio data into text data,
[1646] A means of summarizing the converted text data,
[1647] A means for comparing the summarized text data with past record data,
[1648] A means for detecting delays or discrepancies in progress based on the aforementioned comparison results,
[1649] Means for visualizing the comparison results,
[1650] Means for notifying the user of the results,
[1651] A method for analyzing user emotions from voice data,
[1652] A means of adding analyzed emotion data to text data,
[1653] A system that includes this.
[1654] (Claim 2)
[1655] The system according to claim 1, further comprising means for processing the aforementioned audio data in real time.
[1656] (Claim 3)
[1657] The system according to claim 1, further comprising means for using a speech recognition engine to convert the aforementioned speech data into text.
[1658] "Application example 2 when combining with an emotional engine"
[1659] (Claim 1)
[1660] Means of recording meeting audio,
[1661] A means of converting recorded audio data into text data,
[1662] A means of summarizing the converted text data,
[1663] A means for comparing the summarized text data with past meeting minutes data,
[1664] A means for detecting delays or discrepancies in progress based on the aforementioned comparison results,
[1665] Means for visualizing the comparison results,
[1666] Means for notifying the user of the results,
[1667] A means of recognizing user emotions based on voice data,
[1668] A means for adding the aforementioned emotional data to the progress evaluation,
[1669] A means of recording instructions between robots and workers in a factory and evaluating progress and emotions,
[1670] A means of displaying voice data and emotion data on a dashboard,
[1671] A system that includes this.
[1672] (Claim 2)
[1673] The system according to claim 1, further comprising means for processing the aforementioned audio data in real time.
[1674] (Claim 3)
[1675] The system according to claim 1, further comprising means for using a speech recognition engine to convert the aforementioned speech data into text. [Explanation of symbols]
[1676] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. Means of recording meeting audio, A means of converting recorded audio data into text data, A means of summarizing the converted text data, A means for comparing the summarized text data with past meeting minutes data, A means for detecting delays or discrepancies in progress based on the aforementioned comparison results, Means for visualizing the comparison results, Means for notifying the user of the results, A system that includes this.
2. The system according to claim 1, further comprising means for processing the aforementioned audio data in real time.
3. The system according to claim 1, further comprising means for using a speech recognition engine to convert the aforementioned speech data into text.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A