System

A system that records, analyzes, and provides feedback on meeting content enhances communication skills and productivity by addressing poor meeting quality issues.

JP2026018015APending Publication Date: 2026-02-05SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024119076
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-24
Publication Date
2026-02-05

AI Technical Summary

Technical Problem

Poor communication skills and inappropriate discussion structure in meetings negatively impact productivity, and participants are often unaware of these issues, making self-improvement difficult.

Method used

A system that records meetings, extracts audio, converts it to text, analyzes the text and audio for vocabulary, syntax, tone, and speaking rate, generates feedback, and presents it to participants for improvement.

Benefits of technology

Improves individual communication skills and overall meeting quality, increasing organizational productivity by providing specific and actionable feedback.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026018015000001_ABST
    Figure 2026018015000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system comprising: means for recording content of a meeting; means for extracting audio from the recorded data; means for converting the extracted audio to text; means for analyzing the text to evaluate linguistics, syntax, and structure of discussion; means for analyzing audio data to evaluate voice tone and speaking speed; means for generating feedback based on the analysis; and means for presenting the generated feedback to a user.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] Meetings play an important role in corporate activities, but if the quality of meetings is not improved, it often has a negative impact on productivity. Specifically, poor communication skills among participants and inappropriate discussion structure can be factors. In many cases, participants themselves are unaware of these problems, making self-improvement difficult. Therefore, there is a need for effective tools to improve the quality of meetings and increase productivity throughout the organization. [Means for solving the problem]

[0005] This invention solves this problem with a system that includes a means for recording the contents of a meeting, a means for extracting audio from the recorded data, a means for converting the extracted audio into text, a means for analyzing the text to evaluate vocabulary, syntax, and argument structure, a means for analyzing the audio data to evaluate tone of voice and speaking rate, a means for generating feedback based on the analysis results, and a means for presenting the generated feedback to a user. This system can provide specific and useful feedback to each meeting participant, improving their individual communication skills and logical thinking abilities. This is expected to improve the overall quality of meetings and increase productivity throughout the organization.

[0006] "Means for recording the contents of a meeting" refers to equipment or a system that records the video and audio of a meeting.

[0007] "Means for extracting audio from recorded data" refers to a method or device for separating and acquiring the audio portion from the recorded conference content.

[0008] "Means for converting speech to text" refers to a method or device that converts speech data into text data using speech recognition technology.

[0009] "Means for analyzing text" refers to methods or systems that analyze converted text data using natural language processing (NLP) techniques to evaluate phrasing, syntax, and argument structure.

[0010] "Means for analyzing audio data" refers to a method or system for analyzing audio signals to evaluate acoustic characteristics such as tone of voice and speaking rate.

[0011] "Means for generating feedback" refers to a method or device that generates feedback including suggestions for improvement and advice for participants based on the results of text and speech analysis.

[0012] "Means for presenting feedback to the user" refers to an interface or system for providing the generated feedback to the user visually or audibly. [Brief explanation of the drawings]

[0013] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14]FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0014] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0015] First, the terms used in the following description will be explained.

[0016] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0017] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0018] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0019] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0020] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0021] [First embodiment]

[0022] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0023] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0024] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0025] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0026] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0027] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0028] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0029] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0030] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0031] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0032] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0033] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0034] The present invention relates to a system for improving the quality of meetings by utilizing AI technology. A specific embodiment of this system will be described.

[0035] 1. Recording meetings

[0036] The device captures video and audio as soon as the meeting starts. Specifically, the entire meeting is recorded using the camera and microphone of a laptop or smartphone, for example. The recorded data is uploaded to a server in real time.

[0037] 2. Audio Extraction

[0038] The server extracts the audio portion from the received recording. This extraction can be done using common audio processing techniques, including, for example, separating the audio stream from the combined video and audio of the meeting.

[0039] 3. Speech-to-text

[0040] The server then requests a speech recognition engine to convert the extracted voice data. Specifically, it calls a service such as the Google Cloud Speech-to-Text API to convert the voice into text data. As a result, the spoken content is obtained in text format.

[0041] 4. Text Analysis

[0042] The server analyzes the text data. Natural language processing (NLP) techniques are used for text analysis. For example, Python NLP libraries such as spaCy and BERT can be used to tokenize the text, tag parts of speech, and perform semantic analysis. This analysis makes it possible to evaluate the vocabulary, syntax, and argument structure of the speech.

[0043] 5. Evaluation of audio data

[0044] The server analyzes the audio data and evaluates the tone of voice and speaking rate. This includes analyzing acoustic features. For example, it uses Praat or an open-source audio analysis library to measure pitch (highness of voice) and intensity (power of voice) and makes an evaluation based on these.

[0045] 6. Generate feedback

[0046] The server generates feedback based on the results of text and speech analysis. The feedback is provided to each participant in the form of an evaluation result and recommendations for improvement. For example, it may include specific comments such as "Mr. Sato speaks too quickly" and suggestions for improvement such as "Try to speak a little more slowly next time."

[0047] 7. Providing Feedback

[0048] The device presents the generated feedback to the user. The feedback is displayed visually in a web browser or mobile app, allowing the user to intuitively identify problems and areas for improvement in their comments. For example, feedback could be displayed separately for each comment, and a visually easy-to-understand format could be used, using color coding or graphs.

[0049] Using this system, meeting participants can improve their communication skills based on specific feedback, which is expected to improve the overall quality of meetings and increase productivity across the organization.

[0050] The processing flow will be explained below.

[0051] Step 1:

[0052] The device captures video and audio as soon as the meeting starts. For example, the entire meeting can be recorded using the camera and microphone of a laptop or smartphone, and the data can be sent to a server in real time.

[0053] Step 2:

[0054] The server extracts the audio portion from the recorded data it receives. Specifically, it uses audio processing technology to separate the audio stream from the combined video and audio data.

[0055] Step 3:

[0056] The server inputs the extracted voice data into a speech recognition engine and converts the voice to text, for example, using the Google Cloud Speech-to-Text API.

[0057] Step 4:

[0058] The server analyzes the text data, using natural language processing (NLP) techniques to tokenize, tag parts of speech, and analyze the semantics of the text. This analysis evaluates the phrasing, syntax, and argument structure of the speech.

[0059] Step 5:

[0060] The server uses the audio data to evaluate the tone of voice and speaking rate. Specifically, it analyzes acoustic features to measure pitch (highness of voice) and intensity (power of voice).

[0061] Step 6:

[0062] The server generates feedback based on the results of text and speech analysis, for example, using a specific template to create feedback for each participant, including specific evaluations and suggestions for improvement.

[0063] Step 7:

[0064] The device then presents the generated feedback to the user, which is displayed visually via a smartphone, PC web browser, or dedicated app. The user can then review the displayed feedback and understand their areas for improvement.

[0065] Example 1

[0066] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0067] While conventional conference systems can record conference content and convert audio into text, they lack the functionality to evaluate the quality of the conference and provide specific feedback. As a result, users have no way to objectively evaluate their own comments or the way the conference proceeded, which hinders the improvement of their communication skills. The present invention aims to solve these problems and provide a system that improves the quality of conferences.

[0068] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0069] In this invention, the server includes means for recording the contents of the conference, means for uploading the recorded data to the server in real time, means for extracting voice from the recorded data, means for transmitting the extracted voice to a voice recognition engine and converting it into text data, means for analyzing the text data using natural language processing technology, means for analyzing the voice data using acoustic analysis technology and evaluating the tone of voice and speaking rate, means for generating feedback based on the results of the text analysis and the voice analysis, and means for presenting the generated feedback to the user. This allows conference participants to receive evaluations of their own remarks and voice and know specific areas for improvement, thereby enabling them to improve their communication skills.

[0070] A "meeting recording device" is a device or method that captures and stores video and audio of a meeting in digital form.

[0071] "Means for uploading recorded data to a server in real time" refers to the function of instantly transmitting video and audio data generated during the progress of a conference to a server.

[0072] "Means for extracting audio from recorded data" refers to a technology for separating and extracting the audio portion from recorded video data.

[0073] "Means for transmitting the extracted speech to a speech recognition engine and converting it into text data" refers to the process of transmitting the extracted speech data to speech recognition software and using that software to convert the speech into text information.

[0074] "Means for analyzing text data using natural language processing technology" refers to a technology that analyzes acquired text data using a natural language processing algorithm and evaluates the wording, syntax, and structure of the argument.

[0075] "Means for analyzing voice data using acoustic analysis technology to evaluate voice tone and speaking rate" refers to technology for analyzing voice data using an acoustic analysis tool and measuring pitch and intensity to evaluate voice tone and speaking rate.

[0076] The "means for generating feedback based on text analysis results and speech analysis results" refers to a process of combining the analysis results of text data and speech data to create feedback that includes specific evaluations and improvement measures.

[0077] The "means for presenting the generated feedback to the user" is a device or method that visually or audibly conveys the generated feedback to the user.

[0078] The present invention relates to a system for improving the quality of meetings by utilizing AI technology. Specific embodiments of this system are described in detail below.

[0079] First, at the start of a meeting, the user launches the meeting application and presses the "Start Meeting" button. Devices (e.g., laptops, smartphones) that detect the start of the meeting begin capturing video and audio using their built-in cameras and microphones. This allows the entire meeting to be recorded.

[0080] The captured video and audio data is uploaded to a server in real time using HTTP streaming or WebRTC technology, for example. The server processes the received recording data and separates the video and audio data. Specifically, audio processing technology such as FFmpeg is used to extract the audio stream from the recording data.

[0081] The extracted voice data is then sent by the server to the Google Cloud Speech-to-Text API, where it is converted into text data. This step allows the voice utterance to be obtained in text format.

[0082] The captured text data is then analyzed by the server using natural language processing (NLP) techniques, such as tokenizing, part-of-speech tagging, and semantic analysis using Python NLP libraries (e.g., spaCy, BERT), to evaluate the vocabulary, syntax, and argument structure of the speech.

[0083] Additionally, the voice data is analyzed using acoustic analysis technology. The server analyzes the voice data using an acoustic analysis library (e.g., Librosa) to measure pitch and intensity. This evaluation evaluates the tone of voice and speaking rate.

[0084] The server generates feedback based on these analysis results. The generated feedback includes specific suggestions for improvement and evaluations, such as "You speak too quickly" or "Try to speak more slowly next time." Finally, the generated feedback is presented to the user via the device. The feedback is displayed visually in a web browser or mobile app, and is presented in an easily understandable format using color-coded graphs and icons.

[0085] For example, when a user holds a meeting using video conferencing software, the laptop's camera and microphone capture video and audio and upload them to a server in real time. The server then extracts the audio using FFmpeg and converts it to text using the Google Cloud Speech-to-Text API. The text is then analyzed using spaCy, and the audio is evaluated using Librosa. Finally, feedback generated based on the analysis results is displayed in the web browser for the user to review.

[0086] An example of a prompt might be, "Please use the Google Cloud Speech-to-Text API to convert the following audio data into text and generate meeting feedback based on the results." By using this system, meeting participants can receive evaluations of their own remarks and voices, learn specific areas for improvement, and improve their communication skills.

[0087] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0088] Step 1:

[0089] The user launches the conference application and presses the "Start Conference" button. This causes the device to detect the start of the conference. The input is the user's operation action, and the output is the start of the conference recording process.

[0090] Step 2:

[0091] The device will begin capturing video and audio using the built-in camera and microphone. The captured data will be saved as video and audio files. The input is real-time data from the built-in camera and microphone, and the output is locally saved video and audio data.

[0092] Step 3:

[0093] The device uploads the captured video and audio data to the server in real time. This process uses HTTP streaming or WebRTC technology. The input is the locally stored video and audio data, and the output is the data uploaded to the server.

[0094] Step 4:

[0095] The server processes the received recording data and separates the video and audio data using audio processing technology such as FFmpeg. The input is the uploaded recording data, and the output is the separated video and audio data.

[0096] Step 5:

[0097] The server sends the extracted speech data to a speech recognition engine and converts it into text data. This step uses the Google Cloud Speech-to-Text API. The extracted speech data is the input, and the text data is the output.

[0098] Step 6:

[0099] The server analyzes the acquired text data using natural language processing (NLP) techniques. Specifically, it uses Python NLP libraries (e.g., spaCy, BERT) to perform tokenization, part-of-speech tagging, and semantic analysis. The converted text data is input, and the analysis results are obtained as output.

[0100] Step 7:

[0101] The server analyzes the voice data using acoustic analysis technology to evaluate the tone of voice and speaking rate. Specifically, it uses the Librosa library to measure pitch and intensity. The input is the voice data, and the output is the result of the acoustic analysis.

[0102] Step 8:

[0103] The server generates feedback based on the results of text analysis and speech analysis. The feedback includes specific evaluations and suggestions for improvement. The inputs are the results of text analysis and speech analysis, and the generated feedback is obtained as the output.

[0104] Step 9:

[0105] The device presents the generated feedback to the user, which is then visually displayed in a web browser or mobile app. The input is the generated feedback, and the output is the feedback presented to the user.

[0106] (Application example 1)

[0107] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0108] Conventional conference systems lack specific feedback to improve the quality of meetings, making it difficult for participants to improve their own comments and attitudes. Furthermore, in certain environments, such as project review meetings in factories, recording and analysis methods and tools are often not properly installed, hindering the effective running of meetings. The present invention aims to provide a system that solves these problems and improves the quality of meetings.

[0109] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0110] In this invention, the server includes a means for recording the contents of the conference, a means for extracting audio from the recorded data, and a means for converting the extracted audio into text, thereby enabling the server to accurately record the overall picture of the conference and provide specific feedback based on the analysis results.

[0111] "Means for recording the contents of a conference" refers to a device or program that captures and records video and audio as soon as the conference begins.

[0112] The "means for extracting audio from recorded data" refers to a processing device or program for separating the audio portion from the recorded conference data and extracting it as audio data.

[0113] "Means for converting extracted speech into text" refers to a program or service that converts extracted speech data into text format using speech recognition technology.

[0114] A "means for analyzing text to evaluate language, syntax, and argument structure" is a program that uses natural language processing techniques on text data to evaluate and analyze the content of the text.

[0115] The "means for analyzing voice data to evaluate voice tone and speaking rate" refers to a device or program that uses voice data to analyze acoustic features such as voice pitch and speaking rate and make an evaluation.

[0116] The "means for generating feedback based on the analysis results" is a program that provides participants with specific suggestions for improvement and evaluations based on the results of text and voice analysis.

[0117] The "means for presenting the generated feedback to the user" refers to a device or program that visually displays the generated feedback so that the user can easily confirm it.

[0118] "Means to be installed on factory robots for project review meetings in a factory work environment" refers to programs and devices that are introduced into robots in factories for the purpose of recording meetings and analyzing data within the factory.

[0119] The "means for uploading recorded audio and video data to a server" refers to a communication device or program for transmitting recorded audio and video data to a server and performing analysis processing.

[0120] The "means for providing the generated feedback to participants within the factory" refers to a device or program that notifies each member who participates in the meeting within the factory of feedback based on the analysis results and suggests areas for improvement.

[0121] The present invention relates to a system for improving the quality of project review meetings in a factory. Specific embodiments of the system will be described below.

[0122] System Overview

[0123] This system is used in a factory work environment and is intended for project review meetings. It is composed mainly of a program installed on a factory robot, which records meetings, analyzes data, and generates and provides feedback. Each process is executed by a server, and feedback is provided to participants via their terminals.

[0124] Meeting Recording

[0125] First, a camera and microphone mounted on a factory robot are used to record the contents of the meeting as soon as it starts, and the recorded data is uploaded to a server in real time as video and audio.

[0126] Audio extraction and analysis

[0127] The server extracts the audio portion from the received video recordings. This extraction uses common audio processing techniques, such as separating the audio stream from the combined video and audio data. This process is performed using a programming language such as Python.

[0128] Speech-to-text

[0129] The server converts the extracted voice data into text using the Google Cloud Speech-to-Text API. This API is called and speech recognition technology is used to convert the voice into text data. As a result, the voice utterances are obtained in text format.

[0130] Text analytics

[0131] The server then analyzes the acquired text data using natural language processing techniques. Specifically, it uses the Python NLP libraries spaCy and BERT to tokenize the text, tag parts of speech, and perform semantic analysis. This analysis allows it to evaluate the vocabulary, syntax, and structure of the discussions of participants' comments.

[0132] Audio analysis

[0133] Similarly, the server analyzes the audio data and performs acoustic feature analysis to assess the tone of voice and speaking rate, using Praat or other open-source speech analysis libraries. This analysis measures the pitch and intensity of the speech and evaluates the quality of the speech.

[0134] Generating and Presenting Feedback

[0135] The server generates specific feedback based on the text and speech analysis results. The feedback is generated for each participant in a format that includes evaluation results and recommended improvements. The generated feedback is presented to the user visually on a web browser or mobile app.

[0136] Examples and prompts

[0137] For example, in a project review meeting held in a factory, the system analyzes the speaking style and content of each participant and provides feedback such as, "Yamada-san, you use too much technical jargon, so try using more understandable words." If the tone of your voice is too high, the system also provides specific examples such as, "By lowering your voice a little, you can give the impression of being more calm."

[0138] Prompt Sentence Examples

[0139] Analysis of statements made during factory meetings:

[0140] 1. Upload the recorded audio data to the following URL:

[0141] 2. Please analyze the uploaded audio data and generate feedback in the following format:

[0142] 3. Provide specific suggestions for improving each speaker's delivery.

[0143] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0144] Step 1:

[0145] Recording of meeting content

[0146] As soon as the meeting starts, the device uses its camera and microphone to capture video and audio, recording in real time. The input is the video and audio of the meeting, and the output is recorded data. Specifically, the device's camera captures images of the meeting participants, and the microphone picks up audio. The recorded data is temporarily stored in the device's memory.

[0147] Step 2:

[0148] Uploading recording data

[0149] The device uploads the recorded data to the server. The input is the recorded data, and the output is the data uploaded to the server. Specifically, after the device finishes recording, it transfers the data to the server via Wi-Fi or Ethernet. For security reasons, the SSL / TLS protocol is used.

[0150] Step 3:

[0151] Audio Extraction

[0152] The server extracts the audio portion from the uploaded video data. The input is the video data, and the output is the audio data. Specifically, an audio processing module on the server separates the audio stream from the combined video and audio data and saves it as audio data. Libraries such as FFmpeg are used.

[0153] Step 4:

[0154] Speech-to-text

[0155] The server converts the extracted voice data into text using the Google Cloud Speech-to-Text API. The input is voice data, and the output is text data. Specifically, the voice data is sent to the API, which uses speech recognition technology to convert it into text, and the result is returned to the server.

[0156] Step 5:

[0157] Text Analysis

[0158] The server analyzes the acquired text data using natural language processing technology. The input is the text data, and the output is the analysis results. Specifically, it uses Python NLP libraries (e.g., spaCy and BERT) to perform tokenization, part-of-speech tagging, and semantic analysis. This analysis evaluates the text's vocabulary, syntax, and argument structure.

[0159] Step 6:

[0160] Analysis of audio data

[0161] The server analyzes the audio data and evaluates the tone of voice and speaking rate. The input is the audio data, and the output is the analysis results. Specifically, it uses Praat and other open source audio analysis libraries to measure pitch (voice height) and intensity (voice strength), and makes an evaluation based on the results.

[0162] Step 7:

[0163] Generate feedback

[0164] The server generates feedback based on the text analysis results and speech analysis results. The input is the text analysis results and speech analysis results, and the output is the generated feedback. Specifically, feedback is generated for each participant that includes evaluation results and recommended improvements. For example, it includes a comment such as "A-san speaks too quickly" and a suggestion for improvement such as "Try to speak more slowly next time."

[0165] Step 8:

[0166] Providing feedback

[0167] The device presents the generated feedback to the user via a web browser or mobile app. The input is the generated feedback, and the output is what is presented to the user. Specifically, the feedback is displayed separately for each comment, and is provided in a visually easy-to-understand format using color coding and graphs. For example, the feedback is displayed in a dashboard format, allowing the user to intuitively identify problems and areas for improvement in their own comments.

[0168] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0169] The present invention relates to a system for improving the quality of meetings using AI technology. This system records the contents of a meeting, extracts audio from the recorded data, converts the audio to text, analyzes the text, analyzes the audio data, and generates feedback. Furthermore, by combining it with an emotion engine, it is possible to detect the user's emotions and provide feedback based on those emotions. Specific embodiments are described below.

[0170] 1. Recording meetings

[0171] The device captures video and audio as soon as the meeting starts. For example, the entire meeting can be recorded using the camera and microphone of a laptop or smartphone. This recorded data is then sent to a server in real time.

[0172] 2. Audio Extraction

[0173] The server extracts the audio portion from the recorded data it receives. Specifically, it uses audio processing technology to separate the audio stream from the combined video and audio data.

[0174] 3. Speech-to-text

[0175] The server inputs the extracted voice data into a speech recognition engine and converts the speech into text. For example, it uses a speech recognition service such as Google Cloud Speech-to-Text API. As a result, the spoken content is obtained in text format.

[0176] 4. Text Analysis

[0177] The server analyzes the text data, specifically using natural language processing (NLP) techniques to tokenize, tag parts of speech, and perform semantic analysis of the text, evaluating the phrasing, syntax, and argument structure of what is being said.

[0178] 5. Evaluation of audio data

[0179] The server uses the audio data to evaluate the tone of voice and speaking speed, for example by analyzing acoustic features and measuring pitch (voice height) and intensity (voice strength).

[0180] 6. Emotion Recognition by Emotion Engine

[0181] The server analyzes the audio and video data to recognize the user's emotions. For example, it uses voice tone and facial expression analysis technology to detect whether the user is happy, angry, or sad. To do this, it uses a machine learning model to classify emotions.

[0182] 7. Generate feedback

[0183] The server generates feedback based on the results of text analysis, speech analysis, and emotion recognition. The feedback is provided to each participant in the form of an evaluation result and recommendations for improvement. For example, specific feedback such as, "Mr. Sato, you spoke too quickly. You also seemed a little nervous. Try to speak more relaxedly next time."

[0184] 8. Providing Feedback

[0185] The device presents the generated feedback to the user. The feedback is displayed visually in a web browser or mobile app. The user can review the feedback and understand their areas for improvement. For example, the feedback is displayed separately for each utterance and presented in an easy-to-understand format using visualization techniques such as color coding and graphs.

[0186] In this way, the present invention allows meeting participants to improve their communication skills based on specific feedback. Furthermore, by adding emotion recognition functionality, feedback can be provided that takes psychological aspects into account, further improving the overall quality of meetings and increasing productivity across the organization.

[0187] The processing flow will be explained below.

[0188] Step 1:

[0189] The device captures video and audio as soon as the meeting starts. Specifically, it uses the laptop's camera and microphone to record the entire meeting. The recorded data is sent to the server in real time.

[0190] Step 2:

[0191] The server extracts the audio portion from the recorded data it receives. Audio processing technology is used to separate the audio stream from the combined video and audio data. This step also involves format conversion of the audio file and noise filtering.

[0192] Step 3:

[0193] The server sends the extracted voice data to a speech recognition engine and converts the voice into text, for example, using the Google Cloud Speech-to-Text API, which then stores the converted text in a database.

[0194] Step 4:

[0195] The server analyzes the text data, using natural language processing (NLP) techniques to perform tokenization, part-of-speech tagging, and semantic analysis. This analysis evaluates the phrasing, syntax, and argument structure of what is being said.

[0196] Step 5:

[0197] The server analyzes the voice data to evaluate the tone of voice and speaking rate. For example, it uses tools such as Praat to measure pitch, intensity, and speaking rate. Based on this, it generates a voice evaluation result.

[0198] Step 6:

[0199] The server uses an emotion engine to recognize the user's emotions from audio and video data. It analyzes the tone of the voice and facial expressions to classify the emotional state. It uses a machine learning model to determine whether the user is happy, angry, or sad.

[0200] Step 7:

[0201] The server generates feedback based on the results of text analysis, speech evaluation, and emotion recognition. For example, it generates specific feedback such as, "Mr. Sato, you spoke too quickly. You also seemed a little nervous. Please try to speak more relaxedly next time."

[0202] Step 8:

[0203] The device presents the generated feedback to the user. The feedback is displayed visually through a web browser or mobile app. The user can review the feedback and understand their areas for improvement. For example, the feedback could be displayed in a color-coded format for each comment, and could be formatted in a way that makes it easy to understand visually using graphs and charts.

[0204] Example 2

[0205] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0206] Conventional meeting support systems did not accurately record or analyze what was said during meetings, which resulted in vague feedback and made it difficult to provide specific suggestions for improvement. Furthermore, there was no way to provide feedback that took into account participants' emotions and psychological states, resulting in a lack of specific guidance for improving the quality of communication.

[0207] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for recording the contents of the conference, means for extracting audio from the recorded data, means for converting the extracted audio into text, means for analyzing the text to evaluate the wording, syntax, and structure of the discussion, means for analyzing the audio data to evaluate the tone of voice and speaking rate, means for analyzing the audio data and video data to recognize the user's emotions, means for generating feedback based on the analysis results, and means for presenting the generated feedback to the user. This makes it possible to accurately record and analyze the content of comments made in the conference and provide specific feedback that takes into account the emotions of the participants.

[0208] "Meeting content" refers to the entire information, materials, and presentations spoken or shared by participants at the meeting.

[0209] "Recording Data" means the digital files of video and audio captured during a meeting.

[0210] "Means for extracting audio" refers to a function that performs the process of separating and extracting the audio portion from the recorded data.

[0211] "Means for converting speech to text" refers to technology, particularly a speech recognition engine, that automatically converts extracted speech data into text data.

[0212] "Means of analyzing text to evaluate language, syntax, and argument structure" refers to the application of natural language processing techniques to converted text data to analyze word choice, grammar, and argument flow.

[0213] "Means for analyzing voice data to evaluate voice tone and speaking rate" refers to technology that uses acoustic analysis technology to measure the features of voice data and evaluate the pitch, strength, and speaking rate of the voice.

[0214] "Means for recognizing a user's emotions by analyzing audio and video data" refers to a process that uses voice tone and facial expression recognition technology to identify a user's emotional state.

[0215] "Means for generating feedback" refers to a function that automatically creates feedback including specific advice and areas for improvement for the user based on the results of text analysis, speech analysis, and emotion recognition.

[0216] "Means for presenting the generated feedback to the user" refers to an interface, such as a web browser or mobile app, that displays the generated feedback so that the user can review it.

[0217] The present invention relates to a system for improving the quality of meetings by utilizing AI technology. This system records the contents of meetings, extracts audio from the recorded data, converts the audio to text, analyzes the text, analyzes the audio data, and generates feedback. Furthermore, by combining it with an emotion engine, it is possible to detect the user's emotions and provide feedback based on those emotions. Specific embodiments for implementing this system are described below.

[0218] Hardware and software used

[0219] The device uses the camera and microphone of a laptop or smartphone. For example, a typical laptop or smartphone has a built-in camera and microphone for capturing video and audio.

[0220] The data transfer and processing is carried out by servers using the following software and technologies:

[0221] Audio processing library: FFmpeg

[0222] Speech recognition service: Google Cloud Speech-to-Text API

[0223] Natural language processing library: SpaCy

[0224] Acoustic analysis library: Librosa

[0225] Machine learning frameworks: TensorFlow, PyTorch

[0226] Browser display: D3.js

[0227] System processing flow

[0228] 1. Meeting recording: The device captures video and audio as soon as the meeting starts. For example, you can use Zoom or other video conferencing software to record the meeting. The recording data is sent to the server in real time.

[0229] 2. Audio Extraction: Extract the audio portion from the recording data received by the server. Use FFmpeg to separate the audio stream from the combined video and audio data.

[0230] 3. Speech-to-text conversion: The server sends the extracted voice data to the Google Cloud Speech-to-Text API to convert the voice to text, thereby obtaining the content of the meeting in text format.

[0231] 4. Text Analysis: The server analyzes the acquired text data using natural language processing (NLP) techniques, including tokenization, part-of-speech tagging, and semantic analysis using the SpaCy library.

[0232] 5. Evaluation of speech data: The server uses the Librosa library to analyze the acoustic features of the speech data, measure pitch and intensity from the speech data, and evaluate the speaking rate.

[0233] 6. Emotion recognition using an emotion engine: The server analyzes audio and video data to recognize the user's emotions. Using TensorFlow and other tools, it analyzes the tone of the voice and recognizes facial expressions to classify the user's emotions.

[0234] 7. Feedback Generation: The server generates feedback based on the results of text analysis, speech analysis, and emotion recognition. It uses natural language generation technology to create feedback that includes specific improvements.

[0235] 8. Present feedback: The device presents the feedback generated by the server to the user. This is visually displayed in the web browser. D3.js is used to present the feedback to the user using graphs and color coding to make it easy to understand.

[0236] Specific examples and prompts

[0237] As a concrete example, the following prompt sentence is input to the generative AI model:

[0238] "Extract audio from meeting recordings and convert it to text. Analyze the text content and perform emotion recognition to generate feedback."

[0239] This prompt sentence executes each processing step and generates specific feedback, for example, "Mr. Suzuki, you speak too quickly. Please try to speak a little more slowly."

[0240] In this way, the present invention can help meeting participants improve their communication skills based on specific feedback, potentially improving the overall quality of meetings. Furthermore, by adding emotion recognition functionality, feedback can be provided that takes into account psychological aspects during meetings, which is expected to improve productivity across the organization.

[0241] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0242] Step 1:

[0243] Recording a meeting

[0244] Input: When a meeting starts, the device captures the video and audio of the meeting.

[0245] Data processing: Recorded data that combines video and audio data is generated.

[0246] Output: The generated recording data is sent to the server in real time.

[0247] Specific operation: The device starts recording using video conferencing software (e.g., Zoom or Microsoft Teams) and transfers the data to the server.

[0248] Step 2:

[0249] Audio Extraction

[0250] Input: Recording data received by the server.

[0251] Data processing: Use FFmpeg to separate the audio stream from the combined video and audio data.

[0252] Output: Separated audio data (e.g., .wav format) is generated.

[0253] Specific operation: The server executes the FFmpeg command to extract and save the audio file from the recorded data.

[0254] Step 3:

[0255] Speech-to-text

[0256] Input: Audio data extracted by the server.

[0257] Data processing: Send the audio data to the Google Cloud Speech-to-Text API for speech recognition.

[0258] Output: The spoken words are captured in text format.

[0259] Specific operation: The server uploads the audio file to the Google Cloud Speech-to-Text API and stores the returned text data.

[0260] Step 4:

[0261] Text Analysis

[0262] Input: Speech-to-text data.

[0263] Data processing: Use the SpaCy library to tokenize the text, tag it with parts of speech, and perform semantic analysis.

[0264] Output: The analyzed text data is generated and the evaluation results are obtained.

[0265] What it does: The server uses the SpaCy library to analyze the text data and evaluate the phrasing, syntax, and argument structure of the comments.

[0266] Step 5:

[0267] Evaluation of audio data

[0268] Input: Extracted audio data.

[0269] Data processing: The Librosa library is used to analyze acoustic features and evaluate voice tone and speaking rate.

[0270] Output: Generates an assessment of tone of voice and speaking rate.

[0271] How it works: The server processes the audio data using Librosa, measures pitch and intensity, and evaluates the speaker's vocal characteristics.

[0272] Step 6:

[0273] Emotion recognition by emotion engine

[0274] Input: Audio and video data.

[0275] Data processing: Use TensorFlow and PyTorch to perform voice tone analysis and facial expression recognition to classify emotions.

[0276] Output: The user's emotional state (e.g., joy, anger, sadness, or happiness) is recognized.

[0277] How it works: The server analyzes the audio data and applies an emotion classification model to detect the user's emotional state, for example, whether they are angry, sad, happy, etc.

[0278] Step 7:

[0279] Generate feedback

[0280] Input: Text analysis results, speech analysis results, and emotion recognition results.

[0281] Data processing: The results of each analysis are integrated and feedback is automatically generated using natural language generation technology.

[0282] Output: A specific feedback text is generated for the user.

[0283] Specific behavior: The server uses a natural language generation engine (e.g., GPT-3) to generate feedback based on the analysis results, including specific suggestions for improvement. For example, it generates feedback like, "Mr. Sato, you spoke too quickly. You also seemed a little nervous. Please try to speak more relaxedly next time."

[0284] Step 8:

[0285] Providing feedback

[0286] Input: The generated feedback text.

[0287] Data processing: Converting JSON-formatted feedback data into a visually understandable format.

[0288] Output: Visual feedback that the user can see.

[0289] What it does: The device displays feedback in a web browser or mobile app, and uses D3.js to visually present the feedback as graphs and color-coded text.

[0290] (Application example 2)

[0291] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0292] When operating robots in factories, there is a need to evaluate the technical quality and emotional state of operators and provide specific feedback to improve operational efficiency and safety. However, with current systems, it is difficult to evaluate the emotional state and quality of operators' operations in real time and generate and provide feedback based on that evaluation.

[0293] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0294] In this invention, the server includes means for recording the contents of the conference, means for extracting audio from the recorded data, means for converting the extracted audio into text, means for analyzing the text to evaluate the vocabulary, syntax, and structure of the discussion, means for analyzing the audio data to evaluate the tone of voice and speaking rate, means for generating feedback based on the analysis results, means for presenting the generated feedback to the user, means for acquiring video and audio to record the operation status, means for analyzing the collected data to evaluate the efficiency and safety of the operation and the emotional state of the operator, and means for generating feedback based on the evaluation results, thereby enabling improvement of the operator's skills, ensuring safety, and providing appropriate guidance based on the emotional state.

[0295] "Means for recording meeting content" refers to devices and methods for recording video and audio of meetings.

[0296] "Means for extracting audio from recorded data" refers to a technology for extracting only the audio portion from recorded video and audio data.

[0297] The "means for converting extracted speech into text" is a speech recognition technology that converts speech data into character data.

[0298] "Methods for analyzing text to evaluate language, syntax, and argument structure" refers to techniques for analyzing text data and judging and evaluating its linguistic expression, sentence structure, and argument flow.

[0299] "Means for analyzing voice data to evaluate voice tone and speaking rate" refers to technology that analyzes and evaluates voice characteristics such as pitch, strength, and speaking rate.

[0300] "Means for generating feedback based on the analysis results" refers to a method for providing users with advice and suggestions for improvement based on the analyzed data.

[0301] "Means for presenting the generated feedback to the user" refers to technology for displaying and providing the generated feedback information in a form that the user can understand.

[0302] "Means for acquiring video and audio for recording the operation status" refers to a device or method for recording the operation status with video and audio.

[0303] "Means for analyzing collected data to evaluate the efficiency, safety, and emotional state of the operator" refers to technology that analyzes acquired data to judge and evaluate the efficiency, safety, and emotional state of the operator.

[0304] The "means for generating feedback based on the evaluation results" is a method for providing useful advice and suggestions for improvement to users based on the evaluation results.

[0305] This invention relates to an AI assistance system for improving the quality of robot operation in factories. This system records the situation during operation, generates feedback from the data, and provides it to the operator to help improve their technical skills and manage their emotions.

[0306] Hardware Configuration

[0307] The system includes the following hardware:

[0308] Smart glasses (e.g., Google Glass Enterprise Edition): worn by operators, record the operation status in real time.

[0309] Cameras and microphones: Installed within the factory to capture additional video and audio data.

[0310] Server: For data processing and analysis (e.g., Google Cloud Platform).

[0311] Software Configuration

[0312] The software used is as follows:

[0313] Speech recognition engine (e.g. Google Cloud Speech-to-Text API): Converts extracted speech into text.

[0314] Natural language processing (NLP) techniques (e.g., Google Cloud Natural Language API): Analyze text data to evaluate phrasing, syntax, and argument structure.

[0315] Emotion recognition engine (e.g., Microsoft Azure Face API): Recognizes the operator's emotional state from video and audio data.

[0316] Data analysis programs (e.g., Python, TensorFlow): used to evaluate data and generate feedback.

[0317] Data processing and feedback generation

[0318] The server receives the video and audio data sent from the smart glasses and the camera, and then processes the data in the following steps:

[0319] 1. Audio extraction: The audio is separated from the video data and input into a speech recognition engine. This process converts the audio into text.

[0320] 2. Text analysis: The obtained text data is analyzed using natural language processing technology to evaluate the content of the statements.

[0321] 3. Voice analysis: Analyzes the characteristics of voice data and evaluates the tone of voice and speaking rate, thereby extracting how efficient and safe the operation is.

[0322] 4. Emotion Recognition: Analyze the operator's emotional state using video and audio data and incorporate the results into the evaluation.

[0323] Generating and Presenting Feedback

[0324] The server generates and provides feedback to the user (operator) based on the analysis results, including specific advice on operational efficiency, safety, and emotional state.

[0325] Specific examples

[0326] An operator wearing smart glasses starts recording while operating a robot. The recorded data is sent to a server, where a speech recognition engine converts the speech into text. Natural language processing technology then analyzes the text to evaluate the accuracy and efficiency of the operation. At the same time, an emotion recognition engine analyzes the operator's emotional state from the video data. Ultimately, based on the results of these analyses, feedback is generated, such as, "The order of operations is efficient, but you appear to be a little impatient. Please operate more slowly next time."

[0327] Prompt Sentence Examples

[0328] Voice recognition prompts

[0329] Use the speech recognition API to convert the following audio data to text:

[0330] <path to audio data>

[0331] Prompts for Natural Language Processing

[0332] Use the recognized text data to extract the following information:

[0333] Key Points

[0334] Utterance Syntax

[0335] Semantic analysis

[0336] Emotion recognition prompts

[0337] Detect operator emotions from video data:

[0338] <Video data path>

[0339] Prompts for generating feedback

[0340] Use the data you collect to generate feedback that includes the following metrics:

[0341] Evaluating operational efficiency

[0342] Safety evaluation

[0343] Emotion-based feedback

[0344] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0345] Step 1:

[0346] Video and audio acquisition

[0347] The device uses smart glasses or cameras and microphones in the factory to record video and audio of the operator's operations. This video and audio data is sent to the server in real time. The input is video and audio data, and the output is stream data transferred to the server.

[0348] Step 2:

[0349] Audio Extraction

[0350] The server extracts the audio portion from the transmitted video recording data. Specifically, it uses audio processing technology to separate the audio stream from the combined video and audio data. The input is the video and audio stream data, and the output is the extracted audio data.

[0351] Step 3:

[0352] Speech-to-text

[0353] The server converts the extracted audio data into text using a speech recognition engine (e.g., Google Cloud Speech-to-Text API). In this process, the speech recognition engine analyzes the audio waveform and generates a corresponding string using a language model. The input is the audio data, and the output is the generated text data.

[0354] Step 4:

[0355] Text Analysis

[0356] The server analyzes the generated text data using natural language processing technology (e.g., Google Cloud Natural Language API). This analysis includes tokenization, part-of-speech tagging, and semantic analysis, and evaluates the content of the speech. The input is text data, and the output is the analysis results (key points, sentence structure, flow of discussion, etc.).

[0357] Step 5:

[0358] Evaluation of audio data

[0359] The server analyzes the features of the voice data and evaluates the tone of voice and speaking speed. Here, acoustic features (e.g., pitch, intensity) are extracted and evaluation is performed based on that information. The input is the voice data, and the output is the voice features (voice pitch, speed, strength, etc.) and the evaluation results.

[0360] Step 6:

[0361] Emotion recognition

[0362] The server uses an emotion recognition engine (e.g., Microsoft Azure Face API) to analyze the operator's emotional state from the video and audio data. This includes facial expression analysis technology and voice tone analysis. The input is video and audio data, and the output is a classification result of the operator's emotional state.

[0363] Step 7:

[0364] Generate feedback

[0365] The server generates specific feedback based on the analysis results. This feedback includes evaluations and advice on operational efficiency, safety, and emotional state. The inputs are the text analysis results, speech analysis results, and emotion recognition results, and the generated feedback is obtained as the output.

[0366] Step 8:

[0367] Providing feedback

[0368] The terminal presents the generated feedback to the operator, using a web browser or mobile app to display the feedback in a visually understandable format. The input is the generated feedback, and the output is the visualized feedback that can be viewed by the user.

[0369] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0370] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0371] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0372] [Second embodiment]

[0373] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0374] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0375] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0376] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0377] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0378] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0379] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0380] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0381] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0382] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0383] In the smart glasses 214, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0384] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0385] The present invention relates to a system for improving the quality of meetings by utilizing AI technology. A specific embodiment of this system will be described.

[0386] 1. Recording meetings

[0387] The device captures video and audio as soon as the meeting starts. Specifically, the entire meeting is recorded using the camera and microphone of a laptop or smartphone, for example. The recorded data is uploaded to a server in real time.

[0388] 2. Audio Extraction

[0389] The server extracts the audio portion from the received recording. This extraction can be done using common audio processing techniques, including, for example, separating the audio stream from the combined video and audio of the meeting.

[0390] 3. Speech-to-text

[0391] The server then requests a speech recognition engine to convert the extracted voice data. Specifically, it calls a service such as the Google Cloud Speech-to-Text API to convert the voice into text data. As a result, the spoken content is obtained in text format.

[0392] 4. Text Analysis

[0393] The server analyzes the text data. Natural language processing (NLP) techniques are used for text analysis. For example, Python NLP libraries such as spaCy and BERT can be used to tokenize the text, tag parts of speech, and perform semantic analysis. This analysis makes it possible to evaluate the vocabulary, syntax, and argument structure of the speech.

[0394] 5. Evaluation of audio data

[0395] The server analyzes the audio data and evaluates the tone of voice and speaking rate. This includes analyzing acoustic features. For example, it uses Praat or an open-source audio analysis library to measure pitch (highness of voice) and intensity (power of voice) and makes an evaluation based on these.

[0396] 6. Generate feedback

[0397] The server generates feedback based on the results of text and speech analysis. The feedback is provided to each participant in the form of an evaluation result and recommendations for improvement. For example, it may include specific comments such as "Mr. Sato speaks too quickly" and suggestions for improvement such as "Try to speak a little more slowly next time."

[0398] 7. Providing Feedback

[0399] The device presents the generated feedback to the user. The feedback is displayed visually in a web browser or mobile app, allowing the user to intuitively identify problems and areas for improvement in their comments. For example, feedback could be displayed separately for each comment, and a visually easy-to-understand format could be used, using color coding or graphs.

[0400] Using this system, meeting participants can improve their communication skills based on specific feedback, which is expected to improve the overall quality of meetings and increase productivity across the organization.

[0401] The processing flow will be explained below.

[0402] Step 1:

[0403] The device captures video and audio as soon as the meeting starts. For example, the entire meeting can be recorded using the camera and microphone of a laptop or smartphone, and the data can be sent to a server in real time.

[0404] Step 2:

[0405] The server extracts the audio portion from the recorded data it receives. Specifically, it uses audio processing technology to separate the audio stream from the combined video and audio data.

[0406] Step 3:

[0407] The server inputs the extracted voice data into a speech recognition engine and converts the voice to text, for example, using the Google Cloud Speech-to-Text API.

[0408] Step 4:

[0409] The server analyzes the text data, using natural language processing (NLP) techniques to tokenize, tag parts of speech, and analyze the semantics of the text. This analysis evaluates the phrasing, syntax, and argument structure of the speech.

[0410] Step 5:

[0411] The server uses the audio data to evaluate the tone of voice and speaking rate. Specifically, it analyzes acoustic features to measure pitch (highness of voice) and intensity (power of voice).

[0412] Step 6:

[0413] The server generates feedback based on the results of text and speech analysis, for example, using a specific template to create feedback for each participant, including specific evaluations and suggestions for improvement.

[0414] Step 7:

[0415] The device then presents the generated feedback to the user, which is displayed visually via a smartphone, PC web browser, or dedicated app. The user can then review the displayed feedback and understand their areas for improvement.

[0416] Example 1

[0417] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0418] While conventional conference systems can record conference content and convert audio into text, they lack the functionality to evaluate the quality of the conference and provide specific feedback. As a result, users have no way to objectively evaluate their own comments or the way the conference proceeded, which hinders the improvement of their communication skills. The present invention aims to solve these problems and provide a system that improves the quality of conferences.

[0419] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0420] In this invention, the server includes means for recording the contents of the conference, means for uploading the recorded data to the server in real time, means for extracting voice from the recorded data, means for transmitting the extracted voice to a voice recognition engine and converting it into text data, means for analyzing the text data using natural language processing technology, means for analyzing the voice data using acoustic analysis technology and evaluating the tone of voice and speaking rate, means for generating feedback based on the results of the text analysis and the voice analysis, and means for presenting the generated feedback to the user. This allows conference participants to receive evaluations of their own remarks and voice and know specific areas for improvement, thereby enabling them to improve their communication skills.

[0421] A "meeting recording device" is a device or method that captures and stores video and audio of a meeting in digital form.

[0422] "Means for uploading recorded data to a server in real time" refers to the function of instantly transmitting video and audio data generated during the progress of a conference to a server.

[0423] "Means for extracting audio from recorded data" refers to a technology for separating and extracting the audio portion from recorded video data.

[0424] "Means for transmitting the extracted speech to a speech recognition engine and converting it into text data" refers to the process of transmitting the extracted speech data to speech recognition software and using that software to convert the speech into text information.

[0425] "Means for analyzing text data using natural language processing technology" refers to a technology that analyzes acquired text data using a natural language processing algorithm and evaluates the wording, syntax, and structure of the argument.

[0426] "Means for analyzing voice data using acoustic analysis technology to evaluate voice tone and speaking rate" refers to technology for analyzing voice data using an acoustic analysis tool and measuring pitch and intensity to evaluate voice tone and speaking rate.

[0427] The "means for generating feedback based on text analysis results and speech analysis results" refers to a process of combining the analysis results of text data and speech data to create feedback that includes specific evaluations and improvement measures.

[0428] The "means for presenting the generated feedback to the user" is a device or method that visually or audibly conveys the generated feedback to the user.

[0429] The present invention relates to a system for improving the quality of meetings by utilizing AI technology. Specific embodiments of this system are described in detail below.

[0430] First, at the start of a meeting, the user launches the meeting application and presses the "Start Meeting" button. Devices (e.g., laptops, smartphones) that detect the start of the meeting begin capturing video and audio using their built-in cameras and microphones. This allows the entire meeting to be recorded.

[0431] The captured video and audio data is uploaded to a server in real time using HTTP streaming or WebRTC technology, for example. The server processes the received recording data and separates the video and audio data. Specifically, audio processing technology such as FFmpeg is used to extract the audio stream from the recording data.

[0432] The extracted voice data is then sent by the server to the Google Cloud Speech-to-Text API, where it is converted into text data. This step allows the voice utterance to be obtained in text format.

[0433] The captured text data is then analyzed by the server using natural language processing (NLP) techniques, such as tokenizing, part-of-speech tagging, and semantic analysis using Python NLP libraries (e.g., spaCy, BERT), to evaluate the vocabulary, syntax, and argument structure of the speech.

[0434] Additionally, the voice data is analyzed using acoustic analysis technology. The server analyzes the voice data using an acoustic analysis library (e.g., Librosa) to measure pitch and intensity. This evaluation evaluates the tone of voice and speaking rate.

[0435] The server generates feedback based on these analysis results. The generated feedback includes specific suggestions for improvement and evaluations, such as "You speak too quickly" or "Try to speak more slowly next time." Finally, the generated feedback is presented to the user via the device. The feedback is displayed visually in a web browser or mobile app, and is presented in an easily understandable format using color-coded graphs and icons.

[0436] For example, when a user holds a meeting using video conferencing software, the laptop's camera and microphone capture video and audio and upload them to a server in real time. The server then extracts the audio using FFmpeg and converts it to text using the Google Cloud Speech-to-Text API. The text is then analyzed using spaCy, and the audio is evaluated using Librosa. Finally, feedback generated based on the analysis results is displayed in the web browser for the user to review.

[0437] An example of a prompt might be, "Please use the Google Cloud Speech-to-Text API to convert the following audio data into text and generate meeting feedback based on the results." By using this system, meeting participants can receive evaluations of their own remarks and voices, learn specific areas for improvement, and improve their communication skills.

[0438] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0439] Step 1:

[0440] The user launches the conference application and presses the "Start Conference" button. This causes the device to detect the start of the conference. The input is the user's operation action, and the output is the start of the conference recording process.

[0441] Step 2:

[0442] The device will begin capturing video and audio using the built-in camera and microphone. The captured data will be saved as video and audio files. The input is real-time data from the built-in camera and microphone, and the output is locally saved video and audio data.

[0443] Step 3:

[0444] The device uploads the captured video and audio data to the server in real time. This process uses HTTP streaming or WebRTC technology. The input is the locally stored video and audio data, and the output is the data uploaded to the server.

[0445] Step 4:

[0446] The server processes the received recording data and separates the video and audio data using audio processing technology such as FFmpeg. The input is the uploaded recording data, and the output is the separated video and audio data.

[0447] Step 5:

[0448] The server sends the extracted speech data to a speech recognition engine and converts it into text data. This step uses the Google Cloud Speech-to-Text API. The extracted speech data is the input, and the text data is the output.

[0449] Step 6:

[0450] The server analyzes the acquired text data using natural language processing (NLP) techniques. Specifically, it uses Python NLP libraries (e.g., spaCy, BERT) to perform tokenization, part-of-speech tagging, and semantic analysis. The converted text data is input, and the analysis results are obtained as output.

[0451] Step 7:

[0452] The server analyzes the voice data using acoustic analysis technology to evaluate the tone of voice and speaking rate. Specifically, it uses the Librosa library to measure pitch and intensity. The input is the voice data, and the output is the result of the acoustic analysis.

[0453] Step 8:

[0454] The server generates feedback based on the results of text analysis and speech analysis. The feedback includes specific evaluations and suggestions for improvement. The inputs are the results of text analysis and speech analysis, and the generated feedback is obtained as the output.

[0455] Step 9:

[0456] The device presents the generated feedback to the user, which is then visually displayed in a web browser or mobile app. The input is the generated feedback, and the output is the feedback presented to the user.

[0457] (Application example 1)

[0458] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0459] Conventional conference systems lack specific feedback to improve the quality of meetings, making it difficult for participants to improve their own comments and attitudes. Furthermore, in certain environments, such as project review meetings in factories, recording and analysis methods and tools are often not properly installed, hindering the effective running of meetings. The present invention aims to provide a system that solves these problems and improves the quality of meetings.

[0460] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0461] In this invention, the server includes a means for recording the contents of the conference, a means for extracting audio from the recorded data, and a means for converting the extracted audio into text, thereby enabling the server to accurately record the overall picture of the conference and provide specific feedback based on the analysis results.

[0462] "Means for recording the contents of a conference" refers to a device or program that captures and records video and audio as soon as the conference begins.

[0463] The "means for extracting audio from recorded data" refers to a processing device or program for separating the audio portion from the recorded conference data and extracting it as audio data.

[0464] "Means for converting extracted speech into text" refers to a program or service that converts extracted speech data into text format using speech recognition technology.

[0465] A "means for analyzing text to evaluate language, syntax, and argument structure" is a program that uses natural language processing techniques on text data to evaluate and analyze the content of the text.

[0466] The "means for analyzing voice data to evaluate voice tone and speaking rate" refers to a device or program that uses voice data to analyze acoustic features such as voice pitch and speaking rate and make an evaluation.

[0467] The "means for generating feedback based on the analysis results" is a program that provides participants with specific suggestions for improvement and evaluations based on the results of text and voice analysis.

[0468] The "means for presenting the generated feedback to the user" refers to a device or program that visually displays the generated feedback so that the user can easily confirm it.

[0469] "Means to be installed on factory robots for project review meetings in a factory work environment" refers to programs and devices that are introduced into robots in factories for the purpose of recording meetings and analyzing data within the factory.

[0470] The "means for uploading recorded audio and video data to a server" refers to a communication device or program for transmitting recorded audio and video data to a server and performing analysis processing.

[0471] The "means for providing the generated feedback to participants within the factory" refers to a device or program that notifies each member who participates in the meeting within the factory of feedback based on the analysis results and suggests areas for improvement.

[0472] The present invention relates to a system for improving the quality of project review meetings in a factory. Specific embodiments of the system will be described below.

[0473] System Overview

[0474] This system is used in a factory work environment and is intended for project review meetings. It is composed mainly of a program installed on a factory robot, which records meetings, analyzes data, and generates and provides feedback. Each process is executed by a server, and feedback is provided to participants via their terminals.

[0475] Meeting Recording

[0476] First, a camera and microphone mounted on a factory robot are used to record the contents of the meeting as soon as it starts, and the recorded data is uploaded to a server in real time as video and audio.

[0477] Audio extraction and analysis

[0478] The server extracts the audio portion from the received video recordings. This extraction uses common audio processing techniques, such as separating the audio stream from the combined video and audio data. This process is performed using a programming language such as Python.

[0479] Speech-to-text

[0480] The server converts the extracted voice data into text using the Google Cloud Speech-to-Text API. This API is called and speech recognition technology is used to convert the voice into text data. As a result, the voice utterances are obtained in text format.

[0481] Text analytics

[0482] The server then analyzes the acquired text data using natural language processing techniques. Specifically, it uses the Python NLP libraries spaCy and BERT to tokenize the text, tag parts of speech, and perform semantic analysis. This analysis allows it to evaluate the vocabulary, syntax, and structure of the discussions of participants' comments.

[0483] Audio analysis

[0484] Similarly, the server analyzes the audio data and performs acoustic feature analysis to assess the tone of voice and speaking rate, using Praat or other open-source speech analysis libraries. This analysis measures the pitch and intensity of the speech and evaluates the quality of the speech.

[0485] Generating and Presenting Feedback

[0486] The server generates specific feedback based on the text and speech analysis results. The feedback is generated for each participant in a format that includes evaluation results and recommended improvements. The generated feedback is presented to the user visually on a web browser or mobile app.

[0487] Examples and prompts

[0488] For example, in a project review meeting held in a factory, the system analyzes the speaking style and content of each participant and provides feedback such as, "Yamada-san, you use too much technical jargon, so try using more understandable words." If the tone of your voice is too high, the system also provides specific examples such as, "By lowering your voice a little, you can give the impression of being more calm."

[0489] Prompt Sentence Examples

[0490] Analysis of statements made during factory meetings:

[0491] 1. Upload the recorded audio data to the following URL:

[0492] 2. Please analyze the uploaded audio data and generate feedback in the following format:

[0493] 3. Provide specific suggestions for improving each speaker's delivery.

[0494] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0495] Step 1:

[0496] Recording of meeting content

[0497] As soon as the meeting starts, the device uses its camera and microphone to capture video and audio, recording in real time. The input is the video and audio of the meeting, and the output is recorded data. Specifically, the device's camera captures images of the meeting participants, and the microphone picks up audio. The recorded data is temporarily stored in the device's memory.

[0498] Step 2:

[0499] Uploading recording data

[0500] The device uploads the recorded data to the server. The input is the recorded data, and the output is the data uploaded to the server. Specifically, after the device finishes recording, it transfers the data to the server via Wi-Fi or Ethernet. For security reasons, the SSL / TLS protocol is used.

[0501] Step 3:

[0502] Audio Extraction

[0503] The server extracts the audio portion from the uploaded video data. The input is the video data, and the output is the audio data. Specifically, an audio processing module on the server separates the audio stream from the combined video and audio data and saves it as audio data. Libraries such as FFmpeg are used.

[0504] Step 4:

[0505] Speech-to-text

[0506] The server converts the extracted voice data into text using the Google Cloud Speech-to-Text API. The input is voice data, and the output is text data. Specifically, the voice data is sent to the API, which uses speech recognition technology to convert it into text, and the result is returned to the server.

[0507] Step 5:

[0508] Text Analysis

[0509] The server analyzes the acquired text data using natural language processing technology. The input is the text data, and the output is the analysis results. Specifically, it uses Python NLP libraries (e.g., spaCy and BERT) to perform tokenization, part-of-speech tagging, and semantic analysis. This analysis evaluates the text's vocabulary, syntax, and argument structure.

[0510] Step 6:

[0511] Analysis of audio data

[0512] The server analyzes the audio data and evaluates the tone of voice and speaking rate. The input is the audio data, and the output is the analysis results. Specifically, it uses Praat and other open source audio analysis libraries to measure pitch (voice height) and intensity (voice strength), and makes an evaluation based on the results.

[0513] Step 7:

[0514] Generate feedback

[0515] The server generates feedback based on the text analysis results and speech analysis results. The input is the text analysis results and speech analysis results, and the output is the generated feedback. Specifically, feedback is generated for each participant that includes evaluation results and recommended improvements. For example, it includes a comment such as "A-san speaks too quickly" and a suggestion for improvement such as "Try to speak more slowly next time."

[0516] Step 8:

[0517] Providing feedback

[0518] The device presents the generated feedback to the user via a web browser or mobile app. The input is the generated feedback, and the output is what is presented to the user. Specifically, the feedback is displayed separately for each comment, and is provided in a visually easy-to-understand format using color coding and graphs. For example, the feedback is displayed in a dashboard format, allowing the user to intuitively identify problems and areas for improvement in their own comments.

[0519] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0520] The present invention relates to a system for improving the quality of meetings using AI technology. This system records the contents of a meeting, extracts audio from the recorded data, converts the audio to text, analyzes the text, analyzes the audio data, and generates feedback. Furthermore, by combining it with an emotion engine, it is possible to detect the user's emotions and provide feedback based on those emotions. Specific embodiments are described below.

[0521] 1. Recording meetings

[0522] The device captures video and audio as soon as the meeting starts. For example, the entire meeting can be recorded using the camera and microphone of a laptop or smartphone. This recorded data is then sent to a server in real time.

[0523] 2. Audio Extraction

[0524] The server extracts the audio portion from the recorded data it receives. Specifically, it uses audio processing technology to separate the audio stream from the combined video and audio data.

[0525] 3. Speech-to-text

[0526] The server inputs the extracted voice data into a speech recognition engine and converts the speech into text. For example, it uses a speech recognition service such as Google Cloud Speech-to-Text API. As a result, the spoken content is obtained in text format.

[0527] 4. Text Analysis

[0528] The server analyzes the text data, specifically using natural language processing (NLP) techniques to tokenize, tag parts of speech, and perform semantic analysis of the text, evaluating the phrasing, syntax, and argument structure of what is being said.

[0529] 5. Evaluation of audio data

[0530] The server uses the audio data to evaluate the tone of voice and speaking speed, for example by analyzing acoustic features and measuring pitch (voice height) and intensity (voice strength).

[0531] 6. Emotion Recognition by Emotion Engine

[0532] The server analyzes the audio and video data to recognize the user's emotions. For example, it uses voice tone and facial expression analysis technology to detect whether the user is happy, angry, or sad. To do this, it uses a machine learning model to classify emotions.

[0533] 7. Generate feedback

[0534] The server generates feedback based on the results of text analysis, speech analysis, and emotion recognition. The feedback is provided to each participant in the form of an evaluation result and recommendations for improvement. For example, specific feedback such as, "Mr. Sato, you spoke too quickly. You also seemed a little nervous. Try to speak more relaxedly next time."

[0535] 8. Providing Feedback

[0536] The device presents the generated feedback to the user. The feedback is displayed visually in a web browser or mobile app. The user can review the feedback and understand their areas for improvement. For example, the feedback is displayed separately for each utterance and presented in an easy-to-understand format using visualization techniques such as color coding and graphs.

[0537] In this way, the present invention allows meeting participants to improve their communication skills based on specific feedback. Furthermore, by adding emotion recognition functionality, feedback can be provided that takes psychological aspects into account, further improving the overall quality of meetings and increasing productivity across the organization.

[0538] The processing flow will be explained below.

[0539] Step 1:

[0540] The device captures video and audio as soon as the meeting starts. Specifically, it uses the laptop's camera and microphone to record the entire meeting. The recorded data is sent to the server in real time.

[0541] Step 2:

[0542] The server extracts the audio portion from the recorded data it receives. Audio processing technology is used to separate the audio stream from the combined video and audio data. This step also involves format conversion of the audio file and noise filtering.

[0543] Step 3:

[0544] The server sends the extracted voice data to a speech recognition engine and converts the voice into text, for example, using the Google Cloud Speech-to-Text API, which then stores the converted text in a database.

[0545] Step 4:

[0546] The server analyzes the text data, using natural language processing (NLP) techniques to perform tokenization, part-of-speech tagging, and semantic analysis. This analysis evaluates the phrasing, syntax, and argument structure of what is being said.

[0547] Step 5:

[0548] The server analyzes the voice data to evaluate the tone of voice and speaking rate. For example, it uses tools such as Praat to measure pitch, intensity, and speaking rate. Based on this, it generates a voice evaluation result.

[0549] Step 6:

[0550] The server uses an emotion engine to recognize the user's emotions from audio and video data. It analyzes the tone of the voice and facial expressions to classify the emotional state. It uses a machine learning model to determine whether the user is happy, angry, or sad.

[0551] Step 7:

[0552] The server generates feedback based on the results of text analysis, speech evaluation, and emotion recognition. For example, it generates specific feedback such as, "Mr. Sato, you spoke too quickly. You also seemed a little nervous. Please try to speak more relaxedly next time."

[0553] Step 8:

[0554] The device presents the generated feedback to the user. The feedback is displayed visually through a web browser or mobile app. The user can review the feedback and understand their areas for improvement. For example, the feedback could be displayed in a color-coded format for each comment, and could be formatted in a way that makes it easy to understand visually using graphs and charts.

[0555] Example 2

[0556] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0557] Conventional meeting support systems did not accurately record or analyze what was said during meetings, which resulted in vague feedback and made it difficult to provide specific suggestions for improvement. Furthermore, there was no way to provide feedback that took into account participants' emotions and psychological states, resulting in a lack of specific guidance for improving the quality of communication.

[0558] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for recording the contents of the conference, means for extracting audio from the recorded data, means for converting the extracted audio into text, means for analyzing the text to evaluate the wording, syntax, and structure of the discussion, means for analyzing the audio data to evaluate the tone of voice and speaking rate, means for analyzing the audio data and video data to recognize the user's emotions, means for generating feedback based on the analysis results, and means for presenting the generated feedback to the user. This makes it possible to accurately record and analyze the content of comments made in the conference and provide specific feedback that takes into account the emotions of the participants.

[0559] "Meeting content" refers to the entire information, materials, and presentations spoken or shared by participants at the meeting.

[0560] "Recording Data" means the digital files of video and audio captured during a meeting.

[0561] "Means for extracting audio" refers to a function that performs the process of separating and extracting the audio portion from the recorded data.

[0562] "Means for converting speech to text" refers to technology, particularly a speech recognition engine, that automatically converts extracted speech data into text data.

[0563] "Means of analyzing text to evaluate language, syntax, and argument structure" refers to the application of natural language processing techniques to converted text data to analyze word choice, grammar, and argument flow.

[0564] "Means for analyzing voice data to evaluate voice tone and speaking rate" refers to technology that uses acoustic analysis technology to measure the features of voice data and evaluate the pitch, strength, and speaking rate of the voice.

[0565] "Means for recognizing a user's emotions by analyzing audio and video data" refers to a process that uses voice tone and facial expression recognition technology to identify a user's emotional state.

[0566] "Means for generating feedback" refers to a function that automatically creates feedback including specific advice and areas for improvement for the user based on the results of text analysis, speech analysis, and emotion recognition.

[0567] "Means for presenting the generated feedback to the user" refers to an interface, such as a web browser or mobile app, that displays the generated feedback so that the user can review it.

[0568] The present invention relates to a system for improving the quality of meetings by utilizing AI technology. This system records the contents of meetings, extracts audio from the recorded data, converts the audio to text, analyzes the text, analyzes the audio data, and generates feedback. Furthermore, by combining it with an emotion engine, it is possible to detect the user's emotions and provide feedback based on those emotions. Specific embodiments for implementing this system are described below.

[0569] Hardware and software used

[0570] The device uses the camera and microphone of a laptop or smartphone. For example, a typical laptop or smartphone has a built-in camera and microphone for capturing video and audio.

[0571] The data transfer and processing is carried out by servers using the following software and technologies:

[0572] Audio processing library: FFmpeg

[0573] Speech recognition service: Google Cloud Speech-to-Text API

[0574] Natural language processing library: SpaCy

[0575] Acoustic analysis library: Librosa

[0576] Machine learning frameworks: TensorFlow, PyTorch

[0577] Browser display: D3.js

[0578] System processing flow

[0579] 1. Meeting recording: The device captures video and audio as soon as the meeting starts. For example, you can use Zoom or other video conferencing software to record the meeting. The recording data is sent to the server in real time.

[0580] 2. Audio Extraction: Extract the audio portion from the recording data received by the server. Use FFmpeg to separate the audio stream from the combined video and audio data.

[0581] 3. Speech-to-text conversion: The server sends the extracted voice data to the Google Cloud Speech-to-Text API to convert the voice to text, thereby obtaining the content of the meeting in text format.

[0582] 4. Text Analysis: The server analyzes the acquired text data using natural language processing (NLP) techniques, including tokenization, part-of-speech tagging, and semantic analysis using the SpaCy library.

[0583] 5. Evaluation of speech data: The server uses the Librosa library to analyze the acoustic features of the speech data, measure pitch and intensity from the speech data, and evaluate the speaking rate.

[0584] 6. Emotion recognition using an emotion engine: The server analyzes audio and video data to recognize the user's emotions. Using TensorFlow and other tools, it analyzes the tone of the voice and recognizes facial expressions to classify the user's emotions.

[0585] 7. Feedback Generation: The server generates feedback based on the results of text analysis, speech analysis, and emotion recognition. It uses natural language generation technology to create feedback that includes specific improvements.

[0586] 8. Present feedback: The device presents the feedback generated by the server to the user. This is visually displayed in the web browser. D3.js is used to present the feedback to the user using graphs and color coding to make it easy to understand.

[0587] Specific examples and prompts

[0588] As a concrete example, the following prompt sentence is input to the generative AI model:

[0589] "Extract audio from meeting recordings and convert it to text. Analyze the text content and perform emotion recognition to generate feedback."

[0590] This prompt sentence executes each processing step and generates specific feedback, for example, "Mr. Suzuki, you speak too quickly. Please try to speak a little more slowly."

[0591] In this way, the present invention can help meeting participants improve their communication skills based on specific feedback, potentially improving the overall quality of meetings. Furthermore, by adding emotion recognition functionality, feedback can be provided that takes into account psychological aspects during meetings, which is expected to improve productivity across the organization.

[0592] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0593] Step 1:

[0594] Recording a meeting

[0595] Input: When a meeting starts, the device captures the video and audio of the meeting.

[0596] Data processing: Recorded data that combines video and audio data is generated.

[0597] Output: The generated recording data is sent to the server in real time.

[0598] Specific operation: The device starts recording using video conferencing software (e.g., Zoom or Microsoft Teams) and transfers the data to the server.

[0599] Step 2:

[0600] Audio Extraction

[0601] Input: Recording data received by the server.

[0602] Data processing: Use FFmpeg to separate the audio stream from the combined video and audio data.

[0603] Output: Separated audio data (e.g., .wav format) is generated.

[0604] Specific operation: The server executes the FFmpeg command to extract and save the audio file from the recorded data.

[0605] Step 3:

[0606] Speech-to-text

[0607] Input: Audio data extracted by the server.

[0608] Data processing: Send the audio data to the Google Cloud Speech-to-Text API for speech recognition.

[0609] Output: The spoken words are captured in text format.

[0610] Specific operation: The server uploads the audio file to the Google Cloud Speech-to-Text API and stores the returned text data.

[0611] Step 4:

[0612] Text Analysis

[0613] Input: Speech-to-text data.

[0614] Data processing: Use the SpaCy library to tokenize the text, tag it with parts of speech, and perform semantic analysis.

[0615] Output: The analyzed text data is generated and the evaluation results are obtained.

[0616] What it does: The server uses the SpaCy library to analyze the text data and evaluate the phrasing, syntax, and argument structure of the comments.

[0617] Step 5:

[0618] Evaluation of audio data

[0619] Input: Extracted audio data.

[0620] Data processing: The Librosa library is used to analyze acoustic features and evaluate voice tone and speaking rate.

[0621] Output: Generates an assessment of tone of voice and speaking rate.

[0622] How it works: The server processes the audio data using Librosa, measures pitch and intensity, and evaluates the speaker's vocal characteristics.

[0623] Step 6:

[0624] Emotion recognition by emotion engine

[0625] Input: Audio and video data.

[0626] Data processing: Use TensorFlow and PyTorch to perform voice tone analysis and facial expression recognition to classify emotions.

[0627] Output: The user's emotional state (e.g., joy, anger, sadness, or happiness) is recognized.

[0628] How it works: The server analyzes the audio data and applies an emotion classification model to detect the user's emotional state, for example, whether they are angry, sad, happy, etc.

[0629] Step 7:

[0630] Generate feedback

[0631] Input: Text analysis results, speech analysis results, and emotion recognition results.

[0632] Data processing: The results of each analysis are integrated and feedback is automatically generated using natural language generation technology.

[0633] Output: A specific feedback text is generated for the user.

[0634] Specific behavior: The server uses a natural language generation engine (e.g., GPT-3) to generate feedback based on the analysis results, including specific suggestions for improvement. For example, it generates feedback like, "Mr. Sato, you spoke too quickly. You also seemed a little nervous. Please try to speak more relaxedly next time."

[0635] Step 8:

[0636] Providing feedback

[0637] Input: The generated feedback text.

[0638] Data processing: Converting JSON-formatted feedback data into a visually understandable format.

[0639] Output: Visual feedback that the user can see.

[0640] What it does: The device displays feedback in a web browser or mobile app, and uses D3.js to visually present the feedback as graphs and color-coded text.

[0641] (Application example 2)

[0642] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0643] When operating robots in factories, there is a need to evaluate the technical quality and emotional state of operators and provide specific feedback to improve operational efficiency and safety. However, with current systems, it is difficult to evaluate the emotional state and quality of operators' operations in real time and generate and provide feedback based on that evaluation.

[0644] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0645] In this invention, the server includes means for recording the contents of the conference, means for extracting audio from the recorded data, means for converting the extracted audio into text, means for analyzing the text to evaluate the vocabulary, syntax, and structure of the discussion, means for analyzing the audio data to evaluate the tone of voice and speaking rate, means for generating feedback based on the analysis results, means for presenting the generated feedback to the user, means for acquiring video and audio to record the operation status, means for analyzing the collected data to evaluate the efficiency and safety of the operation and the emotional state of the operator, and means for generating feedback based on the evaluation results, thereby enabling improvement of the operator's skills, ensuring safety, and providing appropriate guidance based on the emotional state.

[0646] "Means for recording meeting content" refers to devices and methods for recording video and audio of meetings.

[0647] "Means for extracting audio from recorded data" refers to a technology for extracting only the audio portion from recorded video and audio data.

[0648] The "means for converting extracted speech into text" is a speech recognition technology that converts speech data into character data.

[0649] "Methods for analyzing text to evaluate language, syntax, and argument structure" refers to techniques for analyzing text data and judging and evaluating its linguistic expression, sentence structure, and argument flow.

[0650] "Means for analyzing voice data to evaluate voice tone and speaking rate" refers to technology that analyzes and evaluates voice characteristics such as pitch, strength, and speaking rate.

[0651] "Means for generating feedback based on the analysis results" refers to a method for providing users with advice and suggestions for improvement based on the analyzed data.

[0652] "Means for presenting the generated feedback to the user" refers to technology for displaying and providing the generated feedback information in a form that the user can understand.

[0653] "Means for acquiring video and audio for recording the operation status" refers to a device or method for recording the operation status with video and audio.

[0654] "Means for analyzing collected data to evaluate the efficiency, safety, and emotional state of the operator" refers to technology that analyzes acquired data to judge and evaluate the efficiency, safety, and emotional state of the operator.

[0655] The "means for generating feedback based on the evaluation results" is a method for providing useful advice and suggestions for improvement to users based on the evaluation results.

[0656] This invention relates to an AI assistance system for improving the quality of robot operation in factories. This system records the situation during operation, generates feedback from the data, and provides it to the operator to help improve their technical skills and manage their emotions.

[0657] Hardware Configuration

[0658] The system includes the following hardware:

[0659] Smart glasses (e.g., Google Glass Enterprise Edition): worn by operators, record the operation status in real time.

[0660] Cameras and microphones: Installed within the factory to capture additional video and audio data.

[0661] Server: For data processing and analysis (e.g., Google Cloud Platform).

[0662] Software Configuration

[0663] The software used is as follows:

[0664] Speech recognition engine (e.g. Google Cloud Speech-to-Text API): Converts extracted speech into text.

[0665] Natural language processing (NLP) techniques (e.g., Google Cloud Natural Language API): Analyze text data to evaluate phrasing, syntax, and argument structure.

[0666] Emotion recognition engine (e.g., Microsoft Azure Face API): Recognizes the operator's emotional state from video and audio data.

[0667] Data analysis programs (e.g., Python, TensorFlow): used to evaluate data and generate feedback.

[0668] Data processing and feedback generation

[0669] The server receives the video and audio data sent from the smart glasses and the camera, and then processes the data in the following steps:

[0670] 1. Audio extraction: The audio is separated from the video data and input into a speech recognition engine. This process converts the audio into text.

[0671] 2. Text analysis: The obtained text data is analyzed using natural language processing technology to evaluate the content of the statements.

[0672] 3. Voice analysis: Analyzes the characteristics of voice data and evaluates the tone of voice and speaking rate, thereby extracting how efficient and safe the operation is.

[0673] 4. Emotion Recognition: Analyze the operator's emotional state using video and audio data and incorporate the results into the evaluation.

[0674] Generating and Presenting Feedback

[0675] The server generates and provides feedback to the user (operator) based on the analysis results, including specific advice on operational efficiency, safety, and emotional state.

[0676] Specific examples

[0677] An operator wearing smart glasses starts recording while operating a robot. The recorded data is sent to a server, where a speech recognition engine converts the speech into text. Natural language processing technology then analyzes the text to evaluate the accuracy and efficiency of the operation. At the same time, an emotion recognition engine analyzes the operator's emotional state from the video data. Ultimately, based on the results of these analyses, feedback is generated, such as, "The order of operations is efficient, but you appear to be a little impatient. Please operate more slowly next time."

[0678] Prompt Sentence Examples

[0679] Voice recognition prompts

[0680] Use the speech recognition API to convert the following audio data to text:

[0681] <path to audio data>

[0682] Prompts for Natural Language Processing

[0683] Use the recognized text data to extract the following information:

[0684] Key Points

[0685] Utterance Syntax

[0686] Semantic analysis

[0687] Emotion recognition prompts

[0688] Detect operator emotions from video data:

[0689] <Video data path>

[0690] Prompts for generating feedback

[0691] Use the data you collect to generate feedback that includes the following metrics:

[0692] Evaluating operational efficiency

[0693] Safety evaluation

[0694] Emotion-based feedback

[0695] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0696] Step 1:

[0697] Video and audio acquisition

[0698] The device uses smart glasses or cameras and microphones in the factory to record video and audio of the operator's operations. This video and audio data is sent to the server in real time. The input is video and audio data, and the output is stream data transferred to the server.

[0699] Step 2:

[0700] Audio Extraction

[0701] The server extracts the audio portion from the transmitted video recording data. Specifically, it uses audio processing technology to separate the audio stream from the combined video and audio data. The input is the video and audio stream data, and the output is the extracted audio data.

[0702] Step 3:

[0703] Speech-to-text

[0704] The server converts the extracted audio data into text using a speech recognition engine (e.g., Google Cloud Speech-to-Text API). In this process, the speech recognition engine analyzes the audio waveform and generates a corresponding string using a language model. The input is the audio data, and the output is the generated text data.

[0705] Step 4:

[0706] Text Analysis

[0707] The server analyzes the generated text data using natural language processing technology (e.g., Google Cloud Natural Language API). This analysis includes tokenization, part-of-speech tagging, and semantic analysis, and evaluates the content of the speech. The input is text data, and the output is the analysis results (key points, sentence structure, flow of discussion, etc.).

[0708] Step 5:

[0709] Evaluation of audio data

[0710] The server analyzes the features of the voice data and evaluates the tone of voice and speaking speed. Here, acoustic features (e.g., pitch, intensity) are extracted and evaluation is performed based on that information. The input is the voice data, and the output is the voice features (voice pitch, speed, strength, etc.) and the evaluation results.

[0711] Step 6:

[0712] Emotion recognition

[0713] The server uses an emotion recognition engine (e.g., Microsoft Azure Face API) to analyze the operator's emotional state from the video and audio data. This includes facial expression analysis technology and voice tone analysis. The input is video and audio data, and the output is a classification result of the operator's emotional state.

[0714] Step 7:

[0715] Generate feedback

[0716] The server generates specific feedback based on the analysis results. This feedback includes evaluations and advice on operational efficiency, safety, and emotional state. The inputs are the text analysis results, speech analysis results, and emotion recognition results, and the generated feedback is obtained as the output.

[0717] Step 8:

[0718] Providing feedback

[0719] The terminal presents the generated feedback to the operator, using a web browser or mobile app to display the feedback in a visually understandable format. The input is the generated feedback, and the output is the visualized feedback that can be viewed by the user.

[0720] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0721] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0722] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0723] [Third embodiment]

[0724] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0725] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0726] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0727] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0728] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0729] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0730] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0731] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0732] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0733] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0734] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0735] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0736] The present invention relates to a system for improving the quality of meetings by utilizing AI technology. A specific embodiment of this system will be described.

[0737] 1. Recording meetings

[0738] The device captures video and audio as soon as the meeting starts. Specifically, the entire meeting is recorded using the camera and microphone of a laptop or smartphone, for example. The recorded data is uploaded to a server in real time.

[0739] 2. Audio Extraction

[0740] The server extracts the audio portion from the received recording. This extraction can be done using common audio processing techniques, including, for example, separating the audio stream from the combined video and audio of the meeting.

[0741] 3. Speech-to-text

[0742] The server then requests a speech recognition engine to convert the extracted voice data. Specifically, it calls a service such as the Google Cloud Speech-to-Text API to convert the voice into text data. As a result, the spoken content is obtained in text format.

[0743] 4. Text Analysis

[0744] The server analyzes the text data. Natural language processing (NLP) techniques are used for text analysis. For example, Python NLP libraries such as spaCy and BERT can be used to tokenize the text, tag parts of speech, and perform semantic analysis. This analysis makes it possible to evaluate the vocabulary, syntax, and argument structure of the speech.

[0745] 5. Evaluation of audio data

[0746] The server analyzes the audio data and evaluates the tone of voice and speaking rate. This includes analyzing acoustic features. For example, it uses Praat or an open-source audio analysis library to measure pitch (highness of voice) and intensity (power of voice) and makes an evaluation based on these.

[0747] 6. Generate feedback

[0748] The server generates feedback based on the results of text and speech analysis. The feedback is provided to each participant in the form of an evaluation result and recommendations for improvement. For example, it may include specific comments such as "Mr. Sato speaks too quickly" and suggestions for improvement such as "Try to speak a little more slowly next time."

[0749] 7. Providing Feedback

[0750] The device presents the generated feedback to the user. The feedback is displayed visually in a web browser or mobile app, allowing the user to intuitively identify problems and areas for improvement in their comments. For example, feedback could be displayed separately for each comment, and a visually easy-to-understand format could be used, using color coding or graphs.

[0751] Using this system, meeting participants can improve their communication skills based on specific feedback, which is expected to improve the overall quality of meetings and increase productivity across the organization.

[0752] The processing flow will be explained below.

[0753] Step 1:

[0754] The device captures video and audio as soon as the meeting starts. For example, the entire meeting can be recorded using the camera and microphone of a laptop or smartphone, and the data can be sent to a server in real time.

[0755] Step 2:

[0756] The server extracts the audio portion from the recorded data it receives. Specifically, it uses audio processing technology to separate the audio stream from the combined video and audio data.

[0757] Step 3:

[0758] The server inputs the extracted voice data into a speech recognition engine and converts the voice to text, for example, using the Google Cloud Speech-to-Text API.

[0759] Step 4:

[0760] The server analyzes the text data, using natural language processing (NLP) techniques to tokenize, tag parts of speech, and analyze the semantics of the text. This analysis evaluates the phrasing, syntax, and argument structure of the speech.

[0761] Step 5:

[0762] The server uses the audio data to evaluate the tone of voice and speaking rate. Specifically, it analyzes acoustic features to measure pitch (highness of voice) and intensity (power of voice).

[0763] Step 6:

[0764] The server generates feedback based on the results of text and speech analysis, for example, using a specific template to create feedback for each participant, including specific evaluations and suggestions for improvement.

[0765] Step 7:

[0766] The device then presents the generated feedback to the user, which is displayed visually via a smartphone, PC web browser, or dedicated app. The user can then review the displayed feedback and understand their areas for improvement.

[0767] Example 1

[0768] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0769] While conventional conference systems can record conference content and convert audio into text, they lack the functionality to evaluate the quality of the conference and provide specific feedback. As a result, users have no way to objectively evaluate their own comments or the way the conference proceeded, which hinders the improvement of their communication skills. The present invention aims to solve these problems and provide a system that improves the quality of conferences.

[0770] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0771] In this invention, the server includes means for recording the contents of the conference, means for uploading the recorded data to the server in real time, means for extracting voice from the recorded data, means for transmitting the extracted voice to a voice recognition engine and converting it into text data, means for analyzing the text data using natural language processing technology, means for analyzing the voice data using acoustic analysis technology and evaluating the tone of voice and speaking rate, means for generating feedback based on the results of the text analysis and the voice analysis, and means for presenting the generated feedback to the user. This allows conference participants to receive evaluations of their own remarks and voice and know specific areas for improvement, thereby enabling them to improve their communication skills.

[0772] A "meeting recording device" is a device or method that captures and stores video and audio of a meeting in digital form.

[0773] "Means for uploading recorded data to a server in real time" refers to the function of instantly transmitting video and audio data generated during the progress of a conference to a server.

[0774] "Means for extracting audio from recorded data" refers to a technology for separating and extracting the audio portion from recorded video data.

[0775] "Means for transmitting the extracted speech to a speech recognition engine and converting it into text data" refers to the process of transmitting the extracted speech data to speech recognition software and using that software to convert the speech into text information.

[0776] "Means for analyzing text data using natural language processing technology" refers to a technology that analyzes acquired text data using a natural language processing algorithm and evaluates the wording, syntax, and structure of the argument.

[0777] "Means for analyzing voice data using acoustic analysis technology to evaluate voice tone and speaking rate" refers to technology for analyzing voice data using an acoustic analysis tool and measuring pitch and intensity to evaluate voice tone and speaking rate.

[0778] The "means for generating feedback based on text analysis results and speech analysis results" refers to a process of combining the analysis results of text data and speech data to create feedback that includes specific evaluations and improvement measures.

[0779] The "means for presenting the generated feedback to the user" is a device or method that visually or audibly conveys the generated feedback to the user.

[0780] The present invention relates to a system for improving the quality of meetings by utilizing AI technology. Specific embodiments of this system are described in detail below.

[0781] First, at the start of a meeting, the user launches the meeting application and presses the "Start Meeting" button. Devices (e.g., laptops, smartphones) that detect the start of the meeting begin capturing video and audio using their built-in cameras and microphones. This allows the entire meeting to be recorded.

[0782] The captured video and audio data is uploaded to a server in real time using HTTP streaming or WebRTC technology, for example. The server processes the received recording data and separates the video and audio data. Specifically, audio processing technology such as FFmpeg is used to extract the audio stream from the recording data.

[0783] The extracted voice data is then sent by the server to the Google Cloud Speech-to-Text API, where it is converted into text data. This step allows the voice utterance to be obtained in text format.

[0784] The captured text data is then analyzed by the server using natural language processing (NLP) techniques, such as tokenizing, part-of-speech tagging, and semantic analysis using Python NLP libraries (e.g., spaCy, BERT), to evaluate the vocabulary, syntax, and argument structure of the speech.

[0785] Additionally, the voice data is analyzed using acoustic analysis technology. The server analyzes the voice data using an acoustic analysis library (e.g., Librosa) to measure pitch and intensity. This evaluation evaluates the tone of voice and speaking rate.

[0786] The server generates feedback based on these analysis results. The generated feedback includes specific suggestions for improvement and evaluations, such as "You speak too quickly" or "Try to speak more slowly next time." Finally, the generated feedback is presented to the user via the device. The feedback is displayed visually in a web browser or mobile app, and is presented in an easily understandable format using color-coded graphs and icons.

[0787] For example, when a user holds a meeting using video conferencing software, the laptop's camera and microphone capture video and audio and upload them to a server in real time. The server then extracts the audio using FFmpeg and converts it to text using the Google Cloud Speech-to-Text API. The text is then analyzed using spaCy, and the audio is evaluated using Librosa. Finally, feedback generated based on the analysis results is displayed in the web browser for the user to review.

[0788] An example of a prompt might be, "Please use the Google Cloud Speech-to-Text API to convert the following audio data into text and generate meeting feedback based on the results." By using this system, meeting participants can receive evaluations of their own remarks and voices, learn specific areas for improvement, and improve their communication skills.

[0789] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0790] Step 1:

[0791] The user launches the conference application and presses the "Start Conference" button. This causes the device to detect the start of the conference. The input is the user's operation action, and the output is the start of the conference recording process.

[0792] Step 2:

[0793] The device will begin capturing video and audio using the built-in camera and microphone. The captured data will be saved as video and audio files. The input is real-time data from the built-in camera and microphone, and the output is locally saved video and audio data.

[0794] Step 3:

[0795] The device uploads the captured video and audio data to the server in real time. This process uses HTTP streaming or WebRTC technology. The input is the locally stored video and audio data, and the output is the data uploaded to the server.

[0796] Step 4:

[0797] The server processes the received recording data and separates the video and audio data using audio processing technology such as FFmpeg. The input is the uploaded recording data, and the output is the separated video and audio data.

[0798] Step 5:

[0799] The server sends the extracted speech data to a speech recognition engine and converts it into text data. This step uses the Google Cloud Speech-to-Text API. The extracted speech data is the input, and the text data is the output.

[0800] Step 6:

[0801] The server analyzes the acquired text data using natural language processing (NLP) techniques. Specifically, it uses Python NLP libraries (e.g., spaCy, BERT) to perform tokenization, part-of-speech tagging, and semantic analysis. The converted text data is input, and the analysis results are obtained as output.

[0802] Step 7:

[0803] The server analyzes the voice data using acoustic analysis technology to evaluate the tone of voice and speaking rate. Specifically, it uses the Librosa library to measure pitch and intensity. The input is the voice data, and the output is the result of the acoustic analysis.

[0804] Step 8:

[0805] The server generates feedback based on the results of text analysis and speech analysis. The feedback includes specific evaluations and suggestions for improvement. The inputs are the results of text analysis and speech analysis, and the generated feedback is obtained as the output.

[0806] Step 9:

[0807] The device presents the generated feedback to the user, which is then visually displayed in a web browser or mobile app. The input is the generated feedback, and the output is the feedback presented to the user.

[0808] (Application example 1)

[0809] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0810] Conventional conference systems lack specific feedback to improve the quality of meetings, making it difficult for participants to improve their own comments and attitudes. Furthermore, in certain environments, such as project review meetings in factories, recording and analysis methods and tools are often not properly installed, hindering the effective running of meetings. The present invention aims to provide a system that solves these problems and improves the quality of meetings.

[0811] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0812] In this invention, the server includes a means for recording the contents of the conference, a means for extracting audio from the recorded data, and a means for converting the extracted audio into text, thereby enabling the server to accurately record the overall picture of the conference and provide specific feedback based on the analysis results.

[0813] "Means for recording the contents of a conference" refers to a device or program that captures and records video and audio as soon as the conference begins.

[0814] The "means for extracting audio from recorded data" refers to a processing device or program for separating the audio portion from the recorded conference data and extracting it as audio data.

[0815] "Means for converting extracted speech into text" refers to a program or service that converts extracted speech data into text format using speech recognition technology.

[0816] A "means for analyzing text to evaluate language, syntax, and argument structure" is a program that uses natural language processing techniques on text data to evaluate and analyze the content of the text.

[0817] The "means for analyzing voice data to evaluate voice tone and speaking rate" refers to a device or program that uses voice data to analyze acoustic features such as voice pitch and speaking rate and make an evaluation.

[0818] The "means for generating feedback based on the analysis results" is a program that provides participants with specific suggestions for improvement and evaluations based on the results of text and voice analysis.

[0819] The "means for presenting the generated feedback to the user" refers to a device or program that visually displays the generated feedback so that the user can easily confirm it.

[0820] "Means to be installed on factory robots for project review meetings in a factory work environment" refers to programs and devices that are introduced into robots in factories for the purpose of recording meetings and analyzing data within the factory.

[0821] The "means for uploading recorded audio and video data to a server" refers to a communication device or program for transmitting recorded audio and video data to a server and performing analysis processing.

[0822] The "means for providing the generated feedback to participants within the factory" refers to a device or program that notifies each member who participates in the meeting within the factory of feedback based on the analysis results and suggests areas for improvement.

[0823] The present invention relates to a system for improving the quality of project review meetings in a factory. Specific embodiments of the system will be described below.

[0824] System Overview

[0825] This system is used in a factory work environment and is intended for project review meetings. It is composed mainly of a program installed on a factory robot, which records meetings, analyzes data, and generates and provides feedback. Each process is executed by a server, and feedback is provided to participants via their terminals.

[0826] Meeting Recording

[0827] First, a camera and microphone mounted on a factory robot are used to record the contents of the meeting as soon as it starts, and the recorded data is uploaded to a server in real time as video and audio.

[0828] Audio extraction and analysis

[0829] The server extracts the audio portion from the received video recordings. This extraction uses common audio processing techniques, such as separating the audio stream from the combined video and audio data. This process is performed using a programming language such as Python.

[0830] Speech-to-text

[0831] The server converts the extracted voice data into text using the Google Cloud Speech-to-Text API. This API is called and speech recognition technology is used to convert the voice into text data. As a result, the voice utterances are obtained in text format.

[0832] Text analytics

[0833] The server then analyzes the acquired text data using natural language processing techniques. Specifically, it uses the Python NLP libraries spaCy and BERT to tokenize the text, tag parts of speech, and perform semantic analysis. This analysis allows it to evaluate the vocabulary, syntax, and structure of the discussions of participants' comments.

[0834] Audio analysis

[0835] Similarly, the server analyzes the audio data and performs acoustic feature analysis to assess the tone of voice and speaking rate, using Praat or other open-source speech analysis libraries. This analysis measures the pitch and intensity of the speech and evaluates the quality of the speech.

[0836] Generating and Presenting Feedback

[0837] The server generates specific feedback based on the text and speech analysis results. The feedback is generated for each participant in a format that includes evaluation results and recommended improvements. The generated feedback is presented to the user visually on a web browser or mobile app.

[0838] Examples and prompts

[0839] For example, in a project review meeting held in a factory, the system analyzes the speaking style and content of each participant and provides feedback such as, "Yamada-san, you use too much technical jargon, so try using more understandable words." If the tone of your voice is too high, the system also provides specific examples such as, "By lowering your voice a little, you can give the impression of being more calm."

[0840] Prompt Sentence Examples

[0841] Analysis of statements made during factory meetings:

[0842] 1. Upload the recorded audio data to the following URL:

[0843] 2. Please analyze the uploaded audio data and generate feedback in the following format:

[0844] 3. Provide specific suggestions for improving each speaker's delivery.

[0845] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0846] Step 1:

[0847] Recording of meeting content

[0848] As soon as the meeting starts, the device uses its camera and microphone to capture video and audio, recording in real time. The input is the video and audio of the meeting, and the output is recorded data. Specifically, the device's camera captures images of the meeting participants, and the microphone picks up audio. The recorded data is temporarily stored in the device's memory.

[0849] Step 2:

[0850] Uploading recording data

[0851] The device uploads the recorded data to the server. The input is the recorded data, and the output is the data uploaded to the server. Specifically, after the device finishes recording, it transfers the data to the server via Wi-Fi or Ethernet. For security reasons, the SSL / TLS protocol is used.

[0852] Step 3:

[0853] Audio Extraction

[0854] The server extracts the audio portion from the uploaded video data. The input is the video data, and the output is the audio data. Specifically, an audio processing module on the server separates the audio stream from the combined video and audio data and saves it as audio data. Libraries such as FFmpeg are used.

[0855] Step 4:

[0856] Speech-to-text

[0857] The server converts the extracted voice data into text using the Google Cloud Speech-to-Text API. The input is voice data, and the output is text data. Specifically, the voice data is sent to the API, which uses speech recognition technology to convert it into text, and the result is returned to the server.

[0858] Step 5:

[0859] Text Analysis

[0860] The server analyzes the acquired text data using natural language processing technology. The input is the text data, and the output is the analysis results. Specifically, it uses Python NLP libraries (e.g., spaCy and BERT) to perform tokenization, part-of-speech tagging, and semantic analysis. This analysis evaluates the text's vocabulary, syntax, and argument structure.

[0861] Step 6:

[0862] Analysis of audio data

[0863] The server analyzes the audio data and evaluates the tone of voice and speaking rate. The input is the audio data, and the output is the analysis results. Specifically, it uses Praat and other open source audio analysis libraries to measure pitch (voice height) and intensity (voice strength), and makes an evaluation based on the results.

[0864] Step 7:

[0865] Generate feedback

[0866] The server generates feedback based on the text analysis results and speech analysis results. The input is the text analysis results and speech analysis results, and the output is the generated feedback. Specifically, feedback is generated for each participant that includes evaluation results and recommended improvements. For example, it includes a comment such as "A-san speaks too quickly" and a suggestion for improvement such as "Try to speak more slowly next time."

[0867] Step 8:

[0868] Providing feedback

[0869] The device presents the generated feedback to the user via a web browser or mobile app. The input is the generated feedback, and the output is what is presented to the user. Specifically, the feedback is displayed separately for each comment, and is provided in a visually easy-to-understand format using color coding and graphs. For example, the feedback is displayed in a dashboard format, allowing the user to intuitively identify problems and areas for improvement in their own comments.

[0870] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0871] The present invention relates to a system for improving the quality of meetings using AI technology. This system records the contents of a meeting, extracts audio from the recorded data, converts the audio to text, analyzes the text, analyzes the audio data, and generates feedback. Furthermore, by combining it with an emotion engine, it is possible to detect the user's emotions and provide feedback based on those emotions. Specific embodiments are described below.

[0872] 1. Recording meetings

[0873] The device captures video and audio as soon as the meeting starts. For example, the entire meeting can be recorded using the camera and microphone of a laptop or smartphone. This recorded data is then sent to a server in real time.

[0874] 2. Audio Extraction

[0875] The server extracts the audio portion from the recorded data it receives. Specifically, it uses audio processing technology to separate the audio stream from the combined video and audio data.

[0876] 3. Speech-to-text

[0877] The server inputs the extracted voice data into a speech recognition engine and converts the speech into text. For example, it uses a speech recognition service such as Google Cloud Speech-to-Text API. As a result, the spoken content is obtained in text format.

[0878] 4. Text Analysis

[0879] The server analyzes the text data, specifically using natural language processing (NLP) techniques to tokenize, tag parts of speech, and perform semantic analysis of the text, evaluating the phrasing, syntax, and argument structure of what is being said.

[0880] 5. Evaluation of audio data

[0881] The server uses the audio data to evaluate the tone of voice and speaking speed, for example by analyzing acoustic features and measuring pitch (voice height) and intensity (voice strength).

[0882] 6. Emotion Recognition by Emotion Engine

[0883] The server analyzes the audio and video data to recognize the user's emotions. For example, it uses voice tone and facial expression analysis technology to detect whether the user is happy, angry, or sad. To do this, it uses a machine learning model to classify emotions.

[0884] 7. Generate feedback

[0885] The server generates feedback based on the results of text analysis, speech analysis, and emotion recognition. The feedback is provided to each participant in the form of an evaluation result and recommendations for improvement. For example, specific feedback such as, "Mr. Sato, you spoke too quickly. You also seemed a little nervous. Try to speak more relaxedly next time."

[0886] 8. Providing Feedback

[0887] The device presents the generated feedback to the user. The feedback is displayed visually in a web browser or mobile app. The user can review the feedback and understand their areas for improvement. For example, the feedback is displayed separately for each utterance and presented in an easy-to-understand format using visualization techniques such as color coding and graphs.

[0888] In this way, the present invention allows meeting participants to improve their communication skills based on specific feedback. Furthermore, by adding emotion recognition functionality, feedback can be provided that takes psychological aspects into account, further improving the overall quality of meetings and increasing productivity across the organization.

[0889] The processing flow will be explained below.

[0890] Step 1:

[0891] The device captures video and audio as soon as the meeting starts. Specifically, it uses the laptop's camera and microphone to record the entire meeting. The recorded data is sent to the server in real time.

[0892] Step 2:

[0893] The server extracts the audio portion from the recorded data it receives. Audio processing technology is used to separate the audio stream from the combined video and audio data. This step also involves format conversion of the audio file and noise filtering.

[0894] Step 3:

[0895] The server sends the extracted voice data to a speech recognition engine and converts the voice into text, for example, using the Google Cloud Speech-to-Text API, which then stores the converted text in a database.

[0896] Step 4:

[0897] The server analyzes the text data, using natural language processing (NLP) techniques to perform tokenization, part-of-speech tagging, and semantic analysis. This analysis evaluates the phrasing, syntax, and argument structure of what is being said.

[0898] Step 5:

[0899] The server analyzes the voice data to evaluate the tone of voice and speaking rate. For example, it uses tools such as Praat to measure pitch, intensity, and speaking rate. Based on this, it generates a voice evaluation result.

[0900] Step 6:

[0901] The server uses an emotion engine to recognize the user's emotions from audio and video data. It analyzes the tone of the voice and facial expressions to classify the emotional state. It uses a machine learning model to determine whether the user is happy, angry, or sad.

[0902] Step 7:

[0903] The server generates feedback based on the results of text analysis, speech evaluation, and emotion recognition. For example, it generates specific feedback such as, "Mr. Sato, you spoke too quickly. You also seemed a little nervous. Please try to speak more relaxedly next time."

[0904] Step 8:

[0905] The device presents the generated feedback to the user. The feedback is displayed visually through a web browser or mobile app. The user can review the feedback and understand their areas for improvement. For example, the feedback could be displayed in a color-coded format for each comment, and could be formatted in a way that makes it easy to understand visually using graphs and charts.

[0906] Example 2

[0907] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0908] Conventional meeting support systems did not accurately record or analyze what was said during meetings, which resulted in vague feedback and made it difficult to provide specific suggestions for improvement. Furthermore, there was no way to provide feedback that took into account participants' emotions and psychological states, resulting in a lack of specific guidance for improving the quality of communication.

[0909] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for recording the contents of the conference, means for extracting audio from the recorded data, means for converting the extracted audio into text, means for analyzing the text to evaluate the wording, syntax, and structure of the discussion, means for analyzing the audio data to evaluate the tone of voice and speaking rate, means for analyzing the audio data and video data to recognize the user's emotions, means for generating feedback based on the analysis results, and means for presenting the generated feedback to the user. This makes it possible to accurately record and analyze the content of comments made in the conference and provide specific feedback that takes into account the emotions of the participants.

[0910] "Meeting content" refers to the entire information, materials, and presentations spoken or shared by participants at the meeting.

[0911] "Recording Data" means the digital files of video and audio captured during a meeting.

[0912] "Means for extracting audio" refers to a function that performs the process of separating and extracting the audio portion from the recorded data.

[0913] "Means for converting speech to text" refers to technology, particularly a speech recognition engine, that automatically converts extracted speech data into text data.

[0914] "Means of analyzing text to evaluate language, syntax, and argument structure" refers to the application of natural language processing techniques to converted text data to analyze word choice, grammar, and argument flow.

[0915] "Means for analyzing voice data to evaluate voice tone and speaking rate" refers to technology that uses acoustic analysis technology to measure the features of voice data and evaluate the pitch, strength, and speaking rate of the voice.

[0916] "Means for recognizing a user's emotions by analyzing audio and video data" refers to a process that uses voice tone and facial expression recognition technology to identify a user's emotional state.

[0917] "Means for generating feedback" refers to a function that automatically creates feedback including specific advice and areas for improvement for the user based on the results of text analysis, speech analysis, and emotion recognition.

[0918] "Means for presenting the generated feedback to the user" refers to an interface, such as a web browser or mobile app, that displays the generated feedback so that the user can review it.

[0919] The present invention relates to a system for improving the quality of meetings by utilizing AI technology. This system records the contents of meetings, extracts audio from the recorded data, converts the audio to text, analyzes the text, analyzes the audio data, and generates feedback. Furthermore, by combining it with an emotion engine, it is possible to detect the user's emotions and provide feedback based on those emotions. Specific embodiments for implementing this system are described below.

[0920] Hardware and software used

[0921] The device uses the camera and microphone of a laptop or smartphone. For example, a typical laptop or smartphone has a built-in camera and microphone for capturing video and audio.

[0922] The data transfer and processing is carried out by servers using the following software and technologies:

[0923] Audio processing library: FFmpeg

[0924] Speech recognition service: Google Cloud Speech-to-Text API

[0925] Natural language processing library: SpaCy

[0926] Acoustic analysis library: Librosa

[0927] Machine learning frameworks: TensorFlow, PyTorch

[0928] Browser display: D3.js

[0929] System processing flow

[0930] 1. Meeting recording: The device captures video and audio as soon as the meeting starts. For example, you can use Zoom or other video conferencing software to record the meeting. The recording data is sent to the server in real time.

[0931] 2. Audio Extraction: Extract the audio portion from the recording data received by the server. Use FFmpeg to separate the audio stream from the combined video and audio data.

[0932] 3. Speech-to-text conversion: The server sends the extracted voice data to the Google Cloud Speech-to-Text API to convert the voice to text, thereby obtaining the content of the meeting in text format.

[0933] 4. Text Analysis: The server analyzes the acquired text data using natural language processing (NLP) techniques, including tokenization, part-of-speech tagging, and semantic analysis using the SpaCy library.

[0934] 5. Evaluation of speech data: The server uses the Librosa library to analyze the acoustic features of the speech data, measure pitch and intensity from the speech data, and evaluate the speaking rate.

[0935] 6. Emotion recognition using an emotion engine: The server analyzes audio and video data to recognize the user's emotions. Using TensorFlow and other tools, it analyzes the tone of the voice and recognizes facial expressions to classify the user's emotions.

[0936] 7. Feedback Generation: The server generates feedback based on the results of text analysis, speech analysis, and emotion recognition. It uses natural language generation technology to create feedback that includes specific improvements.

[0937] 8. Present feedback: The device presents the feedback generated by the server to the user. This is visually displayed in the web browser. D3.js is used to present the feedback to the user using graphs and color coding to make it easy to understand.

[0938] Specific examples and prompts

[0939] As a concrete example, the following prompt sentence is input to the generative AI model:

[0940] "Extract audio from meeting recordings and convert it to text. Analyze the text content and perform emotion recognition to generate feedback."

[0941] This prompt sentence executes each processing step and generates specific feedback, for example, "Mr. Suzuki, you speak too quickly. Please try to speak a little more slowly."

[0942] In this way, the present invention can help meeting participants improve their communication skills based on specific feedback, potentially improving the overall quality of meetings. Furthermore, by adding emotion recognition functionality, feedback can be provided that takes into account psychological aspects during meetings, which is expected to improve productivity across the organization.

[0943] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0944] Step 1:

[0945] Recording a meeting

[0946] Input: When a meeting starts, the device captures the video and audio of the meeting.

[0947] Data processing: Recorded data that combines video and audio data is generated.

[0948] Output: The generated recording data is sent to the server in real time.

[0949] Specific operation: The device starts recording using video conferencing software (e.g., Zoom or Microsoft Teams) and transfers the data to the server.

[0950] Step 2:

[0951] Audio Extraction

[0952] Input: Recording data received by the server.

[0953] Data processing: Use FFmpeg to separate the audio stream from the combined video and audio data.

[0954] Output: Separated audio data (e.g., .wav format) is generated.

[0955] Specific operation: The server executes the FFmpeg command to extract and save the audio file from the recorded data.

[0956] Step 3:

[0957] Speech-to-text

[0958] Input: Audio data extracted by the server.

[0959] Data processing: Send the audio data to the Google Cloud Speech-to-Text API for speech recognition.

[0960] Output: The spoken words are captured in text format.

[0961] Specific operation: The server uploads the audio file to the Google Cloud Speech-to-Text API and stores the returned text data.

[0962] Step 4:

[0963] Text Analysis

[0964] Input: Speech-to-text data.

[0965] Data processing: Use the SpaCy library to tokenize the text, tag it with parts of speech, and perform semantic analysis.

[0966] Output: The analyzed text data is generated and the evaluation results are obtained.

[0967] What it does: The server uses the SpaCy library to analyze the text data and evaluate the phrasing, syntax, and argument structure of the comments.

[0968] Step 5:

[0969] Evaluation of audio data

[0970] Input: Extracted audio data.

[0971] Data processing: The Librosa library is used to analyze acoustic features and evaluate voice tone and speaking rate.

[0972] Output: Generates an assessment of tone of voice and speaking rate.

[0973] How it works: The server processes the audio data using Librosa, measures pitch and intensity, and evaluates the speaker's vocal characteristics.

[0974] Step 6:

[0975] Emotion recognition by emotion engine

[0976] Input: Audio and video data.

[0977] Data processing: Use TensorFlow and PyTorch to perform voice tone analysis and facial expression recognition to classify emotions.

[0978] Output: The user's emotional state (e.g., joy, anger, sadness, or happiness) is recognized.

[0979] How it works: The server analyzes the audio data and applies an emotion classification model to detect the user's emotional state, for example, whether they are angry, sad, happy, etc.

[0980] Step 7:

[0981] Generate feedback

[0982] Input: Text analysis results, speech analysis results, and emotion recognition results.

[0983] Data processing: The results of each analysis are integrated and feedback is automatically generated using natural language generation technology.

[0984] Output: A specific feedback text is generated for the user.

[0985] Specific behavior: The server uses a natural language generation engine (e.g., GPT-3) to generate feedback based on the analysis results, including specific suggestions for improvement. For example, it generates feedback like, "Mr. Sato, you spoke too quickly. You also seemed a little nervous. Please try to speak more relaxedly next time."

[0986] Step 8:

[0987] Providing feedback

[0988] Input: The generated feedback text.

[0989] Data processing: Converting JSON-formatted feedback data into a visually understandable format.

[0990] Output: Visual feedback that the user can see.

[0991] What it does: The device displays feedback in a web browser or mobile app, and uses D3.js to visually present the feedback as graphs and color-coded text.

[0992] (Application example 2)

[0993] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0994] When operating robots in factories, there is a need to evaluate the technical quality and emotional state of operators and provide specific feedback to improve operational efficiency and safety. However, with current systems, it is difficult to evaluate the emotional state and quality of operators' operations in real time and generate and provide feedback based on that evaluation.

[0995] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0996] In this invention, the server includes means for recording the contents of the conference, means for extracting audio from the recorded data, means for converting the extracted audio into text, means for analyzing the text to evaluate the vocabulary, syntax, and structure of the discussion, means for analyzing the audio data to evaluate the tone of voice and speaking rate, means for generating feedback based on the analysis results, means for presenting the generated feedback to the user, means for acquiring video and audio to record the operation status, means for analyzing the collected data to evaluate the efficiency and safety of the operation and the emotional state of the operator, and means for generating feedback based on the evaluation results, thereby enabling improvement of the operator's skills, ensuring safety, and providing appropriate guidance based on the emotional state.

[0997] "Means for recording meeting content" refers to devices and methods for recording video and audio of meetings.

[0998] "Means for extracting audio from recorded data" refers to a technology for extracting only the audio portion from recorded video and audio data.

[0999] The "means for converting extracted speech into text" is a speech recognition technology that converts speech data into character data.

[1000] "Methods for analyzing text to evaluate language, syntax, and argument structure" refers to techniques for analyzing text data and judging and evaluating its linguistic expression, sentence structure, and argument flow.

[1001] "Means for analyzing voice data to evaluate voice tone and speaking rate" refers to technology that analyzes and evaluates voice characteristics such as pitch, strength, and speaking rate.

[1002] "Means for generating feedback based on the analysis results" refers to a method for providing users with advice and suggestions for improvement based on the analyzed data.

[1003] "Means for presenting the generated feedback to the user" refers to technology for displaying and providing the generated feedback information in a form that the user can understand.

[1004] "Means for acquiring video and audio for recording the operation status" refers to a device or method for recording the operation status with video and audio.

[1005] "Means for analyzing collected data to evaluate the efficiency, safety, and emotional state of the operator" refers to technology that analyzes acquired data to judge and evaluate the efficiency, safety, and emotional state of the operator.

[1006] The "means for generating feedback based on the evaluation results" is a method for providing useful advice and suggestions for improvement to users based on the evaluation results.

[1007] This invention relates to an AI assistance system for improving the quality of robot operation in factories. This system records the situation during operation, generates feedback from the data, and provides it to the operator to help improve their technical skills and manage their emotions.

[1008] Hardware Configuration

[1009] The system includes the following hardware:

[1010] Smart glasses (e.g., Google Glass Enterprise Edition): worn by operators, record the operation status in real time.

[1011] Cameras and microphones: Installed within the factory to capture additional video and audio data.

[1012] Server: For data processing and analysis (e.g., Google Cloud Platform).

[1013] Software Configuration

[1014] The software used is as follows:

[1015] Speech recognition engine (e.g. Google Cloud Speech-to-Text API): Converts extracted speech into text.

[1016] Natural language processing (NLP) techniques (e.g., Google Cloud Natural Language API): Analyze text data to evaluate phrasing, syntax, and argument structure.

[1017] Emotion recognition engine (e.g., Microsoft Azure Face API): Recognizes the operator's emotional state from video and audio data.

[1018] Data analysis programs (e.g., Python, TensorFlow): used to evaluate data and generate feedback.

[1019] Data processing and feedback generation

[1020] The server receives the video and audio data sent from the smart glasses and the camera, and then processes the data in the following steps:

[1021] 1. Audio extraction: The audio is separated from the video data and input into a speech recognition engine. This process converts the audio into text.

[1022] 2. Text analysis: The obtained text data is analyzed using natural language processing technology to evaluate the content of the statements.

[1023] 3. Voice analysis: Analyzes the characteristics of voice data and evaluates the tone of voice and speaking rate, thereby extracting how efficient and safe the operation is.

[1024] 4. Emotion Recognition: Analyze the operator's emotional state using video and audio data and incorporate the results into the evaluation.

[1025] Generating and Presenting Feedback

[1026] The server generates and provides feedback to the user (operator) based on the analysis results, including specific advice on operational efficiency, safety, and emotional state.

[1027] Specific examples

[1028] An operator wearing smart glasses starts recording while operating a robot. The recorded data is sent to a server, where a speech recognition engine converts the speech into text. Natural language processing technology then analyzes the text to evaluate the accuracy and efficiency of the operation. At the same time, an emotion recognition engine analyzes the operator's emotional state from the video data. Ultimately, based on the results of these analyses, feedback is generated, such as, "The order of operations is efficient, but you appear to be a little impatient. Please operate more slowly next time."

[1029] Prompt Sentence Examples

[1030] Voice recognition prompts

[1031] Use the speech recognition API to convert the following audio data to text:

[1032] <path to audio data>

[1033] Prompts for Natural Language Processing

[1034] Use the recognized text data to extract the following information:

[1035] Key Points

[1036] Utterance Syntax

[1037] Semantic analysis

[1038] Emotion recognition prompts

[1039] Detect operator emotions from video data:

[1040] <Video data path>

[1041] Prompts for generating feedback

[1042] Use the data you collect to generate feedback that includes the following metrics:

[1043] Evaluating operational efficiency

[1044] Safety evaluation

[1045] Emotion-based feedback

[1046] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1047] Step 1:

[1048] Video and audio acquisition

[1049] The device uses smart glasses or cameras and microphones in the factory to record video and audio of the operator's operations. This video and audio data is sent to the server in real time. The input is video and audio data, and the output is stream data transferred to the server.

[1050] Step 2:

[1051] Audio Extraction

[1052] The server extracts the audio portion from the transmitted video recording data. Specifically, it uses audio processing technology to separate the audio stream from the combined video and audio data. The input is the video and audio stream data, and the output is the extracted audio data.

[1053] Step 3:

[1054] Speech-to-text

[1055] The server converts the extracted audio data into text using a speech recognition engine (e.g., Google Cloud Speech-to-Text API). In this process, the speech recognition engine analyzes the audio waveform and generates a corresponding string using a language model. The input is the audio data, and the output is the generated text data.

[1056] Step 4:

[1057] Text Analysis

[1058] The server analyzes the generated text data using natural language processing technology (e.g., Google Cloud Natural Language API). This analysis includes tokenization, part-of-speech tagging, and semantic analysis, and evaluates the content of the speech. The input is text data, and the output is the analysis results (key points, sentence structure, flow of discussion, etc.).

[1059] Step 5:

[1060] Evaluation of audio data

[1061] The server analyzes the features of the voice data and evaluates the tone of voice and speaking speed. Here, acoustic features (e.g., pitch, intensity) are extracted and evaluation is performed based on that information. The input is the voice data, and the output is the voice features (voice pitch, speed, strength, etc.) and the evaluation results.

[1062] Step 6:

[1063] Emotion recognition

[1064] The server uses an emotion recognition engine (e.g., Microsoft Azure Face API) to analyze the operator's emotional state from the video and audio data. This includes facial expression analysis technology and voice tone analysis. The input is video and audio data, and the output is a classification result of the operator's emotional state.

[1065] Step 7:

[1066] Generate feedback

[1067] The server generates specific feedback based on the analysis results. This feedback includes evaluations and advice on operational efficiency, safety, and emotional state. The inputs are the text analysis results, speech analysis results, and emotion recognition results, and the generated feedback is obtained as the output.

[1068] Step 8:

[1069] Providing feedback

[1070] The terminal presents the generated feedback to the operator, using a web browser or mobile app to display the feedback in a visually understandable format. The input is the generated feedback, and the output is the visualized feedback that can be viewed by the user.

[1071] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1072] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1073] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1074] [Fourth embodiment]

[1075] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1076] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1077] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1078] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1079] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1080] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1081] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1082] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1083] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1084] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1085] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1086] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1087] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1088] The present invention relates to a system for improving the quality of meetings by utilizing AI technology. A specific embodiment of this system will be described.

[1089] 1. Recording meetings

[1090] The device captures video and audio as soon as the meeting starts. Specifically, the entire meeting is recorded using the camera and microphone of a laptop or smartphone, for example. The recorded data is uploaded to a server in real time.

[1091] 2. Audio Extraction

[1092] The server extracts the audio portion from the received recording. This extraction can be done using common audio processing techniques, including, for example, separating the audio stream from the combined video and audio of the meeting.

[1093] 3. Speech-to-text

[1094] The server then requests a speech recognition engine to convert the extracted voice data. Specifically, it calls a service such as the Google Cloud Speech-to-Text API to convert the voice into text data. As a result, the spoken content is obtained in text format.

[1095] 4. Text Analysis

[1096] The server analyzes the text data. Natural language processing (NLP) techniques are used for text analysis. For example, Python NLP libraries such as spaCy and BERT can be used to tokenize the text, tag parts of speech, and perform semantic analysis. This analysis makes it possible to evaluate the vocabulary, syntax, and argument structure of the speech.

[1097] 5. Evaluation of audio data

[1098] The server analyzes the audio data and evaluates the tone of voice and speaking rate. This includes analyzing acoustic features. For example, it uses Praat or an open-source audio analysis library to measure pitch (highness of voice) and intensity (power of voice) and makes an evaluation based on these.

[1099] 6. Generate feedback

[1100] The server generates feedback based on the results of text and speech analysis. The feedback is provided to each participant in the form of an evaluation result and recommendations for improvement. For example, it may include specific comments such as "Mr. Sato speaks too quickly" and suggestions for improvement such as "Try to speak a little more slowly next time."

[1101] 7. Providing Feedback

[1102] The device presents the generated feedback to the user. The feedback is displayed visually in a web browser or mobile app, allowing the user to intuitively identify problems and areas for improvement in their comments. For example, feedback could be displayed separately for each comment, and a visually easy-to-understand format could be used, using color coding or graphs.

[1103] Using this system, meeting participants can improve their communication skills based on specific feedback, which is expected to improve the overall quality of meetings and increase productivity across the organization.

[1104] The processing flow will be explained below.

[1105] Step 1:

[1106] The device captures video and audio as soon as the meeting starts. For example, the entire meeting can be recorded using the camera and microphone of a laptop or smartphone, and the data can be sent to a server in real time.

[1107] Step 2:

[1108] The server extracts the audio portion from the recorded data it receives. Specifically, it uses audio processing technology to separate the audio stream from the combined video and audio data.

[1109] Step 3:

[1110] The server inputs the extracted voice data into a speech recognition engine and converts the voice to text, for example, using the Google Cloud Speech-to-Text API.

[1111] Step 4:

[1112] The server analyzes the text data, using natural language processing (NLP) techniques to tokenize, tag parts of speech, and analyze the semantics of the text. This analysis evaluates the phrasing, syntax, and argument structure of the speech.

[1113] Step 5:

[1114] The server uses the audio data to evaluate the tone of voice and speaking rate. Specifically, it analyzes acoustic features to measure pitch (highness of voice) and intensity (power of voice).

[1115] Step 6:

[1116] The server generates feedback based on the results of text and speech analysis, for example, using a specific template to create feedback for each participant, including specific evaluations and suggestions for improvement.

[1117] Step 7:

[1118] The device then presents the generated feedback to the user, which is displayed visually via a smartphone, PC web browser, or dedicated app. The user can then review the displayed feedback and understand their areas for improvement.

[1119] Example 1

[1120] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1121] While conventional conference systems can record conference content and convert audio into text, they lack the functionality to evaluate the quality of the conference and provide specific feedback. As a result, users have no way to objectively evaluate their own comments or the way the conference proceeded, which hinders the improvement of their communication skills. The present invention aims to solve these problems and provide a system that improves the quality of conferences.

[1122] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1123] In this invention, the server includes means for recording the contents of the conference, means for uploading the recorded data to the server in real time, means for extracting voice from the recorded data, means for transmitting the extracted voice to a voice recognition engine and converting it into text data, means for analyzing the text data using natural language processing technology, means for analyzing the voice data using acoustic analysis technology and evaluating the tone of voice and speaking rate, means for generating feedback based on the results of the text analysis and the voice analysis, and means for presenting the generated feedback to the user. This allows conference participants to receive evaluations of their own remarks and voice and know specific areas for improvement, thereby enabling them to improve their communication skills.

[1124] A "meeting recording device" is a device or method that captures and stores video and audio of a meeting in digital form.

[1125] "Means for uploading recorded data to a server in real time" refers to the function of instantly transmitting video and audio data generated during the progress of a conference to a server.

[1126] "Means for extracting audio from recorded data" refers to a technology for separating and extracting the audio portion from recorded video data.

[1127] "Means for transmitting the extracted speech to a speech recognition engine and converting it into text data" refers to the process of transmitting the extracted speech data to speech recognition software and using that software to convert the speech into text information.

[1128] "Means for analyzing text data using natural language processing technology" refers to a technology that analyzes acquired text data using a natural language processing algorithm and evaluates the wording, syntax, and structure of the argument.

[1129] "Means for analyzing voice data using acoustic analysis technology to evaluate voice tone and speaking rate" refers to technology for analyzing voice data using an acoustic analysis tool and measuring pitch and intensity to evaluate voice tone and speaking rate.

[1130] The "means for generating feedback based on text analysis results and speech analysis results" refers to a process of combining the analysis results of text data and speech data to create feedback that includes specific evaluations and improvement measures.

[1131] The "means for presenting the generated feedback to the user" is a device or method that visually or audibly conveys the generated feedback to the user.

[1132] The present invention relates to a system for improving the quality of meetings by utilizing AI technology. Specific embodiments of this system are described in detail below.

[1133] First, at the start of a meeting, the user launches the meeting application and presses the "Start Meeting" button. Devices (e.g., laptops, smartphones) that detect the start of the meeting begin capturing video and audio using their built-in cameras and microphones. This allows the entire meeting to be recorded.

[1134] The captured video and audio data is uploaded to a server in real time using HTTP streaming or WebRTC technology, for example. The server processes the received recording data and separates the video and audio data. Specifically, audio processing technology such as FFmpeg is used to extract the audio stream from the recording data.

[1135] The extracted voice data is then sent by the server to the Google Cloud Speech-to-Text API, where it is converted into text data. This step allows the voice utterance to be obtained in text format.

[1136] The captured text data is then analyzed by the server using natural language processing (NLP) techniques, such as tokenizing, part-of-speech tagging, and semantic analysis using Python NLP libraries (e.g., spaCy, BERT), to evaluate the vocabulary, syntax, and argument structure of the speech.

[1137] Additionally, the voice data is analyzed using acoustic analysis technology. The server analyzes the voice data using an acoustic analysis library (e.g., Librosa) to measure pitch and intensity. This evaluation evaluates the tone of voice and speaking rate.

[1138] The server generates feedback based on these analysis results. The generated feedback includes specific suggestions for improvement and evaluations, such as "You speak too quickly" or "Try to speak more slowly next time." Finally, the generated feedback is presented to the user via the device. The feedback is displayed visually in a web browser or mobile app, and is presented in an easily understandable format using color-coded graphs and icons.

[1139] For example, when a user holds a meeting using video conferencing software, the laptop's camera and microphone capture video and audio and upload them to a server in real time. The server then extracts the audio using FFmpeg and converts it to text using the Google Cloud Speech-to-Text API. The text is then analyzed using spaCy, and the audio is evaluated using Librosa. Finally, feedback generated based on the analysis results is displayed in the web browser for the user to review.

[1140] An example of a prompt might be, "Please use the Google Cloud Speech-to-Text API to convert the following audio data into text and generate meeting feedback based on the results." By using this system, meeting participants can receive evaluations of their own remarks and voices, learn specific areas for improvement, and improve their communication skills.

[1141] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1142] Step 1:

[1143] The user launches the conference application and presses the "Start Conference" button. This causes the device to detect the start of the conference. The input is the user's operation action, and the output is the start of the conference recording process.

[1144] Step 2:

[1145] The device will begin capturing video and audio using the built-in camera and microphone. The captured data will be saved as video and audio files. The input is real-time data from the built-in camera and microphone, and the output is locally saved video and audio data.

[1146] Step 3:

[1147] The device uploads the captured video and audio data to the server in real time. This process uses HTTP streaming or WebRTC technology. The input is the locally stored video and audio data, and the output is the data uploaded to the server.

[1148] Step 4:

[1149] The server processes the received recording data and separates the video and audio data using audio processing technology such as FFmpeg. The input is the uploaded recording data, and the output is the separated video and audio data.

[1150] Step 5:

[1151] The server sends the extracted speech data to a speech recognition engine and converts it into text data. This step uses the Google Cloud Speech-to-Text API. The extracted speech data is the input, and the text data is the output.

[1152] Step 6:

[1153] The server analyzes the acquired text data using natural language processing (NLP) techniques. Specifically, it uses Python NLP libraries (e.g., spaCy, BERT) to perform tokenization, part-of-speech tagging, and semantic analysis. The converted text data is input, and the analysis results are obtained as output.

[1154] Step 7:

[1155] The server analyzes the voice data using acoustic analysis technology to evaluate the tone of voice and speaking rate. Specifically, it uses the Librosa library to measure pitch and intensity. The input is the voice data, and the output is the result of the acoustic analysis.

[1156] Step 8:

[1157] The server generates feedback based on the results of text analysis and speech analysis. The feedback includes specific evaluations and suggestions for improvement. The inputs are the results of text analysis and speech analysis, and the generated feedback is obtained as the output.

[1158] Step 9:

[1159] The device presents the generated feedback to the user, which is then visually displayed in a web browser or mobile app. The input is the generated feedback, and the output is the feedback presented to the user.

[1160] (Application example 1)

[1161] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1162] Conventional conference systems lack specific feedback to improve the quality of meetings, making it difficult for participants to improve their own comments and attitudes. Furthermore, in certain environments, such as project review meetings in factories, recording and analysis methods and tools are often not properly installed, hindering the effective running of meetings. The present invention aims to provide a system that solves these problems and improves the quality of meetings.

[1163] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1164] In this invention, the server includes a means for recording the contents of the conference, a means for extracting audio from the recorded data, and a means for converting the extracted audio into text, thereby enabling the server to accurately record the overall picture of the conference and provide specific feedback based on the analysis results.

[1165] "Means for recording the contents of a conference" refers to a device or program that captures and records video and audio as soon as the conference begins.

[1166] The "means for extracting audio from recorded data" refers to a processing device or program for separating the audio portion from the recorded conference data and extracting it as audio data.

[1167] "Means for converting extracted speech into text" refers to a program or service that converts extracted speech data into text format using speech recognition technology.

[1168] A "means for analyzing text to evaluate language, syntax, and argument structure" is a program that uses natural language processing techniques on text data to evaluate and analyze the content of the text.

[1169] The "means for analyzing voice data to evaluate voice tone and speaking rate" refers to a device or program that uses voice data to analyze acoustic features such as voice pitch and speaking rate and make an evaluation.

[1170] The "means for generating feedback based on the analysis results" is a program that provides participants with specific suggestions for improvement and evaluations based on the results of text and voice analysis.

[1171] The "means for presenting the generated feedback to the user" refers to a device or program that visually displays the generated feedback so that the user can easily confirm it.

[1172] "Means to be installed on factory robots for project review meetings in a factory work environment" refers to programs and devices that are introduced into robots in factories for the purpose of recording meetings and analyzing data within the factory.

[1173] The "means for uploading recorded audio and video data to a server" refers to a communication device or program for transmitting recorded audio and video data to a server and performing analysis processing.

[1174] The "means for providing the generated feedback to participants within the factory" refers to a device or program that notifies each member who participates in the meeting within the factory of feedback based on the analysis results and suggests areas for improvement.

[1175] The present invention relates to a system for improving the quality of project review meetings in a factory. Specific embodiments of the system will be described below.

[1176] System Overview

[1177] This system is used in a factory work environment and is intended for project review meetings. It is composed mainly of a program installed on a factory robot, which records meetings, analyzes data, and generates and provides feedback. Each process is executed by a server, and feedback is provided to participants via their terminals.

[1178] Meeting Recording

[1179] First, a camera and microphone mounted on a factory robot are used to record the contents of the meeting as soon as it starts, and the recorded data is uploaded to a server in real time as video and audio.

[1180] Audio extraction and analysis

[1181] The server extracts the audio portion from the received video recordings. This extraction uses common audio processing techniques, such as separating the audio stream from the combined video and audio data. This process is performed using a programming language such as Python.

[1182] Speech-to-text

[1183] The server converts the extracted voice data into text using the Google Cloud Speech-to-Text API. This API is called and speech recognition technology is used to convert the voice into text data. As a result, the voice utterances are obtained in text format.

[1184] Text analytics

[1185] The server then analyzes the acquired text data using natural language processing techniques. Specifically, it uses the Python NLP libraries spaCy and BERT to tokenize the text, tag parts of speech, and perform semantic analysis. This analysis allows it to evaluate the vocabulary, syntax, and structure of the discussions of participants' comments.

[1186] Audio analysis

[1187] Similarly, the server analyzes the audio data and performs acoustic feature analysis to assess the tone of voice and speaking rate, using Praat or other open-source speech analysis libraries. This analysis measures the pitch and intensity of the speech and evaluates the quality of the speech.

[1188] Generating and Presenting Feedback

[1189] The server generates specific feedback based on the text and speech analysis results. The feedback is generated for each participant in a format that includes evaluation results and recommended improvements. The generated feedback is presented to the user visually on a web browser or mobile app.

[1190] Examples and prompts

[1191] For example, in a project review meeting held in a factory, the system analyzes the speaking style and content of each participant and provides feedback such as, "Yamada-san, you use too much technical jargon, so try using more understandable words." If the tone of your voice is too high, the system also provides specific examples such as, "By lowering your voice a little, you can give the impression of being more calm."

[1192] Prompt Sentence Examples

[1193] Analysis of statements made during factory meetings:

[1194] 1. Upload the recorded audio data to the following URL:

[1195] 2. Please analyze the uploaded audio data and generate feedback in the following format:

[1196] 3. Provide specific suggestions for improving each speaker's delivery.

[1197] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1198] Step 1:

[1199] Recording of meeting content

[1200] As soon as the meeting starts, the device uses its camera and microphone to capture video and audio, recording in real time. The input is the video and audio of the meeting, and the output is recorded data. Specifically, the device's camera captures images of the meeting participants, and the microphone picks up audio. The recorded data is temporarily stored in the device's memory.

[1201] Step 2:

[1202] Uploading recording data

[1203] The device uploads the recorded data to the server. The input is the recorded data, and the output is the data uploaded to the server. Specifically, after the device finishes recording, it transfers the data to the server via Wi-Fi or Ethernet. For security reasons, the SSL / TLS protocol is used.

[1204] Step 3:

[1205] Audio Extraction

[1206] The server extracts the audio portion from the uploaded video data. The input is the video data, and the output is the audio data. Specifically, an audio processing module on the server separates the audio stream from the combined video and audio data and saves it as audio data. Libraries such as FFmpeg are used.

[1207] Step 4:

[1208] Speech-to-text

[1209] The server converts the extracted voice data into text using the Google Cloud Speech-to-Text API. The input is voice data, and the output is text data. Specifically, the voice data is sent to the API, which uses speech recognition technology to convert it into text, and the result is returned to the server.

[1210] Step 5:

[1211] Text Analysis

[1212] The server analyzes the acquired text data using natural language processing technology. The input is the text data, and the output is the analysis results. Specifically, it uses Python NLP libraries (e.g., spaCy and BERT) to perform tokenization, part-of-speech tagging, and semantic analysis. This analysis evaluates the text's vocabulary, syntax, and argument structure.

[1213] Step 6:

[1214] Analysis of audio data

[1215] The server analyzes the audio data and evaluates the tone of voice and speaking rate. The input is the audio data, and the output is the analysis results. Specifically, it uses Praat and other open source audio analysis libraries to measure pitch (voice height) and intensity (voice strength), and makes an evaluation based on the results.

[1216] Step 7:

[1217] Generate feedback

[1218] The server generates feedback based on the text analysis results and speech analysis results. The input is the text analysis results and speech analysis results, and the output is the generated feedback. Specifically, feedback is generated for each participant that includes evaluation results and recommended improvements. For example, it includes a comment such as "A-san speaks too quickly" and a suggestion for improvement such as "Try to speak more slowly next time."

[1219] Step 8:

[1220] Providing feedback

[1221] The device presents the generated feedback to the user via a web browser or mobile app. The input is the generated feedback, and the output is what is presented to the user. Specifically, the feedback is displayed separately for each comment, and is provided in a visually easy-to-understand format using color coding and graphs. For example, the feedback is displayed in a dashboard format, allowing the user to intuitively identify problems and areas for improvement in their own comments.

[1222] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1223] The present invention relates to a system for improving the quality of meetings using AI technology. This system records the contents of a meeting, extracts audio from the recorded data, converts the audio to text, analyzes the text, analyzes the audio data, and generates feedback. Furthermore, by combining it with an emotion engine, it is possible to detect the user's emotions and provide feedback based on those emotions. Specific embodiments are described below.

[1224] 1. Recording meetings

[1225] The device captures video and audio as soon as the meeting starts. For example, the entire meeting can be recorded using the camera and microphone of a laptop or smartphone. This recorded data is then sent to a server in real time.

[1226] 2. Audio Extraction

[1227] The server extracts the audio portion from the recorded data it receives. Specifically, it uses audio processing technology to separate the audio stream from the combined video and audio data.

[1228] 3. Speech-to-text

[1229] The server inputs the extracted voice data into a speech recognition engine and converts the speech into text. For example, it uses a speech recognition service such as Google Cloud Speech-to-Text API. As a result, the spoken content is obtained in text format.

[1230] 4. Text Analysis

[1231] The server analyzes the text data, specifically using natural language processing (NLP) techniques to tokenize, tag parts of speech, and perform semantic analysis of the text, evaluating the phrasing, syntax, and argument structure of what is being said.

[1232] 5. Evaluation of audio data

[1233] The server uses the audio data to evaluate the tone of voice and speaking speed, for example by analyzing acoustic features and measuring pitch (voice height) and intensity (voice strength).

[1234] 6. Emotion Recognition by Emotion Engine

[1235] The server analyzes the audio and video data to recognize the user's emotions. For example, it uses voice tone and facial expression analysis technology to detect whether the user is happy, angry, or sad. To do this, it uses a machine learning model to classify emotions.

[1236] 7. Generate feedback

[1237] The server generates feedback based on the results of text analysis, speech analysis, and emotion recognition. The feedback is provided to each participant in the form of an evaluation result and recommendations for improvement. For example, specific feedback such as, "Mr. Sato, you spoke too quickly. You also seemed a little nervous. Try to speak more relaxedly next time."

[1238] 8. Providing Feedback

[1239] The device presents the generated feedback to the user. The feedback is displayed visually in a web browser or mobile app. The user can review the feedback and understand their areas for improvement. For example, the feedback is displayed separately for each utterance and presented in an easy-to-understand format using visualization techniques such as color coding and graphs.

[1240] In this way, the present invention allows meeting participants to improve their communication skills based on specific feedback. Furthermore, by adding emotion recognition functionality, feedback can be provided that takes psychological aspects into account, further improving the overall quality of meetings and increasing productivity across the organization.

[1241] The processing flow will be explained below.

[1242] Step 1:

[1243] The device captures video and audio as soon as the meeting starts. Specifically, it uses the laptop's camera and microphone to record the entire meeting. The recorded data is sent to the server in real time.

[1244] Step 2:

[1245] The server extracts the audio portion from the recorded data it receives. Audio processing technology is used to separate the audio stream from the combined video and audio data. This step also involves format conversion of the audio file and noise filtering.

[1246] Step 3:

[1247] The server sends the extracted voice data to a speech recognition engine and converts the voice into text, for example, using the Google Cloud Speech-to-Text API, which then stores the converted text in a database.

[1248] Step 4:

[1249] The server analyzes the text data, using natural language processing (NLP) techniques to perform tokenization, part-of-speech tagging, and semantic analysis. This analysis evaluates the phrasing, syntax, and argument structure of what is being said.

[1250] Step 5:

[1251] The server analyzes the voice data to evaluate the tone of voice and speaking rate. For example, it uses tools such as Praat to measure pitch, intensity, and speaking rate. Based on this, it generates a voice evaluation result.

[1252] Step 6:

[1253] The server uses an emotion engine to recognize the user's emotions from audio and video data. It analyzes the tone of the voice and facial expressions to classify the emotional state. It uses a machine learning model to determine whether the user is happy, angry, or sad.

[1254] Step 7:

[1255] The server generates feedback based on the results of text analysis, speech evaluation, and emotion recognition. For example, it generates specific feedback such as, "Mr. Sato, you spoke too quickly. You also seemed a little nervous. Please try to speak more relaxedly next time."

[1256] Step 8:

[1257] The device presents the generated feedback to the user. The feedback is displayed visually through a web browser or mobile app. The user can review the feedback and understand their areas for improvement. For example, the feedback could be displayed in a color-coded format for each comment, and could be formatted in a way that makes it easy to understand visually using graphs and charts.

[1258] Example 2

[1259] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1260] Conventional meeting support systems did not accurately record or analyze what was said during meetings, which resulted in vague feedback and made it difficult to provide specific suggestions for improvement. Furthermore, there was no way to provide feedback that took into account participants' emotions and psychological states, resulting in a lack of specific guidance for improving the quality of communication.

[1261] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for recording the contents of the conference, means for extracting audio from the recorded data, means for converting the extracted audio into text, means for analyzing the text to evaluate the wording, syntax, and structure of the discussion, means for analyzing the audio data to evaluate the tone of voice and speaking rate, means for analyzing the audio data and video data to recognize the user's emotions, means for generating feedback based on the analysis results, and means for presenting the generated feedback to the user. This makes it possible to accurately record and analyze the content of comments made in the conference and provide specific feedback that takes into account the emotions of the participants.

[1262] "Meeting content" refers to the entire information, materials, and presentations spoken or shared by participants at the meeting.

[1263] "Recording Data" means the digital files of video and audio captured during a meeting.

[1264] "Means for extracting audio" refers to a function that performs the process of separating and extracting the audio portion from the recorded data.

[1265] "Means for converting speech to text" refers to technology, particularly a speech recognition engine, that automatically converts extracted speech data into text data.

[1266] "Means of analyzing text to evaluate language, syntax, and argument structure" refers to the application of natural language processing techniques to converted text data to analyze word choice, grammar, and argument flow.

[1267] "Means for analyzing voice data to evaluate voice tone and speaking rate" refers to technology that uses acoustic analysis technology to measure the features of voice data and evaluate the pitch, strength, and speaking rate of the voice.

[1268] "Means for recognizing a user's emotions by analyzing audio and video data" refers to a process that uses voice tone and facial expression recognition technology to identify a user's emotional state.

[1269] "Means for generating feedback" refers to a function that automatically creates feedback including specific advice and areas for improvement for the user based on the results of text analysis, speech analysis, and emotion recognition.

[1270] "Means for presenting the generated feedback to the user" refers to an interface, such as a web browser or mobile app, that displays the generated feedback so that the user can review it.

[1271] The present invention relates to a system for improving the quality of meetings by utilizing AI technology. This system records the contents of meetings, extracts audio from the recorded data, converts the audio to text, analyzes the text, analyzes the audio data, and generates feedback. Furthermore, by combining it with an emotion engine, it is possible to detect the user's emotions and provide feedback based on those emotions. Specific embodiments for implementing this system are described below.

[1272] Hardware and software used

[1273] The device uses the camera and microphone of a laptop or smartphone. For example, a typical laptop or smartphone has a built-in camera and microphone for capturing video and audio.

[1274] The data transfer and processing is carried out by servers using the following software and technologies:

[1275] Audio processing library: FFmpeg

[1276] Speech recognition service: Google Cloud Speech-to-Text API

[1277] Natural language processing library: SpaCy

[1278] Acoustic analysis library: Librosa

[1279] Machine learning frameworks: TensorFlow, PyTorch

[1280] Browser display: D3.js

[1281] System processing flow

[1282] 1. Meeting recording: The device captures video and audio as soon as the meeting starts. For example, you can use Zoom or other video conferencing software to record the meeting. The recording data is sent to the server in real time.

[1283] 2. Audio Extraction: Extract the audio portion from the recording data received by the server. Use FFmpeg to separate the audio stream from the combined video and audio data.

[1284] 3. Speech-to-text conversion: The server sends the extracted voice data to the Google Cloud Speech-to-Text API to convert the voice to text, thereby obtaining the content of the meeting in text format.

[1285] 4. Text Analysis: The server analyzes the acquired text data using natural language processing (NLP) techniques, including tokenization, part-of-speech tagging, and semantic analysis using the SpaCy library.

[1286] 5. Evaluation of speech data: The server uses the Librosa library to analyze the acoustic features of the speech data, measure pitch and intensity from the speech data, and evaluate the speaking rate.

[1287] 6. Emotion recognition using an emotion engine: The server analyzes audio and video data to recognize the user's emotions. Using TensorFlow and other tools, it analyzes the tone of the voice and recognizes facial expressions to classify the user's emotions.

[1288] 7. Feedback Generation: The server generates feedback based on the results of text analysis, speech analysis, and emotion recognition. It uses natural language generation technology to create feedback that includes specific improvements.

[1289] 8. Present feedback: The device presents the feedback generated by the server to the user. This is visually displayed in the web browser. D3.js is used to present the feedback to the user using graphs and color coding to make it easy to understand.

[1290] Specific examples and prompts

[1291] As a concrete example, the following prompt sentence is input to the generative AI model:

[1292] "Extract audio from meeting recordings and convert it to text. Analyze the text content and perform emotion recognition to generate feedback."

[1293] This prompt sentence executes each processing step and generates specific feedback, for example, "Mr. Suzuki, you speak too quickly. Please try to speak a little more slowly."

[1294] In this way, the present invention can help meeting participants improve their communication skills based on specific feedback, potentially improving the overall quality of meetings. Furthermore, by adding emotion recognition functionality, feedback can be provided that takes into account psychological aspects during meetings, which is expected to improve productivity across the organization.

[1295] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1296] Step 1:

[1297] Recording a meeting

[1298] Input: When a meeting starts, the device captures the video and audio of the meeting.

[1299] Data processing: Recorded data that combines video and audio data is generated.

[1300] Output: The generated recording data is sent to the server in real time.

[1301] Specific operation: The device starts recording using video conferencing software (e.g., Zoom or Microsoft Teams) and transfers the data to the server.

[1302] Step 2:

[1303] Audio Extraction

[1304] Input: Recording data received by the server.

[1305] Data processing: Use FFmpeg to separate the audio stream from the combined video and audio data.

[1306] Output: Separated audio data (e.g., .wav format) is generated.

[1307] Specific operation: The server executes the FFmpeg command to extract and save the audio file from the recorded data.

[1308] Step 3:

[1309] Speech-to-text

[1310] Input: Audio data extracted by the server.

[1311] Data processing: Send the audio data to the Google Cloud Speech-to-Text API for speech recognition.

[1312] Output: The spoken words are captured in text format.

[1313] Specific operation: The server uploads the audio file to the Google Cloud Speech-to-Text API and stores the returned text data.

[1314] Step 4:

[1315] Text Analysis

[1316] Input: Speech-to-text data.

[1317] Data processing: Use the SpaCy library to tokenize the text, tag it with parts of speech, and perform semantic analysis.

[1318] Output: The analyzed text data is generated and the evaluation results are obtained.

[1319] What it does: The server uses the SpaCy library to analyze the text data and evaluate the phrasing, syntax, and argument structure of the comments.

[1320] Step 5:

[1321] Evaluation of audio data

[1322] Input: Extracted audio data.

[1323] Data processing: The Librosa library is used to analyze acoustic features and evaluate voice tone and speaking rate.

[1324] Output: Generates an assessment of tone of voice and speaking rate.

[1325] How it works: The server processes the audio data using Librosa, measures pitch and intensity, and evaluates the speaker's vocal characteristics.

[1326] Step 6:

[1327] Emotion recognition by emotion engine

[1328] Input: Audio and video data.

[1329] Data processing: Use TensorFlow and PyTorch to perform voice tone analysis and facial expression recognition to classify emotions.

[1330] Output: The user's emotional state (e.g., joy, anger, sadness, or happiness) is recognized.

[1331] How it works: The server analyzes the audio data and applies an emotion classification model to detect the user's emotional state, for example, whether they are angry, sad, happy, etc.

[1332] Step 7:

[1333] Generate feedback

[1334] Input: Text analysis results, speech analysis results, and emotion recognition results.

[1335] Data processing: The results of each analysis are integrated and feedback is automatically generated using natural language generation technology.

[1336] Output: A specific feedback text is generated for the user.

[1337] Specific behavior: The server uses a natural language generation engine (e.g., GPT-3) to generate feedback based on the analysis results, including specific suggestions for improvement. For example, it generates feedback like, "Mr. Sato, you spoke too quickly. You also seemed a little nervous. Please try to speak more relaxedly next time."

[1338] Step 8:

[1339] Providing feedback

[1340] Input: The generated feedback text.

[1341] Data processing: Converting JSON-formatted feedback data into a visually understandable format.

[1342] Output: Visual feedback that the user can see.

[1343] What it does: The device displays feedback in a web browser or mobile app, and uses D3.js to visually present the feedback as graphs and color-coded text.

[1344] (Application example 2)

[1345] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1346] When operating robots in factories, there is a need to evaluate the technical quality and emotional state of operators and provide specific feedback to improve operational efficiency and safety. However, with current systems, it is difficult to evaluate the emotional state and quality of operators' operations in real time and generate and provide feedback based on that evaluation.

[1347] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1348] In this invention, the server includes means for recording the contents of the conference, means for extracting audio from the recorded data, means for converting the extracted audio into text, means for analyzing the text to evaluate the vocabulary, syntax, and structure of the discussion, means for analyzing the audio data to evaluate the tone of voice and speaking rate, means for generating feedback based on the analysis results, means for presenting the generated feedback to the user, means for acquiring video and audio to record the operation status, means for analyzing the collected data to evaluate the efficiency and safety of the operation and the emotional state of the operator, and means for generating feedback based on the evaluation results, thereby enabling improvement of the operator's skills, ensuring safety, and providing appropriate guidance based on the emotional state.

[1349] "Means for recording meeting content" refers to devices and methods for recording video and audio of meetings.

[1350] "Means for extracting audio from recorded data" refers to a technology for extracting only the audio portion from recorded video and audio data.

[1351] The "means for converting extracted speech into text" is a speech recognition technology that converts speech data into character data.

[1352] "Methods for analyzing text to evaluate language, syntax, and argument structure" refers to techniques for analyzing text data and judging and evaluating its linguistic expression, sentence structure, and argument flow.

[1353] "Means for analyzing voice data to evaluate voice tone and speaking rate" refers to technology that analyzes and evaluates voice characteristics such as pitch, strength, and speaking rate.

[1354] "Means for generating feedback based on the analysis results" refers to a method for providing users with advice and suggestions for improvement based on the analyzed data.

[1355] "Means for presenting the generated feedback to the user" refers to technology for displaying and providing the generated feedback information in a form that the user can understand.

[1356] "Means for acquiring video and audio for recording the operation status" refers to a device or method for recording the operation status with video and audio.

[1357] "Means for analyzing collected data to evaluate the efficiency, safety, and emotional state of the operator" refers to technology that analyzes acquired data to judge and evaluate the efficiency, safety, and emotional state of the operator.

[1358] The "means for generating feedback based on the evaluation results" is a method for providing useful advice and suggestions for improvement to users based on the evaluation results.

[1359] This invention relates to an AI assistance system for improving the quality of robot operation in factories. This system records the situation during operation, generates feedback from the data, and provides it to the operator to help improve their technical skills and manage their emotions.

[1360] Hardware Configuration

[1361] The system includes the following hardware:

[1362] Smart glasses (e.g., Google Glass Enterprise Edition): worn by operators, record the operation status in real time.

[1363] Cameras and microphones: Installed within the factory to capture additional video and audio data.

[1364] Server: For data processing and analysis (e.g., Google Cloud Platform).

[1365] Software Configuration

[1366] The software used is as follows:

[1367] Speech recognition engine (e.g. Google Cloud Speech-to-Text API): Converts extracted speech into text.

[1368] Natural language processing (NLP) techniques (e.g., Google Cloud Natural Language API): Analyze text data to evaluate phrasing, syntax, and argument structure.

[1369] Emotion recognition engine (e.g., Microsoft Azure Face API): Recognizes the operator's emotional state from video and audio data.

[1370] Data analysis programs (e.g., Python, TensorFlow): used to evaluate data and generate feedback.

[1371] Data processing and feedback generation

[1372] The server receives the video and audio data sent from the smart glasses and the camera, and then processes the data in the following steps:

[1373] 1. Audio extraction: The audio is separated from the video data and input into a speech recognition engine. This process converts the audio into text.

[1374] 2. Text analysis: The obtained text data is analyzed using natural language processing technology to evaluate the content of the statements.

[1375] 3. Voice analysis: Analyzes the characteristics of voice data and evaluates the tone of voice and speaking rate, thereby extracting how efficient and safe the operation is.

[1376] 4. Emotion Recognition: Analyze the operator's emotional state using video and audio data and incorporate the results into the evaluation.

[1377] Generating and Presenting Feedback

[1378] The server generates and provides feedback to the user (operator) based on the analysis results, including specific advice on operational efficiency, safety, and emotional state.

[1379] Specific examples

[1380] An operator wearing smart glasses starts recording while operating a robot. The recorded data is sent to a server, where a speech recognition engine converts the speech into text. Natural language processing technology then analyzes the text to evaluate the accuracy and efficiency of the operation. At the same time, an emotion recognition engine analyzes the operator's emotional state from the video data. Ultimately, based on the results of these analyses, feedback is generated, such as, "The order of operations is efficient, but you appear to be a little impatient. Please operate more slowly next time."

[1381] Prompt Sentence Examples

[1382] Voice recognition prompts

[1383] Use the speech recognition API to convert the following audio data to text:

[1384] <path to audio data>

[1385] Prompts for Natural Language Processing

[1386] Use the recognized text data to extract the following information:

[1387] Key Points

[1388] Utterance Syntax

[1389] Semantic analysis

[1390] Emotion recognition prompts

[1391] Detect operator emotions from video data:

[1392] <Video data path>

[1393] Prompts for generating feedback

[1394] Use the data you collect to generate feedback that includes the following metrics:

[1395] Evaluating operational efficiency

[1396] Safety evaluation

[1397] Emotion-based feedback

[1398] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1399] Step 1:

[1400] Video and audio acquisition

[1401] The device uses smart glasses or cameras and microphones in the factory to record video and audio of the operator's operations. This video and audio data is sent to the server in real time. The input is video and audio data, and the output is stream data transferred to the server.

[1402] Step 2:

[1403] Audio Extraction

[1404] The server extracts the audio portion from the transmitted video recording data. Specifically, it uses audio processing technology to separate the audio stream from the combined video and audio data. The input is the video and audio stream data, and the output is the extracted audio data.

[1405] Step 3:

[1406] Speech-to-text

[1407] The server converts the extracted audio data into text using a speech recognition engine (e.g., Google Cloud Speech-to-Text API). In this process, the speech recognition engine analyzes the audio waveform and generates a corresponding string using a language model. The input is the audio data, and the output is the generated text data.

[1408] Step 4:

[1409] Text Analysis

[1410] The server analyzes the generated text data using natural language processing technology (e.g., Google Cloud Natural Language API). This analysis includes tokenization, part-of-speech tagging, and semantic analysis, and evaluates the content of the speech. The input is text data, and the output is the analysis results (key points, sentence structure, flow of discussion, etc.).

[1411] Step 5:

[1412] Evaluation of audio data

[1413] The server analyzes the features of the voice data and evaluates the tone of voice and speaking speed. Here, acoustic features (e.g., pitch, intensity) are extracted and evaluation is performed based on that information. The input is the voice data, and the output is the voice features (voice pitch, speed, strength, etc.) and the evaluation results.

[1414] Step 6:

[1415] Emotion recognition

[1416] The server uses an emotion recognition engine (e.g., Microsoft Azure Face API) to analyze the operator's emotional state from the video and audio data. This includes facial expression analysis technology and voice tone analysis. The input is video and audio data, and the output is a classification result of the operator's emotional state.

[1417] Step 7:

[1418] Generate feedback

[1419] The server generates specific feedback based on the analysis results. This feedback includes evaluations and advice on operational efficiency, safety, and emotional state. The inputs are the text analysis results, speech analysis results, and emotion recognition results, and the generated feedback is obtained as the output.

[1420] Step 8:

[1421] Providing feedback

[1422] The terminal presents the generated feedback to the operator, using a web browser or mobile app to display the feedback in a visually understandable format. The input is the generated feedback, and the output is the visualized feedback that can be viewed by the user.

[1423] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1424] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1425] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1426] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1427] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1428] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1429] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1430] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1431] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1432] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1433] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1434] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1435] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1436] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1437] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1438] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1439] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1440] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1441] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1442] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1443] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1444] The following is further disclosed regarding the above embodiment.

[1445] (Claim 1)

[1446] A means of recording the meeting content;

[1447] A means for extracting audio from the recording data;

[1448] means for converting the extracted speech into text;

[1449] a means of analyzing text to evaluate language, syntax, and argument structure;

[1450] means for analyzing the audio data to assess tone of voice and speaking rate;

[1451] means for generating feedback based on the analysis results;

[1452] The system includes a means for presenting the generated feedback to a user.

[1453] (Claim 2)

[1454] 10. The system of claim 1, wherein the means for analyzing the text uses natural language processing techniques.

[1455] (Claim 3)

[1456] 2. The system of claim 1, wherein the means for analyzing the speech data uses acoustic features to evaluate tone of voice and speaking rate.

[1457] "Example 1"

[1458] (Claim 1)

[1459] A means of recording the meeting content;

[1460] A means for uploading recorded data to a server in real time;

[1461] A means for extracting audio from the recording data;

[1462] means for transmitting the extracted speech to a speech recognition engine and converting it into text data;

[1463] A means for analyzing text data using natural language processing technology;

[1464] means for analyzing the speech data using acoustic analysis techniques to assess tone of voice and speaking rate;

[1465] means for generating feedback based on the text analysis results and the speech analysis results;

[1466] The system includes a means for presenting the generated feedback to a user.

[1467] (Claim 2)

[1468] 2. The system according to claim 1, wherein the means for analyzing the text data using natural language processing techniques performs tokenization, part-of-speech tagging, and semantic analysis.

[1469] (Claim 3)

[1470] 2. The system of claim 1, wherein the speech data is analyzed using acoustic analysis techniques, and the means for assessing tone of voice and speaking rate measures pitch and intensity using acoustic features.

[1471] "Application Example 1"

[1472] (Claim 1)

[1473] A means of recording the meeting content;

[1474] A means for extracting audio from the recording data;

[1475] means for converting the extracted speech into text;

[1476] a means of analyzing text to evaluate language, syntax, and argument structure;

[1477] means for analyzing the audio data to assess tone of voice and speaking rate;

[1478] means for generating feedback based on the analysis results;

[1479] means for presenting the generated feedback to the user;

[1480] a means for project review meetings in a factory work environment and installed on the factory robot;

[1481] A means for uploading recorded audio and video data to a server;

[1482] a means of providing the generated feedback to participants within the factory;

[1483] A system including:

[1484] (Claim 2)

[1485] 10. The system of claim 1, wherein the means for analyzing the text uses natural language processing techniques.

[1486] (Claim 3)

[1487] 2. The system of claim 1, wherein the means for analyzing the speech data uses acoustic features to evaluate tone of voice and speaking rate.

[1488] "Example 2: Combining Emotion Engines"

[1489] (Claim 1)

[1490] A means of recording the meeting content;

[1491] A means for extracting audio from the recording data;

[1492] means for converting the extracted speech into text;

[1493] a means of analyzing text to evaluate language, syntax, and argument structure;

[1494] means for analyzing the audio data to assess tone of voice and speaking rate;

[1495] means for analyzing audio data and video data to recognize user emotions;

[1496] means for generating feedback based on the analysis results;

[1497] The system includes a means for presenting the generated feedback to a user.

[1498] (Claim 2)

[1499] 10. The system of claim 1, wherein the means for analyzing the text uses natural language processing techniques.

[1500] (Claim 3)

[1501] 2. The system of claim 1, wherein the means for analyzing the speech data uses acoustic features to evaluate tone of voice and speaking rate.

[1502] "Application example 2 when combining emotion engines"

[1503] (Claim 1)

[1504] A means of recording the meeting content;

[1505] A means for extracting audio from the recording data;

[1506] means for converting the extracted speech into text;

[1507] a means of analyzing text to evaluate language, syntax, and argument structure;

[1508] means for analyzing the audio data to assess tone of voice and speaking rate;

[1509] means for generating feedback based on the analysis results;

[1510] means for presenting the generated feedback to the user;

[1511] A means for acquiring video and audio for recording the operation status;

[1512] a means of analyzing the collected data to assess the efficiency, safety, and emotional state of the operation;

[1513] The system includes a means for generating feedback based on the evaluation results.

[1514] (Claim 2)

[1515] 10. The system of claim 1, wherein the means for analyzing the text uses natural language processing techniques.

[1516] (Claim 3)

[1517] 2. The system of claim 1, wherein the means for analyzing the speech data uses acoustic features to evaluate tone of voice and speaking rate. [Explanation of symbols]

[1518] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. A means of recording the meeting content; A means for extracting audio from the recording data; means for converting the extracted speech into text; a means of analyzing text to evaluate language, syntax, and argument structure; means for analyzing the audio data to assess tone of voice and speaking rate; means for generating feedback based on the analysis results; The system includes a means for presenting the generated feedback to a user.

2. 10. The system of claim 1, wherein the means for analyzing the text uses natural language processing techniques.

3. 2. The system of claim 1, wherein the means for analyzing the speech data uses acoustic features to evaluate tone of voice and speaking rate.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A