System

A system that collects, analyzes, and translates meeting voices, extracts key information, and manages time effectively addresses inefficiencies in modern meetings, ensuring smooth communication and transparency across languages.

JP2026017931APending Publication Date: 2026-02-05SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024118992
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-24
Publication Date
2026-02-05

AI Technical Summary

Technical Problem

Modern meetings, particularly in local neighborhood associations and international or online settings, suffer from inefficiencies such as chaotic discussions, language barriers, and labor-intensive information management, leading to difficulties in sharing and ensuring transparency among participants.

Method used

A system that collects participant voices, analyzes and translates them into text, extracts key information, visualizes it, and manages time, providing real-time support and multilingual capabilities to ensure smooth communication and efficient meeting progress.

Benefits of technology

The system enhances meeting efficiency by enabling real-time information sharing, transparency, and automatic minute generation, supporting participants with diverse languages and improving overall meeting management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026017931000001_ABST
    Figure 2026017931000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: The system includes a means for collecting voices of conference participants, a means for analyzing the collected voice data and identifying statements of the conference participants, a means for converting the identified statement contents into text data, a means for extracting important information from the text data, and a means for visualizing the extracted information.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] Modern meetings require the skills to proceed efficiently and clearly understand what participants are saying. However, local neighborhood associations and PTAs often lack these skills, leading to chaotic meetings. Furthermore, language barriers can hinder the communication of information in international and online meetings that support multiple languages. Furthermore, creating meeting minutes, managing time, and extracting and visualizing important information are labor-intensive, making it difficult to share information and ensure transparency among participants. A system that can solve these issues is needed. [Means for solving the problem]

[0005] The present invention solves these problems by providing a system that includes a means for collecting the voices of conference participants, a means for analyzing the collected voice data and identifying the participants' comments, a means for converting the identified comments into text data, a means for extracting important information from the text data, and a means for visualizing the extracted information. Furthermore, by further including a means for translating the collected voice data into multiple languages ​​and a means for displaying the translated content, smooth information sharing is achieved even among participants who speak different languages. Furthermore, by adding a means for monitoring the progress of the conference, managing time, and a means for notifying participants when the scheduled time has been exceeded, the system supports the efficient progress of the conference. This allows all conference participants to grasp information in real time, eliminating information gaps and improving transparency.

[0006] "Meeting participant" refers to a person who participates in a meeting and speaks or exchanges opinions.

[0007] "Audio Data" refers to audio information in digital form captured through an audio collection device such as a microphone.

[0008] "Analysis" refers to a method of performing specific processing on collected voice data to recognize and identify its content.

[0009] "Identification" refers to the process of distinguishing and individually recognizing specific conference participants from the analyzed voice data.

[0010] "Text data" refers to data in a format in which voice data is converted into character information.

[0011] "Extraction" refers to the process of selecting specific information or important points from a large amount of data.

[0012] "Visualization" refers to a method of displaying extracted information in a visually easy-to-understand format.

[0013] "Multilingual translation" refers to the process of converting content written in a specific language into another language.

[0014] "Monitoring" refers to a method of observing the progress of a meeting and the passage of time to grasp the situation.

[0015] "Notification" refers to the process of sending information to notify of a specific situation, time limit, etc.

[0016] "Minutes" refers to a document that records and organizes what was said and decided during a meeting. [Brief explanation of the drawings]

[0017] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram illustrating a processing flow of the data processing system according to the first embodiment. [Figure 12]FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0018] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0019] First, the terms used in the following description will be explained.

[0020] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0021] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0022] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0023] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0024] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0025] [First embodiment]

[0026] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0027] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0028] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0029] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0030] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0031] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0032] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0033] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0034] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0035] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0036] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0037] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0038] This invention is a system that supports efficient conference management by collecting and analyzing the comments of conference participants, and organizing and visualizing the content. This system is realized through cooperation between "terminals," "servers," and "users."

[0039] System Configuration

[0040] 1. Collection of voice data (device)

[0041] The terminals collect the voices of the conference participants through microphones. The collected voice data is recorded in real time and sent to a server. The terminals are usually devices such as smartphones or PCs.

[0042] 2. Sending audio data (terminal)

[0043] The device transmits the collected voice data over the network to a server, where it is compressed and encrypted to ensure security and efficiency.

[0044] 3. Analysis of voice data (server)

[0045] The server uses voice recognition technology to analyze the received voice data, converting it into text data and identifying the speaker.

[0046] 4. Extraction of comment content (server)

[0047] The server uses natural language processing (NLP) techniques to extract key information from the text data, such as the purpose of the meeting and related agenda items.

[0048] 5. Content organization and visualization (server)

[0049] The server organizes the extracted information, summarizes important points, and visualizes them in the form of mind maps or lists. The visualized information is sent to the terminal in real time and displayed to the user.

[0050] 6. Multilingual Translation (Server)

[0051] The server uses a multilingual translation engine to translate speech into multiple languages ​​as needed, and the translation results are also displayed on the device in real time.

[0052] 7. Time Management and Notification (Server)

[0053] The server monitors the progress of the conference and notifies the user if the scheduled time is exceeded or if the progress is delayed. This notification is sent to the terminal and displayed to the user.

[0054] 8. Generating minutes (server)

[0055] After the meeting, the server automatically generates minutes based on the organized and visualized information. The minutes are then saved in cloud storage so that users can access them later.

[0056] Specific examples

[0057] Example 1: Neighborhood Association Meeting

[0058] During neighborhood association meetings, devices (smartphones and PCs) collect participants' voices and send them to a server. The server analyzes the voice data, converts comments into text, and extracts and visualizes information related to important agenda items. The extracted information is displayed in real time on the device in mind map format, helping to guide the meeting. After the meeting, the server automatically generates minutes, which users can access and check by accessing cloud storage.

[0059] Example 2: International online conference

[0060] In international online conferences, devices (PCs) collect participants' voices through online conference tools and send them to a server. The server analyzes the voice data, identifies the speaker, and converts what is being said into text. It also translates what is being said into multiple languages ​​and displays the translated content on the device in real time. The server monitors the progress and notifies users in the event of delays or exceeding the scheduled time. After the conference ends, minutes are automatically generated in multiple languages, and users can access them via cloud storage.

[0061] With the system configuration and specific example described above, the present invention can support the efficient progress of a conference and information sharing among participants, and ensure information transparency.

[0062] The processing flow will be explained below.

[0063] Step 1:

[0064] When a meeting starts, the device activates the built-in or connected microphone device to collect participants' speech, and the collected audio data is converted into a digital format.

[0065] Step 2:

[0066] The device compresses and encrypts the collected voice data in real time, ensuring security and efficient transmission of the voice data.

[0067] Step 3:

[0068] The terminal transmits the encrypted voice data to the server via the network.

[0069] Step 4:

[0070] The server receives the received voice data and inputs it into the analysis engine, which uses voice recognition technology to convert the voice data into text data.

[0071] Step 5:

[0072] The server performs speaker identification analysis on the text data, where it identifies who spoke based on voice characteristics and other identifying information.

[0073] Step 6:

[0074] The server analyzes the text data using natural language processing (NLP) techniques to extract key information about the purpose of the meeting and related agenda items.

[0075] Step 7:

[0076] The server visualizes the extracted information in the form of a mind map or list, and this visualization data is sent to the device in real time.

[0077] Step 8:

[0078] The terminal displays the received visualization data on a user interface so that the user can refer to it.

[0079] Step 9:

[0080] The server monitors the progress of the conference and generates an alert if the conference is running behind schedule or exceeds the scheduled time.

[0081] Step 10:

[0082] The server sends the generated alert to the terminal, and the user receives the alert on the terminal and adjusts the progress of the conference.

[0083] Step 11:

[0084] The server will then launch a translation engine to translate the speech into the specified languages ​​as needed, and the translated content will be sent to the device in real time.

[0085] Step 12:

[0086] The terminal displays the translated content to the user, facilitating understanding between participants who speak different languages.

[0087] Step 13:

[0088] At the end of the meeting, the server automatically generates minutes based on all collected and analyzed information.

[0089] Step 14:

[0090] The server stores the generated minutes in cloud storage for users to access later.

[0091] Step 15:

[0092] After the meeting, users can access the cloud storage to view the saved minutes and other related information.

[0093] With this specific processing flow, the system provides support at every stage of meeting management, creating an environment in which all participants can collaborate efficiently.

[0094] Example 1

[0095] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0096] To ensure efficient meeting management and information transparency, a system is needed that can quickly and accurately collect and analyze participants' comments and organize and visualize related information. However, current systems have fragmented processes from audio collection to analysis and visualization, and lack real-time capabilities. Furthermore, they lack multilingual support and meeting progress management, making them difficult to use in international or multilingual meetings. Furthermore, the automatic generation of meeting minutes and their storage in cloud storage are not very practical.

[0097] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0098] In this invention, the server includes means for collecting the voices of conference participants, means for compressing and encrypting the collected voice data and transmitting it, means for analyzing the transmitted voice data and converting the voices into text data, natural language processing means for extracting important information from the text data, means for organizing the extracted information and visualizing it in the form of a mind map or list, and means for displaying the visualized information in real time. This enables real-time support for the progress of the conference, multilingual support, progress monitoring, automatic generation of minutes, and storage in cloud storage.

[0099] "Conference Participant" means a person who attends and speaks at a conference.

[0100] "Audio" refers to the words and voices of conference participants.

[0101] "Audio Data" means collected audio in digital form.

[0102] "Means for collecting" refers to a method or device for capturing sound using a microphone.

[0103] "Compression" refers to the technique of reducing the size of data.

[0104] "Encryption" refers to the technology of converting data into a form that is not easily understandable by third parties.

[0105] "Transmitting means" refers to a method or device for transferring data to another device over a network.

[0106] "Analysis" refers to the process of examining data in detail to understand its structure and content.

[0107] "Speech recognition" refers to the technology of analyzing voice data and converting it into text data.

[0108] "Text data" refers to data that expresses information using characters.

[0109] "Natural language processing" refers to the technology that allows computers to understand and process the language that humans use naturally.

[0110] "Extraction" refers to the process of extracting the desired information from the data.

[0111] "Organizing information" refers to classifying acquired information, making it relevant, and arranging it in a form that is easy to understand.

[0112] "Visualization" refers to displaying information graphically to make it easier to understand visually.

[0113] A "mind map" is a diagram that shows related items radiating out from a central concept.

[0114] "List format" refers to a format in which information is organized in bullet points and displayed in an orderly manner.

[0115] "Real-time" refers to reporting current events as they occur with almost no delay.

[0116] "Display means" refers to a method or device for visually displaying information using a screen, display, etc.

[0117] "Multilingual translation" refers to the technology of converting one language into another.

[0118] "Progress monitoring" refers to the process of overseeing and managing the progress of a meeting.

[0119] "Notification" means sending a message informing you about a particular condition or event.

[0120] "Minutes" refers to a document that records the discussions and decisions made at a meeting.

[0121] "Cloud storage" refers to online storage services that store and make data accessible over the internet.

[0122] The present invention is a system that supports efficient conference management by collecting, analyzing, organizing, and visualizing the comments of conference participants. Details of how to implement this system are described below.

[0123] This system includes a "terminal" used by a user, a "server" that performs analysis and management, and a "user" who is the user of the system.

[0124] Hardware and software used

[0125] The terminals used are devices with the ability to collect audio, such as smartphones, tablets, and PCs, with a meeting collection application installed.

[0126] The server incorporates a speech recognition engine (e.g., Google Cloud Speech-to-Text), a natural language processing engine (e.g., spaCy, NLTK), and a multilingual translation engine (e.g., Google Translate API).

[0127] System configuration

[0128] 1. Collection of audio data

[0129] The terminal uses a microphone to collect the voices of the conference participants. At the start of the conference, the user launches the conference collection application and presses the "Start Conference" button. This operation causes the terminal to collect voices in real time and record the collected voice data.

[0130] 2. Sending audio data

[0131] The collected voice data is compressed and encrypted on the device, ensuring security and efficiency, and then transmitted over the network to a server.

[0132] 3. Analysis of audio data

[0133] The server receives the voice data sent from the device and analyzes it using a voice recognition engine. It converts the voice into text data and identifies the speaker. Specifically, it extracts features from the patterns of the voice data and converts the speech into text.

[0134] 4. Extraction of speech content

[0135] The server then performs natural language processing on the text data to extract key information and identify key agenda items and action points for the meeting.

[0136] 5. Organizing and visualizing information

[0137] The extracted information is organized on the server and visualized in mind maps or lists. For example, subtopics and related items for each agenda item are organized and displayed in a format that is easy for users to understand.

[0138] 6. Multilingual translation and notification management

[0139] If necessary, the server translates the speech into other languages ​​using a multilingual translation engine. The translated information is displayed on the user's device in real time. In addition, the server monitors the progress of the meeting and notifies the device if the meeting goes over the scheduled time or if the progress is behind schedule.

[0140] 7. Generate and save minutes

[0141] After the meeting, the server automatically generates minutes based on the organized and visualized information. The minutes are then saved in cloud storage, allowing users to access them later.

[0142] Specific examples

[0143] Example 1: Neighborhood Association Meeting

[0144] During neighborhood association meetings, users use their smartphones to collect audio. During the meeting, the device sends the collected audio data to a server, which converts it into text in real time and identifies and visualizes important topics. After the meeting, the automatically generated minutes can be viewed from cloud storage.

[0145] Example 2: International online conference

[0146] In international online conferences, participants use their computers to collect audio and send it to a server via an online conference tool. The server analyzes the audio data and translates what is being said into multiple languages. The translation results are displayed on the device in real time, and progress management and notification functions are also supported. After the conference ends, multilingual minutes can be viewed from cloud storage.

[0147] Prompt Sentence Examples

[0148] "Please explain a program that creates real-time, multilingual minutes for international conferences."

[0149] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0150] Step 1:

[0151] The user starts up the terminal, opens the conference collection application, and presses the "Start Conference" button. This starts the collection of audio data. The user's operation leads to the start of audio capture using the microphone.

[0152] Input: User's instruction to start a meeting

[0153] Output: Trigger to start collection

[0154] Step 2:

[0155] The device collects the voices of the meeting participants in real time through a microphone, and each participant's speech is recorded as a separate audio stream.

[0156] Input: Conference room audio

[0157] Output: Raw audio data

[0158] Step 3:

[0159] The device compresses the collected audio data and then encrypts it, which is then sent to a server using a secure communication protocol.

[0160] Input: Raw audio data

[0161] Output: Encrypted audio data

[0162] Step 4:

[0163] The server receives the voice data sent from the terminal and temporarily stores the received data in a buffer.

[0164] Input: Encrypted audio data

[0165] Output: Buffered audio data

[0166] Step 5:

[0167] The server uses a speech recognition engine to analyze the voice data and convert it into text data. Specifically, it extracts features from the voice pattern, analyzes them, and converts them into text.

[0168] Input: Buffered audio data

[0169] Output: Text data

[0170] Step 6:

[0171] The server uses natural language processing on the text data to extract key information, a process that identifies key topics and action points for the meeting.

[0172] Input: Text data

[0173] Output: Important information (extracted agenda items, action points, etc.)

[0174] Step 7:

[0175] The server organizes and visualizes the extracted information in the form of a mind map or list. Specifically, it arranges related information for each agenda item in a radial pattern, making it visually easy to understand.

[0176] Input: Important Information

[0177] Output: Visualized data

[0178] Step 8:

[0179] The server transmits the visualized information to the terminal in real time, and the terminal displays this information and provides it to the user in a timely manner.

[0180] Input: Visualization data

[0181] Output: Information displayed on the terminal

[0182] Step 9:

[0183] The server uses a multilingual translation engine to translate the speech into other languages ​​as needed, and the translated information is also displayed on the device in real time.

[0184] Input: Important Information

[0185] Output: Translated information

[0186] Step 10:

[0187] The server monitors the progress of the conference and sends notifications to the terminals when the scheduled time has passed or when the progress is delayed. This notification is used to adjust the progress.

[0188] Input: Meeting progress data

[0189] Output: Notification of overdue time

[0190] Step 11:

[0191] After the meeting, the server automatically generates minutes based on the organized and visualized information, and the generated minutes are saved in cloud storage.

[0192] Input: Organized and visualized information

[0193] Output: Meeting minutes

[0194] Step 12:

[0195] Users can access cloud storage and check and download the generated minutes, making it easy to review them after the meeting.

[0196] Input: Meeting minutes stored in cloud storage

[0197] Output: User-accessible transcript

[0198] (Application example 1)

[0199] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0200] Ensuring safety and streamlining operations are important issues in modern factory environments. However, addressing these challenges requires fast and accurate communication and information sharing. However, it is difficult to quickly generate and display important information and guidelines during real-time meetings. Furthermore, language barriers exist in multilingual meetings, making it difficult for all participants to share the same information. There is a need to solve these issues and improve the efficiency of meeting management within factories and information transparency.

[0201] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0202] In this invention, the server includes means for collecting the voices of conference participants, means for analyzing the collected voice data and identifying what the conference participants are saying, means for converting the identified content of the speech into text data, means for extracting important information from the text data, means for visualizing the extracted information, means for displaying the extracted information on a factory work robot, means for automatically generating a safety manual and efficiency guidelines for the conference in real time, and means for displaying the generated manuals and guidelines using smart glasses or a head-mounted display. This enables quick and accurate information sharing in factory meetings, ensuring safety and improving work efficiency.

[0203] A "conference participant" is someone who actually attends a conference and takes part in speeches and discussions.

[0204] "Audio data" refers to information that has been collected from the voices of conference participants and converted into digital format.

[0205] "Analysis" is the process of processing collected audio data using algorithms to extract useful information.

[0206] "Speech identification" refers to recognizing the content of speech made by each individual participant in a conference.

[0207] "Text data" is digital data that has been converted from voice data into character information.

[0208] "Important information" is information that is particularly significant to the progress of the meeting and has a direct impact on decision-making and discussion.

[0209] "Visualization" refers to visually displaying extracted information in the form of charts, graphs, etc., to make it easier to understand.

[0210] A "factory robot" is a mechanical device designed to perform automated tasks in a factory.

[0211] "Real-time" means that information is collected and processed with little to no delay, providing immediate results.

[0212] A "safety manual" is a guide to ensuring safety at work sites.

[0213] "Efficiency guidelines" are documents that contain specific procedures and instructions for improving work efficiency.

[0214] "Smart glasses" are a wearable eyeglass-type device equipped with a display function.

[0215] A "head-mounted display" is a display device that is worn on the head and provides visual information.

[0216] This invention is a system that supports the efficient management of factory meetings by collecting and analyzing the comments of meeting participants, and organizing and visualizing the content. This system is realized through cooperation between "terminals," "servers," and "users."

[0217] System Configuration

[0218] 1. Collection of voice data (device)

[0219] The terminal collects the voices of the conference participants through a microphone. The collected voice data is recorded in real time and sent to a server. The terminal is typically a device such as smart glasses, a PC, or a factory robot.

[0220] 2. Sending audio data (terminal)

[0221] The device transmits the collected voice data over the network to a server, where it is compressed and encrypted to ensure security and efficiency.

[0222] 3. Analysis of voice data (server)

[0223] The server uses speech recognition technology to analyze the received voice data. Specifically, it converts the collected voice data into text data using the speech_recognition library. High-accuracy speech recognition is also possible by using Google's speech recognition API.

[0224] 4. Extraction of comment content (server)

[0225] The server uses natural language processing (NLP) techniques to extract important information from text data, leveraging the SpaCy library to extract important noun phrases and key phrases from the text.

[0226] 5. Content organization and visualization (server)

[0227] The server organizes the extracted information and visualizes the important content. It uses NetworkX and Matplotlib to display key phrases in graph form, allowing users to grasp the progress of the meeting and important topics at a glance.

[0228] 6. Multilingual Translation (Server)

[0229] The server uses the googletrans library to translate speech into multiple languages ​​as needed, and the translation results are displayed on the device in real time.

[0230] 7. Time Management and Notification (Server)

[0231] The server monitors the progress of the conference and notifies the user if the scheduled time is exceeded or if the progress is delayed. This notification is sent to the terminal and displayed to the user.

[0232] 8. Multi-purpose display (server)

[0233] The server automatically generates safety manuals and efficiency guidelines for meetings in real time, and displays the generated manuals and guidelines using smart glasses or a head-mounted display.

[0234] Specific examples

[0235] During meetings on safety and efficiency within a factory, devices (smart glasses or head-mounted displays) collect the voices of participants and send them to a server. The server analyzes the voice data, converts comments into text, and extracts and visualizes information related to important topics. Furthermore, it automatically generates safety manuals and efficiency guidelines in real time, and visualizes this data for participants via the smart glasses or head-mounted displays.

[0236] The server also translates multiple languages ​​as the meeting progresses, providing information in a format that is easy for all participants to understand, facilitating smooth information sharing during meetings within the factory and ensuring that important information on safety and other matters is communicated promptly.

[0237] Example prompts to input to the generative AI model:

[0238] I am trying to build a "meeting support system" for use in factory meetings. This system collects meeting audio in real time, extracts important keywords and phrases, and supports efficient meeting management. Please refer to the code below and help with additional improvements and enhancements.

[0239] Code fragment:

[0240] def record_audio(duration=60):

[0241] recognizer = sr.Recognizer()

[0242] mic = sr.Microphone()

[0243] with mic as source:

[0244] print("Meeting recording begins...")

[0245] audio = recognizer.record(source, duration=duration)

[0246] print("Recording complete")

[0247] return audio

[0248] Based on this, please give me some advice.

[0249] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0250] Step 1:

[0251] The terminal collects the voices of the conference participants. Specifically, it records the voices through a microphone and captures the voice data in real time. The input is the voices of the conference participants, and the output is the collected voice data.

[0252] Step 2:

[0253] The terminal sends the collected voice data to the server. When sending the data, the voice data is compressed and encrypted. The voice data is input, and the compressed and encrypted voice data is sent to the server as output.

[0254] Step 3:

[0255] The server analyzes the received voice data. Specifically, it converts the voice data into text data using the speech_recognition library. The input is compressed and encrypted voice data, and the output is text data.

[0256] Step 4:

[0257] The server extracts important information from the text data. Specifically, it uses the SpaCy library to extract noun phrases and key phrases from the text. The input is the text data, and the output is the extracted noun phrases and key phrases.

[0258] Step 5:

[0259] The server visualizes the extracted information in a graph format using NetworkX and Matplotlib. The input is the extracted key phrases, and the output is the visualized graph.

[0260] Step 6:

[0261] The server translates the speech into multiple languages ​​as needed. Specifically, it uses the GoogleTrans library for translation. The input is text data, and the output is translated text data.

[0262] Step 7:

[0263] The server monitors the progress of the conference and manages the time. It notifies participants if the conference is delayed or exceeds the scheduled time. The input is the start time of the conference, and the output is notifications according to the progress of the conference.

[0264] Step 8:

[0265] The server automatically generates safety manuals and efficiency guidelines and displays them on the terminal. The generated manuals and guidelines are displayed in real time through smart glasses or a head-mounted display. The input is extracted information and key phrases, and the generated safety manuals and efficiency guidelines are displayed as output.

[0266] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0267] This invention is a system that collects and analyzes the comments of meeting participants, organizes and visualizes the content, and supports efficient meeting management. By combining this with emotion recognition functionality, it is possible to grasp the emotional state of participants and adjust the progress of the meeting. This system is realized through cooperation between "terminals," "servers," and "users."

[0268] System Configuration

[0269] 1. Collection of voice data (device)

[0270] The terminal collects the voices of the participants through a microphone. The collected voice data is converted into a digital format in real time and sent to a server. The terminal is usually a device such as a smartphone or PC.

[0271] 2. Sending audio data (terminal)

[0272] The device compresses and encrypts the collected audio data and sends it to a server over the network, ensuring data security and efficient transfer.

[0273] 3. Analysis of voice data (server)

[0274] The server uses speech recognition technology to analyze the received voice data, converts the voice data into text data, identifies the speaker, and then uses natural language processing (NLP) technology to extract important information from the text data.

[0275] 4. Emotion Analysis (Server)

[0276] The server uses an emotion engine for emotion recognition to analyze participants' emotions from the voice data. The emotion engine uses an algorithm to estimate emotions based on voice characteristics such as intonation, tempo, and volume.

[0277] 5. Information visualization (server)

[0278] The server organizes the extracted information and emotional data and visualizes them in mind maps and lists. The emotional data is displayed in graphs and color codes, allowing the progress of the meeting to be understood at a glance. The visualized information is sent to the device in real time and displayed to the user.

[0279] 6. Multilingual Translation (Server)

[0280] The server translates the speech into multiple languages ​​as needed, and the translation results are sent to the device in real time and displayed.

[0281] 7. Time Management and Notification (Server)

[0282] The server monitors the progress of the conference and notifies the user if the conference exceeds the scheduled time or if the conference is behind schedule. The notification is sent to the terminal and displayed to the user.

[0283] 8. Generating minutes (server)

[0284] After the meeting, the server automatically generates minutes based on all collected and analyzed information, and the minutes are saved in cloud storage for users to access later.

[0285] Specific examples

[0286] Example 1: Neighborhood Association Meeting

[0287] At neighborhood association meetings, devices (smartphones and PCs) collect the voices of participants and send them to a server. The server analyzes the voice data, converts comments into text, and extracts and visualizes information related to important topics. It also uses an emotion engine to analyze the emotional state of participants and visualizes this at the same time. The emotional data is displayed in graphs and colors to show how participants are feeling when they speak, and is used to adjust the progress of the meeting. After the meeting ends, the server automatically generates minutes, which users can access and view by accessing cloud storage.

[0288] Example 2: International online conference

[0289] In international online conferences, devices (PCs) collect participants' voices through online conference tools and send them to a server. The server analyzes the voice data, identifies the speaker, and converts what was said into text. Natural language processing technology is used to extract important information, which is then translated using a multilingual translation engine. Emotions are then analyzed using an emotion engine and displayed in real time. The server monitors the progress and issues notifications in the event of delays or overruns. Minutes generated after the conference end contain emotional data along with the content of what was said, and users can access them via cloud storage.

[0290] With this specific configuration, the system provides comprehensive support at every stage of meeting management, ensuring effective progress while also taking into account the emotional states of participants.

[0291] The processing flow will be explained below.

[0292] Step 1:

[0293] When a meeting starts, the device activates the built-in or connected microphone device to collect participants' speech, and the collected audio data is converted into a digital format.

[0294] Step 2:

[0295] The device compresses and encrypts the collected voice data in real time, ensuring security and efficient transmission of the voice data.

[0296] Step 3:

[0297] The terminal transmits the encrypted voice data to the server via the network.

[0298] Step 4:

[0299] The server receives the received voice data and inputs it into a voice recognition engine, which converts the voice data into text data.

[0300] Step 5:

[0301] The server performs speaker identification analysis on the text data, where it identifies who spoke based on voice characteristics and other identifying information.

[0302] Step 6:

[0303] The server analyzes the text data using natural language processing (NLP) techniques to extract key information about the purpose of the meeting and related agenda items.

[0304] Step 7:

[0305] The server inputs the voice data into an emotion engine for emotion analysis, which estimates the participants' emotions based on voice characteristics such as intonation, tempo, and volume.

[0306] Step 8:

[0307] The server organizes the extracted information and emotional data and visualizes them in mind maps and lists, while the emotional data is displayed in graphs and color codes.

[0308] Step 9:

[0309] The server transmits the visualized information to the terminal in real time.

[0310] Step 10:

[0311] The terminal displays the received visualization data on a user interface so that the user can refer to it.

[0312] Step 11:

[0313] The server monitors the progress of the conference and generates an alert if the conference is running behind schedule or exceeds the scheduled time.

[0314] Step 12:

[0315] The server sends the generated alert to the terminal, and the user receives the alert on the terminal and adjusts the progress of the conference.

[0316] Step 13:

[0317] The server will then launch a translation engine to translate the speech into the specified languages ​​as needed, and the translated content will be sent to the device in real time.

[0318] Step 14:

[0319] The terminal displays the translated content to the user, facilitating understanding between participants who speak different languages.

[0320] Step 15:

[0321] At the end of the meeting, the server automatically generates minutes based on all collected and analyzed information.

[0322] Step 16:

[0323] The server stores the generated minutes in cloud storage for users to access later.

[0324] Step 17:

[0325] After the meeting, users can access the cloud storage to view the saved minutes and other related information.

[0326] With this specific processing flow, the system provides support at every stage of meeting management, creating an environment where meetings can proceed efficiently while taking into account the emotional state of participants.

[0327] Example 2

[0328] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0329] Although conventional meeting support systems had the ability to organize and visualize participants' comments, they were inadequate in terms of grasping participants' emotional states and supporting multiple languages. As a result, the progress of the meeting could be stalled due to emotional factors, and communication in a multilingual environment could not proceed smoothly. In addition, there were limited means for effectively monitoring the progress of the meeting and managing time. This resulted in problems that reduced the efficiency of meeting management.

[0330] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0331] In this invention, the server includes means for analyzing voice data and converting speech content into text data, means for analyzing emotional states from the voice data, means for visualizing important information and emotional data, means for translating into multiple languages, and means for monitoring progress and managing time. This makes it possible to grasp the emotional states of conference participants and efficiently manage conferences in multiple languages.

[0332] A "means for capturing participant voices" is a device and method for recording participant speech during a conference and converting it into a digital format.

[0333] "Means for compressing, encrypting and transmitting collected voice data" refers to a device and method that reduces the data size of collected voice data for efficient storage and transmission, encrypts the data to ensure security, and transmits it to a server via a network.

[0334] "Means for analyzing collected voice data and identifying speech of conference participants" refers to speech recognition technology and methods used to convert collected voice data into text and identify speakers.

[0335] The "means for converting the identified speech content into text data" refers to a method and apparatus for analyzing speech data and converting it into text format using speech recognition technology.

[0336] The "means for extracting important information from text data" refers to an apparatus and method for extracting important topics and key points in a meeting from text data using natural language processing technology.

[0337] The "means for analyzing the emotional state of participants from voice data" refers to an algorithm and method for estimating the emotional state of conference participants by analyzing the intonation, tempo, volume, etc. of the voice in the voice data.

[0338] "Means for visualizing extracted information and emotional data" refers to software and methods for visually displaying information and emotional data collected during a meeting in the form of mind maps, graphs, or lists.

[0339] A "means for translating into multiple languages" is an apparatus and method for translating text data into multiple different languages ​​in real time.

[0340] The "means for monitoring progress and managing time" refers to a device and method for monitoring the progress of a meeting and managing the meeting so that it proceeds within the scheduled time.

[0341] MODE FOR CARRYING OUT THE INVENTION

[0342] This invention is a system that supports efficient meeting management by collecting and analyzing the comments of meeting participants, and organizing and visualizing the content. By further combining this system with emotion recognition functionality, it is possible to grasp the emotional state of participants and adjust the progress of the meeting. This system is realized through cooperation between "terminals," "servers," and "users."

[0343] The device uses a microphone to collect the audio of the meeting participants. For example, a voice recording app or software on a smartphone or PC is used. The device converts the collected audio data into a digital format in real time and prepares it for transmission to the server. At this time, the audio data is compressed using "FFmpeg" and encrypted using "OpenSSL." The collected audio data is then sent to the server via the network.

[0344] The server converts the received voice data into text using speech recognition APIs such as Google Cloud Speech-to-Text and IBM Watson Speech to Text. It then uses natural language processing (NLP) libraries such as spaCy and NLTK to extract important information from the text data. The server also analyzes participants' emotional states from the voice data using emotion recognition technologies such as Microsoft Azure Emotion API and Affectiva. The analyzed emotion data is estimated based on voice characteristics such as intonation, tempo, and volume.

[0345] The server organizes the extracted information and emotion data and visualizes it using D3.js and Chart.js. This allows the information to be organized in mind maps and list formats, and the emotion data to be displayed in graphs and color codes. The visualized information is sent to the device in real time and displayed to the user. For example, the progress of a meeting can be displayed using graphs and tables so that it can be seen at a glance.

[0346] If necessary, the server translates the speech into multiple languages ​​using Google Translate API or Microsoft Translator. The translation results are also sent to the device in real time, allowing users to understand what is being said in different languages. The server monitors the progress of the meeting and uses Twilio or Slack API to notify users if the meeting goes over the scheduled time or if there are delays. Notifications are sent to the device and displayed to the user.

[0347] After the meeting, the server automatically generates minutes based on all collected and analyzed information. The documents are generated using Microsoft Word API and Google Docs API. The generated minutes are saved in cloud storage services such as Google Drive and Dropbox, allowing users to access them later.

[0348] Specific examples

[0349] Neighborhood Association Meeting

[0350] During neighborhood association meetings, devices (smartphones and PCs) collect participants' audio, compress it with FFmpeg, encrypt it with OpenSSL, and then send it to a server. The server converts the audio data into text using Google Cloud Speech-to-Text and extracts important information using spaCy. The server then analyzes the participants' emotional states using Affectiva and visualizes the emotional data using graphs and colors. After the meeting ends, the server automatically generates minutes, which users can view via cloud storage (e.g., Google Drive).

[0351] International Online Conference

[0352] In international online conferences, devices (PCs) collect participants' voices through online conference tools (e.g., Zoom), compress them with FFmpeg, encrypt them with OpenSSL, and send them to a server. The server analyzes the voice data using IBM Watson Speech to Text and converts what is being said into text. Important information is extracted using spaCy and translated using Google Translate API. The server further analyzes emotions using Microsoft Azure Emotion API and displays them on the device in real time. After the conference ends, the server generates minutes including the content of the speech and emotion data, which users can access via Dropbox.

[0353] The system provides comprehensive support at every stage of the meeting, enabling effective meeting management that takes into account the emotional state of participants. Multilingual support also enables smooth communication even in international meetings.

[0354] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0355] Step 1:

[0356] The device collects the voices of the meeting participants through a microphone. The input is the live audio from the meeting, and converts it into a digital format (e.g., WAV or MP3). Specifically, a dedicated application (e.g., audio recording software) on a smartphone or PC is used. The output digital audio data is stored in temporary memory for subsequent processing.

[0357] Step 2:

[0358] The terminal compresses the collected audio data using "FFmpeg" and encrypts it with "OpenSSL." The input is the digital audio data obtained in step 1, which is compressed to reduce the data size and encrypted to ensure data security. Specifically, the terminal executes the FFmpeg compression command and the OpenSSL encryption command. The output is compressed and encrypted audio data, which is sent to the server via the network.

[0359] Step 3:

[0360] The server receives the audio data sent from the device, decrypts it, and decompresses it. The input is compressed and encrypted audio data, and OpenSSL and FFmpeg are used to decompress it and return it to plain text. Specifically, the server executes the OpenSSL decryption command and the FFmpeg decompression command. The output is the original digital audio data.

[0361] Step 4:

[0362] The server converts the voice data into text data using speech recognition technology. The input is the original digital voice data, which is converted into text using APIs such as "Google Cloud Speech-to-Text" and "IBM Watson Speech to Text." The API is called to analyze the voice data and convert the utterances into text format. The output is text data for each speaker.

[0363] Step 5:

[0364] The server uses natural language processing (NLP) technology to extract important information from text data. The input is text data obtained through speech recognition, and information extraction is performed using tools such as "spaCy" and "NLTK." Specifically, the text data is input into an NLP library, and related keywords and topics are extracted. The output is text data extracted as important information.

[0365] Step 6:

[0366] The server analyzes the participants' emotional states from the voice data. The input is the voice data obtained in step 3, and emotion analysis is performed using tools such as the Microsoft Azure Emotion API. Specifically, the voice data is passed to the emotion engine, which estimates the emotional state (e.g., joy, anger, sadness). The output is data indicating the emotional state of each participant.

[0367] Step 7:

[0368] The server visualizes the extracted information and emotion data. The input is the important information obtained in step 5 and the emotion data obtained in step 6, and these are displayed visually using "D3.js" and "Chart.js." Specifically, it generates mind maps and graphs and sends them to the user's device in real time. The output is the visualized data.

[0369] Step 8:

[0370] The server translates the speech into multiple languages ​​as needed. The input is text data obtained through speech recognition, which is translated using the Google Translate API or Microsoft Translator. Specifically, it calls a translation engine to convert the text into another language, and the output is the translated text data.

[0371] Step 9:

[0372] The server monitors the progress of the meeting and manages the time. The input is meeting data updated in real time, and a notification is generated if the scheduled time is exceeded or progress is delayed. For example, a notification message is created using "Twilio" or "Slack API" and sent to the user's device. The output is a notification to the user.

[0373] Step 10:

[0374] After the meeting ends, the server automatically generates minutes based on all collected and analyzed information. The input is the voice data collected during the meeting, analyzed text data, and emotion data, and minutes are generated using the Microsoft Word API or Google Docs API. Specifically, the data is passed to the minutes generation engine to generate a document, which is then saved in cloud storage. The output is the generated minutes.

[0375] These processing steps effectively organize the remarks of the conference participants and realize efficient conference proceedings that take into account their emotional states.

[0376] (Application example 2)

[0377] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0378] While conventional meeting support systems can collect participants' voices, convert what is being said into text, and extract and visualize important information, they are unable to grasp the participants' emotional state and adjust the progress of the meeting. Furthermore, when making advertising presentations, there is no mechanism for providing real-time feedback on participants' reactions and emotions, making it difficult to effectively present advertising campaigns.

[0379] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0380] In this invention, the server includes a means for analyzing the emotional states of conference participants from their comments, a means for visualizing the analyzed emotional states in real time, and a means for automatically generating a record including the comments and emotional states after the conference ends. This allows participants' reactions and emotional states to be grasped in real time, enabling effective progress and feedback of the advertising presentation.

[0381] The "means for collecting the voices of the conference participants" refers to a device or system that converts the voices of the conference participants into a digital format in real time via a microphone and collects them as data.

[0382] "Means for analyzing collected voice data and identifying what is being said by conference participants" refers to algorithms or software for identifying conference participants and specifying what is being said based on collected voice data.

[0383] The "means for converting the identified speech content into text data" is a process for converting the identified speech content into text data in sentence format using speech recognition technology.

[0384] The "means for extracting important information from text data" is a system that uses natural language processing technology to extract key topics and important information discussed in meetings and presentations from text data.

[0385] "Means for visualizing extracted information" refers to software or tools that visually display extracted important information and data in the form of graphs, lists, mind maps, etc.

[0386] "Means for analyzing emotional state from speech content" refers to algorithms or emotion recognition engines for estimating emotional state based on characteristics of the speaker's voice data, such as intonation, tempo, and volume.

[0387] The "means for visualizing analyzed emotional states in real time" is a system for displaying emotional data estimated by an emotion recognition engine in real time using graphs and colors.

[0388] "Means for automatically generating records including remarks and emotional states after a meeting" refers to software that automatically creates and saves minutes and records that integrate remarks and emotional states based on information collected and analyzed after a meeting.

[0389] The present invention is a conference support system used in advertising presentations that analyzes and visualizes participants' comments and emotional states to effectively support the progress of the conference. Specific embodiments are described below.

[0390] System Configuration

[0391] 1. Collection of voice data (device)

[0392] The terminals, typically digital devices such as smartphones and personal computers, collect the voices of conference participants in real time through microphones and convert them into digital format. This data is then compressed, encrypted, and sent to the server.

[0393] 2. Analysis of voice data (server)

[0394] The server receives the collected voice data and converts it into text using speech recognition technology, identifies the speaker, analyzes the speech, and extracts important information. Natural language processing (NLP) technology is used here.

[0395] 3. Emotion Analysis (Server)

[0396] The server uses an emotion recognition engine to estimate participants' emotional states from the voice data. The emotion recognition engine uses an algorithm to analyze voice intonation, tempo, volume, etc. to estimate emotions.

[0397] 4. Real-time visualization (server)

[0398] The server organizes the extracted text information and emotional data and visualizes it in real time in graph and list format. The emotional data is displayed using colors and graphs, allowing participants' emotional states to be understood at a glance. The visualized information is sent to the device in real time and displayed to the user.

[0399] 5. Generation of minutes (server)

[0400] After the meeting, the server automatically generates minutes based on all collected and analyzed information, which are then saved in cloud storage for users to access later.

[0401] Usage example

[0402] Specific examples

[0403] During an advertising presentation meeting, when an advertising agency is presenting a new advertising campaign, the comments made by participants and the emotional state of those comments are analyzed and visualized in real time. For example, it is possible to instantly understand whether participants' reactions are positive or negative during the presentation. This allows the presenter to flexibly adjust the progress of the meeting and deliver an effective presentation.

[0404] Prompt Sentence Examples

[0405] Create a system that analyzes participants' comments and emotions in real time during advertising campaign meetings and visualizes the results. Include a function to grasp participants' emotional state and automatically generate and save minutes after the meeting.

[0406] Hardware and Software Configuration

[0407] Microphone: for collecting audio data

[0408] Personal computer or smartphone: for analysis and data storage

[0409] Python: a programming language

[0410] SpeechRecognition: A speech recognition library

[0411] transformers: Hugging Face's NLP library for emotion analysis

[0412] Google Speech API: Convert speech to text

[0413] Using this hardware and software, it is possible to collect and analyze voice data, analyze emotional states, visualize information, and generate meeting minutes.

[0414] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0415] Step 1: Collecting voice data (device)

[0416] The terminal collects the voices of the conference participants in real time through a microphone and converts them into a digital format. The terminal then compresses and encrypts the collected voice data and sends it to the server. The input is raw voice, and the output is digitized encrypted voice data.

[0417] Step 2: Analyzing the voice data (server)

[0418] The server converts the received voice data into text data using speech recognition technology. This process uses the Google Speech API to analyze the voice data and convert it into text. It also identifies the speaker and identifies each utterance. The input is encrypted voice data, and the output is the identified utterance text data.

[0419] Step 3: Extracting important information (server)

[0420] The server extracts important information from the identified text data using natural language processing (NLP) techniques. This process uses the Hugging Face transformers library. The input is the text data, and the output is the extracted important information.

[0421] Step 4: Sentiment Analysis (Server)

[0422] The server analyzes the emotional state of the identified speech using an emotion recognition engine. This process estimates emotions based on voice intonation, tempo, volume, etc. The input is text data, and the output is emotion analysis data.

[0423] Step 5: Real-time visualization (server)

[0424] The server visualizes the extracted information and emotional data in graphs and lists, allowing participants' emotional states to be visually displayed in real time. The input is important information and emotional analysis data, and the output is visualized graphs and lists.

[0425] Step 6: Real-time display (terminal)

[0426] The terminal displays the visualized data sent from the server in real time, allowing users to instantly check the emotional state and important information during a meeting. The input is the visualized data, and the output is the displayed graph or list.

[0427] Step 7: Generate minutes (server)

[0428] After the meeting ends, the server automatically generates minutes based on all collected and analyzed information. These minutes include the content of remarks and the emotional state of the participants. The generated minutes are saved in cloud storage and can be accessed by users later. The input is important information and emotional analysis data, and the output is the minutes data.

[0429] Step 8: Saving and Accessing the Minutes (Devices and Users)

[0430] Users can access the cloud storage through their devices and check the generated minutes. The input is the minutes data on the cloud, and the output is the minutes that users can view.

[0431] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0432] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0433] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0434] [Second embodiment]

[0435] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0436] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0437] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0438] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0439] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0440] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0441] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0442] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0443] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0444] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0445] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0446] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0447] This invention is a system that supports efficient conference management by collecting and analyzing the comments of conference participants, and organizing and visualizing the content. This system is realized through cooperation between "terminals," "servers," and "users."

[0448] System Configuration

[0449] 1. Collection of voice data (device)

[0450] The terminals collect the voices of the conference participants through microphones. The collected voice data is recorded in real time and sent to a server. The terminals are usually devices such as smartphones or PCs.

[0451] 2. Sending audio data (terminal)

[0452] The device transmits the collected voice data over the network to a server, where it is compressed and encrypted to ensure security and efficiency.

[0453] 3. Analysis of voice data (server)

[0454] The server uses voice recognition technology to analyze the received voice data, converting it into text data and identifying the speaker.

[0455] 4. Extraction of comment content (server)

[0456] The server uses natural language processing (NLP) techniques to extract key information from the text data, such as the purpose of the meeting and related agenda items.

[0457] 5. Content organization and visualization (server)

[0458] The server organizes the extracted information, summarizes important points, and visualizes them in the form of mind maps or lists. The visualized information is sent to the terminal in real time and displayed to the user.

[0459] 6. Multilingual Translation (Server)

[0460] The server uses a multilingual translation engine to translate speech into multiple languages ​​as needed, and the translation results are also displayed on the device in real time.

[0461] 7. Time Management and Notification (Server)

[0462] The server monitors the progress of the conference and notifies the user if the scheduled time is exceeded or if the progress is delayed. This notification is sent to the terminal and displayed to the user.

[0463] 8. Generating minutes (server)

[0464] After the meeting, the server automatically generates minutes based on the organized and visualized information. The minutes are then saved in cloud storage so that users can access them later.

[0465] Specific examples

[0466] Example 1: Neighborhood Association Meeting

[0467] During neighborhood association meetings, devices (smartphones and PCs) collect participants' voices and send them to a server. The server analyzes the voice data, converts comments into text, and extracts and visualizes information related to important agenda items. The extracted information is displayed in real time on the device in mind map format, helping to guide the meeting. After the meeting, the server automatically generates minutes, which users can access and check by accessing cloud storage.

[0468] Example 2: International online conference

[0469] In international online conferences, devices (PCs) collect participants' voices through online conference tools and send them to a server. The server analyzes the voice data, identifies the speaker, and converts what is being said into text. It also translates what is being said into multiple languages ​​and displays the translated content on the device in real time. The server monitors the progress and notifies users in the event of delays or exceeding the scheduled time. After the conference ends, minutes are automatically generated in multiple languages, and users can access them via cloud storage.

[0470] With the system configuration and specific example described above, the present invention can support the efficient progress of a conference and information sharing among participants, and ensure information transparency.

[0471] The processing flow will be explained below.

[0472] Step 1:

[0473] When a meeting starts, the device activates the built-in or connected microphone device to collect participants' speech, and the collected audio data is converted into a digital format.

[0474] Step 2:

[0475] The device compresses and encrypts the collected voice data in real time, ensuring security and efficient transmission of the voice data.

[0476] Step 3:

[0477] The terminal transmits the encrypted voice data to the server via the network.

[0478] Step 4:

[0479] The server receives the received voice data and inputs it into the analysis engine, which uses voice recognition technology to convert the voice data into text data.

[0480] Step 5:

[0481] The server performs speaker identification analysis on the text data, where it identifies who spoke based on voice characteristics and other identifying information.

[0482] Step 6:

[0483] The server analyzes the text data using natural language processing (NLP) techniques to extract key information about the purpose of the meeting and related agenda items.

[0484] Step 7:

[0485] The server visualizes the extracted information in the form of a mind map or list, and this visualization data is sent to the device in real time.

[0486] Step 8:

[0487] The terminal displays the received visualization data on a user interface so that the user can refer to it.

[0488] Step 9:

[0489] The server monitors the progress of the conference and generates an alert if the conference is running behind schedule or exceeds the scheduled time.

[0490] Step 10:

[0491] The server sends the generated alert to the terminal, and the user receives the alert on the terminal and adjusts the progress of the conference.

[0492] Step 11:

[0493] The server will then launch a translation engine to translate the speech into the specified languages ​​as needed, and the translated content will be sent to the device in real time.

[0494] Step 12:

[0495] The terminal displays the translated content to the user, facilitating understanding between participants who speak different languages.

[0496] Step 13:

[0497] At the end of the meeting, the server automatically generates minutes based on all collected and analyzed information.

[0498] Step 14:

[0499] The server stores the generated minutes in cloud storage for users to access later.

[0500] Step 15:

[0501] After the meeting, users can access the cloud storage to view the saved minutes and other related information.

[0502] With this specific processing flow, the system provides support at every stage of meeting management, creating an environment in which all participants can collaborate efficiently.

[0503] Example 1

[0504] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0505] To ensure efficient meeting management and information transparency, a system is needed that can quickly and accurately collect and analyze participants' comments and organize and visualize related information. However, current systems have fragmented processes from audio collection to analysis and visualization, and lack real-time capabilities. Furthermore, they lack multilingual support and meeting progress management, making them difficult to use in international or multilingual meetings. Furthermore, the automatic generation of meeting minutes and their storage in cloud storage are not very practical.

[0506] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0507] In this invention, the server includes means for collecting the voices of conference participants, means for compressing and encrypting the collected voice data and transmitting it, means for analyzing the transmitted voice data and converting the voices into text data, natural language processing means for extracting important information from the text data, means for organizing the extracted information and visualizing it in the form of a mind map or list, and means for displaying the visualized information in real time. This enables real-time support for the progress of the conference, multilingual support, progress monitoring, automatic generation of minutes, and storage in cloud storage.

[0508] "Conference Participant" means a person who attends and speaks at a conference.

[0509] "Audio" refers to the words and voices of conference participants.

[0510] "Audio Data" means collected audio in digital form.

[0511] "Means for collecting" refers to a method or device for capturing sound using a microphone.

[0512] "Compression" refers to the technique of reducing the size of data.

[0513] "Encryption" refers to the technology of converting data into a form that is not easily understandable by third parties.

[0514] "Transmitting means" refers to a method or device for transferring data to another device over a network.

[0515] "Analysis" refers to the process of examining data in detail to understand its structure and content.

[0516] "Speech recognition" refers to the technology of analyzing voice data and converting it into text data.

[0517] "Text data" refers to data that expresses information using characters.

[0518] "Natural language processing" refers to the technology that allows computers to understand and process the language that humans use naturally.

[0519] "Extraction" refers to the process of extracting the desired information from the data.

[0520] "Organizing information" refers to classifying acquired information, making it relevant, and arranging it in a form that is easy to understand.

[0521] "Visualization" refers to displaying information graphically to make it easier to understand visually.

[0522] A "mind map" is a diagram that shows related items radiating out from a central concept.

[0523] "List format" refers to a format in which information is organized in bullet points and displayed in an orderly manner.

[0524] "Real-time" refers to reporting current events as they occur with almost no delay.

[0525] "Display means" refers to a method or device for visually displaying information using a screen, display, etc.

[0526] "Multilingual translation" refers to the technology of converting one language into another.

[0527] "Progress monitoring" refers to the process of overseeing and managing the progress of a meeting.

[0528] "Notification" means sending a message informing you about a particular condition or event.

[0529] "Minutes" refers to a document that records the discussions and decisions made at a meeting.

[0530] "Cloud storage" refers to online storage services that store and make data accessible over the internet.

[0531] The present invention is a system that supports efficient conference management by collecting, analyzing, organizing, and visualizing the comments of conference participants. Details of how to implement this system are described below.

[0532] This system includes a "terminal" used by a user, a "server" that performs analysis and management, and a "user" who is the user of the system.

[0533] Hardware and software used

[0534] The terminals used are devices with the ability to collect audio, such as smartphones, tablets, and PCs, with a meeting collection application installed.

[0535] The server incorporates a speech recognition engine (e.g., Google Cloud Speech-to-Text), a natural language processing engine (e.g., spaCy, NLTK), and a multilingual translation engine (e.g., Google Translate API).

[0536] System configuration

[0537] 1. Collection of audio data

[0538] The terminal uses a microphone to collect the voices of the conference participants. At the start of the conference, the user launches the conference collection application and presses the "Start Conference" button. This operation causes the terminal to collect voices in real time and record the collected voice data.

[0539] 2. Sending audio data

[0540] The collected voice data is compressed and encrypted on the device, ensuring security and efficiency, and then transmitted over the network to a server.

[0541] 3. Analysis of audio data

[0542] The server receives the voice data sent from the device and analyzes it using a voice recognition engine. It converts the voice into text data and identifies the speaker. Specifically, it extracts features from the patterns of the voice data and converts the speech into text.

[0543] 4. Extraction of speech content

[0544] The server then performs natural language processing on the text data to extract key information and identify key agenda items and action points for the meeting.

[0545] 5. Organizing and visualizing information

[0546] The extracted information is organized on the server and visualized in mind maps or lists. For example, subtopics and related items for each agenda item are organized and displayed in a format that is easy for users to understand.

[0547] 6. Multilingual translation and notification management

[0548] If necessary, the server translates the speech into other languages ​​using a multilingual translation engine. The translated information is displayed on the user's device in real time. In addition, the server monitors the progress of the meeting and notifies the device if the meeting goes over the scheduled time or if the progress is behind schedule.

[0549] 7. Generate and save minutes

[0550] After the meeting, the server automatically generates minutes based on the organized and visualized information. The minutes are then saved in cloud storage, allowing users to access them later.

[0551] Specific examples

[0552] Example 1: Neighborhood Association Meeting

[0553] During neighborhood association meetings, users use their smartphones to collect audio. During the meeting, the device sends the collected audio data to a server, which converts it into text in real time and identifies and visualizes important topics. After the meeting, the automatically generated minutes can be viewed from cloud storage.

[0554] Example 2: International online conference

[0555] In international online conferences, participants use their computers to collect audio and send it to a server via an online conference tool. The server analyzes the audio data and translates what is being said into multiple languages. The translation results are displayed on the device in real time, and progress management and notification functions are also supported. After the conference ends, multilingual minutes can be viewed from cloud storage.

[0556] Prompt Sentence Examples

[0557] "Please explain a program that creates real-time, multilingual minutes for international conferences."

[0558] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0559] Step 1:

[0560] The user starts up the terminal, opens the conference collection application, and presses the "Start Conference" button. This starts the collection of audio data. The user's operation leads to the start of audio capture using the microphone.

[0561] Input: User's instruction to start a meeting

[0562] Output: Trigger to start collection

[0563] Step 2:

[0564] The device collects the voices of the meeting participants in real time through a microphone, and each participant's speech is recorded as a separate audio stream.

[0565] Input: Conference room audio

[0566] Output: Raw audio data

[0567] Step 3:

[0568] The device compresses the collected audio data and then encrypts it, which is then sent to a server using a secure communication protocol.

[0569] Input: Raw audio data

[0570] Output: Encrypted audio data

[0571] Step 4:

[0572] The server receives the voice data sent from the terminal and temporarily stores the received data in a buffer.

[0573] Input: Encrypted audio data

[0574] Output: Buffered audio data

[0575] Step 5:

[0576] The server uses a speech recognition engine to analyze the voice data and convert it into text data. Specifically, it extracts features from the voice pattern, analyzes them, and converts them into text.

[0577] Input: Buffered audio data

[0578] Output: Text data

[0579] Step 6:

[0580] The server uses natural language processing on the text data to extract key information, a process that identifies key topics and action points for the meeting.

[0581] Input: Text data

[0582] Output: Important information (extracted agenda items, action points, etc.)

[0583] Step 7:

[0584] The server organizes and visualizes the extracted information in the form of a mind map or list. Specifically, it arranges related information for each agenda item in a radial pattern, making it visually easy to understand.

[0585] Input: Important Information

[0586] Output: Visualized data

[0587] Step 8:

[0588] The server transmits the visualized information to the terminal in real time, and the terminal displays this information and provides it to the user in a timely manner.

[0589] Input: Visualization data

[0590] Output: Information displayed on the terminal

[0591] Step 9:

[0592] The server uses a multilingual translation engine to translate the speech into other languages ​​as needed, and the translated information is also displayed on the device in real time.

[0593] Input: Important Information

[0594] Output: Translated information

[0595] Step 10:

[0596] The server monitors the progress of the conference and sends notifications to the terminals when the scheduled time has passed or when the progress is delayed. This notification is used to adjust the progress.

[0597] Input: Meeting progress data

[0598] Output: Notification of overdue time

[0599] Step 11:

[0600] After the meeting, the server automatically generates minutes based on the organized and visualized information, and the generated minutes are saved in cloud storage.

[0601] Input: Organized and visualized information

[0602] Output: Meeting minutes

[0603] Step 12:

[0604] Users can access cloud storage and check and download the generated minutes, making it easy to review them after the meeting.

[0605] Input: Meeting minutes stored in cloud storage

[0606] Output: User-accessible transcript

[0607] (Application example 1)

[0608] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0609] Ensuring safety and streamlining operations are important issues in modern factory environments. However, addressing these challenges requires fast and accurate communication and information sharing. However, it is difficult to quickly generate and display important information and guidelines during real-time meetings. Furthermore, language barriers exist in multilingual meetings, making it difficult for all participants to share the same information. There is a need to solve these issues and improve the efficiency of meeting management within factories and information transparency.

[0610] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0611] In this invention, the server includes means for collecting the voices of conference participants, means for analyzing the collected voice data and identifying what the conference participants are saying, means for converting the identified content of the speech into text data, means for extracting important information from the text data, means for visualizing the extracted information, means for displaying the extracted information on a factory work robot, means for automatically generating a safety manual and efficiency guidelines for the conference in real time, and means for displaying the generated manuals and guidelines using smart glasses or a head-mounted display. This enables quick and accurate information sharing in factory meetings, ensuring safety and improving work efficiency.

[0612] A "conference participant" is someone who actually attends a conference and takes part in speeches and discussions.

[0613] "Audio data" refers to information that has been collected from the voices of conference participants and converted into digital format.

[0614] "Analysis" is the process of processing collected audio data using algorithms to extract useful information.

[0615] "Speech identification" refers to recognizing the content of speech made by each individual participant in a conference.

[0616] "Text data" is digital data that has been converted from voice data into character information.

[0617] "Important information" is information that is particularly significant to the progress of the meeting and has a direct impact on decision-making and discussion.

[0618] "Visualization" refers to visually displaying extracted information in the form of charts, graphs, etc., to make it easier to understand.

[0619] A "factory robot" is a mechanical device designed to perform automated tasks in a factory.

[0620] "Real-time" means that information is collected and processed with little to no delay, providing immediate results.

[0621] A "safety manual" is a guide to ensuring safety at work sites.

[0622] "Efficiency guidelines" are documents that contain specific procedures and instructions for improving work efficiency.

[0623] "Smart glasses" are a wearable eyeglass-type device equipped with a display function.

[0624] A "head-mounted display" is a display device that is worn on the head and provides visual information.

[0625] This invention is a system that supports the efficient management of factory meetings by collecting and analyzing the comments of meeting participants, and organizing and visualizing the content. This system is realized through cooperation between "terminals," "servers," and "users."

[0626] System Configuration

[0627] 1. Collection of voice data (device)

[0628] The terminal collects the voices of the conference participants through a microphone. The collected voice data is recorded in real time and sent to a server. The terminal is typically a device such as smart glasses, a PC, or a factory robot.

[0629] 2. Sending audio data (terminal)

[0630] The device transmits the collected voice data over the network to a server, where it is compressed and encrypted to ensure security and efficiency.

[0631] 3. Analysis of voice data (server)

[0632] The server uses speech recognition technology to analyze the received voice data. Specifically, it converts the collected voice data into text data using the speech_recognition library. High-accuracy speech recognition is also possible by using Google's speech recognition API.

[0633] 4. Extraction of comment content (server)

[0634] The server uses natural language processing (NLP) techniques to extract important information from text data, leveraging the SpaCy library to extract important noun phrases and key phrases from the text.

[0635] 5. Content organization and visualization (server)

[0636] The server organizes the extracted information and visualizes the important content. It uses NetworkX and Matplotlib to display key phrases in graph form, allowing users to grasp the progress of the meeting and important topics at a glance.

[0637] 6. Multilingual Translation (Server)

[0638] The server uses the googletrans library to translate speech into multiple languages ​​as needed, and the translation results are displayed on the device in real time.

[0639] 7. Time Management and Notification (Server)

[0640] The server monitors the progress of the conference and notifies the user if the scheduled time is exceeded or if the progress is delayed. This notification is sent to the terminal and displayed to the user.

[0641] 8. Multi-purpose display (server)

[0642] The server automatically generates safety manuals and efficiency guidelines for meetings in real time, and displays the generated manuals and guidelines using smart glasses or a head-mounted display.

[0643] Specific examples

[0644] During meetings on safety and efficiency within a factory, devices (smart glasses or head-mounted displays) collect the voices of participants and send them to a server. The server analyzes the voice data, converts comments into text, and extracts and visualizes information related to important topics. Furthermore, it automatically generates safety manuals and efficiency guidelines in real time, and visualizes this data for participants via the smart glasses or head-mounted displays.

[0645] The server also translates multiple languages ​​as the meeting progresses, providing information in a format that is easy for all participants to understand, facilitating smooth information sharing during meetings within the factory and ensuring that important information on safety and other matters is communicated promptly.

[0646] Example prompts to input to the generative AI model:

[0647] I am trying to build a "meeting support system" for use in factory meetings. This system collects meeting audio in real time, extracts important keywords and phrases, and supports efficient meeting management. Please refer to the code below and help with additional improvements and enhancements.

[0648] Code fragment:

[0649] def record_audio(duration=60):

[0650] recognizer = sr.Recognizer()

[0651] mic = sr.Microphone()

[0652] with mic as source:

[0653] print("Meeting recording begins...")

[0654] audio = recognizer.record(source, duration=duration)

[0655] print("Recording complete")

[0656] return audio

[0657] Based on this, please give me some advice.

[0658] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0659] Step 1:

[0660] The terminal collects the voices of the conference participants. Specifically, it records the voices through a microphone and captures the voice data in real time. The input is the voices of the conference participants, and the output is the collected voice data.

[0661] Step 2:

[0662] The terminal sends the collected voice data to the server. When sending the data, the voice data is compressed and encrypted. The voice data is input, and the compressed and encrypted voice data is sent to the server as output.

[0663] Step 3:

[0664] The server analyzes the received voice data. Specifically, it converts the voice data into text data using the speech_recognition library. The input is compressed and encrypted voice data, and the output is text data.

[0665] Step 4:

[0666] The server extracts important information from the text data. Specifically, it uses the SpaCy library to extract noun phrases and key phrases from the text. The input is the text data, and the output is the extracted noun phrases and key phrases.

[0667] Step 5:

[0668] The server visualizes the extracted information in a graph format using NetworkX and Matplotlib. The input is the extracted key phrases, and the output is the visualized graph.

[0669] Step 6:

[0670] The server translates the speech into multiple languages ​​as needed. Specifically, it uses the GoogleTrans library for translation. The input is text data, and the output is translated text data.

[0671] Step 7:

[0672] The server monitors the progress of the conference and manages the time. It notifies participants if the conference is delayed or exceeds the scheduled time. The input is the start time of the conference, and the output is notifications according to the progress of the conference.

[0673] Step 8:

[0674] The server automatically generates safety manuals and efficiency guidelines and displays them on the terminal. The generated manuals and guidelines are displayed in real time through smart glasses or a head-mounted display. The input is extracted information and key phrases, and the generated safety manuals and efficiency guidelines are displayed as output.

[0675] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0676] This invention is a system that collects and analyzes the comments of meeting participants, organizes and visualizes the content, and supports efficient meeting management. By combining this with emotion recognition functionality, it is possible to grasp the emotional state of participants and adjust the progress of the meeting. This system is realized through cooperation between "terminals," "servers," and "users."

[0677] System Configuration

[0678] 1. Collection of voice data (device)

[0679] The terminal collects the voices of the participants through a microphone. The collected voice data is converted into a digital format in real time and sent to a server. The terminal is usually a device such as a smartphone or PC.

[0680] 2. Sending audio data (terminal)

[0681] The device compresses and encrypts the collected audio data and sends it to a server over the network, ensuring data security and efficient transfer.

[0682] 3. Analysis of voice data (server)

[0683] The server uses speech recognition technology to analyze the received voice data, converts the voice data into text data, identifies the speaker, and then uses natural language processing (NLP) technology to extract important information from the text data.

[0684] 4. Emotion Analysis (Server)

[0685] The server uses an emotion engine for emotion recognition to analyze participants' emotions from the voice data. The emotion engine uses an algorithm to estimate emotions based on voice characteristics such as intonation, tempo, and volume.

[0686] 5. Information visualization (server)

[0687] The server organizes the extracted information and emotional data and visualizes them in mind maps and lists. The emotional data is displayed in graphs and color codes, allowing the progress of the meeting to be understood at a glance. The visualized information is sent to the device in real time and displayed to the user.

[0688] 6. Multilingual Translation (Server)

[0689] The server translates the speech into multiple languages ​​as needed, and the translation results are sent to the device in real time and displayed.

[0690] 7. Time Management and Notification (Server)

[0691] The server monitors the progress of the conference and notifies the user if the conference exceeds the scheduled time or if the conference is behind schedule. The notification is sent to the terminal and displayed to the user.

[0692] 8. Generating minutes (server)

[0693] After the meeting, the server automatically generates minutes based on all collected and analyzed information, and the minutes are saved in cloud storage for users to access later.

[0694] Specific examples

[0695] Example 1: Neighborhood Association Meeting

[0696] At neighborhood association meetings, devices (smartphones and PCs) collect the voices of participants and send them to a server. The server analyzes the voice data, converts comments into text, and extracts and visualizes information related to important topics. It also uses an emotion engine to analyze the emotional state of participants and visualizes this at the same time. The emotional data is displayed in graphs and colors to show how participants are feeling when they speak, and is used to adjust the progress of the meeting. After the meeting ends, the server automatically generates minutes, which users can access and view by accessing cloud storage.

[0697] Example 2: International online conference

[0698] In international online conferences, devices (PCs) collect participants' voices through online conference tools and send them to a server. The server analyzes the voice data, identifies the speaker, and converts what was said into text. Natural language processing technology is used to extract important information, which is then translated using a multilingual translation engine. Emotions are then analyzed using an emotion engine and displayed in real time. The server monitors the progress and issues notifications in the event of delays or overruns. Minutes generated after the conference end contain emotional data along with the content of what was said, and users can access them via cloud storage.

[0699] With this specific configuration, the system provides comprehensive support at every stage of meeting management, ensuring effective progress while also taking into account the emotional states of participants.

[0700] The processing flow will be explained below.

[0701] Step 1:

[0702] When a meeting starts, the device activates the built-in or connected microphone device to collect participants' speech, and the collected audio data is converted into a digital format.

[0703] Step 2:

[0704] The device compresses and encrypts the collected voice data in real time, ensuring security and efficient transmission of the voice data.

[0705] Step 3:

[0706] The terminal transmits the encrypted voice data to the server via the network.

[0707] Step 4:

[0708] The server receives the received voice data and inputs it into a voice recognition engine, which converts the voice data into text data.

[0709] Step 5:

[0710] The server performs speaker identification analysis on the text data, where it identifies who spoke based on voice characteristics and other identifying information.

[0711] Step 6:

[0712] The server analyzes the text data using natural language processing (NLP) techniques to extract key information about the purpose of the meeting and related agenda items.

[0713] Step 7:

[0714] The server inputs the voice data into an emotion engine for emotion analysis, which estimates the participants' emotions based on voice characteristics such as intonation, tempo, and volume.

[0715] Step 8:

[0716] The server organizes the extracted information and emotional data and visualizes them in mind maps and lists, while the emotional data is displayed in graphs and color codes.

[0717] Step 9:

[0718] The server transmits the visualized information to the terminal in real time.

[0719] Step 10:

[0720] The terminal displays the received visualization data on a user interface so that the user can refer to it.

[0721] Step 11:

[0722] The server monitors the progress of the conference and generates an alert if the conference is running behind schedule or exceeds the scheduled time.

[0723] Step 12:

[0724] The server sends the generated alert to the terminal, and the user receives the alert on the terminal and adjusts the progress of the conference.

[0725] Step 13:

[0726] The server will then launch a translation engine to translate the speech into the specified languages ​​as needed, and the translated content will be sent to the device in real time.

[0727] Step 14:

[0728] The terminal displays the translated content to the user, facilitating understanding between participants who speak different languages.

[0729] Step 15:

[0730] At the end of the meeting, the server automatically generates minutes based on all collected and analyzed information.

[0731] Step 16:

[0732] The server stores the generated minutes in cloud storage for users to access later.

[0733] Step 17:

[0734] After the meeting, users can access the cloud storage to view the saved minutes and other related information.

[0735] With this specific processing flow, the system provides support at every stage of meeting management, creating an environment where meetings can proceed efficiently while taking into account the emotional state of participants.

[0736] Example 2

[0737] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0738] Although conventional meeting support systems had the ability to organize and visualize participants' comments, they were inadequate in terms of grasping participants' emotional states and supporting multiple languages. As a result, the progress of the meeting could be stalled due to emotional factors, and communication in a multilingual environment could not proceed smoothly. In addition, there were limited means for effectively monitoring the progress of the meeting and managing time. This resulted in problems that reduced the efficiency of meeting management.

[0739] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0740] In this invention, the server includes means for analyzing voice data and converting speech content into text data, means for analyzing emotional states from the voice data, means for visualizing important information and emotional data, means for translating into multiple languages, and means for monitoring progress and managing time. This makes it possible to grasp the emotional states of conference participants and efficiently manage conferences in multiple languages.

[0741] A "means for capturing participant voices" is a device and method for recording participant speech during a conference and converting it into a digital format.

[0742] "Means for compressing, encrypting and transmitting collected voice data" refers to a device and method that reduces the data size of collected voice data for efficient storage and transmission, encrypts the data to ensure security, and transmits it to a server via a network.

[0743] "Means for analyzing collected voice data and identifying speech of conference participants" refers to speech recognition technology and methods used to convert collected voice data into text and identify speakers.

[0744] The "means for converting the identified speech content into text data" refers to a method and apparatus for analyzing speech data and converting it into text format using speech recognition technology.

[0745] The "means for extracting important information from text data" refers to an apparatus and method for extracting important topics and key points in a meeting from text data using natural language processing technology.

[0746] The "means for analyzing the emotional state of participants from voice data" refers to an algorithm and method for estimating the emotional state of conference participants by analyzing the intonation, tempo, volume, etc. of the voice in the voice data.

[0747] "Means for visualizing extracted information and emotional data" refers to software and methods for visually displaying information and emotional data collected during a meeting in the form of mind maps, graphs, or lists.

[0748] A "means for translating into multiple languages" is an apparatus and method for translating text data into multiple different languages ​​in real time.

[0749] The "means for monitoring progress and managing time" refers to a device and method for monitoring the progress of a meeting and managing the meeting so that it proceeds within the scheduled time.

[0750] MODE FOR CARRYING OUT THE INVENTION

[0751] This invention is a system that supports efficient meeting management by collecting and analyzing the comments of meeting participants, and organizing and visualizing the content. By further combining this system with emotion recognition functionality, it is possible to grasp the emotional state of participants and adjust the progress of the meeting. This system is realized through cooperation between "terminals," "servers," and "users."

[0752] The device uses a microphone to collect the audio of the meeting participants. For example, a voice recording app or software on a smartphone or PC is used. The device converts the collected audio data into a digital format in real time and prepares it for transmission to the server. At this time, the audio data is compressed using "FFmpeg" and encrypted using "OpenSSL." The collected audio data is then sent to the server via the network.

[0753] The server converts the received voice data into text using speech recognition APIs such as Google Cloud Speech-to-Text and IBM Watson Speech to Text. It then uses natural language processing (NLP) libraries such as spaCy and NLTK to extract important information from the text data. The server also analyzes participants' emotional states from the voice data using emotion recognition technologies such as Microsoft Azure Emotion API and Affectiva. The analyzed emotion data is estimated based on voice characteristics such as intonation, tempo, and volume.

[0754] The server organizes the extracted information and emotion data and visualizes it using D3.js and Chart.js. This allows the information to be organized in mind maps and list formats, and the emotion data to be displayed in graphs and color codes. The visualized information is sent to the device in real time and displayed to the user. For example, the progress of a meeting can be displayed using graphs and tables so that it can be seen at a glance.

[0755] If necessary, the server translates the speech into multiple languages ​​using Google Translate API or Microsoft Translator. The translation results are also sent to the device in real time, allowing users to understand what is being said in different languages. The server monitors the progress of the meeting and uses Twilio or Slack API to notify users if the meeting goes over the scheduled time or if there are delays. Notifications are sent to the device and displayed to the user.

[0756] After the meeting, the server automatically generates minutes based on all collected and analyzed information. The documents are generated using Microsoft Word API and Google Docs API. The generated minutes are saved in cloud storage services such as Google Drive and Dropbox, allowing users to access them later.

[0757] Specific examples

[0758] Neighborhood Association Meeting

[0759] During neighborhood association meetings, devices (smartphones and PCs) collect participants' audio, compress it with FFmpeg, encrypt it with OpenSSL, and then send it to a server. The server converts the audio data into text using Google Cloud Speech-to-Text and extracts important information using spaCy. The server then analyzes the participants' emotional states using Affectiva and visualizes the emotional data using graphs and colors. After the meeting ends, the server automatically generates minutes, which users can view via cloud storage (e.g., Google Drive).

[0760] International Online Conference

[0761] In international online conferences, devices (PCs) collect participants' voices through online conference tools (e.g., Zoom), compress them with FFmpeg, encrypt them with OpenSSL, and send them to a server. The server analyzes the voice data using IBM Watson Speech to Text and converts what is being said into text. Important information is extracted using spaCy and translated using Google Translate API. The server further analyzes emotions using Microsoft Azure Emotion API and displays them on the device in real time. After the conference ends, the server generates minutes including the content of the speech and emotion data, which users can access via Dropbox.

[0762] The system provides comprehensive support at every stage of the meeting, enabling effective meeting management that takes into account the emotional state of participants. Multilingual support also enables smooth communication even in international meetings.

[0763] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0764] Step 1:

[0765] The device collects the voices of the meeting participants through a microphone. The input is the live audio from the meeting, and converts it into a digital format (e.g., WAV or MP3). Specifically, a dedicated application (e.g., audio recording software) on a smartphone or PC is used. The output digital audio data is stored in temporary memory for subsequent processing.

[0766] Step 2:

[0767] The terminal compresses the collected audio data using "FFmpeg" and encrypts it with "OpenSSL." The input is the digital audio data obtained in step 1, which is compressed to reduce the data size and encrypted to ensure data security. Specifically, the terminal executes the FFmpeg compression command and the OpenSSL encryption command. The output is compressed and encrypted audio data, which is sent to the server via the network.

[0768] Step 3:

[0769] The server receives the audio data sent from the device, decrypts it, and decompresses it. The input is compressed and encrypted audio data, and OpenSSL and FFmpeg are used to decompress it and return it to plain text. Specifically, the server executes the OpenSSL decryption command and the FFmpeg decompression command. The output is the original digital audio data.

[0770] Step 4:

[0771] The server converts the voice data into text data using speech recognition technology. The input is the original digital voice data, which is converted into text using APIs such as "Google Cloud Speech-to-Text" and "IBM Watson Speech to Text." The API is called to analyze the voice data and convert the utterances into text format. The output is text data for each speaker.

[0772] Step 5:

[0773] The server uses natural language processing (NLP) technology to extract important information from text data. The input is text data obtained through speech recognition, and information extraction is performed using tools such as "spaCy" and "NLTK." Specifically, the text data is input into an NLP library, and related keywords and topics are extracted. The output is text data extracted as important information.

[0774] Step 6:

[0775] The server analyzes the participants' emotional states from the voice data. The input is the voice data obtained in step 3, and emotion analysis is performed using tools such as the Microsoft Azure Emotion API. Specifically, the voice data is passed to the emotion engine, which estimates the emotional state (e.g., joy, anger, sadness). The output is data indicating the emotional state of each participant.

[0776] Step 7:

[0777] The server visualizes the extracted information and emotion data. The input is the important information obtained in step 5 and the emotion data obtained in step 6, and these are displayed visually using "D3.js" and "Chart.js." Specifically, it generates mind maps and graphs and sends them to the user's device in real time. The output is the visualized data.

[0778] Step 8:

[0779] The server translates the speech into multiple languages ​​as needed. The input is text data obtained through speech recognition, which is translated using the Google Translate API or Microsoft Translator. Specifically, it calls a translation engine to convert the text into another language, and the output is the translated text data.

[0780] Step 9:

[0781] The server monitors the progress of the meeting and manages the time. The input is meeting data updated in real time, and a notification is generated if the scheduled time is exceeded or progress is delayed. For example, a notification message is created using "Twilio" or "Slack API" and sent to the user's device. The output is a notification to the user.

[0782] Step 10:

[0783] After the meeting ends, the server automatically generates minutes based on all collected and analyzed information. The input is the voice data collected during the meeting, analyzed text data, and emotion data, and minutes are generated using the Microsoft Word API or Google Docs API. Specifically, the data is passed to the minutes generation engine to generate a document, which is then saved in cloud storage. The output is the generated minutes.

[0784] These processing steps effectively organize the remarks of the conference participants and realize efficient conference proceedings that take into account their emotional states.

[0785] (Application example 2)

[0786] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0787] While conventional meeting support systems can collect participants' voices, convert what is being said into text, and extract and visualize important information, they are unable to grasp the participants' emotional state and adjust the progress of the meeting. Furthermore, when making advertising presentations, there is no mechanism for providing real-time feedback on participants' reactions and emotions, making it difficult to effectively present advertising campaigns.

[0788] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0789] In this invention, the server includes a means for analyzing the emotional states of conference participants from their comments, a means for visualizing the analyzed emotional states in real time, and a means for automatically generating a record including the comments and emotional states after the conference ends. This allows participants' reactions and emotional states to be grasped in real time, enabling effective progress and feedback of the advertising presentation.

[0790] The "means for collecting the voices of the conference participants" refers to a device or system that converts the voices of the conference participants into a digital format in real time via a microphone and collects them as data.

[0791] "Means for analyzing collected voice data and identifying what is being said by conference participants" refers to algorithms or software for identifying conference participants and specifying what is being said based on collected voice data.

[0792] The "means for converting the identified speech content into text data" is a process for converting the identified speech content into text data in sentence format using speech recognition technology.

[0793] The "means for extracting important information from text data" is a system that uses natural language processing technology to extract key topics and important information discussed in meetings and presentations from text data.

[0794] "Means for visualizing extracted information" refers to software or tools that visually display extracted important information and data in the form of graphs, lists, mind maps, etc.

[0795] "Means for analyzing emotional state from speech content" refers to algorithms or emotion recognition engines for estimating emotional state based on characteristics of the speaker's voice data, such as intonation, tempo, and volume.

[0796] The "means for visualizing analyzed emotional states in real time" is a system for displaying emotional data estimated by an emotion recognition engine in real time using graphs and colors.

[0797] "Means for automatically generating records including remarks and emotional states after a meeting" refers to software that automatically creates and saves minutes and records that integrate remarks and emotional states based on information collected and analyzed after a meeting.

[0798] The present invention is a conference support system used in advertising presentations that analyzes and visualizes participants' comments and emotional states to effectively support the progress of the conference. Specific embodiments are described below.

[0799] System Configuration

[0800] 1. Collection of voice data (device)

[0801] The terminals, typically digital devices such as smartphones and personal computers, collect the voices of conference participants in real time through microphones and convert them into digital format. This data is then compressed, encrypted, and sent to the server.

[0802] 2. Analysis of voice data (server)

[0803] The server receives the collected voice data and converts it into text using speech recognition technology, identifies the speaker, analyzes the speech, and extracts important information. Natural language processing (NLP) technology is used here.

[0804] 3. Emotion Analysis (Server)

[0805] The server uses an emotion recognition engine to estimate participants' emotional states from the voice data. The emotion recognition engine uses an algorithm to analyze voice intonation, tempo, volume, etc. to estimate emotions.

[0806] 4. Real-time visualization (server)

[0807] The server organizes the extracted text information and emotional data and visualizes it in real time in graph and list format. The emotional data is displayed using colors and graphs, allowing participants' emotional states to be understood at a glance. The visualized information is sent to the device in real time and displayed to the user.

[0808] 5. Generation of minutes (server)

[0809] After the meeting, the server automatically generates minutes based on all collected and analyzed information, which are then saved in cloud storage for users to access later.

[0810] Usage example

[0811] Specific examples

[0812] During an advertising presentation meeting, when an advertising agency is presenting a new advertising campaign, the comments made by participants and the emotional state of those comments are analyzed and visualized in real time. For example, it is possible to instantly understand whether participants' reactions are positive or negative during the presentation. This allows the presenter to flexibly adjust the progress of the meeting and deliver an effective presentation.

[0813] Prompt Sentence Examples

[0814] Create a system that analyzes participants' comments and emotions in real time during advertising campaign meetings and visualizes the results. Include a function to grasp participants' emotional state and automatically generate and save minutes after the meeting.

[0815] Hardware and Software Configuration

[0816] Microphone: for collecting audio data

[0817] Personal computer or smartphone: for analysis and data storage

[0818] Python: a programming language

[0819] SpeechRecognition: A speech recognition library

[0820] transformers: Hugging Face's NLP library for emotion analysis

[0821] Google Speech API: Convert speech to text

[0822] Using this hardware and software, it is possible to collect and analyze voice data, analyze emotional states, visualize information, and generate meeting minutes.

[0823] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0824] Step 1: Collecting voice data (device)

[0825] The terminal collects the voices of the conference participants in real time through a microphone and converts them into a digital format. The terminal then compresses and encrypts the collected voice data and sends it to the server. The input is raw voice, and the output is digitized encrypted voice data.

[0826] Step 2: Analyzing the voice data (server)

[0827] The server converts the received voice data into text data using speech recognition technology. This process uses the Google Speech API to analyze the voice data and convert it into text. It also identifies the speaker and identifies each utterance. The input is encrypted voice data, and the output is the identified utterance text data.

[0828] Step 3: Extracting important information (server)

[0829] The server extracts important information from the identified text data using natural language processing (NLP) techniques. This process uses the Hugging Face transformers library. The input is the text data, and the output is the extracted important information.

[0830] Step 4: Sentiment Analysis (Server)

[0831] The server analyzes the emotional state of the identified speech using an emotion recognition engine. This process estimates emotions based on voice intonation, tempo, volume, etc. The input is text data, and the output is emotion analysis data.

[0832] Step 5: Real-time visualization (server)

[0833] The server visualizes the extracted information and emotional data in graphs and lists, allowing participants' emotional states to be visually displayed in real time. The input is important information and emotional analysis data, and the output is visualized graphs and lists.

[0834] Step 6: Real-time display (terminal)

[0835] The terminal displays the visualized data sent from the server in real time, allowing users to instantly check the emotional state and important information during a meeting. The input is the visualized data, and the output is the displayed graph or list.

[0836] Step 7: Generate minutes (server)

[0837] After the meeting ends, the server automatically generates minutes based on all collected and analyzed information. These minutes include the content of remarks and the emotional state of the participants. The generated minutes are saved in cloud storage and can be accessed by users later. The input is important information and emotional analysis data, and the output is the minutes data.

[0838] Step 8: Saving and Accessing the Minutes (Devices and Users)

[0839] Users can access the cloud storage through their devices and check the generated minutes. The input is the minutes data on the cloud, and the output is the minutes that users can view.

[0840] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0841] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0842] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0843] [Third embodiment]

[0844] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0845] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[0846] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0847] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0848] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0849] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0850] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0851] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0852] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0853] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0854] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0855] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0856] This invention is a system that supports efficient conference management by collecting and analyzing the comments of conference participants, and organizing and visualizing the content. This system is realized through cooperation between "terminals," "servers," and "users."

[0857] System Configuration

[0858] 1. Collection of voice data (device)

[0859] The terminals collect the voices of the conference participants through microphones. The collected voice data is recorded in real time and sent to a server. The terminals are usually devices such as smartphones or PCs.

[0860] 2. Sending audio data (terminal)

[0861] The device transmits the collected voice data over the network to a server, where it is compressed and encrypted to ensure security and efficiency.

[0862] 3. Analysis of voice data (server)

[0863] The server uses voice recognition technology to analyze the received voice data, converting it into text data and identifying the speaker.

[0864] 4. Extraction of comment content (server)

[0865] The server uses natural language processing (NLP) techniques to extract key information from the text data, such as the purpose of the meeting and related agenda items.

[0866] 5. Content organization and visualization (server)

[0867] The server organizes the extracted information, summarizes important points, and visualizes them in the form of mind maps or lists. The visualized information is sent to the terminal in real time and displayed to the user.

[0868] 6. Multilingual Translation (Server)

[0869] The server uses a multilingual translation engine to translate speech into multiple languages ​​as needed, and the translation results are also displayed on the device in real time.

[0870] 7. Time Management and Notification (Server)

[0871] The server monitors the progress of the conference and notifies the user if the scheduled time is exceeded or if the progress is delayed. This notification is sent to the terminal and displayed to the user.

[0872] 8. Generating minutes (server)

[0873] After the meeting, the server automatically generates minutes based on the organized and visualized information. The minutes are then saved in cloud storage so that users can access them later.

[0874] Specific examples

[0875] Example 1: Neighborhood Association Meeting

[0876] During neighborhood association meetings, devices (smartphones and PCs) collect participants' voices and send them to a server. The server analyzes the voice data, converts comments into text, and extracts and visualizes information related to important agenda items. The extracted information is displayed in real time on the device in mind map format, helping to guide the meeting. After the meeting, the server automatically generates minutes, which users can access and check by accessing cloud storage.

[0877] Example 2: International online conference

[0878] In international online conferences, devices (PCs) collect participants' voices through online conference tools and send them to a server. The server analyzes the voice data, identifies the speaker, and converts what is being said into text. It also translates what is being said into multiple languages ​​and displays the translated content on the device in real time. The server monitors the progress and notifies users in the event of delays or exceeding the scheduled time. After the conference ends, minutes are automatically generated in multiple languages, and users can access them via cloud storage.

[0879] With the system configuration and specific example described above, the present invention can support the efficient progress of a conference and information sharing among participants, and ensure information transparency.

[0880] The processing flow will be explained below.

[0881] Step 1:

[0882] When a meeting starts, the device activates the built-in or connected microphone device to collect participants' speech, and the collected audio data is converted into a digital format.

[0883] Step 2:

[0884] The device compresses and encrypts the collected voice data in real time, ensuring security and efficient transmission of the voice data.

[0885] Step 3:

[0886] The terminal transmits the encrypted voice data to the server via the network.

[0887] Step 4:

[0888] The server receives the received voice data and inputs it into the analysis engine, which uses voice recognition technology to convert the voice data into text data.

[0889] Step 5:

[0890] The server performs speaker identification analysis on the text data, where it identifies who spoke based on voice characteristics and other identifying information.

[0891] Step 6:

[0892] The server analyzes the text data using natural language processing (NLP) techniques to extract key information about the purpose of the meeting and related agenda items.

[0893] Step 7:

[0894] The server visualizes the extracted information in the form of a mind map or list, and this visualization data is sent to the device in real time.

[0895] Step 8:

[0896] The terminal displays the received visualization data on a user interface so that the user can refer to it.

[0897] Step 9:

[0898] The server monitors the progress of the conference and generates an alert if the conference is running behind schedule or exceeds the scheduled time.

[0899] Step 10:

[0900] The server sends the generated alert to the terminal, and the user receives the alert on the terminal and adjusts the progress of the conference.

[0901] Step 11:

[0902] The server will then launch a translation engine to translate the speech into the specified languages ​​as needed, and the translated content will be sent to the device in real time.

[0903] Step 12:

[0904] The terminal displays the translated content to the user, facilitating understanding between participants who speak different languages.

[0905] Step 13:

[0906] At the end of the meeting, the server automatically generates minutes based on all collected and analyzed information.

[0907] Step 14:

[0908] The server stores the generated minutes in cloud storage for users to access later.

[0909] Step 15:

[0910] After the meeting, users can access the cloud storage to view the saved minutes and other related information.

[0911] With this specific processing flow, the system provides support at every stage of meeting management, creating an environment in which all participants can collaborate efficiently.

[0912] Example 1

[0913] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0914] To ensure efficient meeting management and information transparency, a system is needed that can quickly and accurately collect and analyze participants' comments and organize and visualize related information. However, current systems have fragmented processes from audio collection to analysis and visualization, and lack real-time capabilities. Furthermore, they lack multilingual support and meeting progress management, making them difficult to use in international or multilingual meetings. Furthermore, the automatic generation of meeting minutes and their storage in cloud storage are not very practical.

[0915] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0916] In this invention, the server includes means for collecting the voices of conference participants, means for compressing and encrypting the collected voice data and transmitting it, means for analyzing the transmitted voice data and converting the voices into text data, natural language processing means for extracting important information from the text data, means for organizing the extracted information and visualizing it in the form of a mind map or list, and means for displaying the visualized information in real time. This enables real-time support for the progress of the conference, multilingual support, progress monitoring, automatic generation of minutes, and storage in cloud storage.

[0917] "Conference Participant" means a person who attends and speaks at a conference.

[0918] "Audio" refers to the words and voices of conference participants.

[0919] "Audio Data" means collected audio in digital form.

[0920] "Means for collecting" refers to a method or device for capturing sound using a microphone.

[0921] "Compression" refers to the technique of reducing the size of data.

[0922] "Encryption" refers to the technology of converting data into a form that is not easily understandable by third parties.

[0923] "Transmitting means" refers to a method or device for transferring data to another device over a network.

[0924] "Analysis" refers to the process of examining data in detail to understand its structure and content.

[0925] "Speech recognition" refers to the technology of analyzing voice data and converting it into text data.

[0926] "Text data" refers to data that expresses information using characters.

[0927] "Natural language processing" refers to the technology that allows computers to understand and process the language that humans use naturally.

[0928] "Extraction" refers to the process of extracting the desired information from the data.

[0929] "Organizing information" refers to classifying acquired information, making it relevant, and arranging it in a form that is easy to understand.

[0930] "Visualization" refers to displaying information graphically to make it easier to understand visually.

[0931] A "mind map" is a diagram that shows related items radiating out from a central concept.

[0932] "List format" refers to a format in which information is organized in bullet points and displayed in an orderly manner.

[0933] "Real-time" refers to reporting current events as they occur with almost no delay.

[0934] "Display means" refers to a method or device for visually displaying information using a screen, display, etc.

[0935] "Multilingual translation" refers to the technology of converting one language into another.

[0936] "Progress monitoring" refers to the process of overseeing and managing the progress of a meeting.

[0937] "Notification" means sending a message informing you about a particular condition or event.

[0938] "Minutes" refers to a document that records the discussions and decisions made at a meeting.

[0939] "Cloud storage" refers to online storage services that store and make data accessible over the internet.

[0940] The present invention is a system that supports efficient conference management by collecting, analyzing, organizing, and visualizing the comments of conference participants. Details of how to implement this system are described below.

[0941] This system includes a "terminal" used by a user, a "server" that performs analysis and management, and a "user" who is the user of the system.

[0942] Hardware and software used

[0943] The terminals used are devices with the ability to collect audio, such as smartphones, tablets, and PCs, with a meeting collection application installed.

[0944] The server incorporates a speech recognition engine (e.g., Google Cloud Speech-to-Text), a natural language processing engine (e.g., spaCy, NLTK), and a multilingual translation engine (e.g., Google Translate API).

[0945] System configuration

[0946] 1. Collection of audio data

[0947] The terminal uses a microphone to collect the voices of the conference participants. At the start of the conference, the user launches the conference collection application and presses the "Start Conference" button. This operation causes the terminal to collect voices in real time and record the collected voice data.

[0948] 2. Sending audio data

[0949] The collected voice data is compressed and encrypted on the device, ensuring security and efficiency, and then transmitted over the network to a server.

[0950] 3. Analysis of audio data

[0951] The server receives the voice data sent from the device and analyzes it using a voice recognition engine. It converts the voice into text data and identifies the speaker. Specifically, it extracts features from the patterns of the voice data and converts the speech into text.

[0952] 4. Extraction of speech content

[0953] The server then performs natural language processing on the text data to extract key information and identify key agenda items and action points for the meeting.

[0954] 5. Organizing and visualizing information

[0955] The extracted information is organized on the server and visualized in mind maps or lists. For example, subtopics and related items for each agenda item are organized and displayed in a format that is easy for users to understand.

[0956] 6. Multilingual translation and notification management

[0957] If necessary, the server translates the speech into other languages ​​using a multilingual translation engine. The translated information is displayed on the user's device in real time. In addition, the server monitors the progress of the meeting and notifies the device if the meeting goes over the scheduled time or if the progress is behind schedule.

[0958] 7. Generate and save minutes

[0959] After the meeting, the server automatically generates minutes based on the organized and visualized information. The minutes are then saved in cloud storage, allowing users to access them later.

[0960] Specific examples

[0961] Example 1: Neighborhood Association Meeting

[0962] During neighborhood association meetings, users use their smartphones to collect audio. During the meeting, the device sends the collected audio data to a server, which converts it into text in real time and identifies and visualizes important topics. After the meeting, the automatically generated minutes can be viewed from cloud storage.

[0963] Example 2: International online conference

[0964] In international online conferences, participants use their computers to collect audio and send it to a server via an online conference tool. The server analyzes the audio data and translates what is being said into multiple languages. The translation results are displayed on the device in real time, and progress management and notification functions are also supported. After the conference ends, multilingual minutes can be viewed from cloud storage.

[0965] Prompt Sentence Examples

[0966] "Please explain a program that creates real-time, multilingual minutes for international conferences."

[0967] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0968] Step 1:

[0969] The user starts up the terminal, opens the conference collection application, and presses the "Start Conference" button. This starts the collection of audio data. The user's operation leads to the start of audio capture using the microphone.

[0970] Input: User's instruction to start a meeting

[0971] Output: Trigger to start collection

[0972] Step 2:

[0973] The device collects the voices of the meeting participants in real time through a microphone, and each participant's speech is recorded as a separate audio stream.

[0974] Input: Conference room audio

[0975] Output: Raw audio data

[0976] Step 3:

[0977] The device compresses the collected audio data and then encrypts it, which is then sent to a server using a secure communication protocol.

[0978] Input: Raw audio data

[0979] Output: Encrypted audio data

[0980] Step 4:

[0981] The server receives the voice data sent from the terminal and temporarily stores the received data in a buffer.

[0982] Input: Encrypted audio data

[0983] Output: Buffered audio data

[0984] Step 5:

[0985] The server uses a speech recognition engine to analyze the voice data and convert it into text data. Specifically, it extracts features from the voice pattern, analyzes them, and converts them into text.

[0986] Input: Buffered audio data

[0987] Output: Text data

[0988] Step 6:

[0989] The server uses natural language processing on the text data to extract key information, a process that identifies key topics and action points for the meeting.

[0990] Input: Text data

[0991] Output: Important information (extracted agenda items, action points, etc.)

[0992] Step 7:

[0993] The server organizes and visualizes the extracted information in the form of a mind map or list. Specifically, it arranges related information for each agenda item in a radial pattern, making it visually easy to understand.

[0994] Input: Important Information

[0995] Output: Visualized data

[0996] Step 8:

[0997] The server transmits the visualized information to the terminal in real time, and the terminal displays this information and provides it to the user in a timely manner.

[0998] Input: Visualization data

[0999] Output: Information displayed on the terminal

[1000] Step 9:

[1001] The server uses a multilingual translation engine to translate the speech into other languages ​​as needed, and the translated information is also displayed on the device in real time.

[1002] Input: Important Information

[1003] Output: Translated information

[1004] Step 10:

[1005] The server monitors the progress of the conference and sends notifications to the terminals when the scheduled time has passed or when the progress is delayed. This notification is used to adjust the progress.

[1006] Input: Meeting progress data

[1007] Output: Notification of overdue time

[1008] Step 11:

[1009] After the meeting, the server automatically generates minutes based on the organized and visualized information, and the generated minutes are saved in cloud storage.

[1010] Input: Organized and visualized information

[1011] Output: Meeting minutes

[1012] Step 12:

[1013] Users can access cloud storage and check and download the generated minutes, making it easy to review them after the meeting.

[1014] Input: Meeting minutes stored in cloud storage

[1015] Output: User-accessible transcript

[1016] (Application example 1)

[1017] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1018] Ensuring safety and streamlining operations are important issues in modern factory environments. However, addressing these challenges requires fast and accurate communication and information sharing. However, it is difficult to quickly generate and display important information and guidelines during real-time meetings. Furthermore, language barriers exist in multilingual meetings, making it difficult for all participants to share the same information. There is a need to solve these issues and improve the efficiency of meeting management within factories and information transparency.

[1019] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1020] In this invention, the server includes means for collecting the voices of conference participants, means for analyzing the collected voice data and identifying what the conference participants are saying, means for converting the identified content of the speech into text data, means for extracting important information from the text data, means for visualizing the extracted information, means for displaying the extracted information on a factory work robot, means for automatically generating a safety manual and efficiency guidelines for the conference in real time, and means for displaying the generated manuals and guidelines using smart glasses or a head-mounted display. This enables quick and accurate information sharing in factory meetings, ensuring safety and improving work efficiency.

[1021] A "conference participant" is someone who actually attends a conference and takes part in speeches and discussions.

[1022] "Audio data" refers to information that has been collected from the voices of conference participants and converted into digital format.

[1023] "Analysis" is the process of processing collected audio data using algorithms to extract useful information.

[1024] "Speech identification" refers to recognizing the content of speech made by each individual participant in a conference.

[1025] "Text data" is digital data that has been converted from voice data into character information.

[1026] "Important information" is information that is particularly significant to the progress of the meeting and has a direct impact on decision-making and discussion.

[1027] "Visualization" refers to visually displaying extracted information in the form of charts, graphs, etc., to make it easier to understand.

[1028] A "factory robot" is a mechanical device designed to perform automated tasks in a factory.

[1029] "Real-time" means that information is collected and processed with little to no delay, providing immediate results.

[1030] A "safety manual" is a guide to ensuring safety at work sites.

[1031] "Efficiency guidelines" are documents that contain specific procedures and instructions for improving work efficiency.

[1032] "Smart glasses" are a wearable eyeglass-type device equipped with a display function.

[1033] A "head-mounted display" is a display device that is worn on the head and provides visual information.

[1034] This invention is a system that supports the efficient management of factory meetings by collecting and analyzing the comments of meeting participants, and organizing and visualizing the content. This system is realized through cooperation between "terminals," "servers," and "users."

[1035] System Configuration

[1036] 1. Collection of voice data (device)

[1037] The terminal collects the voices of the conference participants through a microphone. The collected voice data is recorded in real time and sent to a server. The terminal is typically a device such as smart glasses, a PC, or a factory robot.

[1038] 2. Sending audio data (terminal)

[1039] The device transmits the collected voice data over the network to a server, where it is compressed and encrypted to ensure security and efficiency.

[1040] 3. Analysis of voice data (server)

[1041] The server uses speech recognition technology to analyze the received voice data. Specifically, it converts the collected voice data into text data using the speech_recognition library. High-accuracy speech recognition is also possible by using Google's speech recognition API.

[1042] 4. Extraction of comment content (server)

[1043] The server uses natural language processing (NLP) techniques to extract important information from text data, leveraging the SpaCy library to extract important noun phrases and key phrases from the text.

[1044] 5. Content organization and visualization (server)

[1045] The server organizes the extracted information and visualizes the important content. It uses NetworkX and Matplotlib to display key phrases in graph form, allowing users to grasp the progress of the meeting and important topics at a glance.

[1046] 6. Multilingual Translation (Server)

[1047] The server uses the googletrans library to translate speech into multiple languages ​​as needed, and the translation results are displayed on the device in real time.

[1048] 7. Time Management and Notification (Server)

[1049] The server monitors the progress of the conference and notifies the user if the scheduled time is exceeded or if the progress is delayed. This notification is sent to the terminal and displayed to the user.

[1050] 8. Multi-purpose display (server)

[1051] The server automatically generates safety manuals and efficiency guidelines for meetings in real time, and displays the generated manuals and guidelines using smart glasses or a head-mounted display.

[1052] Specific examples

[1053] During meetings on safety and efficiency within a factory, devices (smart glasses or head-mounted displays) collect the voices of participants and send them to a server. The server analyzes the voice data, converts comments into text, and extracts and visualizes information related to important topics. Furthermore, it automatically generates safety manuals and efficiency guidelines in real time, and visualizes this data for participants via the smart glasses or head-mounted displays.

[1054] The server also translates multiple languages ​​as the meeting progresses, providing information in a format that is easy for all participants to understand, facilitating smooth information sharing during meetings within the factory and ensuring that important information on safety and other matters is communicated promptly.

[1055] Example prompts to input to the generative AI model:

[1056] I am trying to build a "meeting support system" for use in factory meetings. This system collects meeting audio in real time, extracts important keywords and phrases, and supports efficient meeting management. Please refer to the code below and help with additional improvements and enhancements.

[1057] Code fragment:

[1058] def record_audio(duration=60):

[1059] recognizer = sr.Recognizer()

[1060] mic = sr.Microphone()

[1061] with mic as source:

[1062] print("Meeting recording begins...")

[1063] audio = recognizer.record(source, duration=duration)

[1064] print("Recording complete")

[1065] return audio

[1066] Based on this, please give me some advice.

[1067] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1068] Step 1:

[1069] The terminal collects the voices of the conference participants. Specifically, it records the voices through a microphone and captures the voice data in real time. The input is the voices of the conference participants, and the output is the collected voice data.

[1070] Step 2:

[1071] The terminal sends the collected voice data to the server. When sending the data, the voice data is compressed and encrypted. The voice data is input, and the compressed and encrypted voice data is sent to the server as output.

[1072] Step 3:

[1073] The server analyzes the received voice data. Specifically, it converts the voice data into text data using the speech_recognition library. The input is compressed and encrypted voice data, and the output is text data.

[1074] Step 4:

[1075] The server extracts important information from the text data. Specifically, it uses the SpaCy library to extract noun phrases and key phrases from the text. The input is the text data, and the output is the extracted noun phrases and key phrases.

[1076] Step 5:

[1077] The server visualizes the extracted information in a graph format using NetworkX and Matplotlib. The input is the extracted key phrases, and the output is the visualized graph.

[1078] Step 6:

[1079] The server translates the speech into multiple languages ​​as needed. Specifically, it uses the GoogleTrans library for translation. The input is text data, and the output is translated text data.

[1080] Step 7:

[1081] The server monitors the progress of the conference and manages the time. It notifies participants if the conference is delayed or exceeds the scheduled time. The input is the start time of the conference, and the output is notifications according to the progress of the conference.

[1082] Step 8:

[1083] The server automatically generates safety manuals and efficiency guidelines and displays them on the terminal. The generated manuals and guidelines are displayed in real time through smart glasses or a head-mounted display. The input is extracted information and key phrases, and the generated safety manuals and efficiency guidelines are displayed as output.

[1084] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1085] This invention is a system that collects and analyzes the comments of meeting participants, organizes and visualizes the content, and supports efficient meeting management. By combining this with emotion recognition functionality, it is possible to grasp the emotional state of participants and adjust the progress of the meeting. This system is realized through cooperation between "terminals," "servers," and "users."

[1086] System Configuration

[1087] 1. Collection of voice data (device)

[1088] The terminal collects the voices of the participants through a microphone. The collected voice data is converted into a digital format in real time and sent to a server. The terminal is usually a device such as a smartphone or PC.

[1089] 2. Sending audio data (terminal)

[1090] The device compresses and encrypts the collected audio data and sends it to a server over the network, ensuring data security and efficient transfer.

[1091] 3. Analysis of voice data (server)

[1092] The server uses speech recognition technology to analyze the received voice data, converts the voice data into text data, identifies the speaker, and then uses natural language processing (NLP) technology to extract important information from the text data.

[1093] 4. Emotion Analysis (Server)

[1094] The server uses an emotion engine for emotion recognition to analyze participants' emotions from the voice data. The emotion engine uses an algorithm to estimate emotions based on voice characteristics such as intonation, tempo, and volume.

[1095] 5. Information visualization (server)

[1096] The server organizes the extracted information and emotional data and visualizes them in mind maps and lists. The emotional data is displayed in graphs and color codes, allowing the progress of the meeting to be understood at a glance. The visualized information is sent to the device in real time and displayed to the user.

[1097] 6. Multilingual Translation (Server)

[1098] The server translates the speech into multiple languages ​​as needed, and the translation results are sent to the device in real time and displayed.

[1099] 7. Time Management and Notification (Server)

[1100] The server monitors the progress of the conference and notifies the user if the conference exceeds the scheduled time or if the conference is behind schedule. The notification is sent to the terminal and displayed to the user.

[1101] 8. Generating minutes (server)

[1102] After the meeting, the server automatically generates minutes based on all collected and analyzed information, and the minutes are saved in cloud storage for users to access later.

[1103] Specific examples

[1104] Example 1: Neighborhood Association Meeting

[1105] At neighborhood association meetings, devices (smartphones and PCs) collect the voices of participants and send them to a server. The server analyzes the voice data, converts comments into text, and extracts and visualizes information related to important topics. It also uses an emotion engine to analyze the emotional state of participants and visualizes this at the same time. The emotional data is displayed in graphs and colors to show how participants are feeling when they speak, and is used to adjust the progress of the meeting. After the meeting ends, the server automatically generates minutes, which users can access and view by accessing cloud storage.

[1106] Example 2: International online conference

[1107] In international online conferences, devices (PCs) collect participants' voices through online conference tools and send them to a server. The server analyzes the voice data, identifies the speaker, and converts what was said into text. Natural language processing technology is used to extract important information, which is then translated using a multilingual translation engine. Emotions are then analyzed using an emotion engine and displayed in real time. The server monitors the progress and issues notifications in the event of delays or overruns. Minutes generated after the conference end contain emotional data along with the content of what was said, and users can access them via cloud storage.

[1108] With this specific configuration, the system provides comprehensive support at every stage of meeting management, ensuring effective progress while also taking into account the emotional states of participants.

[1109] The processing flow will be explained below.

[1110] Step 1:

[1111] When a meeting starts, the device activates the built-in or connected microphone device to collect participants' speech, and the collected audio data is converted into a digital format.

[1112] Step 2:

[1113] The device compresses and encrypts the collected voice data in real time, ensuring security and efficient transmission of the voice data.

[1114] Step 3:

[1115] The terminal transmits the encrypted voice data to the server via the network.

[1116] Step 4:

[1117] The server receives the received voice data and inputs it into a voice recognition engine, which converts the voice data into text data.

[1118] Step 5:

[1119] The server performs speaker identification analysis on the text data, where it identifies who spoke based on voice characteristics and other identifying information.

[1120] Step 6:

[1121] The server analyzes the text data using natural language processing (NLP) techniques to extract key information about the purpose of the meeting and related agenda items.

[1122] Step 7:

[1123] The server inputs the voice data into an emotion engine for emotion analysis, which estimates the participants' emotions based on voice characteristics such as intonation, tempo, and volume.

[1124] Step 8:

[1125] The server organizes the extracted information and emotional data and visualizes them in mind maps and lists, while the emotional data is displayed in graphs and color codes.

[1126] Step 9:

[1127] The server transmits the visualized information to the terminal in real time.

[1128] Step 10:

[1129] The terminal displays the received visualization data on a user interface so that the user can refer to it.

[1130] Step 11:

[1131] The server monitors the progress of the conference and generates an alert if the conference is running behind schedule or exceeds the scheduled time.

[1132] Step 12:

[1133] The server sends the generated alert to the terminal, and the user receives the alert on the terminal and adjusts the progress of the conference.

[1134] Step 13:

[1135] The server will then launch a translation engine to translate the speech into the specified languages ​​as needed, and the translated content will be sent to the device in real time.

[1136] Step 14:

[1137] The terminal displays the translated content to the user, facilitating understanding between participants who speak different languages.

[1138] Step 15:

[1139] At the end of the meeting, the server automatically generates minutes based on all collected and analyzed information.

[1140] Step 16:

[1141] The server stores the generated minutes in cloud storage for users to access later.

[1142] Step 17:

[1143] After the meeting, users can access the cloud storage to view the saved minutes and other related information.

[1144] With this specific processing flow, the system provides support at every stage of meeting management, creating an environment where meetings can proceed efficiently while taking into account the emotional state of participants.

[1145] Example 2

[1146] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1147] Although conventional meeting support systems had the ability to organize and visualize participants' comments, they were inadequate in terms of grasping participants' emotional states and supporting multiple languages. As a result, the progress of the meeting could be stalled due to emotional factors, and communication in a multilingual environment could not proceed smoothly. In addition, there were limited means for effectively monitoring the progress of the meeting and managing time. This resulted in problems that reduced the efficiency of meeting management.

[1148] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1149] In this invention, the server includes means for analyzing voice data and converting speech content into text data, means for analyzing emotional states from the voice data, means for visualizing important information and emotional data, means for translating into multiple languages, and means for monitoring progress and managing time. This makes it possible to grasp the emotional states of conference participants and efficiently manage conferences in multiple languages.

[1150] A "means for capturing participant voices" is a device and method for recording participant speech during a conference and converting it into a digital format.

[1151] "Means for compressing, encrypting and transmitting collected voice data" refers to a device and method that reduces the data size of collected voice data for efficient storage and transmission, encrypts the data to ensure security, and transmits it to a server via a network.

[1152] "Means for analyzing collected voice data and identifying speech of conference participants" refers to speech recognition technology and methods used to convert collected voice data into text and identify speakers.

[1153] The "means for converting the identified speech content into text data" refers to a method and apparatus for analyzing speech data and converting it into text format using speech recognition technology.

[1154] The "means for extracting important information from text data" refers to an apparatus and method for extracting important topics and key points in a meeting from text data using natural language processing technology.

[1155] The "means for analyzing the emotional state of participants from voice data" refers to an algorithm and method for estimating the emotional state of conference participants by analyzing the intonation, tempo, volume, etc. of the voice in the voice data.

[1156] "Means for visualizing extracted information and emotional data" refers to software and methods for visually displaying information and emotional data collected during a meeting in the form of mind maps, graphs, or lists.

[1157] A "means for translating into multiple languages" is an apparatus and method for translating text data into multiple different languages ​​in real time.

[1158] The "means for monitoring progress and managing time" refers to a device and method for monitoring the progress of a meeting and managing the meeting so that it proceeds within the scheduled time.

[1159] MODE FOR CARRYING OUT THE INVENTION

[1160] This invention is a system that supports efficient meeting management by collecting and analyzing the comments of meeting participants, and organizing and visualizing the content. By further combining this system with emotion recognition functionality, it is possible to grasp the emotional state of participants and adjust the progress of the meeting. This system is realized through cooperation between "terminals," "servers," and "users."

[1161] The device uses a microphone to collect the audio of the meeting participants. For example, a voice recording app or software on a smartphone or PC is used. The device converts the collected audio data into a digital format in real time and prepares it for transmission to the server. At this time, the audio data is compressed using "FFmpeg" and encrypted using "OpenSSL." The collected audio data is then sent to the server via the network.

[1162] The server converts the received voice data into text using speech recognition APIs such as Google Cloud Speech-to-Text and IBM Watson Speech to Text. It then uses natural language processing (NLP) libraries such as spaCy and NLTK to extract important information from the text data. The server also analyzes participants' emotional states from the voice data using emotion recognition technologies such as Microsoft Azure Emotion API and Affectiva. The analyzed emotion data is estimated based on voice characteristics such as intonation, tempo, and volume.

[1163] The server organizes the extracted information and emotion data and visualizes it using D3.js and Chart.js. This allows the information to be organized in mind maps and list formats, and the emotion data to be displayed in graphs and color codes. The visualized information is sent to the device in real time and displayed to the user. For example, the progress of a meeting can be displayed using graphs and tables so that it can be seen at a glance.

[1164] If necessary, the server translates the speech into multiple languages ​​using Google Translate API or Microsoft Translator. The translation results are also sent to the device in real time, allowing users to understand what is being said in different languages. The server monitors the progress of the meeting and uses Twilio or Slack API to notify users if the meeting goes over the scheduled time or if there are delays. Notifications are sent to the device and displayed to the user.

[1165] After the meeting, the server automatically generates minutes based on all collected and analyzed information. The documents are generated using Microsoft Word API and Google Docs API. The generated minutes are saved in cloud storage services such as Google Drive and Dropbox, allowing users to access them later.

[1166] Specific examples

[1167] Neighborhood Association Meeting

[1168] During neighborhood association meetings, devices (smartphones and PCs) collect participants' audio, compress it with FFmpeg, encrypt it with OpenSSL, and then send it to a server. The server converts the audio data into text using Google Cloud Speech-to-Text and extracts important information using spaCy. The server then analyzes the participants' emotional states using Affectiva and visualizes the emotional data using graphs and colors. After the meeting ends, the server automatically generates minutes, which users can view via cloud storage (e.g., Google Drive).

[1169] International Online Conference

[1170] In international online conferences, devices (PCs) collect participants' voices through online conference tools (e.g., Zoom), compress them with FFmpeg, encrypt them with OpenSSL, and send them to a server. The server analyzes the voice data using IBM Watson Speech to Text and converts what is being said into text. Important information is extracted using spaCy and translated using Google Translate API. The server further analyzes emotions using Microsoft Azure Emotion API and displays them on the device in real time. After the conference ends, the server generates minutes including the content of the speech and emotion data, which users can access via Dropbox.

[1171] The system provides comprehensive support at every stage of the meeting, enabling effective meeting management that takes into account the emotional state of participants. Multilingual support also enables smooth communication even in international meetings.

[1172] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1173] Step 1:

[1174] The device collects the voices of the meeting participants through a microphone. The input is the live audio from the meeting, and converts it into a digital format (e.g., WAV or MP3). Specifically, a dedicated application (e.g., audio recording software) on a smartphone or PC is used. The output digital audio data is stored in temporary memory for subsequent processing.

[1175] Step 2:

[1176] The terminal compresses the collected audio data using "FFmpeg" and encrypts it with "OpenSSL." The input is the digital audio data obtained in step 1, which is compressed to reduce the data size and encrypted to ensure data security. Specifically, the terminal executes the FFmpeg compression command and the OpenSSL encryption command. The output is compressed and encrypted audio data, which is sent to the server via the network.

[1177] Step 3:

[1178] The server receives the audio data sent from the device, decrypts it, and decompresses it. The input is compressed and encrypted audio data, and OpenSSL and FFmpeg are used to decompress it and return it to plain text. Specifically, the server executes the OpenSSL decryption command and the FFmpeg decompression command. The output is the original digital audio data.

[1179] Step 4:

[1180] The server converts the voice data into text data using speech recognition technology. The input is the original digital voice data, which is converted into text using APIs such as "Google Cloud Speech-to-Text" and "IBM Watson Speech to Text." The API is called to analyze the voice data and convert the utterances into text format. The output is text data for each speaker.

[1181] Step 5:

[1182] The server uses natural language processing (NLP) technology to extract important information from text data. The input is text data obtained through speech recognition, and information extraction is performed using tools such as "spaCy" and "NLTK." Specifically, the text data is input into an NLP library, and related keywords and topics are extracted. The output is text data extracted as important information.

[1183] Step 6:

[1184] The server analyzes the participants' emotional states from the voice data. The input is the voice data obtained in step 3, and emotion analysis is performed using tools such as the Microsoft Azure Emotion API. Specifically, the voice data is passed to the emotion engine, which estimates the emotional state (e.g., joy, anger, sadness). The output is data indicating the emotional state of each participant.

[1185] Step 7:

[1186] The server visualizes the extracted information and emotion data. The input is the important information obtained in step 5 and the emotion data obtained in step 6, and these are displayed visually using "D3.js" and "Chart.js." Specifically, it generates mind maps and graphs and sends them to the user's device in real time. The output is the visualized data.

[1187] Step 8:

[1188] The server translates the speech into multiple languages ​​as needed. The input is text data obtained through speech recognition, which is translated using the Google Translate API or Microsoft Translator. Specifically, it calls a translation engine to convert the text into another language, and the output is the translated text data.

[1189] Step 9:

[1190] The server monitors the progress of the meeting and manages the time. The input is meeting data updated in real time, and a notification is generated if the scheduled time is exceeded or progress is delayed. For example, a notification message is created using "Twilio" or "Slack API" and sent to the user's device. The output is a notification to the user.

[1191] Step 10:

[1192] After the meeting ends, the server automatically generates minutes based on all collected and analyzed information. The input is the voice data collected during the meeting, analyzed text data, and emotion data, and minutes are generated using the Microsoft Word API or Google Docs API. Specifically, the data is passed to the minutes generation engine to generate a document, which is then saved in cloud storage. The output is the generated minutes.

[1193] These processing steps effectively organize the remarks of the conference participants and realize efficient conference proceedings that take into account their emotional states.

[1194] (Application example 2)

[1195] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1196] While conventional meeting support systems can collect participants' voices, convert what is being said into text, and extract and visualize important information, they are unable to grasp the participants' emotional state and adjust the progress of the meeting. Furthermore, when making advertising presentations, there is no mechanism for providing real-time feedback on participants' reactions and emotions, making it difficult to effectively present advertising campaigns.

[1197] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1198] In this invention, the server includes a means for analyzing the emotional states of conference participants from their comments, a means for visualizing the analyzed emotional states in real time, and a means for automatically generating a record including the comments and emotional states after the conference ends. This allows participants' reactions and emotional states to be grasped in real time, enabling effective progress and feedback of the advertising presentation.

[1199] The "means for collecting the voices of the conference participants" refers to a device or system that converts the voices of the conference participants into a digital format in real time via a microphone and collects them as data.

[1200] "Means for analyzing collected voice data and identifying what is being said by conference participants" refers to algorithms or software for identifying conference participants and specifying what is being said based on collected voice data.

[1201] The "means for converting the identified speech content into text data" is a process for converting the identified speech content into text data in sentence format using speech recognition technology.

[1202] The "means for extracting important information from text data" is a system that uses natural language processing technology to extract key topics and important information discussed in meetings and presentations from text data.

[1203] "Means for visualizing extracted information" refers to software or tools that visually display extracted important information and data in the form of graphs, lists, mind maps, etc.

[1204] "Means for analyzing emotional state from speech content" refers to algorithms or emotion recognition engines for estimating emotional state based on characteristics of the speaker's voice data, such as intonation, tempo, and volume.

[1205] The "means for visualizing analyzed emotional states in real time" is a system for displaying emotional data estimated by an emotion recognition engine in real time using graphs and colors.

[1206] "Means for automatically generating records including remarks and emotional states after a meeting" refers to software that automatically creates and saves minutes and records that integrate remarks and emotional states based on information collected and analyzed after a meeting.

[1207] The present invention is a conference support system used in advertising presentations that analyzes and visualizes participants' comments and emotional states to effectively support the progress of the conference. Specific embodiments are described below.

[1208] System Configuration

[1209] 1. Collection of voice data (device)

[1210] The terminals, typically digital devices such as smartphones and personal computers, collect the voices of conference participants in real time through microphones and convert them into digital format. This data is then compressed, encrypted, and sent to the server.

[1211] 2. Analysis of voice data (server)

[1212] The server receives the collected voice data and converts it into text using speech recognition technology, identifies the speaker, analyzes the speech, and extracts important information. Natural language processing (NLP) technology is used here.

[1213] 3. Emotion Analysis (Server)

[1214] The server uses an emotion recognition engine to estimate participants' emotional states from the voice data. The emotion recognition engine uses an algorithm to analyze voice intonation, tempo, volume, etc. to estimate emotions.

[1215] 4. Real-time visualization (server)

[1216] The server organizes the extracted text information and emotional data and visualizes it in real time in graph and list format. The emotional data is displayed using colors and graphs, allowing participants' emotional states to be understood at a glance. The visualized information is sent to the device in real time and displayed to the user.

[1217] 5. Generation of minutes (server)

[1218] After the meeting, the server automatically generates minutes based on all collected and analyzed information, which are then saved in cloud storage for users to access later.

[1219] Usage example

[1220] Specific examples

[1221] During an advertising presentation meeting, when an advertising agency is presenting a new advertising campaign, the comments made by participants and the emotional state of those comments are analyzed and visualized in real time. For example, it is possible to instantly understand whether participants' reactions are positive or negative during the presentation. This allows the presenter to flexibly adjust the progress of the meeting and deliver an effective presentation.

[1222] Prompt Sentence Examples

[1223] Create a system that analyzes participants' comments and emotions in real time during advertising campaign meetings and visualizes the results. Include a function to grasp participants' emotional state and automatically generate and save minutes after the meeting.

[1224] Hardware and Software Configuration

[1225] Microphone: for collecting audio data

[1226] Personal computer or smartphone: for analysis and data storage

[1227] Python: a programming language

[1228] SpeechRecognition: A speech recognition library

[1229] transformers: Hugging Face's NLP library for emotion analysis

[1230] Google Speech API: Convert speech to text

[1231] Using this hardware and software, it is possible to collect and analyze voice data, analyze emotional states, visualize information, and generate meeting minutes.

[1232] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1233] Step 1: Collecting voice data (device)

[1234] The terminal collects the voices of the conference participants in real time through a microphone and converts them into a digital format. The terminal then compresses and encrypts the collected voice data and sends it to the server. The input is raw voice, and the output is digitized encrypted voice data.

[1235] Step 2: Analyzing the voice data (server)

[1236] The server converts the received voice data into text data using speech recognition technology. This process uses the Google Speech API to analyze the voice data and convert it into text. It also identifies the speaker and identifies each utterance. The input is encrypted voice data, and the output is the identified utterance text data.

[1237] Step 3: Extracting important information (server)

[1238] The server extracts important information from the identified text data using natural language processing (NLP) techniques. This process uses the Hugging Face transformers library. The input is the text data, and the output is the extracted important information.

[1239] Step 4: Sentiment Analysis (Server)

[1240] The server analyzes the emotional state of the identified speech using an emotion recognition engine. This process estimates emotions based on voice intonation, tempo, volume, etc. The input is text data, and the output is emotion analysis data.

[1241] Step 5: Real-time visualization (server)

[1242] The server visualizes the extracted information and emotional data in graphs and lists, allowing participants' emotional states to be visually displayed in real time. The input is important information and emotional analysis data, and the output is visualized graphs and lists.

[1243] Step 6: Real-time display (terminal)

[1244] The terminal displays the visualized data sent from the server in real time, allowing users to instantly check the emotional state and important information during a meeting. The input is the visualized data, and the output is the displayed graph or list.

[1245] Step 7: Generate minutes (server)

[1246] After the meeting ends, the server automatically generates minutes based on all collected and analyzed information. These minutes include the content of remarks and the emotional state of the participants. The generated minutes are saved in cloud storage and can be accessed by users later. The input is important information and emotional analysis data, and the output is the minutes data.

[1247] Step 8: Saving and Accessing the Minutes (Devices and Users)

[1248] Users can access the cloud storage through their devices and check the generated minutes. The input is the minutes data on the cloud, and the output is the minutes that users can view.

[1249] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1250] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1251] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1252] [Fourth embodiment]

[1253] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1254] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1255] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1256] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1257] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1258] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1259] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1260] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1261] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1262] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1263] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1264] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1265] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1266] This invention is a system that supports efficient conference management by collecting and analyzing the comments of conference participants, and organizing and visualizing the content. This system is realized through cooperation between "terminals," "servers," and "users."

[1267] System Configuration

[1268] 1. Collection of voice data (device)

[1269] The terminals collect the voices of the conference participants through microphones. The collected voice data is recorded in real time and sent to a server. The terminals are usually devices such as smartphones or PCs.

[1270] 2. Sending audio data (terminal)

[1271] The device transmits the collected voice data over the network to a server, where it is compressed and encrypted to ensure security and efficiency.

[1272] 3. Analysis of voice data (server)

[1273] The server uses voice recognition technology to analyze the received voice data, converting it into text data and identifying the speaker.

[1274] 4. Extraction of comment content (server)

[1275] The server uses natural language processing (NLP) techniques to extract key information from the text data, such as the purpose of the meeting and related agenda items.

[1276] 5. Content organization and visualization (server)

[1277] The server organizes the extracted information, summarizes important points, and visualizes them in the form of mind maps or lists. The visualized information is sent to the terminal in real time and displayed to the user.

[1278] 6. Multilingual Translation (Server)

[1279] The server uses a multilingual translation engine to translate speech into multiple languages ​​as needed, and the translation results are also displayed on the device in real time.

[1280] 7. Time Management and Notification (Server)

[1281] The server monitors the progress of the conference and notifies the user if the scheduled time is exceeded or if the progress is delayed. This notification is sent to the terminal and displayed to the user.

[1282] 8. Generating minutes (server)

[1283] After the meeting, the server automatically generates minutes based on the organized and visualized information. The minutes are then saved in cloud storage so that users can access them later.

[1284] Specific examples

[1285] Example 1: Neighborhood Association Meeting

[1286] During neighborhood association meetings, devices (smartphones and PCs) collect participants' voices and send them to a server. The server analyzes the voice data, converts comments into text, and extracts and visualizes information related to important agenda items. The extracted information is displayed in real time on the device in mind map format, helping to guide the meeting. After the meeting, the server automatically generates minutes, which users can access and check by accessing cloud storage.

[1287] Example 2: International online conference

[1288] In international online conferences, devices (PCs) collect participants' voices through online conference tools and send them to a server. The server analyzes the voice data, identifies the speaker, and converts what is being said into text. It also translates what is being said into multiple languages ​​and displays the translated content on the device in real time. The server monitors the progress and notifies users in the event of delays or exceeding the scheduled time. After the conference ends, minutes are automatically generated in multiple languages, and users can access them via cloud storage.

[1289] With the system configuration and specific example described above, the present invention can support the efficient progress of a conference and information sharing among participants, and ensure information transparency.

[1290] The processing flow will be explained below.

[1291] Step 1:

[1292] When a meeting starts, the device activates the built-in or connected microphone device to collect participants' speech, and the collected audio data is converted into a digital format.

[1293] Step 2:

[1294] The device compresses and encrypts the collected voice data in real time, ensuring security and efficient transmission of the voice data.

[1295] Step 3:

[1296] The terminal transmits the encrypted voice data to the server via the network.

[1297] Step 4:

[1298] The server receives the received voice data and inputs it into the analysis engine, which uses voice recognition technology to convert the voice data into text data.

[1299] Step 5:

[1300] The server performs speaker identification analysis on the text data, where it identifies who spoke based on voice characteristics and other identifying information.

[1301] Step 6:

[1302] The server analyzes the text data using natural language processing (NLP) techniques to extract key information about the purpose of the meeting and related agenda items.

[1303] Step 7:

[1304] The server visualizes the extracted information in the form of a mind map or list, and this visualization data is sent to the device in real time.

[1305] Step 8:

[1306] The terminal displays the received visualization data on a user interface so that the user can refer to it.

[1307] Step 9:

[1308] The server monitors the progress of the conference and generates an alert if the conference is running behind schedule or exceeds the scheduled time.

[1309] Step 10:

[1310] The server sends the generated alert to the terminal, and the user receives the alert on the terminal and adjusts the progress of the conference.

[1311] Step 11:

[1312] The server will then launch a translation engine to translate the speech into the specified languages ​​as needed, and the translated content will be sent to the device in real time.

[1313] Step 12:

[1314] The terminal displays the translated content to the user, facilitating understanding between participants who speak different languages.

[1315] Step 13:

[1316] At the end of the meeting, the server automatically generates minutes based on all collected and analyzed information.

[1317] Step 14:

[1318] The server stores the generated minutes in cloud storage for users to access later.

[1319] Step 15:

[1320] After the meeting, users can access the cloud storage to view the saved minutes and other related information.

[1321] With this specific processing flow, the system provides support at every stage of meeting management, creating an environment in which all participants can collaborate efficiently.

[1322] Example 1

[1323] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1324] To ensure efficient meeting management and information transparency, a system is needed that can quickly and accurately collect and analyze participants' comments and organize and visualize related information. However, current systems have fragmented processes from audio collection to analysis and visualization, and lack real-time capabilities. Furthermore, they lack multilingual support and meeting progress management, making them difficult to use in international or multilingual meetings. Furthermore, the automatic generation of meeting minutes and their storage in cloud storage are not very practical.

[1325] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1326] In this invention, the server includes means for collecting the voices of conference participants, means for compressing and encrypting the collected voice data and transmitting it, means for analyzing the transmitted voice data and converting the voices into text data, natural language processing means for extracting important information from the text data, means for organizing the extracted information and visualizing it in the form of a mind map or list, and means for displaying the visualized information in real time. This enables real-time support for the progress of the conference, multilingual support, progress monitoring, automatic generation of minutes, and storage in cloud storage.

[1327] "Conference Participant" means a person who attends and speaks at a conference.

[1328] "Audio" refers to the words and voices of conference participants.

[1329] "Audio Data" means collected audio in digital form.

[1330] "Means for collecting" refers to a method or device for capturing sound using a microphone.

[1331] "Compression" refers to the technique of reducing the size of data.

[1332] "Encryption" refers to the technology of converting data into a form that is not easily understandable by third parties.

[1333] "Transmitting means" refers to a method or device for transferring data to another device over a network.

[1334] "Analysis" refers to the process of examining data in detail to understand its structure and content.

[1335] "Speech recognition" refers to the technology of analyzing voice data and converting it into text data.

[1336] "Text data" refers to data that expresses information using characters.

[1337] "Natural language processing" refers to the technology that allows computers to understand and process the language that humans use naturally.

[1338] "Extraction" refers to the process of extracting the desired information from the data.

[1339] "Organizing information" refers to classifying acquired information, making it relevant, and arranging it in a form that is easy to understand.

[1340] "Visualization" refers to displaying information graphically to make it easier to understand visually.

[1341] A "mind map" is a diagram that shows related items radiating out from a central concept.

[1342] "List format" refers to a format in which information is organized in bullet points and displayed in an orderly manner.

[1343] "Real-time" refers to reporting current events as they occur with almost no delay.

[1344] "Display means" refers to a method or device for visually displaying information using a screen, display, etc.

[1345] "Multilingual translation" refers to the technology of converting one language into another.

[1346] "Progress monitoring" refers to the process of overseeing and managing the progress of a meeting.

[1347] "Notification" means sending a message informing you about a particular condition or event.

[1348] "Minutes" refers to a document that records the discussions and decisions made at a meeting.

[1349] "Cloud storage" refers to online storage services that store and make data accessible over the internet.

[1350] The present invention is a system that supports efficient conference management by collecting, analyzing, organizing, and visualizing the comments of conference participants. Details of how to implement this system are described below.

[1351] This system includes a "terminal" used by a user, a "server" that performs analysis and management, and a "user" who is the user of the system.

[1352] Hardware and software used

[1353] The terminals used are devices with the ability to collect audio, such as smartphones, tablets, and PCs, with a meeting collection application installed.

[1354] The server incorporates a speech recognition engine (e.g., Google Cloud Speech-to-Text), a natural language processing engine (e.g., spaCy, NLTK), and a multilingual translation engine (e.g., Google Translate API).

[1355] System configuration

[1356] 1. Collection of audio data

[1357] The terminal uses a microphone to collect the voices of the conference participants. At the start of the conference, the user launches the conference collection application and presses the "Start Conference" button. This operation causes the terminal to collect voices in real time and record the collected voice data.

[1358] 2. Sending audio data

[1359] The collected voice data is compressed and encrypted on the device, ensuring security and efficiency, and then transmitted over the network to a server.

[1360] 3. Analysis of audio data

[1361] The server receives the voice data sent from the device and analyzes it using a voice recognition engine. It converts the voice into text data and identifies the speaker. Specifically, it extracts features from the patterns of the voice data and converts the speech into text.

[1362] 4. Extraction of speech content

[1363] The server then performs natural language processing on the text data to extract key information and identify key agenda items and action points for the meeting.

[1364] 5. Organizing and visualizing information

[1365] The extracted information is organized on the server and visualized in mind maps or lists. For example, subtopics and related items for each agenda item are organized and displayed in a format that is easy for users to understand.

[1366] 6. Multilingual translation and notification management

[1367] If necessary, the server translates the speech into other languages ​​using a multilingual translation engine. The translated information is displayed on the user's device in real time. In addition, the server monitors the progress of the meeting and notifies the device if the meeting goes over the scheduled time or if the progress is behind schedule.

[1368] 7. Generate and save minutes

[1369] After the meeting, the server automatically generates minutes based on the organized and visualized information. The minutes are then saved in cloud storage, allowing users to access them later.

[1370] Specific examples

[1371] Example 1: Neighborhood Association Meeting

[1372] During neighborhood association meetings, users use their smartphones to collect audio. During the meeting, the device sends the collected audio data to a server, which converts it into text in real time and identifies and visualizes important topics. After the meeting, the automatically generated minutes can be viewed from cloud storage.

[1373] Example 2: International online conference

[1374] In international online conferences, participants use their computers to collect audio and send it to a server via an online conference tool. The server analyzes the audio data and translates what is being said into multiple languages. The translation results are displayed on the device in real time, and progress management and notification functions are also supported. After the conference ends, multilingual minutes can be viewed from cloud storage.

[1375] Prompt Sentence Examples

[1376] "Please explain a program that creates real-time, multilingual minutes for international conferences."

[1377] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1378] Step 1:

[1379] The user starts up the terminal, opens the conference collection application, and presses the "Start Conference" button. This starts the collection of audio data. The user's operation leads to the start of audio capture using the microphone.

[1380] Input: User's instruction to start a meeting

[1381] Output: Trigger to start collection

[1382] Step 2:

[1383] The device collects the voices of the meeting participants in real time through a microphone, and each participant's speech is recorded as a separate audio stream.

[1384] Input: Conference room audio

[1385] Output: Raw audio data

[1386] Step 3:

[1387] The device compresses the collected audio data and then encrypts it, which is then sent to a server using a secure communication protocol.

[1388] Input: Raw audio data

[1389] Output: Encrypted audio data

[1390] Step 4:

[1391] The server receives the voice data sent from the terminal and temporarily stores the received data in a buffer.

[1392] Input: Encrypted audio data

[1393] Output: Buffered audio data

[1394] Step 5:

[1395] The server uses a speech recognition engine to analyze the voice data and convert it into text data. Specifically, it extracts features from the voice pattern, analyzes them, and converts them into text.

[1396] Input: Buffered audio data

[1397] Output: Text data

[1398] Step 6:

[1399] The server uses natural language processing on the text data to extract key information, a process that identifies key topics and action points for the meeting.

[1400] Input: Text data

[1401] Output: Important information (extracted agenda items, action points, etc.)

[1402] Step 7:

[1403] The server organizes and visualizes the extracted information in the form of a mind map or list. Specifically, it arranges related information for each agenda item in a radial pattern, making it visually easy to understand.

[1404] Input: Important Information

[1405] Output: Visualized data

[1406] Step 8:

[1407] The server transmits the visualized information to the terminal in real time, and the terminal displays this information and provides it to the user in a timely manner.

[1408] Input: Visualization data

[1409] Output: Information displayed on the terminal

[1410] Step 9:

[1411] The server uses a multilingual translation engine to translate the speech into other languages ​​as needed, and the translated information is also displayed on the device in real time.

[1412] Input: Important Information

[1413] Output: Translated information

[1414] Step 10:

[1415] The server monitors the progress of the conference and sends notifications to the terminals when the scheduled time has passed or when the progress is delayed. This notification is used to adjust the progress.

[1416] Input: Meeting progress data

[1417] Output: Notification of overdue time

[1418] Step 11:

[1419] After the meeting, the server automatically generates minutes based on the organized and visualized information, and the generated minutes are saved in cloud storage.

[1420] Input: Organized and visualized information

[1421] Output: Meeting minutes

[1422] Step 12:

[1423] Users can access cloud storage and check and download the generated minutes, making it easy to review them after the meeting.

[1424] Input: Meeting minutes stored in cloud storage

[1425] Output: User-accessible transcript

[1426] (Application example 1)

[1427] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1428] Ensuring safety and streamlining operations are important issues in modern factory environments. However, addressing these challenges requires fast and accurate communication and information sharing. However, it is difficult to quickly generate and display important information and guidelines during real-time meetings. Furthermore, language barriers exist in multilingual meetings, making it difficult for all participants to share the same information. There is a need to solve these issues and improve the efficiency of meeting management within factories and information transparency.

[1429] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1430] In this invention, the server includes means for collecting the voices of conference participants, means for analyzing the collected voice data and identifying what the conference participants are saying, means for converting the identified content of the speech into text data, means for extracting important information from the text data, means for visualizing the extracted information, means for displaying the extracted information on a factory work robot, means for automatically generating a safety manual and efficiency guidelines for the conference in real time, and means for displaying the generated manuals and guidelines using smart glasses or a head-mounted display. This enables quick and accurate information sharing in factory meetings, ensuring safety and improving work efficiency.

[1431] A "conference participant" is someone who actually attends a conference and takes part in speeches and discussions.

[1432] "Audio data" refers to information that has been collected from the voices of conference participants and converted into digital format.

[1433] "Analysis" is the process of processing collected audio data using algorithms to extract useful information.

[1434] "Speech identification" refers to recognizing the content of speech made by each individual participant in a conference.

[1435] "Text data" is digital data that has been converted from voice data into character information.

[1436] "Important information" is information that is particularly significant to the progress of the meeting and has a direct impact on decision-making and discussion.

[1437] "Visualization" refers to visually displaying extracted information in the form of charts, graphs, etc., to make it easier to understand.

[1438] A "factory robot" is a mechanical device designed to perform automated tasks in a factory.

[1439] "Real-time" means that information is collected and processed with little to no delay, providing immediate results.

[1440] A "safety manual" is a guide to ensuring safety at work sites.

[1441] "Efficiency guidelines" are documents that contain specific procedures and instructions for improving work efficiency.

[1442] "Smart glasses" are a wearable eyeglass-type device equipped with a display function.

[1443] A "head-mounted display" is a display device that is worn on the head and provides visual information.

[1444] This invention is a system that supports the efficient management of factory meetings by collecting and analyzing the comments of meeting participants, and organizing and visualizing the content. This system is realized through cooperation between "terminals," "servers," and "users."

[1445] System Configuration

[1446] 1. Collection of voice data (device)

[1447] The terminal collects the voices of the conference participants through a microphone. The collected voice data is recorded in real time and sent to a server. The terminal is typically a device such as smart glasses, a PC, or a factory robot.

[1448] 2. Sending audio data (terminal)

[1449] The device transmits the collected voice data over the network to a server, where it is compressed and encrypted to ensure security and efficiency.

[1450] 3. Analysis of voice data (server)

[1451] The server uses speech recognition technology to analyze the received voice data. Specifically, it converts the collected voice data into text data using the speech_recognition library. High-accuracy speech recognition is also possible by using Google's speech recognition API.

[1452] 4. Extraction of comment content (server)

[1453] The server uses natural language processing (NLP) techniques to extract important information from text data, leveraging the SpaCy library to extract important noun phrases and key phrases from the text.

[1454] 5. Content organization and visualization (server)

[1455] The server organizes the extracted information and visualizes the important content. It uses NetworkX and Matplotlib to display key phrases in graph form, allowing users to grasp the progress of the meeting and important topics at a glance.

[1456] 6. Multilingual Translation (Server)

[1457] The server uses the googletrans library to translate speech into multiple languages ​​as needed, and the translation results are displayed on the device in real time.

[1458] 7. Time Management and Notification (Server)

[1459] The server monitors the progress of the conference and notifies the user if the scheduled time is exceeded or if the progress is delayed. This notification is sent to the terminal and displayed to the user.

[1460] 8. Multi-purpose display (server)

[1461] The server automatically generates safety manuals and efficiency guidelines for meetings in real time, and displays the generated manuals and guidelines using smart glasses or a head-mounted display.

[1462] Specific examples

[1463] During meetings on safety and efficiency within a factory, devices (smart glasses or head-mounted displays) collect the voices of participants and send them to a server. The server analyzes the voice data, converts comments into text, and extracts and visualizes information related to important topics. Furthermore, it automatically generates safety manuals and efficiency guidelines in real time, and visualizes this data for participants via the smart glasses or head-mounted displays.

[1464] The server also translates multiple languages ​​as the meeting progresses, providing information in a format that is easy for all participants to understand, facilitating smooth information sharing during meetings within the factory and ensuring that important information on safety and other matters is communicated promptly.

[1465] Example prompts to input to the generative AI model:

[1466] I am trying to build a "meeting support system" for use in factory meetings. This system collects meeting audio in real time, extracts important keywords and phrases, and supports efficient meeting management. Please refer to the code below and help with additional improvements and enhancements.

[1467] Code fragment:

[1468] def record_audio(duration=60):

[1469] recognizer = sr.Recognizer()

[1470] mic = sr.Microphone()

[1471] with mic as source:

[1472] print("Meeting recording begins...")

[1473] audio = recognizer.record(source, duration=duration)

[1474] print("Recording complete")

[1475] return audio

[1476] Based on this, please give me some advice.

[1477] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1478] Step 1:

[1479] The terminal collects the voices of the conference participants. Specifically, it records the voices through a microphone and captures the voice data in real time. The input is the voices of the conference participants, and the output is the collected voice data.

[1480] Step 2:

[1481] The terminal sends the collected voice data to the server. When sending the data, the voice data is compressed and encrypted. The voice data is input, and the compressed and encrypted voice data is sent to the server as output.

[1482] Step 3:

[1483] The server analyzes the received voice data. Specifically, it converts the voice data into text data using the speech_recognition library. The input is compressed and encrypted voice data, and the output is text data.

[1484] Step 4:

[1485] The server extracts important information from the text data. Specifically, it uses the SpaCy library to extract noun phrases and key phrases from the text. The input is the text data, and the output is the extracted noun phrases and key phrases.

[1486] Step 5:

[1487] The server visualizes the extracted information in a graph format using NetworkX and Matplotlib. The input is the extracted key phrases, and the output is the visualized graph.

[1488] Step 6:

[1489] The server translates the speech into multiple languages ​​as needed. Specifically, it uses the GoogleTrans library for translation. The input is text data, and the output is translated text data.

[1490] Step 7:

[1491] The server monitors the progress of the conference and manages the time. It notifies participants if the conference is delayed or exceeds the scheduled time. The input is the start time of the conference, and the output is notifications according to the progress of the conference.

[1492] Step 8:

[1493] The server automatically generates safety manuals and efficiency guidelines and displays them on the terminal. The generated manuals and guidelines are displayed in real time through smart glasses or a head-mounted display. The input is extracted information and key phrases, and the generated safety manuals and efficiency guidelines are displayed as output.

[1494] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1495] This invention is a system that collects and analyzes the comments of meeting participants, organizes and visualizes the content, and supports efficient meeting management. By combining this with emotion recognition functionality, it is possible to grasp the emotional state of participants and adjust the progress of the meeting. This system is realized through cooperation between "terminals," "servers," and "users."

[1496] System Configuration

[1497] 1. Collection of voice data (device)

[1498] The terminal collects the voices of the participants through a microphone. The collected voice data is converted into a digital format in real time and sent to a server. The terminal is usually a device such as a smartphone or PC.

[1499] 2. Sending audio data (terminal)

[1500] The device compresses and encrypts the collected audio data and sends it to a server over the network, ensuring data security and efficient transfer.

[1501] 3. Analysis of voice data (server)

[1502] The server uses speech recognition technology to analyze the received voice data, converts the voice data into text data, identifies the speaker, and then uses natural language processing (NLP) technology to extract important information from the text data.

[1503] 4. Emotion Analysis (Server)

[1504] The server uses an emotion engine for emotion recognition to analyze participants' emotions from the voice data. The emotion engine uses an algorithm to estimate emotions based on voice characteristics such as intonation, tempo, and volume.

[1505] 5. Information visualization (server)

[1506] The server organizes the extracted information and emotional data and visualizes them in mind maps and lists. The emotional data is displayed in graphs and color codes, allowing the progress of the meeting to be understood at a glance. The visualized information is sent to the device in real time and displayed to the user.

[1507] 6. Multilingual Translation (Server)

[1508] The server translates the speech into multiple languages ​​as needed, and the translation results are sent to the device in real time and displayed.

[1509] 7. Time Management and Notification (Server)

[1510] The server monitors the progress of the conference and notifies the user if the conference exceeds the scheduled time or if the conference is behind schedule. The notification is sent to the terminal and displayed to the user.

[1511] 8. Generating minutes (server)

[1512] After the meeting, the server automatically generates minutes based on all collected and analyzed information, and the minutes are saved in cloud storage for users to access later.

[1513] Specific examples

[1514] Example 1: Neighborhood Association Meeting

[1515] At neighborhood association meetings, devices (smartphones and PCs) collect the voices of participants and send them to a server. The server analyzes the voice data, converts comments into text, and extracts and visualizes information related to important topics. It also uses an emotion engine to analyze the emotional state of participants and visualizes this at the same time. The emotional data is displayed in graphs and colors to show how participants are feeling when they speak, and is used to adjust the progress of the meeting. After the meeting ends, the server automatically generates minutes, which users can access and view by accessing cloud storage.

[1516] Example 2: International online conference

[1517] In international online conferences, devices (PCs) collect participants' voices through online conference tools and send them to a server. The server analyzes the voice data, identifies the speaker, and converts what was said into text. Natural language processing technology is used to extract important information, which is then translated using a multilingual translation engine. Emotions are then analyzed using an emotion engine and displayed in real time. The server monitors the progress and issues notifications in the event of delays or overruns. Minutes generated after the conference end contain emotional data along with the content of what was said, and users can access them via cloud storage.

[1518] With this specific configuration, the system provides comprehensive support at every stage of meeting management, ensuring effective progress while also taking into account the emotional states of participants.

[1519] The processing flow will be explained below.

[1520] Step 1:

[1521] When a meeting starts, the device activates the built-in or connected microphone device to collect participants' speech, and the collected audio data is converted into a digital format.

[1522] Step 2:

[1523] The device compresses and encrypts the collected voice data in real time, ensuring security and efficient transmission of the voice data.

[1524] Step 3:

[1525] The terminal transmits the encrypted voice data to the server via the network.

[1526] Step 4:

[1527] The server receives the received voice data and inputs it into a voice recognition engine, which converts the voice data into text data.

[1528] Step 5:

[1529] The server performs speaker identification analysis on the text data, where it identifies who spoke based on voice characteristics and other identifying information.

[1530] Step 6:

[1531] The server analyzes the text data using natural language processing (NLP) techniques to extract key information about the purpose of the meeting and related agenda items.

[1532] Step 7:

[1533] The server inputs the voice data into an emotion engine for emotion analysis, which estimates the participants' emotions based on voice characteristics such as intonation, tempo, and volume.

[1534] Step 8:

[1535] The server organizes the extracted information and emotional data and visualizes them in mind maps and lists, while the emotional data is displayed in graphs and color codes.

[1536] Step 9:

[1537] The server transmits the visualized information to the terminal in real time.

[1538] Step 10:

[1539] The terminal displays the received visualization data on a user interface so that the user can refer to it.

[1540] Step 11:

[1541] The server monitors the progress of the conference and generates an alert if the conference is running behind schedule or exceeds the scheduled time.

[1542] Step 12:

[1543] The server sends the generated alert to the terminal, and the user receives the alert on the terminal and adjusts the progress of the conference.

[1544] Step 13:

[1545] The server will then launch a translation engine to translate the speech into the specified languages ​​as needed, and the translated content will be sent to the device in real time.

[1546] Step 14:

[1547] The terminal displays the translated content to the user, facilitating understanding between participants who speak different languages.

[1548] Step 15:

[1549] At the end of the meeting, the server automatically generates minutes based on all collected and analyzed information.

[1550] Step 16:

[1551] The server stores the generated minutes in cloud storage for users to access later.

[1552] Step 17:

[1553] After the meeting, users can access the cloud storage to view the saved minutes and other related information.

[1554] With this specific processing flow, the system provides support at every stage of meeting management, creating an environment where meetings can proceed efficiently while taking into account the emotional state of participants.

[1555] Example 2

[1556] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1557] Although conventional meeting support systems had the ability to organize and visualize participants' comments, they were inadequate in terms of grasping participants' emotional states and supporting multiple languages. As a result, the progress of the meeting could be stalled due to emotional factors, and communication in a multilingual environment could not proceed smoothly. In addition, there were limited means for effectively monitoring the progress of the meeting and managing time. This resulted in problems that reduced the efficiency of meeting management.

[1558] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1559] In this invention, the server includes means for analyzing voice data and converting speech content into text data, means for analyzing emotional states from the voice data, means for visualizing important information and emotional data, means for translating into multiple languages, and means for monitoring progress and managing time. This makes it possible to grasp the emotional states of conference participants and efficiently manage conferences in multiple languages.

[1560] A "means for capturing participant voices" is a device and method for recording participant speech during a conference and converting it into a digital format.

[1561] "Means for compressing, encrypting and transmitting collected voice data" refers to a device and method that reduces the data size of collected voice data for efficient storage and transmission, encrypts the data to ensure security, and transmits it to a server via a network.

[1562] "Means for analyzing collected voice data and identifying speech of conference participants" refers to speech recognition technology and methods used to convert collected voice data into text and identify speakers.

[1563] The "means for converting the identified speech content into text data" refers to a method and apparatus for analyzing speech data and converting it into text format using speech recognition technology.

[1564] The "means for extracting important information from text data" refers to an apparatus and method for extracting important topics and key points in a meeting from text data using natural language processing technology.

[1565] The "means for analyzing the emotional state of participants from voice data" refers to an algorithm and method for estimating the emotional state of conference participants by analyzing the intonation, tempo, volume, etc. of the voice in the voice data.

[1566] "Means for visualizing extracted information and emotional data" refers to software and methods for visually displaying information and emotional data collected during a meeting in the form of mind maps, graphs, or lists.

[1567] A "means for translating into multiple languages" is an apparatus and method for translating text data into multiple different languages ​​in real time.

[1568] The "means for monitoring progress and managing time" refers to a device and method for monitoring the progress of a meeting and managing the meeting so that it proceeds within the scheduled time.

[1569] MODE FOR CARRYING OUT THE INVENTION

[1570] This invention is a system that supports efficient meeting management by collecting and analyzing the comments of meeting participants, and organizing and visualizing the content. By further combining this system with emotion recognition functionality, it is possible to grasp the emotional state of participants and adjust the progress of the meeting. This system is realized through cooperation between "terminals," "servers," and "users."

[1571] The device uses a microphone to collect the audio of the meeting participants. For example, a voice recording app or software on a smartphone or PC is used. The device converts the collected audio data into a digital format in real time and prepares it for transmission to the server. At this time, the audio data is compressed using "FFmpeg" and encrypted using "OpenSSL." The collected audio data is then sent to the server via the network.

[1572] The server converts the received voice data into text using speech recognition APIs such as Google Cloud Speech-to-Text and IBM Watson Speech to Text. It then uses natural language processing (NLP) libraries such as spaCy and NLTK to extract important information from the text data. The server also analyzes participants' emotional states from the voice data using emotion recognition technologies such as Microsoft Azure Emotion API and Affectiva. The analyzed emotion data is estimated based on voice characteristics such as intonation, tempo, and volume.

[1573] The server organizes the extracted information and emotion data and visualizes it using D3.js and Chart.js. This allows the information to be organized in mind maps and list formats, and the emotion data to be displayed in graphs and color codes. The visualized information is sent to the device in real time and displayed to the user. For example, the progress of a meeting can be displayed using graphs and tables so that it can be seen at a glance.

[1574] If necessary, the server translates the speech into multiple languages ​​using Google Translate API or Microsoft Translator. The translation results are also sent to the device in real time, allowing users to understand what is being said in different languages. The server monitors the progress of the meeting and uses Twilio or Slack API to notify users if the meeting goes over the scheduled time or if there are delays. Notifications are sent to the device and displayed to the user.

[1575] After the meeting, the server automatically generates minutes based on all collected and analyzed information. The documents are generated using Microsoft Word API and Google Docs API. The generated minutes are saved in cloud storage services such as Google Drive and Dropbox, allowing users to access them later.

[1576] Specific examples

[1577] Neighborhood Association Meeting

[1578] During neighborhood association meetings, devices (smartphones and PCs) collect participants' audio, compress it with FFmpeg, encrypt it with OpenSSL, and then send it to a server. The server converts the audio data into text using Google Cloud Speech-to-Text and extracts important information using spaCy. The server then analyzes the participants' emotional states using Affectiva and visualizes the emotional data using graphs and colors. After the meeting ends, the server automatically generates minutes, which users can view via cloud storage (e.g., Google Drive).

[1579] International Online Conference

[1580] In international online conferences, devices (PCs) collect participants' voices through online conference tools (e.g., Zoom), compress them with FFmpeg, encrypt them with OpenSSL, and send them to a server. The server analyzes the voice data using IBM Watson Speech to Text and converts what is being said into text. Important information is extracted using spaCy and translated using Google Translate API. The server further analyzes emotions using Microsoft Azure Emotion API and displays them on the device in real time. After the conference ends, the server generates minutes including the content of the speech and emotion data, which users can access via Dropbox.

[1581] The system provides comprehensive support at every stage of the meeting, enabling effective meeting management that takes into account the emotional state of participants. Multilingual support also enables smooth communication even in international meetings.

[1582] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1583] Step 1:

[1584] The device collects the voices of the meeting participants through a microphone. The input is the live audio from the meeting, and converts it into a digital format (e.g., WAV or MP3). Specifically, a dedicated application (e.g., audio recording software) on a smartphone or PC is used. The output digital audio data is stored in temporary memory for subsequent processing.

[1585] Step 2:

[1586] The terminal compresses the collected audio data using "FFmpeg" and encrypts it with "OpenSSL." The input is the digital audio data obtained in step 1, which is compressed to reduce the data size and encrypted to ensure data security. Specifically, the terminal executes the FFmpeg compression command and the OpenSSL encryption command. The output is compressed and encrypted audio data, which is sent to the server via the network.

[1587] Step 3:

[1588] The server receives the audio data sent from the device, decrypts it, and decompresses it. The input is compressed and encrypted audio data, and OpenSSL and FFmpeg are used to decompress it and return it to plain text. Specifically, the server executes the OpenSSL decryption command and the FFmpeg decompression command. The output is the original digital audio data.

[1589] Step 4:

[1590] The server converts the voice data into text data using speech recognition technology. The input is the original digital voice data, which is converted into text using APIs such as "Google Cloud Speech-to-Text" and "IBM Watson Speech to Text." The API is called to analyze the voice data and convert the utterances into text format. The output is text data for each speaker.

[1591] Step 5:

[1592] The server uses natural language processing (NLP) technology to extract important information from text data. The input is text data obtained through speech recognition, and information extraction is performed using tools such as "spaCy" and "NLTK." Specifically, the text data is input into an NLP library, and related keywords and topics are extracted. The output is text data extracted as important information.

[1593] Step 6:

[1594] The server analyzes the participants' emotional states from the voice data. The input is the voice data obtained in step 3, and emotion analysis is performed using tools such as the Microsoft Azure Emotion API. Specifically, the voice data is passed to the emotion engine, which estimates the emotional state (e.g., joy, anger, sadness). The output is data indicating the emotional state of each participant.

[1595] Step 7:

[1596] The server visualizes the extracted information and emotion data. The input is the important information obtained in step 5 and the emotion data obtained in step 6, and these are displayed visually using "D3.js" and "Chart.js." Specifically, it generates mind maps and graphs and sends them to the user's device in real time. The output is the visualized data.

[1597] Step 8:

[1598] The server translates the speech into multiple languages ​​as needed. The input is text data obtained through speech recognition, which is translated using the Google Translate API or Microsoft Translator. Specifically, it calls a translation engine to convert the text into another language, and the output is the translated text data.

[1599] Step 9:

[1600] The server monitors the progress of the meeting and manages the time. The input is meeting data updated in real time, and a notification is generated if the scheduled time is exceeded or progress is delayed. For example, a notification message is created using "Twilio" or "Slack API" and sent to the user's device. The output is a notification to the user.

[1601] Step 10:

[1602] After the meeting ends, the server automatically generates minutes based on all collected and analyzed information. The input is the voice data collected during the meeting, analyzed text data, and emotion data, and minutes are generated using the Microsoft Word API or Google Docs API. Specifically, the data is passed to the minutes generation engine to generate a document, which is then saved in cloud storage. The output is the generated minutes.

[1603] These processing steps effectively organize the remarks of the conference participants and realize efficient conference proceedings that take into account their emotional states.

[1604] (Application example 2)

[1605] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1606] While conventional meeting support systems can collect participants' voices, convert what is being said into text, and extract and visualize important information, they are unable to grasp the participants' emotional state and adjust the progress of the meeting. Furthermore, when making advertising presentations, there is no mechanism for providing real-time feedback on participants' reactions and emotions, making it difficult to effectively present advertising campaigns.

[1607] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1608] In this invention, the server includes a means for analyzing the emotional states of conference participants from their comments, a means for visualizing the analyzed emotional states in real time, and a means for automatically generating a record including the comments and emotional states after the conference ends. This allows participants' reactions and emotional states to be grasped in real time, enabling effective progress and feedback of the advertising presentation.

[1609] The "means for collecting the voices of the conference participants" refers to a device or system that converts the voices of the conference participants into a digital format in real time via a microphone and collects them as data.

[1610] "Means for analyzing collected voice data and identifying what is being said by conference participants" refers to algorithms or software for identifying conference participants and specifying what is being said based on collected voice data.

[1611] The "means for converting the identified speech content into text data" is a process for converting the identified speech content into text data in sentence format using speech recognition technology.

[1612] The "means for extracting important information from text data" is a system that uses natural language processing technology to extract key topics and important information discussed in meetings and presentations from text data.

[1613] "Means for visualizing extracted information" refers to software or tools that visually display extracted important information and data in the form of graphs, lists, mind maps, etc.

[1614] "Means for analyzing emotional state from speech content" refers to algorithms or emotion recognition engines for estimating emotional state based on characteristics of the speaker's voice data, such as intonation, tempo, and volume.

[1615] The "means for visualizing analyzed emotional states in real time" is a system for displaying emotional data estimated by an emotion recognition engine in real time using graphs and colors.

[1616] "Means for automatically generating records including remarks and emotional states after a meeting" refers to software that automatically creates and saves minutes and records that integrate remarks and emotional states based on information collected and analyzed after a meeting.

[1617] The present invention is a conference support system used in advertising presentations that analyzes and visualizes participants' comments and emotional states to effectively support the progress of the conference. Specific embodiments are described below.

[1618] System Configuration

[1619] 1. Collection of voice data (device)

[1620] The terminals, typically digital devices such as smartphones and personal computers, collect the voices of conference participants in real time through microphones and convert them into digital format. This data is then compressed, encrypted, and sent to the server.

[1621] 2. Analysis of voice data (server)

[1622] The server receives the collected voice data and converts it into text using speech recognition technology, identifies the speaker, analyzes the speech, and extracts important information. Natural language processing (NLP) technology is used here.

[1623] 3. Emotion Analysis (Server)

[1624] The server uses an emotion recognition engine to estimate participants' emotional states from the voice data. The emotion recognition engine uses an algorithm to analyze voice intonation, tempo, volume, etc. to estimate emotions.

[1625] 4. Real-time visualization (server)

[1626] The server organizes the extracted text information and emotional data and visualizes it in real time in graph and list format. The emotional data is displayed using colors and graphs, allowing participants' emotional states to be understood at a glance. The visualized information is sent to the device in real time and displayed to the user.

[1627] 5. Generation of minutes (server)

[1628] After the meeting, the server automatically generates minutes based on all collected and analyzed information, which are then saved in cloud storage for users to access later.

[1629] Usage example

[1630] Specific examples

[1631] During an advertising presentation meeting, when an advertising agency is presenting a new advertising campaign, the comments made by participants and the emotional state of those comments are analyzed and visualized in real time. For example, it is possible to instantly understand whether participants' reactions are positive or negative during the presentation. This allows the presenter to flexibly adjust the progress of the meeting and deliver an effective presentation.

[1632] Prompt Sentence Examples

[1633] Create a system that analyzes participants' comments and emotions in real time during advertising campaign meetings and visualizes the results. Include a function to grasp participants' emotional state and automatically generate and save minutes after the meeting.

[1634] Hardware and Software Configuration

[1635] Microphone: for collecting audio data

[1636] Personal computer or smartphone: for analysis and data storage

[1637] Python: a programming language

[1638] SpeechRecognition: A speech recognition library

[1639] transformers: Hugging Face's NLP library for emotion analysis

[1640] Google Speech API: Convert speech to text

[1641] Using this hardware and software, it is possible to collect and analyze voice data, analyze emotional states, visualize information, and generate meeting minutes.

[1642] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1643] Step 1: Collecting voice data (device)

[1644] The terminal collects the voices of the conference participants in real time through a microphone and converts them into a digital format. The terminal then compresses and encrypts the collected voice data and sends it to the server. The input is raw voice, and the output is digitized encrypted voice data.

[1645] Step 2: Analyzing the voice data (server)

[1646] The server converts the received voice data into text data using speech recognition technology. This process uses the Google Speech API to analyze the voice data and convert it into text. It also identifies the speaker and identifies each utterance. The input is encrypted voice data, and the output is the identified utterance text data.

[1647] Step 3: Extracting important information (server)

[1648] The server extracts important information from the identified text data using natural language processing (NLP) techniques. This process uses the Hugging Face transformers library. The input is the text data, and the output is the extracted important information.

[1649] Step 4: Sentiment Analysis (Server)

[1650] The server analyzes the emotional state of the identified speech using an emotion recognition engine. This process estimates emotions based on voice intonation, tempo, volume, etc. The input is text data, and the output is emotion analysis data.

[1651] Step 5: Real-time visualization (server)

[1652] The server visualizes the extracted information and emotional data in graphs and lists, allowing participants' emotional states to be visually displayed in real time. The input is important information and emotional analysis data, and the output is visualized graphs and lists.

[1653] Step 6: Real-time display (terminal)

[1654] The terminal displays the visualized data sent from the server in real time, allowing users to instantly check the emotional state and important information during a meeting. The input is the visualized data, and the output is the displayed graph or list.

[1655] Step 7: Generate minutes (server)

[1656] After the meeting ends, the server automatically generates minutes based on all collected and analyzed information. These minutes include the content of remarks and the emotional state of the participants. The generated minutes are saved in cloud storage and can be accessed by users later. The input is important information and emotional analysis data, and the output is the minutes data.

[1657] Step 8: Saving and Accessing the Minutes (Devices and Users)

[1658] Users can access the cloud storage through their devices and check the generated minutes. The input is the minutes data on the cloud, and the output is the minutes that users can view.

[1659] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1660] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1661] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1662] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1663] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1664] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1665] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1666] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1667] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1668] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1669] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1670] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1671] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1672] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1673] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1674] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1675] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1676] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1677] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1678] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1679] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1680] The following is further disclosed regarding the above embodiment.

[1681] (Claim 1)

[1682] means for collecting audio of conference participants;

[1683] A means for analyzing the collected voice data and identifying the speech of the conference participants;

[1684] A means for converting the identified utterance content into text data;

[1685] A means for extracting important information from text data;

[1686] The system includes a means for visualizing the extracted information.

[1687] (Claim 2)

[1688] A means for translating the collected voice data into multiple languages;

[1689] 10. The system of claim 1, further comprising: means for displaying the translated content.

[1690] (Claim 3)

[1691] A means of monitoring meeting progress and time management;

[1692] 10. The system of claim 1, further comprising means for providing notification if the scheduled time is exceeded.

[1693] "Example 1"

[1694] (Claim 1)

[1695] means for collecting audio of conference participants;

[1696] a means for compressing and encrypting the collected voice data and transmitting the data;

[1697] means for analyzing the transmitted voice data and converting the voice data into text data;

[1698] natural language processing means for extracting important information from text data;

[1699] A means to organize the extracted information and visualize it in mind maps or list formats,

[1700] A system including a means for displaying visualized information in real time.

[1701] (Claim 2)

[1702] A means for translating the collected voice data into multiple languages;

[1703] A means to display the translated content in real time;

[1704] A means of monitoring the progress of the meeting and providing notification if the meeting goes over schedule or is behind schedule; and

[1705] 10. The system according to claim 1, further comprising means for automatically generating minutes after the meeting and storing them in cloud storage.

[1706] (Claim 3)

[1707] 10. The system of claim 1, further comprising means for storing data generated by the system of claim 1 in cloud storage for later access by a user.

[1708] "Application Example 1"

[1709] (Claim 1)

[1710] means for collecting audio of conference participants;

[1711] A means for analyzing the collected voice data and identifying the speech of the conference participants;

[1712] A means for converting the identified utterance content into text data;

[1713] A means for extracting important information from text data;

[1714] a means for visualizing the extracted information;

[1715] a means for displaying the extracted information through a factory work robot;

[1716] A means to automatically generate safety manuals and efficiency guidelines during meetings in real time,

[1717] A system including a means for displaying the generated manuals and guidelines using smart glasses or a head-mounted display.

[1718] (Claim 2)

[1719] A means for translating the collected voice data into multiple languages;

[1720] 10. The system of claim 1, further comprising: means for displaying the translated content.

[1721] (Claim 3)

[1722] A means of monitoring meeting progress and time management;

[1723] 10. The system of claim 1, further comprising means for providing notification if the scheduled time is exceeded.

[1724] "Example 2: Combining Emotion Engines"

[1725] (Claim 1)

[1726] means for collecting audio of conference participants;

[1727] a means for compressing and encrypting the collected voice data and transmitting the data;

[1728] A means for analyzing the collected voice data and identifying the speech of the conference participants;

[1729] A means for converting the identified utterance content into text data;

[1730] A means for extracting important information from text data;

[1731] a means for analyzing the emotional state of participants from the audio data;

[1732] The system includes a means for visualizing the extracted information and emotion data.

[1733] (Claim 2)

[1734] A means for translating the collected voice data into multiple languages;

[1735] 10. The system of claim 1, further comprising: means for displaying the translated content.

[1736] (Claim 3)

[1737] A means of monitoring meeting progress and time management;

[1738] 10. The system of claim 1, further comprising means for providing notification if the scheduled time is exceeded.

[1739] "Application example 2 when combining emotion engines"

[1740] (Claim 1)

[1741] means for collecting audio of conference participants;

[1742] A means for analyzing the collected voice data and identifying the speech of the conference participants;

[1743] A means for converting the identified utterance content into text data;

[1744] A means for extracting important information from text data;

[1745] a means for visualizing the extracted information;

[1746] A means of analyzing the emotional state from the content of the speech,

[1747] A means for visualizing the analyzed emotional state in real time;

[1748] The system includes a means for automatically generating a transcript after a meeting, including what was said and the emotional state of the participants.

[1749] (Claim 2)

[1750] A means for translating the collected voice data into multiple languages;

[1751] 10. The system of claim 1, further comprising: means for displaying the translated content.

[1752] (Claim 3)

[1753] A means of monitoring meeting progress and time management;

[1754] 10. The system of claim 1, further comprising means for providing notification if the scheduled time is exceeded. [Explanation of symbols]

[1755] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. means for collecting audio of conference participants; A means for analyzing the collected voice data and identifying the speech of the conference participants; A means for converting the identified utterance content into text data; A means for extracting important information from text data; The system includes a means for visualizing the extracted information.

2. A means for translating the collected voice data into multiple languages; 10. The system of claim 1, further comprising: means for displaying the translated content.

3. A means of monitoring meeting progress and time management; 10. The system of claim 1, further comprising means for notifying if the scheduled time has been exceeded.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A