system

The system addresses inefficiencies in meeting operations by automating voice input, recognition, classification, and document generation, enhancing productivity and reducing overtime through streamlined meeting management.

JP2026064614APending Publication Date: 2026-04-14SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-10-02
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Modern enterprises face inefficiencies in meeting operations, including reduced actual work time due to numerous meetings, human errors in manual meeting minutes and task management, and inaccurate information sharing, which hinder productivity.

Method used

A system that includes voice acquisition, real-time communication, speech recognition, text classification, automatic document generation, storage, and transmission capabilities to streamline meeting management and document creation, allowing for efficient meeting setup and metadata management.

Benefits of technology

This system automates meeting recording and document creation, reducing overtime and ensuring accurate information sharing, thereby improving employee productivity and meeting efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026064614000001_ABST
    Figure 2026064614000001_ABST
Patent Text Reader

Abstract

We provide a system that improves the efficiency of setting up and managing meetings. [Solution] These problems are solved by providing a system that includes a voice acquisition means for the user to input voice data, a communication means for transmitting voice data to a server in real time, a voice recognition means for the server to convert the received voice data into text data, a text classification means for classifying the text data based on the meeting content, a document generation means for automatically generating meeting documents from the classified text data, a storage means for saving the generated documents, and a transmission means for sending the generated documents to the user. Furthermore, the system also includes a setting means for the user to configure meeting settings and a management means for the server to manage meeting metadata.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, the method including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In modern enterprises, it is common for employees to hold many internal meetings to improve work efficiency. However, due to the large number of meetings, the time employees spend on actual work is reduced, resulting in an increase in overtime hours. In addition, manual operation of meeting minutes and task management is prone to human errors and low efficiency. Furthermore, leakage of important information and sharing of inaccurate information may also become problems. Thus, the current meeting operation is an element that hinders the productivity of employees and the efficiency of the entire company. It is necessary to solve these problems and significantly improve the efficiency of meeting operation.

Means for Solving the Problems

[0005] This invention solves these problems by providing a system that includes a voice acquisition means for the user to input voice data, a communication means for transmitting voice data to a server in real time, a voice recognition means for the server to convert the received voice data into text data, a text classification means for classifying the text data based on the meeting content, a document generation means for automatically generating meeting documents from the classified text data, a storage means for saving the generated documents, and a transmission means for sending the generated documents to the user. Furthermore, by including a setting means for the user to set up the meeting and a management means for the server to manage the meeting metadata, the efficiency of meeting setup and management is improved. As a result, automatic recording of meeting content and document creation become possible, allowing employees to concentrate on their actual work hours, reducing overtime and enabling efficient information sharing.

[0006] 1. "Voice acquisition means" refers to a device or function for a user to perform voice input, and consists of a microphone, a virtual assistant, etc.

[0007] 2. "Communication means" refers to a network-based system for transmitting voice data in real time, such as using internet connectivity or WebSocket technology.

[0008] 3. "Speech recognition means" refers to the technology for converting received speech data into text data, and includes software and systems that incorporate speech recognition algorithms.

[0009] 4. "Text classification means" refers to a function for classifying text data based on meeting content, and is a system that utilizes natural language processing technology to organize data by agenda item or category.

[0010] 5. "Document generation means" refers to software or a system for automatically generating meeting documents from classified text data, such as meeting minutes and task lists.

[0011] 6. "Storage means" refers to functions or systems for securely saving generated meeting documents to databases, cloud storage, etc.

[0012] 7. "Transmission method" refers to the function or system for sending the generated meeting documents to users, and may include using email or file sharing services.

[0013] 8. "Setting means" refers to the interface or tool that users use to set up a meeting, and is used for inputting and managing the meeting date and time, participants, agenda, etc.

[0014] 9. "Management means" refers to the functions and systems that the server uses to manage meeting metadata, such as the meeting name, participant list, and date and time of the meeting, and which are used to record and update this information. [Brief explanation of the drawing]

[0015] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8]It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which a plurality of emotions are mapped. [Figure 10] It shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when an emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when an emotion engine is combined.

Modes for Carrying Out the Invention

[0016] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0017] First, the language used in the following description will be explained.

[0018] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of a plurality of arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of a plurality of types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0019] In the following embodiments, signed RAM (Random Access Memory) is a memory that temporarily stores information and is used as work memory by the processor.

[0020] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.

[0021] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0022] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0023] [First Embodiment]

[0024] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0025] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0026] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0027] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0028] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0029] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0030] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0031] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0032] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0033] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0034] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0035] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0036] This invention is a system that acquires user voice input, transmits the voice data to a server in real time, converts the voice data into text data, classifies it, and automatically generates, saves, and transmits meeting documents. Specific embodiments are described below.

[0037] In this system, the user first performs voice input using a voice acquisition method. For example, the microphone built into the user's terminal (a PC or mobile device) is used. This voice data is transmitted in real time from the terminal to the server via a communication method (such as an internet connection or WebSocket technology via a dedicated application).

[0038] The server, acting as a speech recognition system, converts received audio data into text data using a highly accurate speech recognition algorithm. This text data is then automatically categorized according to the meeting agenda, utilizing natural language processing technology as a text classification system. Classification is performed not only by agenda item but also by category (e.g., shared items, homework, SOW, Q&A).

[0039] Next, based on the classified text data, the server, acting as a document generation mechanism, automatically generates meeting documents (such as meeting minutes and task lists). These documents are generated in the specified format (e.g., PDF or Word document). Furthermore, the generated documents are securely stored in a database or cloud storage.

[0040] After the meeting documents have been generated and saved, the server, acting as the transmission method, sends them to the user. The documents are sent via the specified email address or file-sharing service. The user's device is designed to display the received documents in a highly readable format.

[0041] Furthermore, users can set meeting details (date and time, participants, agenda, etc.) using a dedicated interface (web application or mobile app) before the meeting. This information is sent to the server and managed as metadata by the server's management system. Metadata management streamlines pre-meeting preparations and post-meeting documentation.

[0042] As a concrete example, consider a scenario where a user sets up a new product development meeting. First, the user schedules the meeting using the scheduling tool and notifies participants. On the day of the meeting, the user uses a terminal in the meeting room to send the audio of the conversation to the server using the audio acquisition tool. The server converts the audio to text in real time, categorizes it by agenda item, and automatically generates meeting documents immediately. Once the meeting ends, the generated minutes and task list are saved to cloud storage using the saving tool and quickly sent to the user's terminal via the transmission tool. The user and participants can review these materials and provide feedback as needed, further streamlining the preparation for the next meeting.

[0043] In this way, the system of the present invention dramatically improves the efficiency of meetings and provides an environment in which employees can concentrate on their actual work time. This leads to a reduction in overtime hours and accurate sharing of information.

[0044] The following describes the processing flow.

[0045] Step 1:

[0046] The user schedules a meeting. Using the terminal's settings, they input the meeting date and time, participants, and agenda, and send this information to the server. The server stores the received information in a database using its management system.

[0047] Step 2:

[0048] The user starts a meeting. At the specified date and time of the meeting, the user inputs voice data using the device's audio acquisition method (microphone). The device acquires this voice data in real time and transmits it to the server using a communication method.

[0049] Step 3:

[0050] The server converts the received audio data into text data using speech recognition technology. The speech recognition algorithm analyzes the audio and converts it into text format.

[0051] Step 4:

[0052] The server's text classification system categorizes text data based on meeting content. Using natural language processing technology, it organizes the text into categories such as meeting minutes, shared items, homework, statements of work (SOW), and Q&A.

[0053] Step 5:

[0054] The server's document generation mechanism automatically generates meeting documents from classified text data. The documents are created in the specified format (PDF, Word document, etc.).

[0055] Step 6:

[0056] The server's storage method involves saving the generated meeting documents to a database or cloud storage. The document's metadata (creation date, creator, etc.) is also saved simultaneously.

[0057] Step 7:

[0058] The server's transmission method sends the generated meeting documents to the user's terminal. The documents are sent via email or a file-sharing service.

[0059] Step 8:

[0060] Users review meeting documents received on their devices. They edit the documents as needed, incorporating any corrections or additional information. They also collect feedback to improve future meetings.

[0061] Step 9:

[0062] The user distributes the final edited document to other participants. This can be done by resending it through the server or manually using email or a file-sharing service.

[0063] (Example 1)

[0064] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0065] Traditional methods for accurately and quickly recording discussions during meetings and automatically generating minutes and task lists are time-consuming and laborious. Furthermore, smoothly sharing these materials with all meeting participants is also a cumbersome process. Therefore, there is a need for efficient meeting management and accurate sharing of important information.

[0066] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0067] In this invention, the server includes: voice acquisition means for the user to input voice data; communication means for transmitting voice data to a data processing device in real time; and voice recognition means for the data processing device to convert the received voice data into text data. This makes it possible to instantly transcribe, classify, and automatically generate, save, and share meeting documents based on what the user says during a meeting.

[0068] "Voice acquisition means" refers to devices or functions that allow users to perform voice input.

[0069] "Communication means" refers to technical methods and devices for transmitting voice data to a data processing device in real time.

[0070] "Speech recognition means" refers to algorithms and software used by a data processing device to convert received speech data into text data.

[0071] "Natural language processing means" refers to a technology used by data processing devices to classify text data based on the content of a meeting.

[0072] "Document generation means" refers to software or algorithms used by a data processing device to automatically generate meeting documents from classified text data.

[0073] "Storage means" refers to the technology and equipment used to save generated meeting documents to a storage medium.

[0074] "Transmission means" refers to the technology or functionality used to send generated meeting documents to users.

[0075] "Display means" refers to devices or software that display meeting documents received by a user's terminal in a highly visible format.

[0076] "Configuration means" refers to the interface or software that allows users to configure detailed meeting information.

[0077] "Management means" refers to the functions and technologies that data processing devices use to manage meeting metadata.

[0078] This invention relates to a system in which a user performs voice input, transmits the voice data to a server in real time, converts it into text data, classifies it, and automatically generates and saves meeting documents. Specific embodiments of the system are described in detail below.

[0079] overview

[0080] This system captures user speech and transmits it in real time to a data processing unit (server). The server converts the speech data into text data and classifies it based on the meeting content. The system automatically generates, saves, and transmits meeting documents from the classified text data. The aim of this system is to improve meeting efficiency and ensure accurate information sharing.

[0081] Hardware and software to be used

[0082] The following hardware and software will be used to implement this system.

[0083] Audio acquisition method: User's device (PC, mobile device) equipped with a microphone device.

[0084] Communication methods: Internet connection, WebSocket technology, HTTP POST

[0085] Speech recognition methods: Google® Cloud Speech-to-Text API, Amazon Transcribe

[0086] Natural language processing tools: Google Cloud Natural Language API

[0087] Document generation methods: Python's reportlab library, docx library

[0088] Storage method: Cloud storage (Amazon S3, Google Drive)

[0089] Sending method: SendGrid, Google Drive shared link

[0090] Specific example

[0091] For example, let's consider a new product development meeting. The user sets the meeting date and participants using a web application and sends them to the data processing unit. On the day of the meeting, the user continuously sends audio of the conversation to the data processing unit using the microphone on their device. The data processing unit converts the audio to text in real time and categorizes it by agenda item. After the meeting, the data processing unit automatically generates meeting minutes and a task list and saves them to cloud storage. These documents are sent to the user and displayed on their device in a highly readable format. The following is an example of a prompt to be input into the generating AI model.

[0092] Example of a prompt:

[0093] Please analyze the audio data from the next meeting and generate meeting minutes categorized by agenda item.

[0094] Users can review this document and provide feedback as needed, further streamlining their preparations for the next meeting. The automated generation and sharing of meeting documents provides employees with an environment where they can focus on their actual work time.

[0095] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0096] Step 1: Obtaining voice input

[0097] The device acquires the user's voice input via the microphone. This voice data is input and temporarily stored in a buffer on the device side. This data is then sent to the next step in a short time for real-time processing.

[0098] Step 2: Sending the audio data

[0099] The terminal transmits the acquired audio data to the server in real time. This transmission uses an internet connection and WebSocket technology. The terminal streams the audio data in the buffer to the server via the WebSocket connection. As a specific example, the JavaScript (registered trademark) WebSocket API is used.

[0100] Step 3: Receiving audio data

[0101] The server receives audio data transmitted from the terminal in real time. The server opens a WebSocket connection and receives the audio data as a stream. The received audio data is temporarily stored in internal memory.

[0102] Step 4: Convert speech to text

[0103] The server converts the received audio data into text data using a speech recognition algorithm. The input is audio data, and the output is the corresponding text. Specifically, the server uses the Google Cloud Speech-to-Text API to send audio data to the API and receive the resulting text data.

[0104] Step 5: Classification of text data

[0105] The server classifies text data by topic using natural language processing techniques. In this step, the input is text data, and the output is text data with classification tags. The Google Cloud Natural Language API is used to parse the text and assign appropriate tags, resulting in categorical text classification.

[0106] Step 6: Generating meeting documents

[0107] The server automatically generates meeting documents based on classified text data. The input is classified text data, and the output is meeting documents (PDF or Word documents). The Python reportlab and docx libraries are used to generate documents in the specified format.

[0108] Step 7: Save the document

[0109] The server saves the generated meeting documents to cloud storage. The input is the meeting documents, and the output is the files saved in cloud storage. Specifically, documents are uploaded using Amazon S3 or the Google Drive API.

[0110] Step 8: Send the document

[0111] The server sends the saved meeting documents to the user. The input is the meeting documents, and the output is the documents received by the user. A download link for the documents is sent to the user's email address using a SendGrid or Google Drive sharing link.

[0112] Step 9: Setting up the meeting

[0113] Users configure meeting details (date, time, participants, agenda, etc.) using a dedicated interface (web application or mobile app). The input is the meeting details, and the output is metadata generated based on that information. The metadata is sent to the server and stored and managed by a management system.

[0114] Step 10: Document Review and Feedback

[0115] Users review the received meeting documents and provide feedback as needed. The input is the meeting document, and the output is the revised document or additional feedback information. Users use this to prepare for future meetings.

[0116] (Application Example 1)

[0117] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0118] Traditional factory operations often involved manual processes for instructing workers and managing their progress, resulting in inefficiency and a heavy burden on workers. In particular, generating and managing work instructions and progress reports in real time was difficult, leading to delays and errors. Therefore, there is a growing need for increased efficiency in factory operations and real-time progress management.

[0119] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0120] In this invention, the server includes voice acquisition means for the user to input voice data, communication means for transmitting voice data to the server in real time, voice recognition means for the server to convert the received voice data into text data, text classification means for the server to classify the text data based on work instructions and progress status, document generation means for the server to automatically generate progress reports and task lists from the classified text data, storage means for the server to save the generated progress reports and task lists, and transmission means for the server to send the generated progress reports and task lists to the user. This makes it possible to efficiently provide instructions and manage the progress of factory work in real time.

[0121] "Voice acquisition means" refers to a device or interface for a user to perform voice input.

[0122] "Communication methods" refer to network technologies and protocols for transmitting voice data to a server in real time.

[0123] "Speech recognition means" refers to algorithms and software used by a server to convert received speech data into text data.

[0124] "Text classification means" refers to natural language processing technology used by a server to classify text data based on work instructions and progress status.

[0125] "Document generation means" refers to software or functionality that allows a server to automatically generate progress reports and task lists from classified text data.

[0126] "Storage methods" refer to technologies and systems for saving generated progress reports and task lists to databases or cloud storage.

[0127] "Transmission method" refers to the technology or protocol used by a server to send generated progress reports and task lists to users.

[0128] "Configuration means" refers to the interface or software that allows users to configure settings for their work.

[0129] "Management means" refers to the systems and technologies that servers use to manage metadata for their operations.

[0130] This invention is a system for managing work instructions and progress in real time within a factory, enabling efficient work operations. The elements constituting this system and their specific embodiments are described below.

[0131] First, the user performs voice input using a voice acquisition device. This device can be a head-mounted display or smart glasses equipped with a high-performance microphone. This data is transmitted in real time to a cloud server via a communication device. The communication device can utilize an internet connection or WebSocket technology through a dedicated application.

[0132] The cloud server converts received audio data into text data using speech recognition. High-precision speech recognition APIs such as Google Cloud Speech-to-Text are used for speech recognition. This text data is then classified by text classification tools based on work instructions and progress. Natural language processing libraries such as spaCy and NLTK are utilized for classification.

[0133] Next, progress reports and task lists are automatically generated from the classified text data by a document generation mechanism. The generated documents are stored in cloud storage such as Google Cloud Storage by a storage mechanism. Furthermore, the server transmits the generated documents to the user's head-mounted display via a transmission mechanism.

[0134] Users can check progress and next work instructions in real time through a head-mounted display. This improves work efficiency and reduces the burden on workers.

[0135] As a concrete example, suppose a user says, "I have finished preparing for step 1. Please give me the next instructions." This voice input is sent to a cloud server via a voice acquisition device. The cloud server converts the voice to text, and then uses Natural Language Processing (NLP) technology to classify the text as "Preparation for step 1 complete." Based on the classified text data, a progress report is automatically generated and displayed in real time on a head-mounted display.

[0136] Specific examples of prompt statements are as follows:

[0137] "Voice input: 'Preparation for step 1 is complete. Please give me the next instruction.'"

[0138] By using this system, work instructions and progress management within the factory can be carried out efficiently in real time, reducing delays and errors and improving overall production efficiency.

[0139] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0140] Step 1:

[0141] Users use a head-mounted display or smart glasses to input voice commands.

[0142] Input: User's voice instructions

[0143] Output: The microphone in the device acquires audio data.

[0144] Specific action: The user says, "I've finished preparing for step 1. Please give me the next instructions," and the audio data is captured by the microphone.

[0145] Step 2:

[0146] The device transmits the acquired audio data to the cloud server in real time.

[0147] Input: Acquired audio data

[0148] Output: Audio data sent to the cloud server

[0149] Specific operation: Audio data is sent to a cloud server over the internet using WebSocket technology.

[0150] Step 3:

[0151] The server converts the received audio data into text data using speech recognition technology.

[0152] Input: Audio data received by the server

[0153] Output: Audio data converted to text data

[0154] Specific operation: Using the Google Cloud Speech-to-Text API, the audio data is converted into text data that reads, "Step 1 is ready. Please give me the next instructions."

[0155] Step 4:

[0156] The server uses natural language processing techniques to classify text data based on work instructions and progress.

[0157] Input: Converted text data

[0158] Output: Classified text data

[0159] Specific operation: Text is used to perform a text analysis using spaCy or NLTK, and the progress status is determined as "Ready".

[0160] Step 5:

[0161] The server automatically generates progress reports and task lists from classified text data using document generation methods.

[0162] Input: Classified text data

[0163] Output: Progress report and task list

[0164] Specific operation: Document generation software on the server generates progress reports and task lists, and creates them in PDF or text format.

[0165] Step 6:

[0166] The server saves the generated progress reports and task lists to cloud storage.

[0167] Input: Generated progress reports and task lists

[0168] Output: Stored data in cloud storage

[0169] Specific action: Upload progress reports and task lists to Google Cloud Storage and store them securely.

[0170] Step 7:

[0171] The server sends generated progress reports and task lists to the user in real time.

[0172] Input: Generated progress reports and task lists

[0173] Output: Data sent to the user's head-mounted display

[0174] Specific operation: Progress reports and task lists are sent via the internet to the user's head-mounted display and displayed on the screen.

[0175] In this way, work instructions and progress management within the factory are carried out efficiently in real time through this series of processes. This system automates all steps from voice input to text conversion, progress report generation, saving, and transmission, thereby reducing the burden on workers and improving work efficiency.

[0176] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0177] This invention is a system that acquires user voice input, transmits the voice data to a server in real time, converts the voice data into text data, classifies it, and automatically generates, saves, and transmits meeting documents, in addition to an emotion engine that recognizes the user's emotions. Specific embodiments are described below.

[0178] In this system, the user first inputs voice data using a voice acquisition device. For example, the microphone built into the user's device (a PC or mobile device) is used. This voice data is transmitted in real time from the device to the server via a communication method (such as an internet connection or WebSocket technology via a dedicated application).

[0179] The server, acting as a speech recognition system, converts received audio data into text data using a highly accurate speech recognition algorithm. This text data is then automatically categorized according to the meeting agenda, utilizing natural language processing technology as a text classification system. The data is then organized by agenda item or category (e.g., shared items, homework, SOW, Q&A, etc.).

[0180] Next, to automatically generate meeting documents from text data, the server, acting as the document generation method, creates documents in the specified format (PDF, Word document, etc.). These meeting documents are stored securely in a database or cloud storage. After the generation and saving of the meeting documents is complete, the server sends the generated documents to the user's terminal using a transmission method. The documents are sent via email or a file sharing service.

[0181] Furthermore, this invention utilizes an emotion engine means for analyzing the user's emotions. The emotion engine means analyzes the user's emotions (e.g., joy, anger, sadness, surprise, etc.) in real time from voice input. The results of the emotion analysis are reflected in the meeting document, allowing users and participants to review changes in their emotions during the meeting as a record. This data is transmitted using a server's transmission means and included in meeting minutes, task lists, etc., generated after the meeting.

[0182] For example, in a new product development meeting, the user schedules the meeting and starts it at the designated time. The user inputs voice data using a voice acquisition device, and this voice data is sent to the server in real time. The server converts the voice into text data and categorizes it by topic and category. Furthermore, an emotion engine device analyzes the user's emotions at the time of speaking, and this emotion data is also processed. After the meeting ends, the server automatically generates a meeting document, saves it to cloud storage using a storage device, and sends it to the user using a transmission device. The user reviews the received document, edits it, and gathers feedback. By using this information to prepare for the next meeting, discussions can be conducted more efficiently.

[0183] Thus, the system of the present invention dramatically improves meeting efficiency and provides an environment where employees can concentrate on their actual work time. Furthermore, by collecting and sharing emotional data during meetings, it facilitates smoother communication among participants. This leads to a reduction in overtime hours and accurate sharing of information.

[0184] The following describes the processing flow.

[0185] Step 1:

[0186] The user schedules a meeting. Using the terminal's settings, they input the meeting date and time, participants, and agenda, and send this information to the server. The server stores the received information in a database using its management system.

[0187] Step 2:

[0188] The user initiates a meeting. At the specified date and time, they use the device's audio input device (microphone) to input voice data. The device acquires this voice data in real time and transmits it to the server using a communication device.

[0189] Step 3:

[0190] The server converts the received audio data into text data using speech recognition technology. The speech recognition algorithm analyzes the audio and converts it into text format.

[0191] Step 4:

[0192] The server's emotion engine analyzes the received audio data to identify the user's emotions. The emotion recognition algorithm analyzes the tone, speed, intonation, etc., of the voice to identify the user's emotions (joy, anger, sadness, surprise, etc.).

[0193] Step 5:

[0194] The server's text classification system categorizes text data and sentiment data based on the meeting content. Using natural language processing technology, it organizes the data by agenda item and category (shared items, homework, SOW, Q&A).

[0195] Step 6:

[0196] The server's document generation mechanism automatically generates meeting documents from classified text and sentiment data. The documents are created in the specified format (PDF, Word document, etc.).

[0197] Step 7:

[0198] The server's storage method involves saving generated meeting documents and sentiment data to a database or cloud storage. Document and sentiment metadata (creation date, creator, sentiment progression, etc.) are also saved simultaneously.

[0199] Step 8:

[0200] The server's transmission method sends the generated meeting documents and sentiment data to the user's terminal. The documents are sent via email or a file-sharing service.

[0201] Step 9:

[0202] Users review meeting documents and sentiment data received on their devices. They edit the documents as needed, incorporating revisions and additional information. They also collect feedback to improve future meetings.

[0203] Step 10:

[0204] The user distributes the final edited document to other participants. This can be done by resending it through the server or manually using email or a file-sharing service.

[0205] (Example 2)

[0206] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0207] Conventional meeting recording systems are often limited to functions such as converting voice input to text and classifying that text. They are unable to analyze user emotions and reflect them in meeting documents, making it difficult to accurately grasp changes in participants' emotions and opinions during a meeting. There is a need to provide a system that solves these problems and improves meeting efficiency and facilitates smoother communication among participants.

[0208] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0209] In this invention, the server includes voice acquisition means for the user to input voice data, communication means for transmitting voice data to the server in real time, voice recognition means for converting the received voice data into text data, text classification means for classifying the text data based on the meeting content, document generation means for automatically generating meeting documents from the classified text data, storage means for saving the generated meeting documents, transmission means for sending the generated meeting documents to the user, and emotion analysis means for analyzing emotions from the voice data. This makes it possible to not only transcribe what the user says into text in real time and classify the content, but also analyze emotion data and reflect it in the document.

[0210] "Voice acquisition means" refers to devices or systems that allow users to input voice, and includes microphones, recording devices, and the like.

[0211] "Communication methods" refer to technologies and devices for transmitting voice data to a server in real time, and include technologies using internet connections and dedicated applications.

[0212] "Speech recognition means" refers to algorithms and technologies used by a server to convert received speech data into text data, and includes speech recognition software and APIs.

[0213] "Text classification means" refers to natural language processing technology used by a server to classify text data based on the content of a meeting, and includes specific algorithms and software.

[0214] "Document generation means" refers to technology that enables a server to automatically generate meeting documents from classified text data, and includes software for formatting and layout creation.

[0215] "Storage methods" refer to technologies and devices for securely storing generated meeting documents, and include databases and cloud storage.

[0216] "Transmission means" refers to the technologies and devices used to send generated meeting documents to users, and includes email systems and file sharing services.

[0217] "Emotion analysis means" refers to engines and algorithms for analyzing emotions from voice data, and includes technologies for identifying the user's emotions (e.g., joy, anger, sadness, surprise, etc.).

[0218] This invention is a system that combines a system that acquires user voice input, transmits voice data to a server in real time, converts the voice data into text data, classifies it, and automatically generates, saves, and transmits meeting documents, with an emotion engine that recognizes the user's emotions.

[0219] First, the user uses the microphone built into their device (a PC or mobile device) to input voice data. This voice data is transmitted in real time from the device to the server via WebSocket technology through an internet connection or a dedicated application.

[0220] The server converts the received audio data into text data using a high-precision speech recognition algorithm (e.g., Google Speech-to-Text API or standard speech recognition software). This text data is then automatically categorized according to the meeting agenda, utilizing natural language processing technologies (e.g., spaCy or NLTK). It is then organized by agenda item or category (e.g., shared items, homework, SOW, Q&A, etc.).

[0221] Next, the server automatically generates meeting documents from text data, creating documents in the specified format (PDF, Word document, etc.). It uses LaTeX or the Microsoft® Word API to format the documents. These meeting documents are stored securely in a database or cloud storage (e.g., AWS® S3 or Google Drive). After the generation and saving of the meeting documents is complete, the server sends the generated documents to the user's device via email or a file-sharing service (e.g., Dropbox or Slack).

[0222] Furthermore, the system utilizes an emotion engine to analyze user emotions. This engine uses tools such as Amazon Comprehend and Microsoft Azure® Text Analytics to analyze user emotions (e.g., joy, anger, sadness, surprise, etc.) in real time from voice input. The results of the emotion analysis are reflected in the meeting documents, allowing users and participants to review their emotional changes during the meeting. This data is transmitted via a server and included in meeting minutes, task lists, and other documents generated after the meeting.

[0223] For example, in a new product development meeting, the user schedules the meeting and starts it at the designated time. The user uses a voice input device to input speech, and this speech data is sent to the server in real time. The server uses the Google Speech-to-Text API to convert the speech into text data and classifies it by topic and category using natural language processing technology (such as spaCy). Furthermore, Amazon Comprehend is used to analyze the user's emotions during speech, and this emotion data is also processed. After the meeting ends, the server automatically generates meeting documents using LaTeX or Microsoft Word APIs, saves them to AWS S3 or Google Drive, and sends them to the user via Dropbox or email. The user reviews the received documents, edits them, and gathers feedback. This information can then be used to prepare for the next meeting, allowing for more efficient discussions.

[0224] The following are examples of prompt statements:

[0225] "Please record what users say in meetings in real time, categorize the comments, and automatically generate meeting minutes. Also, record emotional information associated with each comment, save all the information after the meeting, and send it to the participants."

[0226] This system streamlines meetings and facilitates communication by collecting emotional data. This, in turn, leads to accurate information sharing and a reduction in overtime.

[0227] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0228] Specific flow of program processing

[0229] Step 1:

[0230] Acquisition of audio data

[0231] The user uses the device's microphone to input voice commands. For example, they might say, "Let's discuss the features of the new product." This voice data is captured as input by the device. Specifically, the device's microphone collects voice waveform data, which is temporarily stored in storage. The output is the captured raw voice data.

[0232] Step 2:

[0233] Sending audio data

[0234] The device captures audio data and sends it to the server in real time. WebSocket technology is used via an internet connection or a dedicated application. The input is the captured audio data, and the output is the streamed audio data sent to the server. Specifically, the device establishes a WebSocket connection and transfers the audio data to the server in real time.

[0235] Step 3:

[0236] Voice data recognition and conversion

[0237] The server processes the received audio data and converts it into text data using a highly accurate speech recognition algorithm. For example, it can use the Google Speech-to-Text API. The input is real-time streamed audio data, and the output is converted text data. Specifically, the server sends the audio data to the API and receives the text data returned from the API.

[0238] Step 4:

[0239] Classification of text data

[0240] The server uses natural language processing techniques to classify text data based on the meeting agenda. For example, tools such as spaCy are used. The input is text data obtained from speech recognition, and the output is text data classified by agenda item. Specifically, the server analyzes the text data and classifies it based on specific keywords or phrases.

[0241] Step 5:

[0242] Emotion analysis

[0243] The server uses Amazon Comprehend or a similar engine to analyze emotions from text data. The input is classified text data, and the output is analyzed emotion data (e.g., joy, anger, sadness, surprise, etc.). Specifically, text data is input into the emotion analysis engine, and emotion data is extracted.

[0244] Step 6:

[0245] Document generation

[0246] The server automatically generates meeting documents containing text and sentiment data using LaTeX and the Microsoft Word API. The input is categorized text and sentiment data, and the output is a formatted meeting document (PDF, Word document, etc.). Specifically, the server inserts data into a text template and uses a document generation tool to create the final document.

[0247] Step 7:

[0248] Save document

[0249] The server saves the generated meeting documents to cloud storage (e.g., AWS S3 or Google Drive). The input is the generated meeting documents, and the output is the documents saved on cloud storage. Specifically, the server uploads the data to cloud storage using an API.

[0250] Step 8:

[0251] Sending documents

[0252] The server sends meeting documents to the user's device. This is done, for example, through an email system or file sharing service (such as Dropbox or Slack). The input is a document stored in cloud storage, and the output is a document downloaded to the user's device. Specifically, the server sends the document using an email sending API or a file sharing API.

[0253] This allows users to receive meeting documents based on real-time converted text data and sentiment analysis results, leading to more efficient meetings and accurate information sharing.

[0254] (Application Example 2)

[0255] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0256] Current meeting systems have problems such as users having to manually create documents, which prevents them from focusing on the meeting's progress. In particular, conducting meetings efficiently while traveling in autonomous vehicles is difficult, and there is a need for technology to record meeting content in real time. Furthermore, there is a lack of technology to understand changes in emotions during meetings and facilitate smooth communication.

[0257] In Application Example 2, the specific processing performed by the specific processing unit 290 of the data processing device 12 is realized by the following means. In this invention, the server includes voice acquisition means for the user to perform voice input, communication means for transmitting voice data to the server in real time, voice recognition means for the server to convert the received voice data into text data, text classification means for the server to classify the text data based on the meeting content, document generation means for the server to automatically generate meeting documents from the classified text data, storage means for the server to save the generated meeting documents, transmission means for the server to send the generated meeting documents to the user, voice acquisition means for inside the autonomous vehicle to acquire voice input inside the autonomous vehicle, emotion analysis means for analyzing the user's emotions in real time, and display means for displaying the generated meeting documents on the vehicle display. This enables efficient meeting management inside the autonomous vehicle, and since the emotions of participants while they are speaking are analyzed and recorded in real time, communication is facilitated.

[0258] "Voice acquisition means" refers to a microphone or voice capture device used to acquire user voice input.

[0259] "Communication means" refers to equipment including an internet connection and communication protocols for transmitting voice data to a server in real time.

[0260] "Speech recognition means" refers to speech recognition algorithms and related software used by the server to convert received speech data into text data.

[0261] A "text classification means" is a device that utilizes natural language processing technology to allow a server to automatically classify text data into categories based on the content of a meeting.

[0262] "Document generation means" refers to software and tools that allow a server to automatically generate meeting documents based on classified text data.

[0263] "Storage method" refers to a database or cloud storage for securely saving generated meeting documents.

[0264] "Transmission method" refers to email or file-sharing services used to send generated meeting documents to users.

[0265] "Voice acquisition means within an autonomous vehicle" refers to an in-vehicle microphone or voice capture device used to acquire user voice input inside an autonomous vehicle.

[0266] "Emotion analysis means" refers to algorithms and related software for analyzing user emotions in real time from voice data.

[0267] "Display means" refers to a device and interface for displaying the generated meeting documents on a display inside the vehicle.

[0268] This invention is a system in which a user in an autonomous vehicle performs voice input, transmits that voice data to a server in real time, converts it into text data, and generates, saves, and transmits meeting documents. Furthermore, by combining it with an emotion analysis function that recognizes the user's emotions, changes in emotions during the meeting are also recorded.

[0269] System Configuration

[0270] 1. Means of acquiring sound:

[0271] This device uses an in-vehicle microphone installed inside an autonomous vehicle to acquire voice input from the user.

[0272] 2. Means of communication:

[0273] The acquired voice data is transmitted to the server in real time using an Internet connection (e.g., 5G or LTE). The WebSocket technology is used for this communication.

[0274] 3. Voice Recognition Means:

[0275] The server converts the received voice data into text data using a voice recognition algorithm such as Google Cloud Speech-to-Text.

[0276] 4. Text Classification Means:

[0277] The converted text data is automatically classified based on the meeting content using natural language processing technology (e.g., spaCy). The classified data is organized into each topic or category (shared items, tasks, Q&A, etc.).

[0278] 5. Document Generation Means:

[0279] The server generates a meeting document from the classified text data. At this time, a document is created in a specified format (PDF, Word document, etc.) using tools such as PDFKit or docx.

[0280] 6. Saving Means:

[0281] The generated meeting document is securely saved in a database or cloud storage (e.g., AWS S3).

[0282] 7. Sending Means:

[0283] The generated meeting document is sent to the user via email or a file sharing service. For example, email is sent using AWS SES.

[0284] 8. In-Vehicle Voice Acquisition Means for Autonomous Vehicles:

[0285] For voice acquisition inside an autonomous vehicle, specific high-sensitivity microphones are used, enabling clear voice input even while the vehicle is in motion.

[0286] 9. Sentiment analysis means:

[0287] The server uses a sentiment analysis engine such as IBM Watson (registered trademark) Tone Analyzer to analyze the sentiment from the user's voice data in real time. The analysis results are reflected in the meeting document.

[0288] 10. Display means:

[0289] The generated meeting document is displayed on the display inside the vehicle. For example, it is displayed on an in-vehicle display or a tablet terminal.

[0290] Usage example

[0291] In a new product development meeting, the user makes a voice input using the voice acquisition means inside the autonomous vehicle. This voice data is sent to the server in real time through WebSocket. The server uses Google Cloud Speech-to-Text to convert the voice data into text data, and further classifies the text data based on the meeting content using spaCy. The classified data is generated as a meeting document using PDFKit or docx and saved in AWS S3. The saved document is sent to the user using AWS SES. In addition, sentiment analysis is performed on each speech during the meeting using IBM Watson Tone Analyzer, and the analysis results are also included in the meeting document. Finally, the generated meeting document is displayed on the in-vehicle display so that the user can confirm it. An example of a prompt sentence is "What do you think about the price setting of the new product?"

[0292] This system enables efficient meeting management within autonomous vehicles, allowing participants to facilitate smoother communication by utilizing real-time voice input and sentiment analysis.

[0293] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0294] Step 1:

[0295] The user provides voice input within the autonomous vehicle. The input is acquired via the vehicle's microphone and transmitted as voice data to the vehicle's terminal. The output is voice data.

[0296] Step 2:

[0297] Audio data is transmitted in real time from a terminal inside the autonomous vehicle to a server via WebSocket. The input is audio data acquired from the vehicle's microphone, and the output is the audio data sent to the server.

[0298] Step 3:

[0299] The server converts the received audio data into text data using Google Cloud Speech-to-Text. The input is the audio data received via WebSocket, and the output is the converted text data. A speech recognition algorithm is then applied.

[0300] Step 4:

[0301] Text data is classified on the server using natural language processing techniques (e.g., spaCy). The input is the transformed text data, and the output is text data classified into meeting topics and categories. Syntactic and semantic analysis of the text is performed.

[0302] Step 5:

[0303] The classified text data is automatically generated by the server into a meeting document in a specified format, such as PDFKit or docx. The input is the classified text data, and the output is the meeting document (PDF or Word document). Here, a document template is used.

[0304] Step 6:

[0305] The server saves the generated meeting document in a database or cloud storage (e.g., AWS S3). The input is the generated meeting document, and the output is the document saved in cloud storage. The saving process is performed.

[0306] Step 7:

[0307] The server sends the saved meeting document to the user. At this time, email sending is performed using AWS SES. The input is the link information saved in cloud storage and the user's email address, and the output is the sent email.

[0308] Step 8:

[0309] During the meeting, the server performs sentiment analysis on the voice data in real-time using IBM Watson Tone Analyzer. The input is the voice data received via WebSocket, and the output is the analyzed sentiment data.

[0310] Step 9:

[0311] The sentiment analysis results are included in the meeting document together with the classified text data, and the user can check the content after the meeting ends. The input is the sentiment analysis results and the text data, and the output is the meeting document containing the sentiment data.

[0312] Step 10:

[0313] The generated meeting documents are displayed on the in-vehicle display, allowing the user to view them in real time. The input is the generated meeting documents, and the output is the content displayed on the in-vehicle display.

[0314] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0315] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0316] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0317] [Second Embodiment]

[0318] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0319] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0320] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0321] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0322] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0323] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0324] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0325] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0326] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0327] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0328] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0329] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0330] This invention is a system that acquires user voice input, transmits the voice data to a server in real time, converts the voice data into text data, classifies it, and automatically generates, saves, and transmits meeting documents. Specific embodiments are described below.

[0331] In this system, the user first performs voice input using a voice acquisition method. For example, the microphone built into the user's terminal (a PC or mobile device) is used. This voice data is transmitted in real time from the terminal to the server via a communication method (such as an internet connection or WebSocket technology via a dedicated application).

[0332] The server, acting as a speech recognition system, converts received audio data into text data using a highly accurate speech recognition algorithm. This text data is then automatically categorized according to the meeting agenda, utilizing natural language processing technology as a text classification system. Classification is performed not only by agenda item but also by category (e.g., shared items, homework, SOW, Q&A).

[0333] Next, based on the classified text data, the server, acting as a document generation mechanism, automatically generates meeting documents (such as meeting minutes and task lists). These documents are generated in the specified format (e.g., PDF or Word document). Furthermore, the generated documents are securely stored in a database or cloud storage.

[0334] After the meeting documents have been generated and saved, the server, acting as the transmission method, sends them to the user. The documents are sent via the specified email address or file-sharing service. The user's device is designed to display the received documents in a highly readable format.

[0335] Furthermore, users can set meeting details (date and time, participants, agenda, etc.) using a dedicated interface (web application or mobile app) before the meeting. This information is sent to the server and managed as metadata by the server's management system. Metadata management streamlines pre-meeting preparations and post-meeting documentation.

[0336] As a concrete example, consider a scenario where a user sets up a new product development meeting. First, the user schedules the meeting using the scheduling tool and notifies participants. On the day of the meeting, the user uses a terminal in the meeting room to send the audio of the conversation to the server using the audio acquisition tool. The server converts the audio to text in real time, categorizes it by agenda item, and automatically generates meeting documents immediately. Once the meeting ends, the generated minutes and task list are saved to cloud storage using the saving tool and quickly sent to the user's terminal via the transmission tool. The user and participants can review these materials and provide feedback as needed, further streamlining the preparation for the next meeting.

[0337] In this way, the system of the present invention dramatically improves the efficiency of meetings and provides an environment in which employees can concentrate on their actual work time. This leads to a reduction in overtime hours and accurate sharing of information.

[0338] The following describes the processing flow.

[0339] Step 1:

[0340] The user schedules a meeting. Using the terminal's settings, they input the meeting date and time, participants, and agenda, and send this information to the server. The server stores the received information in a database using its management system.

[0341] Step 2:

[0342] The user starts a meeting. At the specified date and time of the meeting, the user inputs voice data using the device's audio acquisition method (microphone). The device acquires this voice data in real time and transmits it to the server using a communication method.

[0343] Step 3:

[0344] The server converts the received audio data into text data using speech recognition technology. The speech recognition algorithm analyzes the audio and converts it into text format.

[0345] Step 4:

[0346] The server's text classification system categorizes text data based on meeting content. Using natural language processing technology, it organizes the text into categories such as meeting minutes, shared items, homework, statements of work (SOW), and Q&A.

[0347] Step 5:

[0348] The server's document generation mechanism automatically generates meeting documents from classified text data. The documents are created in the specified format (PDF, Word document, etc.).

[0349] Step 6:

[0350] The server's storage method involves saving the generated meeting documents to a database or cloud storage. The document's metadata (creation date, creator, etc.) is also saved simultaneously.

[0351] Step 7:

[0352] The server's transmission method sends the generated meeting documents to the user's terminal. The documents are sent via email or a file-sharing service.

[0353] Step 8:

[0354] Users review meeting documents received on their devices. They edit the documents as needed, incorporating any corrections or additional information. They also collect feedback to improve future meetings.

[0355] Step 9:

[0356] The user distributes the final edited document to other participants. This can be done by resending it through the server or manually using email or a file-sharing service.

[0357] (Example 1)

[0358] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0359] Traditional methods for accurately and quickly recording discussions during meetings and automatically generating minutes and task lists are time-consuming and laborious. Furthermore, smoothly sharing these materials with all meeting participants is also a cumbersome process. Therefore, there is a need for efficient meeting management and accurate sharing of important information.

[0360] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0361] In this invention, the server includes: voice acquisition means for the user to input voice data; communication means for transmitting voice data to a data processing device in real time; and voice recognition means for the data processing device to convert the received voice data into text data. This makes it possible to instantly transcribe, classify, and automatically generate, save, and share meeting documents based on what the user says during a meeting.

[0362] "Voice acquisition means" refers to devices or functions that allow users to perform voice input.

[0363] "Communication means" refers to technical methods and devices for transmitting voice data to a data processing device in real time.

[0364] "Speech recognition means" refers to algorithms and software used by a data processing device to convert received speech data into text data.

[0365] "Natural language processing means" refers to a technology used by data processing devices to classify text data based on the content of a meeting.

[0366] "Document generation means" refers to software or algorithms used by a data processing device to automatically generate meeting documents from classified text data.

[0367] "Storage means" refers to the technology and equipment used to save generated meeting documents to a storage medium.

[0368] "Transmission means" refers to the technology or functionality used to send generated meeting documents to users.

[0369] "Display means" refers to devices or software that display meeting documents received by a user's terminal in a highly visible format.

[0370] "Configuration means" refers to the interface or software that allows users to configure detailed meeting information.

[0371] "Management means" refers to the functions and technologies that data processing devices use to manage meeting metadata.

[0372] This invention relates to a system in which a user performs voice input, transmits the voice data to a server in real time, converts it into text data, classifies it, and automatically generates and saves meeting documents. Specific embodiments of the system are described in detail below.

[0373] overview

[0374] This system captures user speech and transmits it in real time to a data processing unit (server). The server converts the speech data into text data and classifies it based on the meeting content. The system automatically generates, saves, and transmits meeting documents from the classified text data. The aim of this system is to improve meeting efficiency and ensure accurate information sharing.

[0375] Hardware and software to use

[0376] The following hardware and software will be used to implement this system.

[0377] Audio acquisition method: User's device (PC, mobile device) equipped with a microphone device.

[0378] Communication methods: Internet connection, WebSocket technology, HTTP POST

[0379] Speech recognition method: Google Cloud Speech-to-Text API, Amazon Transcribe

[0380] Natural language processing tools: Google Cloud Natural Language API

[0381] Document generation methods: Python's reportlab library, docx library

[0382] Storage method: Cloud storage (Amazon S3, Google Drive)

[0383] Sending method: SendGrid, Google Drive shared link

[0384] Specific example

[0385] For example, let's consider a new product development meeting. The user sets the meeting date and participants using a web application and sends them to the data processing unit. On the day of the meeting, the user continuously sends audio of the conversation to the data processing unit using the microphone on their device. The data processing unit converts the audio to text in real time and categorizes it by agenda item. After the meeting, the data processing unit automatically generates meeting minutes and a task list and saves them to cloud storage. These documents are sent to the user and displayed on their device in a highly readable format. The following is an example of a prompt to be input into the generating AI model.

[0386] Example of a prompt:

[0387] Please analyze the audio data from the next meeting and generate meeting minutes categorized by agenda item.

[0388] Users can review this document and provide feedback as needed, further streamlining their preparations for the next meeting. The automated generation and sharing of meeting documents provides employees with an environment where they can focus on their actual work time.

[0389] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0390] Step 1: Obtaining voice input

[0391] The device acquires the user's voice input via the microphone. This voice data is input and temporarily stored in a buffer on the device side. This data is then sent to the next step in a short time for real-time processing.

[0392] Step 2: Sending the audio data

[0393] The terminal transmits the acquired audio data to the server in real time. This transmission uses an internet connection and WebSocket technology. The terminal streams the audio data in a buffer to the server via the WebSocket connection. As a specific example, the JavaScript WebSocket API is used.

[0394] Step 3: Receiving audio data

[0395] The server receives audio data transmitted from the terminal in real time. The server opens a WebSocket connection and receives the audio data as a stream. The received audio data is temporarily stored in internal memory.

[0396] Step 4: Convert speech to text

[0397] The server converts the received audio data into text data using a speech recognition algorithm. The input is audio data, and the output is the corresponding text. Specifically, the server uses the Google Cloud Speech-to-Text API to send audio data to the API and receive the resulting text data.

[0398] Step 5: Classification of text data

[0399] The server classifies text data by topic using natural language processing techniques. In this step, the input is text data, and the output is text data with classification tags. The Google Cloud Natural Language API is used to parse the text and assign appropriate tags, resulting in categorical text classification.

[0400] Step 6: Generating meeting documents

[0401] The server automatically generates meeting documents based on classified text data. The input is classified text data, and the output is meeting documents (PDF or Word documents). The Python reportlab and docx libraries are used to generate documents in the specified format.

[0402] Step 7: Save the document

[0403] The server saves the generated meeting documents to cloud storage. The input is the meeting documents, and the output is the files saved in cloud storage. Specifically, documents are uploaded using Amazon S3 or the Google Drive API.

[0404] Step 8: Send the document

[0405] The server sends the saved meeting documents to the user. The input is the meeting documents, and the output is the documents received by the user. A download link for the documents is sent to the user's email address using a SendGrid or Google Drive sharing link.

[0406] Step 9: Setting up the meeting

[0407] Users configure meeting details (date, time, participants, agenda, etc.) using a dedicated interface (web application or mobile app). The input is the meeting details, and the output is metadata generated based on that information. The metadata is sent to the server and stored and managed by a management system.

[0408] Step 10: Document Review and Feedback

[0409] Users review the received meeting documents and provide feedback as needed. The input is the meeting document, and the output is the revised document or additional feedback information. Users use this to prepare for future meetings.

[0410] (Application Example 1)

[0411] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0412] Traditional factory operations often involved manual processes for instructing workers and managing their progress, resulting in inefficiency and a heavy burden on workers. In particular, generating and managing work instructions and progress reports in real time was difficult, leading to delays and errors. Therefore, there is a growing need for increased efficiency in factory operations and real-time progress management.

[0413] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0414] In this invention, the server includes voice acquisition means for the user to input voice data, communication means for transmitting voice data to the server in real time, voice recognition means for the server to convert the received voice data into text data, text classification means for the server to classify the text data based on work instructions and progress status, document generation means for the server to automatically generate progress reports and task lists from the classified text data, storage means for the server to save the generated progress reports and task lists, and transmission means for the server to send the generated progress reports and task lists to the user. This makes it possible to efficiently provide instructions and manage the progress of factory work in real time.

[0415] "Voice acquisition means" refers to a device or interface for a user to perform voice input.

[0416] "Communication methods" refer to network technologies and protocols for transmitting voice data to a server in real time.

[0417] "Speech recognition means" refers to algorithms and software used by a server to convert received speech data into text data.

[0418] "Text classification means" refers to natural language processing technology used by a server to classify text data based on work instructions and progress status.

[0419] "Document generation means" refers to software or functionality that allows a server to automatically generate progress reports and task lists from classified text data.

[0420] "Storage methods" refer to technologies and systems for saving generated progress reports and task lists to databases or cloud storage.

[0421] "Transmission method" refers to the technology or protocol used by a server to send generated progress reports and task lists to users.

[0422] "Configuration means" refers to the interface or software that allows users to configure settings for their work.

[0423] "Management means" refers to the systems and technologies that servers use to manage metadata for their operations.

[0424] This invention is a system for managing work instructions and progress in real time within a factory, enabling efficient work operations. The elements constituting this system and their specific embodiments are described below.

[0425] First, the user performs voice input using a voice acquisition device. This device can be a head-mounted display or smart glasses equipped with a high-performance microphone. This data is transmitted in real time to a cloud server via a communication device. The communication device can utilize an internet connection or WebSocket technology through a dedicated application.

[0426] The cloud server converts received audio data into text data using speech recognition. High-precision speech recognition APIs such as Google Cloud Speech-to-Text are used for speech recognition. This text data is then classified by text classification tools based on work instructions and progress. Natural language processing libraries such as spaCy and NLTK are utilized for classification.

[0427] Next, progress reports and task lists are automatically generated from the classified text data by a document generation mechanism. The generated documents are stored in cloud storage such as Google Cloud Storage by a storage mechanism. Furthermore, the server transmits the generated documents to the user's head-mounted display via a transmission mechanism.

[0428] Users can check progress and next work instructions in real time through a head-mounted display. This improves work efficiency and reduces the burden on workers.

[0429] As a concrete example, suppose a user says, "I have finished preparing for step 1. Please give me the next instructions." This voice input is sent to a cloud server via a voice acquisition device. The cloud server converts the voice to text, and then uses Natural Language Processing (NLP) technology to classify the text as "Preparation for step 1 complete." Based on the classified text data, a progress report is automatically generated and displayed in real time on a head-mounted display.

[0430] Specific examples of prompt statements are as follows:

[0431] "Voice input: 'Preparation for step 1 is complete. Please give me the next instruction.'"

[0432] By using this system, work instructions and progress management within the factory can be carried out efficiently in real time, reducing delays and errors and improving overall production efficiency.

[0433] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0434] Step 1:

[0435] Users use a head-mounted display or smart glasses to input voice commands.

[0436] Input: User's voice instructions

[0437] Output: The microphone in the device acquires audio data.

[0438] Specific action: The user says, "I've finished preparing for step 1. Please give me the next instructions," and the audio data is captured by the microphone.

[0439] Step 2:

[0440] The device transmits the acquired audio data to the cloud server in real time.

[0441] Input: Acquired audio data

[0442] Output: Audio data sent to the cloud server

[0443] Specific operation: Audio data is sent to a cloud server over the internet using WebSocket technology.

[0444] Step 3:

[0445] The server converts the received audio data into text data using speech recognition technology.

[0446] Input: Audio data received by the server

[0447] Output: Audio data converted to text data

[0448] Specific operation: Using the Google Cloud Speech-to-Text API, the audio data is converted into text data that reads, "Step 1 is ready. Please give me the next instructions."

[0449] Step 4:

[0450] The server uses natural language processing techniques to classify text data based on work instructions and progress.

[0451] Input: Converted text data

[0452] Output: Classified text data

[0453] Specific operation: Text is used to perform a text analysis using spaCy or NLTK, and the progress status is determined as "Ready".

[0454] Step 5:

[0455] The server automatically generates progress reports and task lists from classified text data using document generation methods.

[0456] Input: Classified text data

[0457] Output: Progress report and task list

[0458] Specific operation: Document generation software on the server generates progress reports and task lists, and creates them in PDF or text format.

[0459] Step 6:

[0460] The server saves the generated progress reports and task lists to cloud storage.

[0461] Input: Generated progress reports and task lists

[0462] Output: Stored data in cloud storage

[0463] Specific action: Upload progress reports and task lists to Google Cloud Storage and store them securely.

[0464] Step 7:

[0465] The server sends generated progress reports and task lists to the user in real time.

[0466] Input: Generated progress reports and task lists

[0467] Output: Data sent to the user's head-mounted display

[0468] Specific operation: Progress reports and task lists are sent via the internet to the user's head-mounted display and displayed on the screen.

[0469] In this way, work instructions and progress management within the factory are carried out efficiently in real time through this series of processes. This system automates all steps from voice input to text conversion, progress report generation, saving, and transmission, thereby reducing the burden on workers and improving work efficiency.

[0470] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0471] This invention is a system that acquires user voice input, transmits the voice data to a server in real time, converts the voice data into text data, classifies it, and automatically generates, saves, and transmits meeting documents, in addition to an emotion engine that recognizes the user's emotions. Specific embodiments are described below.

[0472] In this system, the user first inputs voice data using a voice acquisition device. For example, the microphone built into the user's device (a PC or mobile device) is used. This voice data is transmitted in real time from the device to the server via a communication method (such as an internet connection or WebSocket technology via a dedicated application).

[0473] The server, acting as a speech recognition system, converts received audio data into text data using a highly accurate speech recognition algorithm. This text data is then automatically categorized according to the meeting agenda, utilizing natural language processing technology as a text classification system. The data is then organized by agenda item or category (e.g., shared items, homework, SOW, Q&A, etc.).

[0474] Next, to automatically generate meeting documents from text data, the server, acting as the document generation method, creates documents in the specified format (PDF, Word document, etc.). These meeting documents are stored securely in a database or cloud storage. After the generation and saving of the meeting documents is complete, the server sends the generated documents to the user's terminal using a transmission method. The documents are sent via email or a file sharing service.

[0475] Furthermore, this invention utilizes an emotion engine means for analyzing the user's emotions. The emotion engine means analyzes the user's emotions (e.g., joy, anger, sadness, surprise, etc.) in real time from voice input. The results of the emotion analysis are reflected in the meeting document, allowing users and participants to review changes in their emotions during the meeting as a record. This data is transmitted using a server's transmission means and included in meeting minutes, task lists, etc., generated after the meeting.

[0476] For example, in a new product development meeting, the user schedules the meeting and starts it at the designated time. The user inputs voice data using a voice acquisition device, and this voice data is sent to the server in real time. The server converts the voice into text data and categorizes it by topic and category. Furthermore, an emotion engine device analyzes the user's emotions at the time of speaking, and this emotion data is also processed. After the meeting ends, the server automatically generates a meeting document, saves it to cloud storage using a storage device, and sends it to the user using a transmission device. The user reviews the received document, edits it, and gathers feedback. By using this information to prepare for the next meeting, discussions can be conducted more efficiently.

[0477] Thus, the system of the present invention dramatically improves meeting efficiency and provides an environment where employees can concentrate on their actual work time. Furthermore, by collecting and sharing emotional data during meetings, it facilitates smoother communication among participants. This leads to a reduction in overtime hours and accurate sharing of information.

[0478] The following describes the processing flow.

[0479] Step 1:

[0480] The user schedules a meeting. Using the terminal's settings, they input the meeting date and time, participants, and agenda, and send this information to the server. The server stores the received information in a database using its management system.

[0481] Step 2:

[0482] The user initiates a meeting. At the specified date and time, they use the device's audio input device (microphone) to input voice data. The device acquires this voice data in real time and transmits it to the server using a communication device.

[0483] Step 3:

[0484] The server converts the received audio data into text data using speech recognition technology. The speech recognition algorithm analyzes the audio and converts it into text format.

[0485] Step 4:

[0486] The server's emotion engine analyzes the received audio data to identify the user's emotions. The emotion recognition algorithm analyzes the tone, speed, intonation, etc., of the voice to identify the user's emotions (joy, anger, sadness, surprise, etc.).

[0487] Step 5:

[0488] The server's text classification system categorizes text data and sentiment data based on the meeting content. Using natural language processing technology, it organizes the data by agenda item and category (shared items, homework, SOW, Q&A).

[0489] Step 6:

[0490] The server's document generation mechanism automatically generates meeting documents from classified text and sentiment data. The documents are created in the specified format (PDF, Word document, etc.).

[0491] Step 7:

[0492] The server's storage method involves saving generated meeting documents and sentiment data to a database or cloud storage. Document and sentiment metadata (creation date, creator, sentiment progression, etc.) are also saved simultaneously.

[0493] Step 8:

[0494] The server's transmission method sends the generated meeting documents and sentiment data to the user's terminal. The documents are sent via email or a file-sharing service.

[0495] Step 9:

[0496] Users review meeting documents and sentiment data received on their devices. They edit the documents as needed, incorporating revisions and additional information. They also collect feedback to improve future meetings.

[0497] Step 10:

[0498] The user distributes the final edited document to other participants. This can be done by resending it through the server or manually using email or a file-sharing service.

[0499] (Example 2)

[0500] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0501] Conventional meeting recording systems are often limited to functions such as converting voice input to text and classifying that text. They are unable to analyze user emotions and reflect them in meeting documents, making it difficult to accurately grasp changes in participants' emotions and opinions during a meeting. There is a need to provide a system that solves these problems and improves meeting efficiency and facilitates smoother communication among participants.

[0502] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0503] In this invention, the server includes voice acquisition means for the user to input voice data, communication means for transmitting voice data to the server in real time, voice recognition means for converting the received voice data into text data, text classification means for classifying the text data based on the meeting content, document generation means for automatically generating meeting documents from the classified text data, storage means for saving the generated meeting documents, transmission means for sending the generated meeting documents to the user, and emotion analysis means for analyzing emotions from the voice data. This makes it possible to not only transcribe what the user says into text in real time and classify the content, but also analyze emotion data and reflect it in the document.

[0504] "Voice acquisition means" refers to devices or systems that allow users to input voice, and includes microphones, recording devices, and the like.

[0505] "Communication methods" refer to technologies and devices for transmitting voice data to a server in real time, and include technologies using internet connections and dedicated applications.

[0506] "Speech recognition means" refers to algorithms and technologies used by a server to convert received speech data into text data, and includes speech recognition software and APIs.

[0507] "Text classification means" refers to natural language processing technology used by a server to classify text data based on the content of a meeting, and includes specific algorithms and software.

[0508] "Document generation means" refers to technology that enables a server to automatically generate meeting documents from classified text data, and includes software for formatting and layout creation.

[0509] "Storage methods" refer to technologies and devices for securely storing generated meeting documents, and include databases and cloud storage.

[0510] "Transmission means" refers to the technologies and devices used to send generated meeting documents to users, and includes email systems and file sharing services.

[0511] "Emotion analysis means" refers to engines and algorithms for analyzing emotions from voice data, and includes technologies for identifying the user's emotions (e.g., joy, anger, sadness, surprise, etc.).

[0512] This invention is a system that combines a system that acquires user voice input, transmits voice data to a server in real time, converts the voice data into text data, classifies it, and automatically generates, saves, and transmits meeting documents, with an emotion engine that recognizes the user's emotions.

[0513] First, the user uses the microphone built into their device (a PC or mobile device) to input voice data. This voice data is transmitted in real time from the device to the server via WebSocket technology through an internet connection or a dedicated application.

[0514] The server converts the received audio data into text data using a high-precision speech recognition algorithm (e.g., Google Speech-to-Text API or standard speech recognition software). This text data is then automatically categorized according to the meeting agenda, utilizing natural language processing technologies (e.g., spaCy or NLTK). It is then organized by agenda item or category (e.g., shared items, homework, SOW, Q&A, etc.).

[0515] Next, the server automatically generates meeting documents from the text data, creating documents in the specified format (PDF, Word document, etc.). It uses LaTeX or the Microsoft Word API to format the documents. These meeting documents are stored securely in a database or cloud storage (e.g., AWS S3 or Google Drive). After the generation and saving of the meeting documents is complete, the server sends the generated documents to the user's device via email or a file-sharing service (e.g., Dropbox or Slack).

[0516] Furthermore, the system utilizes an emotion engine to analyze user emotions. This engine uses tools such as Amazon Comprehend and Microsoft Azure Text Analytics to analyze user emotions (e.g., joy, anger, sadness, surprise, etc.) in real time from voice input. The results of the emotion analysis are reflected in the meeting documents, allowing users and participants to review their emotional changes during the meeting. This data is then transmitted via a server and included in meeting minutes, task lists, and other documents generated after the meeting.

[0517] For example, in a new product development meeting, the user schedules the meeting and starts it at the designated time. The user uses a voice input device to input speech, and this speech data is sent to the server in real time. The server uses the Google Speech-to-Text API to convert the speech into text data and classifies it by topic and category using natural language processing technology (such as spaCy). Furthermore, Amazon Comprehend is used to analyze the user's emotions during speech, and this emotion data is also processed. After the meeting ends, the server automatically generates meeting documents using LaTeX or Microsoft Word APIs, saves them to AWS S3 or Google Drive, and sends them to the user via Dropbox or email. The user reviews the received documents, edits them, and gathers feedback. This information can then be used to prepare for the next meeting, allowing for more efficient discussions.

[0518] The following are examples of prompt statements:

[0519] "Please record what users say in meetings in real time, categorize the comments, and automatically generate meeting minutes. Also, record emotional information associated with each comment, save all the information after the meeting, and send it to the participants."

[0520] This system streamlines meetings and facilitates communication by collecting emotional data. This, in turn, leads to accurate information sharing and a reduction in overtime.

[0521] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0522] Specific flow of program processing

[0523] Step 1:

[0524] Acquisition of audio data

[0525] The user uses the device's microphone to input voice commands. For example, they might say, "Let's discuss the features of the new product." This voice data is captured as input by the device. Specifically, the device's microphone collects voice waveform data, which is temporarily stored in storage. The output is the captured raw voice data.

[0526] Step 2:

[0527] Sending audio data

[0528] The device captures audio data and sends it to the server in real time. WebSocket technology is used via an internet connection or a dedicated application. The input is the captured audio data, and the output is the streamed audio data sent to the server. Specifically, the device establishes a WebSocket connection and transfers the audio data to the server in real time.

[0529] Step 3:

[0530] Voice data recognition and conversion

[0531] The server processes the received audio data and converts it into text data using a highly accurate speech recognition algorithm. For example, it can use the Google Speech-to-Text API. The input is real-time streamed audio data, and the output is converted text data. Specifically, the server sends the audio data to the API and receives the text data returned from the API.

[0532] Step 4:

[0533] Classification of text data

[0534] The server uses natural language processing techniques to classify text data based on the meeting agenda. For example, tools such as spaCy are used. The input is text data obtained from speech recognition, and the output is text data classified by agenda item. Specifically, the server analyzes the text data and classifies it based on specific keywords or phrases.

[0535] Step 5:

[0536] Emotion analysis

[0537] The server uses Amazon Comprehend or a similar engine to analyze emotions from text data. The input is classified text data, and the output is analyzed emotion data (e.g., joy, anger, sadness, surprise, etc.). Specifically, text data is input into the emotion analysis engine, and emotion data is extracted.

[0538] Step 6:

[0539] Document generation

[0540] The server automatically generates meeting documents containing text and sentiment data using LaTeX and the Microsoft Word API. The input is categorized text and sentiment data, and the output is a formatted meeting document (PDF, Word document, etc.). Specifically, the server inserts data into a text template and uses a document generation tool to create the final document.

[0541] Step 7:

[0542] Save document

[0543] The server saves the generated meeting documents to cloud storage (e.g., AWS S3 or Google Drive). The input is the generated meeting documents, and the output is the documents saved on cloud storage. Specifically, the server uploads the data to cloud storage using an API.

[0544] Step 8:

[0545] Sending documents

[0546] The server sends meeting documents to the user's device. This is done, for example, through an email system or file sharing service (such as Dropbox or Slack). The input is a document stored in cloud storage, and the output is a document downloaded to the user's device. Specifically, the server sends the document using an email sending API or a file sharing API.

[0547] This allows users to receive meeting documents based on real-time converted text data and sentiment analysis results, leading to more efficient meetings and accurate information sharing.

[0548] (Application Example 2)

[0549] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0550] Current meeting systems have problems such as users having to manually create documents, which prevents them from focusing on the meeting's progress. In particular, conducting meetings efficiently while traveling in autonomous vehicles is difficult, and there is a need for technology to record meeting content in real time. Furthermore, there is a lack of technology to understand changes in emotions during meetings and facilitate smooth communication.

[0551] In Application Example 2, the specific processing performed by the specific processing unit 290 of the data processing device 12 is realized by the following means. In this invention, the server includes voice acquisition means for the user to perform voice input, communication means for transmitting voice data to the server in real time, voice recognition means for the server to convert the received voice data into text data, text classification means for the server to classify the text data based on the meeting content, document generation means for the server to automatically generate meeting documents from the classified text data, storage means for the server to save the generated meeting documents, transmission means for the server to send the generated meeting documents to the user, voice acquisition means for inside the autonomous vehicle to acquire voice input inside the autonomous vehicle, emotion analysis means for analyzing the user's emotions in real time, and display means for displaying the generated meeting documents on the vehicle display. This enables efficient meeting management inside the autonomous vehicle, and since the emotions of participants while they are speaking are analyzed and recorded in real time, communication is facilitated.

[0552] "Voice acquisition means" refers to a microphone or voice capture device used to acquire user voice input.

[0553] "Communication means" refers to equipment including an internet connection and communication protocols for transmitting voice data to a server in real time.

[0554] "Speech recognition means" refers to speech recognition algorithms and related software used by the server to convert received speech data into text data.

[0555] A "text classification means" is a device that utilizes natural language processing technology to allow a server to automatically classify text data into categories based on the content of a meeting.

[0556] "Document generation means" refers to software and tools that allow a server to automatically generate meeting documents based on classified text data.

[0557] "Storage method" refers to a database or cloud storage for securely saving generated meeting documents.

[0558] "Transmission method" refers to email or file-sharing services used to send generated meeting documents to users.

[0559] "Voice acquisition means within an autonomous vehicle" refers to an in-vehicle microphone or voice capture device used to acquire user voice input inside an autonomous vehicle.

[0560] "Emotion analysis means" refers to algorithms and related software for analyzing user emotions in real time from voice data.

[0561] "Display means" refers to a device and interface for displaying the generated meeting documents on a display inside the vehicle.

[0562] This invention is a system in which a user in an autonomous vehicle performs voice input, transmits that voice data to a server in real time, converts it into text data, and generates, saves, and transmits meeting documents. Furthermore, by combining it with an emotion analysis function that recognizes the user's emotions, changes in emotions during the meeting are also recorded.

[0563] System Configuration

[0564] 1. Means of acquiring sound:

[0565] This device uses an in-vehicle microphone installed inside an autonomous vehicle to acquire voice input from the user.

[0566] 2. Means of communication:

[0567] The acquired voice data is transmitted to the server in real time using an internet connection (e.g., 5G or LTE). WebSocket technology is used for this communication.

[0568] 3. Voice recognition means:

[0569] The server converts the received audio data into text data using speech recognition algorithms such as Google Cloud Speech-to-Text.

[0570] 4. Text classification means:

[0571] The converted text data is automatically categorized based on the meeting content using natural language processing technology (e.g., spaCy). The categorized data is then organized into various agenda items and categories (e.g., shared items, homework, Q&A).

[0572] 5. Document generation means:

[0573] The server generates meeting documents from the classified text data. During this process, tools such as PDFKit and docx are used to create documents in the specified format (PDF, Word document, etc.).

[0574] 6. Preservation means:

[0575] The generated meeting documents are securely stored in a database or cloud storage (e.g., AWS S3).

[0576] 7. Transmission method:

[0577] The generated meeting documents are sent to users via email or file sharing services. For example, email sending is done using AWS SES.

[0578] 8. Means for acquiring voice data inside an autonomous vehicle:

[0579] A specific high-sensitivity microphone is used to acquire voice data inside the autonomous vehicle, enabling clear voice input even while the vehicle is in motion.

[0580] 9. Emotion analysis means:

[0581] The server uses emotion analysis engines such as IBM Watson Tone Analyzer to analyze the user's emotions in real time from their voice data. The analysis results are reflected in the meeting documents.

[0582] 10. Display means:

[0583] The generated meeting documents will be displayed on the vehicle's display, for example, on an in-car display or a tablet device.

[0584] Usage example

[0585] In a new product development meeting, the user provides voice input using a voice acquisition device inside the autonomous vehicle. This voice data is transmitted to a server in real time via WebSocket. The server converts the voice data into text data using Google Cloud Speech-to-Text, and then classifies the text data based on the meeting content using spaCy. The classified data is generated as a meeting document using PDFKit or docx and stored in AWS S3. The stored document is sent to the user using AWS SES. In addition, sentiment analysis is performed on each statement during the meeting using IBM Watson Tone Analyzer, and the analysis results are also included in the meeting document. Finally, the generated meeting document is displayed on the in-vehicle display for the user to review. An example of a prompt phrase used is, "What are your thoughts on pricing the new product?"

[0586] This system enables efficient meeting management within autonomous vehicles, allowing participants to facilitate smoother communication by utilizing real-time voice input and sentiment analysis.

[0587] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0588] Step 1:

[0589] The user provides voice input within the autonomous vehicle. The input is acquired via the vehicle's microphone and transmitted as voice data to the vehicle's terminal. The output is voice data.

[0590] Step 2:

[0591] Audio data is transmitted in real time from a terminal inside the autonomous vehicle to a server via WebSocket. The input is audio data acquired from the vehicle's microphone, and the output is the audio data sent to the server.

[0592] Step 3:

[0593] The server converts the received audio data into text data using Google Cloud Speech-to-Text. The input is the audio data received via WebSocket, and the output is the converted text data. A speech recognition algorithm is then applied.

[0594] Step 4:

[0595] Text data is classified on the server using natural language processing techniques (e.g., spaCy). The input is the transformed text data, and the output is text data classified into meeting topics and categories. Syntactic and semantic analysis of the text is performed.

[0596] Step 5:

[0597] The classified text data is automatically generated by the server into a meeting document in the specified format using tools such as PDFKit or docx. The input is the classified text data, and the output is a meeting document (PDF or Word document). A document template is used here.

[0598] Step 6:

[0599] The server saves the generated meeting documents to a database or cloud storage (e.g., AWS S3). The input is the generated meeting documents, and the output is the documents saved in cloud storage. The saving process is then performed.

[0600] Step 7:

[0601] The server sends the saved meeting documents to the user. This email transmission is performed using AWS SES. The input is the link information stored in cloud storage and the user's email address, and the output is the sent email.

[0602] Step 8:

[0603] During the meeting, the server performs real-time sentiment analysis on audio data using IBM Watson Tone Analyzer. The input is audio data received via WebSocket, and the output is the analyzed sentiment data.

[0604] Step 9:

[0605] The sentiment analysis results, along with categorized text data, are included in the meeting document, which users can review after the meeting. The input is the sentiment analysis results and text data, and the output is the meeting document containing the sentiment data.

[0606] Step 10:

[0607] The generated meeting documents are displayed on the in-vehicle display, allowing the user to view them in real time. The input is the generated meeting documents, and the output is the content displayed on the in-vehicle display.

[0608] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0609] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0610] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0611] [Third Embodiment]

[0612] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0613] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0614] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0615] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0616] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0617] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0618] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0619] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0620] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0621] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0622] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0623] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0624] This invention is a system that acquires user voice input, transmits the voice data to a server in real time, converts the voice data into text data, classifies it, and automatically generates, saves, and transmits meeting documents. Specific embodiments are described below.

[0625] In this system, the user first performs voice input using a voice acquisition method. For example, the microphone built into the user's terminal (a PC or mobile device) is used. This voice data is transmitted in real time from the terminal to the server via a communication method (such as an internet connection or WebSocket technology via a dedicated application).

[0626] The server, acting as a speech recognition system, converts received audio data into text data using a highly accurate speech recognition algorithm. This text data is then automatically categorized according to the meeting agenda, utilizing natural language processing technology as a text classification system. Classification is performed not only by agenda item but also by category (e.g., shared items, homework, SOW, Q&A).

[0627] Next, based on the classified text data, the server, acting as a document generation mechanism, automatically generates meeting documents (such as meeting minutes and task lists). These documents are generated in the specified format (e.g., PDF or Word document). Furthermore, the generated documents are securely stored in a database or cloud storage.

[0628] After the meeting documents have been generated and saved, the server, acting as the transmission method, sends them to the user. The documents are sent via the specified email address or file-sharing service. The user's device is designed to display the received documents in a highly readable format.

[0629] Furthermore, users can set meeting details (date and time, participants, agenda, etc.) using a dedicated interface (web application or mobile app) before the meeting. This information is sent to the server and managed as metadata by the server's management system. Metadata management streamlines pre-meeting preparations and post-meeting documentation.

[0630] As a concrete example, consider a scenario where a user sets up a new product development meeting. First, the user schedules the meeting using the scheduling tool and notifies participants. On the day of the meeting, the user uses a terminal in the meeting room to send the audio of the conversation to the server using the audio acquisition tool. The server converts the audio to text in real time, categorizes it by agenda item, and automatically generates meeting documents immediately. Once the meeting ends, the generated minutes and task list are saved to cloud storage using the saving tool and quickly sent to the user's terminal via the transmission tool. The user and participants can review these materials and provide feedback as needed, further streamlining the preparation for the next meeting.

[0631] In this way, the system of the present invention dramatically improves the efficiency of meetings and provides an environment in which employees can concentrate on their actual work time. This leads to a reduction in overtime hours and accurate sharing of information.

[0632] The following describes the processing flow.

[0633] Step 1:

[0634] The user schedules a meeting. Using the terminal's settings, they input the meeting date and time, participants, and agenda, and send this information to the server. The server stores the received information in a database using its management system.

[0635] Step 2:

[0636] The user starts a meeting. At the specified date and time of the meeting, the user inputs voice data using the device's audio acquisition method (microphone). The device acquires this voice data in real time and transmits it to the server using a communication method.

[0637] Step 3:

[0638] The server converts the received audio data into text data using speech recognition technology. The speech recognition algorithm analyzes the audio and converts it into text format.

[0639] Step 4:

[0640] The server's text classification system categorizes text data based on meeting content. Using natural language processing technology, it organizes the text into categories such as meeting minutes, shared items, homework, statements of work (SOW), and Q&A.

[0641] Step 5:

[0642] The server's document generation mechanism automatically generates meeting documents from classified text data. The documents are created in the specified format (PDF, Word document, etc.).

[0643] Step 6:

[0644] The server's storage method involves saving the generated meeting documents to a database or cloud storage. The document's metadata (creation date, creator, etc.) is also saved simultaneously.

[0645] Step 7:

[0646] The server's transmission method sends the generated meeting documents to the user's terminal. The documents are sent via email or a file-sharing service.

[0647] Step 8:

[0648] Users review meeting documents received on their devices. They edit the documents as needed, incorporating any corrections or additional information. They also collect feedback to improve future meetings.

[0649] Step 9:

[0650] The user distributes the final edited document to other participants. This can be done by resending it through the server or manually using email or a file-sharing service.

[0651] (Example 1)

[0652] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0653] Traditional methods for accurately and quickly recording discussions during meetings and automatically generating minutes and task lists are time-consuming and laborious. Furthermore, smoothly sharing these materials with all meeting participants is also a cumbersome process. Therefore, there is a need for efficient meeting management and accurate sharing of important information.

[0654] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0655] In this invention, the server includes: voice acquisition means for the user to input voice data; communication means for transmitting voice data to a data processing device in real time; and voice recognition means for the data processing device to convert the received voice data into text data. This makes it possible to instantly transcribe, classify, and automatically generate, save, and share meeting documents based on what the user says during a meeting.

[0656] "Voice acquisition means" refers to devices or functions that allow users to perform voice input.

[0657] "Communication means" refers to technical methods and devices for transmitting voice data to a data processing device in real time.

[0658] "Speech recognition means" refers to algorithms and software used by a data processing device to convert received speech data into text data.

[0659] "Natural language processing means" refers to a technology used by data processing devices to classify text data based on the content of a meeting.

[0660] "Document generation means" refers to software or algorithms used by a data processing device to automatically generate meeting documents from classified text data.

[0661] "Storage means" refers to the technology and equipment used to save generated meeting documents to a storage medium.

[0662] "Transmission means" refers to the technology or functionality used to send generated meeting documents to users.

[0663] "Display means" refers to devices or software that display meeting documents received by a user's terminal in a highly visible format.

[0664] "Configuration means" refers to the interface or software that allows users to configure detailed meeting information.

[0665] "Management means" refers to the functions and technologies that data processing devices use to manage meeting metadata.

[0666] This invention relates to a system in which a user performs voice input, transmits the voice data to a server in real time, converts it into text data, classifies it, and automatically generates and saves meeting documents. Specific embodiments of the system are described in detail below.

[0667] overview

[0668] This system captures user speech and transmits it in real time to a data processing unit (server). The server converts the speech data into text data and classifies it based on the meeting content. The system automatically generates, saves, and transmits meeting documents from the classified text data. The aim of this system is to improve meeting efficiency and ensure accurate information sharing.

[0669] Hardware and software to be used

[0670] The following hardware and software will be used to implement this system.

[0671] Audio acquisition method: User's device (PC, mobile device) equipped with a microphone device.

[0672] Communication methods: Internet connection, WebSocket technology, HTTP POST

[0673] Speech recognition method: Google Cloud Speech-to-Text API, Amazon Transcribe

[0674] Natural language processing tools: Google Cloud Natural Language API

[0675] Document generation methods: Python's reportlab library, docx library

[0676] Storage method: Cloud storage (Amazon S3, Google Drive)

[0677] Sending method: SendGrid, Google Drive shared link

[0678] Specific example

[0679] For example, let's consider a new product development meeting. The user sets the meeting date and participants using a web application and sends them to the data processing unit. On the day of the meeting, the user continuously sends audio of the conversation to the data processing unit using the microphone on their device. The data processing unit converts the audio to text in real time and categorizes it by agenda item. After the meeting, the data processing unit automatically generates meeting minutes and a task list and saves them to cloud storage. These documents are sent to the user and displayed on their device in a highly readable format. The following is an example of a prompt to be input into the generating AI model.

[0680] Example of a prompt:

[0681] Please analyze the audio data from the next meeting and generate meeting minutes categorized by agenda item.

[0682] Users can review this document and provide feedback as needed, further streamlining their preparations for the next meeting. The automated generation and sharing of meeting documents provides employees with an environment where they can focus on their actual work time.

[0683] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0684] Step 1: Obtaining voice input

[0685] The device acquires the user's voice input via the microphone. This voice data is input and temporarily stored in a buffer on the device side. This data is then sent to the next step in a short time for real-time processing.

[0686] Step 2: Sending the audio data

[0687] The terminal transmits the acquired audio data to the server in real time. This transmission uses an internet connection and WebSocket technology. The terminal streams the audio data in a buffer to the server via the WebSocket connection. As a specific example, the JavaScript WebSocket API is used.

[0688] Step 3: Receiving audio data

[0689] The server receives audio data transmitted from the terminal in real time. The server opens a WebSocket connection and receives the audio data as a stream. The received audio data is temporarily stored in internal memory.

[0690] Step 4: Convert speech to text

[0691] The server converts the received audio data into text data using a speech recognition algorithm. The input is audio data, and the output is the corresponding text. Specifically, the server uses the Google Cloud Speech-to-Text API to send audio data to the API and receive the resulting text data.

[0692] Step 5: Classification of text data

[0693] The server classifies text data by topic using natural language processing techniques. In this step, the input is text data, and the output is text data with classification tags. The Google Cloud Natural Language API is used to parse the text and assign appropriate tags, resulting in categorical text classification.

[0694] Step 6: Generating meeting documents

[0695] The server automatically generates meeting documents based on classified text data. The input is classified text data, and the output is meeting documents (PDF or Word documents). The Python reportlab and docx libraries are used to generate documents in the specified format.

[0696] Step 7: Save the document

[0697] The server saves the generated meeting documents to cloud storage. The input is the meeting documents, and the output is the files saved in cloud storage. Specifically, documents are uploaded using Amazon S3 or the Google Drive API.

[0698] Step 8: Send the document

[0699] The server sends the saved meeting documents to the user. The input is the meeting documents, and the output is the documents received by the user. A download link for the documents is sent to the user's email address using a SendGrid or Google Drive sharing link.

[0700] Step 9: Setting up the meeting

[0701] Users configure meeting details (date, time, participants, agenda, etc.) using a dedicated interface (web application or mobile app). The input is the meeting details, and the output is metadata generated based on that information. The metadata is sent to the server and stored and managed by a management system.

[0702] Step 10: Document Review and Feedback

[0703] Users review the received meeting documents and provide feedback as needed. The input is the meeting document, and the output is the revised document or additional feedback information. Users use this to prepare for future meetings.

[0704] (Application Example 1)

[0705] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0706] Traditional factory operations often involved manual processes for instructing workers and managing their progress, resulting in inefficiency and a heavy burden on workers. In particular, generating and managing work instructions and progress reports in real time was difficult, leading to delays and errors. Therefore, there is a growing need for increased efficiency in factory operations and real-time progress management.

[0707] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0708] In this invention, the server includes voice acquisition means for the user to input voice data, communication means for transmitting voice data to the server in real time, voice recognition means for the server to convert the received voice data into text data, text classification means for the server to classify the text data based on work instructions and progress status, document generation means for the server to automatically generate progress reports and task lists from the classified text data, storage means for the server to save the generated progress reports and task lists, and transmission means for the server to send the generated progress reports and task lists to the user. This makes it possible to efficiently provide instructions and manage the progress of factory work in real time.

[0709] "Voice acquisition means" refers to a device or interface for a user to perform voice input.

[0710] "Communication methods" refer to network technologies and protocols for transmitting voice data to a server in real time.

[0711] "Speech recognition means" refers to algorithms and software used by a server to convert received speech data into text data.

[0712] "Text classification means" refers to natural language processing technology used by a server to classify text data based on work instructions and progress status.

[0713] "Document generation means" refers to software or functionality that allows a server to automatically generate progress reports and task lists from classified text data.

[0714] "Storage methods" refer to technologies and systems for saving generated progress reports and task lists to databases or cloud storage.

[0715] "Transmission method" refers to the technology or protocol used by a server to send generated progress reports and task lists to users.

[0716] "Configuration means" refers to the interface or software that allows users to configure settings for their work.

[0717] "Management means" refers to the systems and technologies that servers use to manage metadata for their operations.

[0718] This invention is a system for managing work instructions and progress in real time within a factory, enabling efficient work operations. The elements constituting this system and their specific embodiments are described below.

[0719] First, the user performs voice input using a voice acquisition device. This device can be a head-mounted display or smart glasses equipped with a high-performance microphone. This data is transmitted in real time to a cloud server via a communication device. The communication device can utilize an internet connection or WebSocket technology through a dedicated application.

[0720] The cloud server converts received audio data into text data using speech recognition. High-precision speech recognition APIs such as Google Cloud Speech-to-Text are used for speech recognition. This text data is then classified by text classification tools based on work instructions and progress. Natural language processing libraries such as spaCy and NLTK are utilized for classification.

[0721] Next, progress reports and task lists are automatically generated from the classified text data by a document generation mechanism. The generated documents are stored in cloud storage such as Google Cloud Storage by a storage mechanism. Furthermore, the server transmits the generated documents to the user's head-mounted display via a transmission mechanism.

[0722] Users can check progress and next work instructions in real time through a head-mounted display. This improves work efficiency and reduces the burden on workers.

[0723] As a concrete example, suppose a user says, "I have finished preparing for step 1. Please give me the next instructions." This voice input is sent to a cloud server via a voice acquisition device. The cloud server converts the voice to text, and then uses Natural Language Processing (NLP) technology to classify the text as "Preparation for step 1 complete." Based on the classified text data, a progress report is automatically generated and displayed in real time on a head-mounted display.

[0724] Specific examples of prompt statements are as follows:

[0725] "Voice input: 'Preparation for step 1 is complete. Please give me the next instruction.'"

[0726] By using this system, work instructions and progress management within the factory can be carried out efficiently in real time, reducing delays and errors and improving overall production efficiency.

[0727] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0728] Step 1:

[0729] Users use a head-mounted display or smart glasses to input voice commands.

[0730] Input: User's voice instructions

[0731] Output: The microphone in the device acquires audio data.

[0732] Specific action: The user says, "I've finished preparing for step 1. Please give me the next instructions," and the audio data is captured by the microphone.

[0733] Step 2:

[0734] The device transmits the acquired audio data to the cloud server in real time.

[0735] Input: Acquired audio data

[0736] Output: Audio data sent to the cloud server

[0737] Specific operation: Audio data is sent to a cloud server over the internet using WebSocket technology.

[0738] Step 3:

[0739] The server converts the received audio data into text data using speech recognition technology.

[0740] Input: Audio data received by the server

[0741] Output: Audio data converted to text data

[0742] Specific operation: Using the Google Cloud Speech-to-Text API, the audio data is converted into text data that reads, "Step 1 is ready. Please give me the next instructions."

[0743] Step 4:

[0744] The server uses natural language processing techniques to classify text data based on work instructions and progress.

[0745] Input: Converted text data

[0746] Output: Classified text data

[0747] Specific operation: Text is used to perform a text analysis using spaCy or NLTK, and the progress status is determined as "Ready".

[0748] Step 5:

[0749] The server automatically generates progress reports and task lists from classified text data using document generation methods.

[0750] Input: Classified text data

[0751] Output: Progress report and task list

[0752] Specific operation: Document generation software on the server generates progress reports and task lists, and creates them in PDF or text format.

[0753] Step 6:

[0754] The server saves the generated progress reports and task lists to cloud storage.

[0755] Input: Generated progress reports and task lists

[0756] Output: Stored data in cloud storage

[0757] Specific action: Upload progress reports and task lists to Google Cloud Storage and store them securely.

[0758] Step 7:

[0759] The server sends generated progress reports and task lists to the user in real time.

[0760] Input: Generated progress reports and task lists

[0761] Output: Data sent to the user's head-mounted display

[0762] Specific operation: Progress reports and task lists are sent via the internet to the user's head-mounted display and displayed on the screen.

[0763] In this way, work instructions and progress management within the factory are carried out efficiently in real time through this series of processes. This system automates all steps from voice input to text conversion, progress report generation, saving, and transmission, thereby reducing the burden on workers and improving work efficiency.

[0764] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0765] This invention is a system that acquires user voice input, transmits the voice data to a server in real time, converts the voice data into text data, classifies it, and automatically generates, saves, and transmits meeting documents, in addition to an emotion engine that recognizes the user's emotions. Specific embodiments are described below.

[0766] In this system, the user first inputs voice data using a voice acquisition device. For example, the microphone built into the user's device (a PC or mobile device) is used. This voice data is transmitted in real time from the device to the server via a communication method (such as an internet connection or WebSocket technology via a dedicated application).

[0767] The server, acting as a speech recognition system, converts received audio data into text data using a highly accurate speech recognition algorithm. This text data is then automatically categorized according to the meeting agenda, utilizing natural language processing technology as a text classification system. The data is then organized by agenda item or category (e.g., shared items, homework, SOW, Q&A, etc.).

[0768] Next, to automatically generate meeting documents from text data, the server, acting as the document generation method, creates documents in the specified format (PDF, Word document, etc.). These meeting documents are stored securely in a database or cloud storage. After the generation and saving of the meeting documents is complete, the server sends the generated documents to the user's terminal using a transmission method. The documents are sent via email or a file sharing service.

[0769] Furthermore, this invention utilizes an emotion engine means for analyzing the user's emotions. The emotion engine means analyzes the user's emotions (e.g., joy, anger, sadness, surprise, etc.) in real time from voice input. The results of the emotion analysis are reflected in the meeting document, allowing users and participants to review changes in their emotions during the meeting as a record. This data is transmitted using a server's transmission means and included in meeting minutes, task lists, etc., generated after the meeting.

[0770] For example, in a new product development meeting, the user schedules the meeting and starts it at the designated time. The user inputs voice data using a voice acquisition device, and this voice data is sent to the server in real time. The server converts the voice into text data and categorizes it by topic and category. Furthermore, an emotion engine device analyzes the user's emotions at the time of speaking, and this emotion data is also processed. After the meeting ends, the server automatically generates a meeting document, saves it to cloud storage using a storage device, and sends it to the user using a transmission device. The user reviews the received document, edits it, and gathers feedback. By using this information to prepare for the next meeting, discussions can be conducted more efficiently.

[0771] Thus, the system of the present invention dramatically improves meeting efficiency and provides an environment where employees can concentrate on their actual work time. Furthermore, by collecting and sharing emotional data during meetings, it facilitates smoother communication among participants. This leads to a reduction in overtime hours and accurate sharing of information.

[0772] The following describes the processing flow.

[0773] Step 1:

[0774] The user schedules a meeting. Using the terminal's settings, they input the meeting date and time, participants, and agenda, and send this information to the server. The server stores the received information in a database using its management system.

[0775] Step 2:

[0776] The user initiates a meeting. At the specified date and time, they use the device's audio input device (microphone) to input voice data. The device acquires this voice data in real time and transmits it to the server using a communication device.

[0777] Step 3:

[0778] The server converts the received audio data into text data using speech recognition technology. The speech recognition algorithm analyzes the audio and converts it into text format.

[0779] Step 4:

[0780] The server's emotion engine analyzes the received audio data to identify the user's emotions. The emotion recognition algorithm analyzes the tone, speed, intonation, etc., of the voice to identify the user's emotions (joy, anger, sadness, surprise, etc.).

[0781] Step 5:

[0782] The server's text classification system categorizes text data and sentiment data based on the meeting content. Using natural language processing technology, it organizes the data by agenda item and category (shared items, homework, SOW, Q&A).

[0783] Step 6:

[0784] The server's document generation mechanism automatically generates meeting documents from classified text and sentiment data. The documents are created in the specified format (PDF, Word document, etc.).

[0785] Step 7:

[0786] The server's storage method involves saving generated meeting documents and sentiment data to a database or cloud storage. Document and sentiment metadata (creation date, creator, sentiment progression, etc.) are also saved simultaneously.

[0787] Step 8:

[0788] The server's transmission method sends the generated meeting documents and sentiment data to the user's terminal. The documents are sent via email or a file-sharing service.

[0789] Step 9:

[0790] Users review meeting documents and sentiment data received on their devices. They edit the documents as needed, incorporating revisions and additional information. They also collect feedback to improve future meetings.

[0791] Step 10:

[0792] The user distributes the final edited document to other participants. This can be done by resending it through the server or manually using email or a file-sharing service.

[0793] (Example 2)

[0794] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0795] Conventional meeting recording systems are often limited to functions such as converting voice input to text and classifying that text. They are unable to analyze user emotions and reflect them in meeting documents, making it difficult to accurately grasp changes in participants' emotions and opinions during a meeting. There is a need to provide a system that solves these problems and improves meeting efficiency and facilitates smoother communication among participants.

[0796] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0797] In this invention, the server includes voice acquisition means for the user to input voice data, communication means for transmitting voice data to the server in real time, voice recognition means for converting the received voice data into text data, text classification means for classifying the text data based on the meeting content, document generation means for automatically generating meeting documents from the classified text data, storage means for saving the generated meeting documents, transmission means for sending the generated meeting documents to the user, and emotion analysis means for analyzing emotions from the voice data. This makes it possible to not only transcribe what the user says into text in real time and classify the content, but also analyze emotion data and reflect it in the document.

[0798] "Voice acquisition means" refers to devices or systems that allow users to input voice, and includes microphones, recording devices, and the like.

[0799] "Communication methods" refer to technologies and devices for transmitting voice data to a server in real time, and include technologies using internet connections and dedicated applications.

[0800] "Speech recognition means" refers to algorithms and technologies used by a server to convert received speech data into text data, and includes speech recognition software and APIs.

[0801] "Text classification means" refers to natural language processing technology used by a server to classify text data based on the content of a meeting, and includes specific algorithms and software.

[0802] "Document generation means" refers to technology that enables a server to automatically generate meeting documents from classified text data, and includes software for formatting and layout creation.

[0803] "Storage methods" refer to technologies and devices for securely storing generated meeting documents, and include databases and cloud storage.

[0804] "Transmission means" refers to the technologies and devices used to send generated meeting documents to users, and includes email systems and file sharing services.

[0805] "Emotion analysis means" refers to engines and algorithms for analyzing emotions from voice data, and includes technologies for identifying the user's emotions (e.g., joy, anger, sadness, surprise, etc.).

[0806] This invention is a system that combines a system that acquires user voice input, transmits voice data to a server in real time, converts the voice data into text data, classifies it, and automatically generates, saves, and transmits meeting documents, with an emotion engine that recognizes the user's emotions.

[0807] First, the user uses the microphone built into their device (a PC or mobile device) to input voice data. This voice data is transmitted in real time from the device to the server via WebSocket technology through an internet connection or a dedicated application.

[0808] The server converts the received audio data into text data using a high-precision speech recognition algorithm (e.g., Google Speech-to-Text API or standard speech recognition software). This text data is then automatically categorized according to the meeting agenda, utilizing natural language processing technologies (e.g., spaCy or NLTK). It is then organized by agenda item or category (e.g., shared items, homework, SOW, Q&A, etc.).

[0809] Next, the server automatically generates meeting documents from the text data, creating documents in the specified format (PDF, Word document, etc.). It uses LaTeX or the Microsoft Word API to format the documents. These meeting documents are stored securely in a database or cloud storage (e.g., AWS S3 or Google Drive). After the generation and saving of the meeting documents is complete, the server sends the generated documents to the user's device via email or a file-sharing service (e.g., Dropbox or Slack).

[0810] Furthermore, the system utilizes an emotion engine to analyze user emotions. This engine uses tools such as Amazon Comprehend and Microsoft Azure Text Analytics to analyze user emotions (e.g., joy, anger, sadness, surprise, etc.) in real time from voice input. The results of the emotion analysis are reflected in the meeting documents, allowing users and participants to review their emotional changes during the meeting. This data is then transmitted via a server and included in meeting minutes, task lists, and other documents generated after the meeting.

[0811] For example, in a new product development meeting, the user schedules the meeting and starts it at the designated time. The user uses a voice input device to input speech, and this speech data is sent to the server in real time. The server uses the Google Speech-to-Text API to convert the speech into text data and classifies it by topic and category using natural language processing technology (such as spaCy). Furthermore, Amazon Comprehend is used to analyze the user's emotions during speech, and this emotion data is also processed. After the meeting ends, the server automatically generates meeting documents using LaTeX or Microsoft Word APIs, saves them to AWS S3 or Google Drive, and sends them to the user via Dropbox or email. The user reviews the received documents, edits them, and gathers feedback. This information can then be used to prepare for the next meeting, allowing for more efficient discussions.

[0812] The following are examples of prompt statements:

[0813] "Please record what users say in meetings in real time, categorize the comments, and automatically generate meeting minutes. Also, record emotional information associated with each comment, save all the information after the meeting, and send it to the participants."

[0814] This system streamlines meetings and facilitates communication by collecting emotional data. This, in turn, leads to accurate information sharing and a reduction in overtime.

[0815] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0816] Specific flow of program processing

[0817] Step 1:

[0818] Acquisition of audio data

[0819] The user uses the device's microphone to input voice commands. For example, they might say, "Let's discuss the features of the new product." This voice data is captured as input by the device. Specifically, the device's microphone collects voice waveform data, which is temporarily stored in storage. The output is the captured raw voice data.

[0820] Step 2:

[0821] Sending audio data

[0822] The device captures audio data and sends it to the server in real time. WebSocket technology is used via an internet connection or a dedicated application. The input is the captured audio data, and the output is the streamed audio data sent to the server. Specifically, the device establishes a WebSocket connection and transfers the audio data to the server in real time.

[0823] Step 3:

[0824] Voice data recognition and conversion

[0825] The server processes the received audio data and converts it into text data using a highly accurate speech recognition algorithm. For example, it can use the Google Speech-to-Text API. The input is real-time streamed audio data, and the output is converted text data. Specifically, the server sends the audio data to the API and receives the text data returned from the API.

[0826] Step 4:

[0827] Classification of text data

[0828] The server uses natural language processing techniques to classify text data based on the meeting agenda. For example, tools such as spaCy are used. The input is text data obtained from speech recognition, and the output is text data classified by agenda item. Specifically, the server analyzes the text data and classifies it based on specific keywords or phrases.

[0829] Step 5:

[0830] Emotion analysis

[0831] The server uses Amazon Comprehend or a similar engine to analyze emotions from text data. The input is classified text data, and the output is analyzed emotion data (e.g., joy, anger, sadness, surprise, etc.). Specifically, text data is input into the emotion analysis engine, and emotion data is extracted.

[0832] Step 6:

[0833] Document generation

[0834] The server automatically generates meeting documents containing text and sentiment data using LaTeX and the Microsoft Word API. The input is categorized text and sentiment data, and the output is a formatted meeting document (PDF, Word document, etc.). Specifically, the server inserts data into a text template and uses a document generation tool to create the final document.

[0835] Step 7:

[0836] Save document

[0837] The server saves the generated meeting documents to cloud storage (e.g., AWS S3 or Google Drive). The input is the generated meeting documents, and the output is the documents saved on cloud storage. Specifically, the server uploads the data to cloud storage using an API.

[0838] Step 8:

[0839] Sending documents

[0840] The server sends meeting documents to the user's device. This is done, for example, through an email system or file sharing service (such as Dropbox or Slack). The input is a document stored in cloud storage, and the output is a document downloaded to the user's device. Specifically, the server sends the document using an email sending API or a file sharing API.

[0841] This allows users to receive meeting documents based on real-time converted text data and sentiment analysis results, leading to more efficient meetings and accurate information sharing.

[0842] (Application Example 2)

[0843] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0844] Current meeting systems have problems such as users having to manually create documents, which prevents them from focusing on the meeting's progress. In particular, conducting meetings efficiently while traveling in autonomous vehicles is difficult, and there is a need for technology to record meeting content in real time. Furthermore, there is a lack of technology to understand changes in emotions during meetings and facilitate smooth communication.

[0845] In Application Example 2, the specific processing performed by the specific processing unit 290 of the data processing device 12 is realized by the following means. In this invention, the server includes voice acquisition means for the user to perform voice input, communication means for transmitting voice data to the server in real time, voice recognition means for the server to convert the received voice data into text data, text classification means for the server to classify the text data based on the meeting content, document generation means for the server to automatically generate meeting documents from the classified text data, storage means for the server to save the generated meeting documents, transmission means for the server to send the generated meeting documents to the user, voice acquisition means for inside the autonomous vehicle to acquire voice input inside the autonomous vehicle, emotion analysis means for analyzing the user's emotions in real time, and display means for displaying the generated meeting documents on the vehicle display. This enables efficient meeting management inside the autonomous vehicle, and since the emotions of participants while they are speaking are analyzed and recorded in real time, communication is facilitated.

[0846] "Voice acquisition means" refers to a microphone or voice capture device used to acquire user voice input.

[0847] "Communication means" refers to equipment including an internet connection and communication protocols for transmitting voice data to a server in real time.

[0848] "Speech recognition means" refers to speech recognition algorithms and related software used by the server to convert received speech data into text data.

[0849] A "text classification means" is a device that utilizes natural language processing technology to allow a server to automatically classify text data into categories based on the content of a meeting.

[0850] "Document generation means" refers to software and tools that allow a server to automatically generate meeting documents based on classified text data.

[0851] "Storage method" refers to a database or cloud storage for securely saving generated meeting documents.

[0852] "Transmission method" refers to email or file-sharing services used to send generated meeting documents to users.

[0853] "Voice acquisition means within an autonomous vehicle" refers to an in-vehicle microphone or voice capture device used to acquire user voice input inside an autonomous vehicle.

[0854] "Emotion analysis means" refers to algorithms and related software for analyzing user emotions in real time from voice data.

[0855] "Display means" refers to a device and interface for displaying the generated meeting documents on a display inside the vehicle.

[0856] This invention is a system in which a user in an autonomous vehicle performs voice input, transmits that voice data to a server in real time, converts it into text data, and generates, saves, and transmits meeting documents. Furthermore, by combining it with an emotion analysis function that recognizes the user's emotions, changes in emotions during the meeting are also recorded.

[0857] System Configuration

[0858] 1. Means of acquiring sound:

[0859] This device uses an in-vehicle microphone installed inside an autonomous vehicle to acquire voice input from the user.

[0860] 2. Means of communication:

[0861] The acquired voice data is transmitted to the server in real time using an internet connection (e.g., 5G or LTE). WebSocket technology is used for this communication.

[0862] 3. Voice recognition means:

[0863] The server converts the received audio data into text data using speech recognition algorithms such as Google Cloud Speech-to-Text.

[0864] 4. Text classification means:

[0865] The converted text data is automatically categorized based on the meeting content using natural language processing technology (e.g., spaCy). The categorized data is then organized into various agenda items and categories (e.g., shared items, homework, Q&A).

[0866] 5. Document generation means:

[0867] The server generates meeting documents from the classified text data. During this process, tools such as PDFKit and docx are used to create documents in the specified format (PDF, Word document, etc.).

[0868] 6. Preservation means:

[0869] The generated meeting documents are securely stored in a database or cloud storage (e.g., AWS S3).

[0870] 7. Transmission method:

[0871] The generated meeting documents are sent to users via email or file sharing services. For example, email sending is done using AWS SES.

[0872] 8. Means for acquiring voice data inside an autonomous vehicle:

[0873] A specific high-sensitivity microphone is used to acquire voice data inside the autonomous vehicle, enabling clear voice input even while the vehicle is in motion.

[0874] 9. Emotion analysis means:

[0875] The server uses emotion analysis engines such as IBM Watson Tone Analyzer to analyze the user's emotions in real time from their voice data. The analysis results are reflected in the meeting documents.

[0876] 10. Display means:

[0877] The generated meeting documents will be displayed on the vehicle's display, for example, on an in-car display or a tablet device.

[0878] Usage example

[0879] In a new product development meeting, the user provides voice input using a voice acquisition device inside the autonomous vehicle. This voice data is transmitted to a server in real time via WebSocket. The server converts the voice data into text data using Google Cloud Speech-to-Text, and then classifies the text data based on the meeting content using spaCy. The classified data is generated as a meeting document using PDFKit or docx and stored in AWS S3. The stored document is sent to the user using AWS SES. In addition, sentiment analysis is performed on each statement during the meeting using IBM Watson Tone Analyzer, and the analysis results are also included in the meeting document. Finally, the generated meeting document is displayed on the in-vehicle display for the user to review. An example of a prompt phrase used is, "What are your thoughts on pricing the new product?"

[0880] This system enables efficient meeting management within autonomous vehicles, allowing participants to facilitate smoother communication by utilizing real-time voice input and sentiment analysis.

[0881] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0882] Step 1:

[0883] The user provides voice input within the autonomous vehicle. The input is acquired via the vehicle's microphone and transmitted as voice data to the vehicle's terminal. The output is voice data.

[0884] Step 2:

[0885] Audio data is transmitted in real time from a terminal inside the autonomous vehicle to a server via WebSocket. The input is audio data acquired from the vehicle's microphone, and the output is the audio data sent to the server.

[0886] Step 3:

[0887] The server converts the received audio data into text data using Google Cloud Speech-to-Text. The input is the audio data received via WebSocket, and the output is the converted text data. A speech recognition algorithm is then applied.

[0888] Step 4:

[0889] Text data is classified on the server using natural language processing techniques (e.g., spaCy). The input is the transformed text data, and the output is text data classified into meeting topics and categories. Syntactic and semantic analysis of the text is performed.

[0890] Step 5:

[0891] The classified text data is automatically generated by the server into a meeting document in the specified format using tools such as PDFKit or docx. The input is the classified text data, and the output is a meeting document (PDF or Word document). A document template is used here.

[0892] Step 6:

[0893] The server saves the generated meeting documents to a database or cloud storage (e.g., AWS S3). The input is the generated meeting documents, and the output is the documents saved in cloud storage. The saving process is then performed.

[0894] Step 7:

[0895] The server sends the saved meeting documents to the user. This email transmission is performed using AWS SES. The input is the link information stored in cloud storage and the user's email address, and the output is the sent email.

[0896] Step 8:

[0897] During the meeting, the server performs real-time sentiment analysis on audio data using IBM Watson Tone Analyzer. The input is audio data received via WebSocket, and the output is the analyzed sentiment data.

[0898] Step 9:

[0899] The sentiment analysis results, along with categorized text data, are included in the meeting document, which users can review after the meeting. The input is the sentiment analysis results and text data, and the output is the meeting document containing the sentiment data.

[0900] Step 10:

[0901] The generated meeting documents are displayed on the in-vehicle display, allowing the user to view them in real time. The input is the generated meeting documents, and the output is the content displayed on the in-vehicle display.

[0902] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0903] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0904] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0905] [Fourth Embodiment]

[0906] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0907] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0908] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0909] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0910] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0911] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0912] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0913] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0914] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0915] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0916] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0917] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0918] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0919] This invention is a system that acquires user voice input, transmits the voice data to a server in real time, converts the voice data into text data, classifies it, and automatically generates, saves, and transmits meeting documents. Specific embodiments are described below.

[0920] In this system, the user first performs voice input using a voice acquisition method. For example, the microphone built into the user's terminal (a PC or mobile device) is used. This voice data is transmitted in real time from the terminal to the server via a communication method (such as an internet connection or WebSocket technology via a dedicated application).

[0921] The server, acting as a speech recognition system, converts received audio data into text data using a highly accurate speech recognition algorithm. This text data is then automatically categorized according to the meeting agenda, utilizing natural language processing technology as a text classification system. Classification is performed not only by agenda item but also by category (e.g., shared items, homework, SOW, Q&A).

[0922] Next, based on the classified text data, the server, acting as a document generation mechanism, automatically generates meeting documents (such as meeting minutes and task lists). These documents are generated in the specified format (e.g., PDF or Word document). Furthermore, the generated documents are securely stored in a database or cloud storage.

[0923] After the meeting documents have been generated and saved, the server, acting as the transmission method, sends them to the user. The documents are sent via the specified email address or file-sharing service. The user's device is designed to display the received documents in a highly readable format.

[0924] Furthermore, users can set meeting details (date and time, participants, agenda, etc.) using a dedicated interface (web application or mobile app) before the meeting. This information is sent to the server and managed as metadata by the server's management system. Metadata management streamlines pre-meeting preparations and post-meeting documentation.

[0925] As a concrete example, consider a scenario where a user sets up a new product development meeting. First, the user schedules the meeting using the scheduling tool and notifies participants. On the day of the meeting, the user uses a terminal in the meeting room to send the audio of the conversation to the server using the audio acquisition tool. The server converts the audio to text in real time, categorizes it by agenda item, and automatically generates meeting documents immediately. Once the meeting ends, the generated minutes and task list are saved to cloud storage using the saving tool and quickly sent to the user's terminal via the transmission tool. The user and participants can review these materials and provide feedback as needed, further streamlining the preparation for the next meeting.

[0926] In this way, the system of the present invention dramatically improves the efficiency of meetings and provides an environment in which employees can concentrate on their actual work time. This leads to a reduction in overtime hours and accurate sharing of information.

[0927] The following describes the processing flow.

[0928] Step 1:

[0929] The user schedules a meeting. Using the terminal's settings, they input the meeting date and time, participants, and agenda, and send this information to the server. The server stores the received information in a database using its management system.

[0930] Step 2:

[0931] The user starts a meeting. At the specified date and time of the meeting, the user inputs voice data using the device's audio acquisition method (microphone). The device acquires this voice data in real time and transmits it to the server using a communication method.

[0932] Step 3:

[0933] The server converts the received audio data into text data using speech recognition technology. The speech recognition algorithm analyzes the audio and converts it into text format.

[0934] Step 4:

[0935] The server's text classification system categorizes text data based on meeting content. Using natural language processing technology, it organizes the text into categories such as meeting minutes, shared items, homework, statements of work (SOW), and Q&A.

[0936] Step 5:

[0937] The server's document generation mechanism automatically generates meeting documents from classified text data. The documents are created in the specified format (PDF, Word document, etc.).

[0938] Step 6:

[0939] The server's storage method involves saving the generated meeting documents to a database or cloud storage. The document's metadata (creation date, creator, etc.) is also saved simultaneously.

[0940] Step 7:

[0941] The server's transmission method sends the generated meeting documents to the user's terminal. The documents are sent via email or a file-sharing service.

[0942] Step 8:

[0943] Users review meeting documents received on their devices. They edit the documents as needed, incorporating any corrections or additional information. They also collect feedback to improve future meetings.

[0944] Step 9:

[0945] The user distributes the final edited document to other participants. This can be done by resending it through the server or manually using email or a file-sharing service.

[0946] (Example 1)

[0947] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0948] Traditional methods for accurately and quickly recording discussions during meetings and automatically generating minutes and task lists are time-consuming and laborious. Furthermore, smoothly sharing these materials with all meeting participants is also a cumbersome process. Therefore, there is a need for efficient meeting management and accurate sharing of important information.

[0949] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0950] In this invention, the server includes: voice acquisition means for the user to input voice data; communication means for transmitting voice data to a data processing device in real time; and voice recognition means for the data processing device to convert the received voice data into text data. This makes it possible to instantly transcribe, classify, and automatically generate, save, and share meeting documents based on what the user says during a meeting.

[0951] "Voice acquisition means" refers to devices or functions that allow users to perform voice input.

[0952] "Communication means" refers to technical methods and devices for transmitting voice data to a data processing device in real time.

[0953] "Speech recognition means" refers to algorithms and software used by a data processing device to convert received speech data into text data.

[0954] "Natural language processing means" refers to a technology used by data processing devices to classify text data based on the content of a meeting.

[0955] "Document generation means" refers to software or algorithms used by a data processing device to automatically generate meeting documents from classified text data.

[0956] "Storage means" refers to the technology and equipment used to save generated meeting documents to a storage medium.

[0957] "Transmission means" refers to the technology or functionality used to send generated meeting documents to users.

[0958] "Display means" refers to devices or software that display meeting documents received by a user's terminal in a highly visible format.

[0959] "Configuration means" refers to the interface or software that allows users to configure detailed meeting information.

[0960] "Management means" refers to the functions and technologies that data processing devices use to manage meeting metadata.

[0961] This invention relates to a system in which a user performs voice input, transmits the voice data to a server in real time, converts it into text data, classifies it, and automatically generates and saves meeting documents. Specific embodiments of the system are described in detail below.

[0962] overview

[0963] This system captures user speech and transmits it in real time to a data processing unit (server). The server converts the speech data into text data and classifies it based on the meeting content. The system automatically generates, saves, and transmits meeting documents from the classified text data. The aim of this system is to improve meeting efficiency and ensure accurate information sharing.

[0964] Hardware and software to be used

[0965] The following hardware and software will be used to implement this system.

[0966] Audio acquisition method: User's device (PC, mobile device) equipped with a microphone device.

[0967] Communication methods: Internet connection, WebSocket technology, HTTP POST

[0968] Speech recognition method: Google Cloud Speech-to-Text API, Amazon Transcribe

[0969] Natural language processing tools: Google Cloud Natural Language API

[0970] Document generation methods: Python's reportlab library, docx library

[0971] Storage method: Cloud storage (Amazon S3, Google Drive)

[0972] Sending method: SendGrid, Google Drive shared link

[0973] Specific example

[0974] For example, let's consider a new product development meeting. The user sets the meeting date and participants using a web application and sends them to the data processing unit. On the day of the meeting, the user continuously sends audio of the conversation to the data processing unit using the microphone on their device. The data processing unit converts the audio to text in real time and categorizes it by agenda item. After the meeting, the data processing unit automatically generates meeting minutes and a task list and saves them to cloud storage. These documents are sent to the user and displayed on their device in a highly readable format. The following is an example of a prompt to be input into the generating AI model.

[0975] Example of a prompt:

[0976] Please analyze the audio data from the next meeting and generate meeting minutes categorized by agenda item.

[0977] Users can review this document and provide feedback as needed, further streamlining their preparations for the next meeting. The automated generation and sharing of meeting documents provides employees with an environment where they can focus on their actual work time.

[0978] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0979] Step 1: Obtaining voice input

[0980] The device acquires the user's voice input via the microphone. This voice data is input and temporarily stored in a buffer on the device side. This data is then sent to the next step in a short time for real-time processing.

[0981] Step 2: Sending the audio data

[0982] The terminal transmits the acquired audio data to the server in real time. This transmission uses an internet connection and WebSocket technology. The terminal streams the audio data in a buffer to the server via the WebSocket connection. As a specific example, the JavaScript WebSocket API is used.

[0983] Step 3: Receiving audio data

[0984] The server receives audio data transmitted from the terminal in real time. The server opens a WebSocket connection and receives the audio data as a stream. The received audio data is temporarily stored in internal memory.

[0985] Step 4: Convert speech to text

[0986] The server converts the received audio data into text data using a speech recognition algorithm. The input is audio data, and the output is the corresponding text. Specifically, the server uses the Google Cloud Speech-to-Text API to send audio data to the API and receive the resulting text data.

[0987] Step 5: Classification of text data

[0988] The server classifies text data by topic using natural language processing techniques. In this step, the input is text data, and the output is text data with classification tags. The Google Cloud Natural Language API is used to parse the text and assign appropriate tags, resulting in categorical text classification.

[0989] Step 6: Generating meeting documents

[0990] The server automatically generates meeting documents based on classified text data. The input is classified text data, and the output is meeting documents (PDF or Word documents). The Python reportlab and docx libraries are used to generate documents in the specified format.

[0991] Step 7: Save the document

[0992] The server saves the generated meeting documents to cloud storage. The input is the meeting documents, and the output is the files saved in cloud storage. Specifically, documents are uploaded using Amazon S3 or the Google Drive API.

[0993] Step 8: Send the document

[0994] The server sends the saved meeting documents to the user. The input is the meeting documents, and the output is the documents received by the user. A download link for the documents is sent to the user's email address using a SendGrid or Google Drive sharing link.

[0995] Step 9: Setting up the meeting

[0996] Users configure meeting details (date, time, participants, agenda, etc.) using a dedicated interface (web application or mobile app). The input is the meeting details, and the output is metadata generated based on that information. The metadata is sent to the server and stored and managed by a management system.

[0997] Step 10: Document Review and Feedback

[0998] Users review the received meeting documents and provide feedback as needed. The input is the meeting document, and the output is the revised document or additional feedback information. Users use this to prepare for future meetings.

[0999] (Application Example 1)

[1000] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1001] Traditional factory operations often involved manual processes for instructing workers and managing their progress, resulting in inefficiency and a heavy burden on workers. In particular, generating and managing work instructions and progress reports in real time was difficult, leading to delays and errors. Therefore, there is a growing need for increased efficiency in factory operations and real-time progress management.

[1002] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[1003] In this invention, the server includes voice acquisition means for the user to input voice data, communication means for transmitting voice data to the server in real time, voice recognition means for the server to convert the received voice data into text data, text classification means for the server to classify the text data based on work instructions and progress status, document generation means for the server to automatically generate progress reports and task lists from the classified text data, storage means for the server to save the generated progress reports and task lists, and transmission means for the server to send the generated progress reports and task lists to the user. This makes it possible to efficiently provide instructions and manage the progress of factory work in real time.

[1004] "Voice acquisition means" refers to a device or interface for a user to perform voice input.

[1005] "Communication methods" refer to network technologies and protocols for transmitting voice data to a server in real time.

[1006] "Speech recognition means" refers to algorithms and software used by a server to convert received speech data into text data.

[1007] "Text classification means" refers to natural language processing technology used by a server to classify text data based on work instructions and progress status.

[1008] "Document generation means" refers to software or functionality that allows a server to automatically generate progress reports and task lists from classified text data.

[1009] "Storage methods" refer to technologies and systems for saving generated progress reports and task lists to databases or cloud storage.

[1010] "Transmission method" refers to the technology or protocol used by a server to send generated progress reports and task lists to users.

[1011] "Configuration means" refers to the interface or software that allows users to configure settings for their work.

[1012] "Management means" refers to the systems and technologies that servers use to manage metadata for their operations.

[1013] This invention is a system for managing work instructions and progress in real time within a factory, enabling efficient work operations. The elements constituting this system and their specific embodiments are described below.

[1014] First, the user performs voice input using a voice acquisition device. This device can be a head-mounted display or smart glasses equipped with a high-performance microphone. This data is transmitted in real time to a cloud server via a communication device. The communication device can utilize an internet connection or WebSocket technology through a dedicated application.

[1015] The cloud server converts received audio data into text data using speech recognition. High-precision speech recognition APIs such as Google Cloud Speech-to-Text are used for speech recognition. This text data is then classified by text classification tools based on work instructions and progress. Natural language processing libraries such as spaCy and NLTK are utilized for classification.

[1016] Next, progress reports and task lists are automatically generated from the classified text data by a document generation mechanism. The generated documents are stored in cloud storage such as Google Cloud Storage by a storage mechanism. Furthermore, the server transmits the generated documents to the user's head-mounted display via a transmission mechanism.

[1017] Users can check progress and next work instructions in real time through a head-mounted display. This improves work efficiency and reduces the burden on workers.

[1018] As a concrete example, suppose a user says, "I have finished preparing for step 1. Please give me the next instructions." This voice input is sent to a cloud server via a voice acquisition device. The cloud server converts the voice to text, and then uses Natural Language Processing (NLP) technology to classify the text as "Preparation for step 1 complete." Based on the classified text data, a progress report is automatically generated and displayed in real time on a head-mounted display.

[1019] Specific examples of prompt statements are as follows:

[1020] "Voice input: 'Preparation for step 1 is complete. Please give me the next instruction.'"

[1021] By using this system, work instructions and progress management within the factory can be carried out efficiently in real time, reducing delays and errors and improving overall production efficiency.

[1022] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[1023] Step 1:

[1024] Users use a head-mounted display or smart glasses to input voice commands.

[1025] Input: User's voice instructions

[1026] Output: The microphone in the device acquires audio data.

[1027] Specific action: The user says, "I've finished preparing for step 1. Please give me the next instructions," and the audio data is captured by the microphone.

[1028] Step 2:

[1029] The device transmits the acquired audio data to the cloud server in real time.

[1030] Input: Acquired audio data

[1031] Output: Audio data sent to the cloud server

[1032] Specific operation: Audio data is sent to a cloud server over the internet using WebSocket technology.

[1033] Step 3:

[1034] The server converts the received audio data into text data using speech recognition technology.

[1035] Input: Audio data received by the server

[1036] Output: Audio data converted to text data

[1037] Specific operation: Using the Google Cloud Speech-to-Text API, the audio data is converted into text data that reads, "Step 1 is ready. Please give me the next instructions."

[1038] Step 4:

[1039] The server uses natural language processing techniques to classify text data based on work instructions and progress.

[1040] Input: Converted text data

[1041] Output: Classified text data

[1042] Specific operation: Text is used to perform a text analysis using spaCy or NLTK, and the progress status is determined as "Ready".

[1043] Step 5:

[1044] The server automatically generates progress reports and task lists from classified text data using document generation methods.

[1045] Input: Classified text data

[1046] Output: Progress report and task list

[1047] Specific operation: Document generation software on the server generates progress reports and task lists, and creates them in PDF or text format.

[1048] Step 6:

[1049] The server saves the generated progress reports and task lists to cloud storage.

[1050] Input: Generated progress reports and task lists

[1051] Output: Stored data in cloud storage

[1052] Specific action: Upload progress reports and task lists to Google Cloud Storage and store them securely.

[1053] Step 7:

[1054] The server sends generated progress reports and task lists to the user in real time.

[1055] Input: Generated progress reports and task lists

[1056] Output: Data sent to the user's head-mounted display

[1057] Specific operation: Progress reports and task lists are sent via the internet to the user's head-mounted display and displayed on the screen.

[1058] In this way, work instructions and progress management within the factory are carried out efficiently in real time through this series of processes. This system automates all steps from voice input to text conversion, progress report generation, saving, and transmission, thereby reducing the burden on workers and improving work efficiency.

[1059] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[1060] This invention is a system that acquires user voice input, transmits the voice data to a server in real time, converts the voice data into text data, classifies it, and automatically generates, saves, and transmits meeting documents, in addition to an emotion engine that recognizes the user's emotions. Specific embodiments are described below.

[1061] In this system, the user first inputs voice data using a voice acquisition device. For example, the microphone built into the user's device (a PC or mobile device) is used. This voice data is transmitted in real time from the device to the server via a communication method (such as an internet connection or WebSocket technology via a dedicated application).

[1062] The server, acting as a speech recognition system, converts received audio data into text data using a highly accurate speech recognition algorithm. This text data is then automatically categorized according to the meeting agenda, utilizing natural language processing technology as a text classification system. The data is then organized by agenda item or category (e.g., shared items, homework, SOW, Q&A, etc.).

[1063] Next, to automatically generate meeting documents from text data, the server, acting as the document generation method, creates documents in the specified format (PDF, Word document, etc.). These meeting documents are stored securely in a database or cloud storage. After the generation and saving of the meeting documents is complete, the server sends the generated documents to the user's terminal using a transmission method. The documents are sent via email or a file sharing service.

[1064] Furthermore, this invention utilizes an emotion engine means for analyzing the user's emotions. The emotion engine means analyzes the user's emotions (e.g., joy, anger, sadness, surprise, etc.) in real time from voice input. The results of the emotion analysis are reflected in the meeting document, allowing users and participants to review changes in their emotions during the meeting as a record. This data is transmitted using a server's transmission means and included in meeting minutes, task lists, etc., generated after the meeting.

[1065] For example, in a new product development meeting, the user schedules the meeting and starts it at the designated time. The user inputs voice data using a voice acquisition device, and this voice data is sent to the server in real time. The server converts the voice into text data and categorizes it by topic and category. Furthermore, an emotion engine device analyzes the user's emotions at the time of speaking, and this emotion data is also processed. After the meeting ends, the server automatically generates a meeting document, saves it to cloud storage using a storage device, and sends it to the user using a transmission device. The user reviews the received document, edits it, and gathers feedback. By using this information to prepare for the next meeting, discussions can be conducted more efficiently.

[1066] Thus, the system of the present invention dramatically improves meeting efficiency and provides an environment where employees can concentrate on their actual work time. Furthermore, by collecting and sharing emotional data during meetings, it facilitates smoother communication among participants. This leads to a reduction in overtime hours and accurate sharing of information.

[1067] The following describes the processing flow.

[1068] Step 1:

[1069] The user schedules a meeting. Using the terminal's settings, they input the meeting date and time, participants, and agenda, and send this information to the server. The server stores the received information in a database using its management system.

[1070] Step 2:

[1071] The user initiates a meeting. At the specified date and time, they use the device's audio input device (microphone) to input voice data. The device acquires this voice data in real time and transmits it to the server using a communication device.

[1072] Step 3:

[1073] The server converts the received audio data into text data using speech recognition technology. The speech recognition algorithm analyzes the audio and converts it into text format.

[1074] Step 4:

[1075] The server's emotion engine analyzes the received audio data to identify the user's emotions. The emotion recognition algorithm analyzes the tone, speed, intonation, etc., of the voice to identify the user's emotions (joy, anger, sadness, surprise, etc.).

[1076] Step 5:

[1077] The server's text classification system categorizes text data and sentiment data based on the meeting content. Using natural language processing technology, it organizes the data by agenda item and category (shared items, homework, SOW, Q&A).

[1078] Step 6:

[1079] The server's document generation mechanism automatically generates meeting documents from classified text and sentiment data. The documents are created in the specified format (PDF, Word document, etc.).

[1080] Step 7:

[1081] The server's storage method involves saving generated meeting documents and sentiment data to a database or cloud storage. Document and sentiment metadata (creation date, creator, sentiment progression, etc.) are also saved simultaneously.

[1082] Step 8:

[1083] The server's transmission method sends the generated meeting documents and sentiment data to the user's terminal. The documents are sent via email or a file-sharing service.

[1084] Step 9:

[1085] Users review meeting documents and sentiment data received on their devices. They edit the documents as needed, incorporating revisions and additional information. They also collect feedback to improve future meetings.

[1086] Step 10:

[1087] The user distributes the final edited document to other participants. This can be done by resending it through the server or manually using email or a file-sharing service.

[1088] (Example 2)

[1089] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1090] Conventional meeting recording systems are often limited to functions such as converting voice input to text and classifying that text. They are unable to analyze user emotions and reflect them in meeting documents, making it difficult to accurately grasp changes in participants' emotions and opinions during a meeting. There is a need to provide a system that solves these problems and improves meeting efficiency and facilitates smoother communication among participants.

[1091] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[1092] In this invention, the server includes voice acquisition means for the user to input voice data, communication means for transmitting voice data to the server in real time, voice recognition means for converting the received voice data into text data, text classification means for classifying the text data based on the meeting content, document generation means for automatically generating meeting documents from the classified text data, storage means for saving the generated meeting documents, transmission means for sending the generated meeting documents to the user, and emotion analysis means for analyzing emotions from the voice data. This makes it possible to not only transcribe what the user says into text in real time and classify the content, but also analyze emotion data and reflect it in the document.

[1093] "Voice acquisition means" refers to devices or systems that allow users to input voice, and includes microphones, recording devices, and the like.

[1094] "Communication methods" refer to technologies and devices for transmitting voice data to a server in real time, and include technologies using internet connections and dedicated applications.

[1095] "Speech recognition means" refers to algorithms and technologies used by a server to convert received speech data into text data, and includes speech recognition software and APIs.

[1096] "Text classification means" refers to natural language processing technology used by a server to classify text data based on the content of a meeting, and includes specific algorithms and software.

[1097] "Document generation means" refers to technology that enables a server to automatically generate meeting documents from classified text data, and includes software for formatting and layout creation.

[1098] "Storage methods" refer to technologies and devices for securely storing generated meeting documents, and include databases and cloud storage.

[1099] "Transmission means" refers to the technologies and devices used to send generated meeting documents to users, and includes email systems and file sharing services.

[1100] "Emotion analysis means" refers to engines and algorithms for analyzing emotions from voice data, and includes technologies for identifying the user's emotions (e.g., joy, anger, sadness, surprise, etc.).

[1101] This invention is a system that combines a system that acquires user voice input, transmits voice data to a server in real time, converts the voice data into text data, classifies it, and automatically generates, saves, and transmits meeting documents, with an emotion engine that recognizes the user's emotions.

[1102] First, the user uses the microphone built into their device (a PC or mobile device) to input voice data. This voice data is transmitted in real time from the device to the server via WebSocket technology through an internet connection or a dedicated application.

[1103] The server converts the received audio data into text data using a high-precision speech recognition algorithm (e.g., Google Speech-to-Text API or standard speech recognition software). This text data is then automatically categorized according to the meeting agenda, utilizing natural language processing technologies (e.g., spaCy or NLTK). It is then organized by agenda item or category (e.g., shared items, homework, SOW, Q&A, etc.).

[1104] Next, the server automatically generates meeting documents from the text data, creating documents in the specified format (PDF, Word document, etc.). It uses LaTeX or the Microsoft Word API to format the documents. These meeting documents are stored securely in a database or cloud storage (e.g., AWS S3 or Google Drive). After the generation and saving of the meeting documents is complete, the server sends the generated documents to the user's device via email or a file-sharing service (e.g., Dropbox or Slack).

[1105] Furthermore, the system utilizes an emotion engine to analyze user emotions. This engine uses tools such as Amazon Comprehend and Microsoft Azure Text Analytics to analyze user emotions (e.g., joy, anger, sadness, surprise, etc.) in real time from voice input. The results of the emotion analysis are reflected in the meeting documents, allowing users and participants to review their emotional changes during the meeting. This data is then transmitted via a server and included in meeting minutes, task lists, and other documents generated after the meeting.

[1106] For example, in a new product development meeting, the user schedules the meeting and starts it at the designated time. The user uses a voice input device to input speech, and this speech data is sent to the server in real time. The server uses the Google Speech-to-Text API to convert the speech into text data and classifies it by topic and category using natural language processing technology (such as spaCy). Furthermore, Amazon Comprehend is used to analyze the user's emotions during speech, and this emotion data is also processed. After the meeting ends, the server automatically generates meeting documents using LaTeX or Microsoft Word APIs, saves them to AWS S3 or Google Drive, and sends them to the user via Dropbox or email. The user reviews the received documents, edits them, and gathers feedback. This information can then be used to prepare for the next meeting, allowing for more efficient discussions.

[1107] The following are examples of prompt statements:

[1108] "Please record what users say in meetings in real time, categorize the comments, and automatically generate meeting minutes. Also, record emotional information associated with each comment, save all the information after the meeting, and send it to the participants."

[1109] This system streamlines meetings and facilitates communication by collecting emotional data. This, in turn, leads to accurate information sharing and a reduction in overtime.

[1110] The flow of the specific processing in Example 2 will be explained using Figure 13.

[1111] Specific flow of program processing

[1112] Step 1:

[1113] Acquisition of audio data

[1114] The user uses the device's microphone to input voice commands. For example, they might say, "Let's discuss the features of the new product." This voice data is captured as input by the device. Specifically, the device's microphone collects voice waveform data, which is temporarily stored in storage. The output is the captured raw voice data.

[1115] Step 2:

[1116] Sending audio data

[1117] The device captures audio data and sends it to the server in real time. WebSocket technology is used via an internet connection or a dedicated application. The input is the captured audio data, and the output is the streamed audio data sent to the server. Specifically, the device establishes a WebSocket connection and transfers the audio data to the server in real time.

[1118] Step 3:

[1119] Voice data recognition and conversion

[1120] The server processes the received audio data and converts it into text data using a highly accurate speech recognition algorithm. For example, it can use the Google Speech-to-Text API. The input is real-time streamed audio data, and the output is converted text data. Specifically, the server sends the audio data to the API and receives the text data returned from the API.

[1121] Step 4:

[1122] Classification of text data

[1123] The server uses natural language processing techniques to classify text data based on the meeting agenda. For example, tools such as spaCy are used. The input is text data obtained from speech recognition, and the output is text data classified by agenda item. Specifically, the server analyzes the text data and classifies it based on specific keywords or phrases.

[1124] Step 5:

[1125] Emotion analysis

[1126] The server uses Amazon Comprehend or a similar engine to analyze emotions from text data. The input is classified text data, and the output is analyzed emotion data (e.g., joy, anger, sadness, surprise, etc.). Specifically, text data is input into the emotion analysis engine, and emotion data is extracted.

[1127] Step 6:

[1128] Document generation

[1129] The server automatically generates meeting documents containing text and sentiment data using LaTeX and the Microsoft Word API. The input is categorized text and sentiment data, and the output is a formatted meeting document (PDF, Word document, etc.). Specifically, the server inserts data into a text template and uses a document generation tool to create the final document.

[1130] Step 7:

[1131] Save document

[1132] The server saves the generated meeting documents to cloud storage (e.g., AWS S3 or Google Drive). The input is the generated meeting documents, and the output is the documents saved on cloud storage. Specifically, the server uploads the data to cloud storage using an API.

[1133] Step 8:

[1134] Sending documents

[1135] The server sends meeting documents to the user's device. This is done, for example, through an email system or file sharing service (such as Dropbox or Slack). The input is a document stored in cloud storage, and the output is a document downloaded to the user's device. Specifically, the server sends the document using an email sending API or a file sharing API.

[1136] This allows users to receive meeting documents based on real-time converted text data and sentiment analysis results, leading to more efficient meetings and accurate information sharing.

[1137] (Application Example 2)

[1138] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1139] Current meeting systems have problems such as users having to manually create documents, which prevents them from focusing on the meeting's progress. In particular, conducting meetings efficiently while traveling in autonomous vehicles is difficult, and there is a need for technology to record meeting content in real time. Furthermore, there is a lack of technology to understand changes in emotions during meetings and facilitate smooth communication.

[1140] In Application Example 2, the specific processing performed by the specific processing unit 290 of the data processing device 12 is realized by the following means. In this invention, the server includes voice acquisition means for the user to perform voice input, communication means for transmitting voice data to the server in real time, voice recognition means for the server to convert the received voice data into text data, text classification means for the server to classify the text data based on the meeting content, document generation means for the server to automatically generate meeting documents from the classified text data, storage means for the server to save the generated meeting documents, transmission means for the server to send the generated meeting documents to the user, voice acquisition means for inside the autonomous vehicle to acquire voice input inside the autonomous vehicle, emotion analysis means for analyzing the user's emotions in real time, and display means for displaying the generated meeting documents on the vehicle display. This enables efficient meeting management inside the autonomous vehicle, and since the emotions of participants while they are speaking are analyzed and recorded in real time, communication is facilitated.

[1141] "Voice acquisition means" refers to a microphone or voice capture device used to acquire user voice input.

[1142] "Communication means" refers to equipment including an internet connection and communication protocols for transmitting voice data to a server in real time.

[1143] "Speech recognition means" refers to speech recognition algorithms and related software used by the server to convert received speech data into text data.

[1144] A "text classification means" is a device that utilizes natural language processing technology to allow a server to automatically classify text data into categories based on the content of a meeting.

[1145] "Document generation means" refers to software and tools that allow a server to automatically generate meeting documents based on classified text data.

[1146] "Storage method" refers to a database or cloud storage for securely saving generated meeting documents.

[1147] "Transmission method" refers to email or file-sharing services used to send generated meeting documents to users.

[1148] "Voice acquisition means within an autonomous vehicle" refers to an in-vehicle microphone or voice capture device used to acquire user voice input inside an autonomous vehicle.

[1149] "Emotion analysis means" refers to algorithms and related software for analyzing user emotions in real time from voice data.

[1150] "Display means" refers to a device and interface for displaying the generated meeting documents on a display inside the vehicle.

[1151] This invention is a system in which a user in an autonomous vehicle performs voice input, transmits that voice data to a server in real time, converts it into text data, and generates, saves, and transmits meeting documents. Furthermore, by combining it with an emotion analysis function that recognizes the user's emotions, changes in emotions during the meeting are also recorded.

[1152] System Configuration

[1153] 1. Means of acquiring sound:

[1154] This device uses an in-vehicle microphone installed inside an autonomous vehicle to acquire voice input from the user.

[1155] 2. Means of communication:

[1156] The acquired voice data is transmitted to the server in real time using an internet connection (e.g., 5G or LTE). WebSocket technology is used for this communication.

[1157] 3. Voice recognition means:

[1158] The server converts the received audio data into text data using speech recognition algorithms such as Google Cloud Speech-to-Text.

[1159] 4. Text classification means:

[1160] The converted text data is automatically categorized based on the meeting content using natural language processing technology (e.g., spaCy). The categorized data is then organized into various agenda items and categories (e.g., shared items, homework, Q&A).

[1161] 5. Document generation means:

[1162] The server generates meeting documents from the classified text data. During this process, tools such as PDFKit and docx are used to create documents in the specified format (PDF, Word document, etc.).

[1163] 6. Preservation means:

[1164] The generated meeting documents are securely stored in a database or cloud storage (e.g., AWS S3).

[1165] 7. Transmission method:

[1166] The generated meeting documents are sent to users via email or file sharing services. For example, email sending is done using AWS SES.

[1167] 8. Means for acquiring voice data inside an autonomous vehicle:

[1168] A specific high-sensitivity microphone is used to acquire voice data inside the autonomous vehicle, enabling clear voice input even while the vehicle is in motion.

[1169] 9. Emotion analysis means:

[1170] The server uses emotion analysis engines such as IBM Watson Tone Analyzer to analyze the user's emotions in real time from their voice data. The analysis results are reflected in the meeting documents.

[1171] 10. Display means:

[1172] The generated meeting documents will be displayed on the vehicle's display, for example, on an in-car display or a tablet device.

[1173] Usage example

[1174] In a new product development meeting, the user provides voice input using a voice acquisition device inside the autonomous vehicle. This voice data is transmitted to a server in real time via WebSocket. The server converts the voice data into text data using Google Cloud Speech-to-Text, and then classifies the text data based on the meeting content using spaCy. The classified data is generated as a meeting document using PDFKit or docx and stored in AWS S3. The stored document is sent to the user using AWS SES. In addition, sentiment analysis is performed on each statement during the meeting using IBM Watson Tone Analyzer, and the analysis results are also included in the meeting document. Finally, the generated meeting document is displayed on the in-vehicle display for the user to review. An example of a prompt phrase used is, "What are your thoughts on pricing the new product?"

[1175] This system enables efficient meeting management within autonomous vehicles, allowing participants to facilitate smoother communication by utilizing real-time voice input and sentiment analysis.

[1176] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[1177] Step 1:

[1178] The user provides voice input within the autonomous vehicle. The input is acquired via the vehicle's microphone and transmitted as voice data to the vehicle's terminal. The output is voice data.

[1179] Step 2:

[1180] Audio data is transmitted in real time from a terminal inside the autonomous vehicle to a server via WebSocket. The input is audio data acquired from the vehicle's microphone, and the output is the audio data sent to the server.

[1181] Step 3:

[1182] The server converts the received audio data into text data using Google Cloud Speech-to-Text. The input is the audio data received via WebSocket, and the output is the converted text data. A speech recognition algorithm is then applied.

[1183] Step 4:

[1184] Text data is classified on the server using natural language processing techniques (e.g., spaCy). The input is the transformed text data, and the output is text data classified into meeting topics and categories. Syntactic and semantic analysis of the text is performed.

[1185] Step 5:

[1186] The classified text data is automatically generated by the server into a meeting document in the specified format using tools such as PDFKit or docx. The input is the classified text data, and the output is a meeting document (PDF or Word document). A document template is used here.

[1187] Step 6:

[1188] The server saves the generated meeting documents to a database or cloud storage (e.g., AWS S3). The input is the generated meeting documents, and the output is the documents saved in cloud storage. The saving process is then performed.

[1189] Step 7:

[1190] The server sends the saved meeting documents to the user. This email transmission is performed using AWS SES. The input is the link information stored in cloud storage and the user's email address, and the output is the sent email.

[1191] Step 8:

[1192] During the meeting, the server performs real-time sentiment analysis on audio data using IBM Watson Tone Analyzer. The input is audio data received via WebSocket, and the output is the analyzed sentiment data.

[1193] Step 9:

[1194] The sentiment analysis results, along with categorized text data, are included in the meeting document, which users can review after the meeting. The input is the sentiment analysis results and text data, and the output is the meeting document containing the sentiment data.

[1195] Step 10:

[1196] The generated meeting documents are displayed on the in-vehicle display, allowing the user to view them in real time. The input is the generated meeting documents, and the output is the content displayed on the in-vehicle display.

[1197] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[1198] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1199] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[1200] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1201] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[1202] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[1203] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[1204] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[1205] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[1206] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[1207] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[1208] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[1209] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[1210] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1211] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[1212] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[1213] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[1214] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[1215] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[1216] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[1217] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[1218] The following is further disclosed regarding the embodiments described above.

[1219] (Claim 1)

[1220] [Voice acquisition means for the user to perform voice input,

[1221] [A communication means for transmitting audio data to a server in real time,

[1222] [The server has speech recognition means for converting received audio data into text data,

[1223] [The server includes text classification means for classifying text data based on the content of the meeting,

[1224] [The server provides document generation means for automatically generating meeting documents from classified text data,

[1225] [The server has storage means for saving the generated meeting documents,

[1226] [The server provides a transmission means for sending the generated meeting document to the user,

[1227] A system that includes this.

[1228] (Claim 2)

[1229] The system according to claim 1, further comprising [configuration means for a user to set up a meeting].

[1230] (Claim 3)

[1231] The system according to claim 1 or 2, further comprising [management means for the server to manage meeting metadata].

[1232] "Example 1"

[1233] (Claim 1)

[1234] [Voice acquisition means for the user to perform voice input,

[1235] [Communication means for transmitting audio data to a data processing device in real time,

[1236] [The data processing device includes speech recognition means for converting received audio data into text data,

[1237] [The data processing device includes natural language processing means for classifying text data based on the content of a meeting,

[1238] [A data processing device comprising document generation means for automatically generating meeting documents from classified text data,

[1239] [A data processing device includes storage means for saving generated meeting documents to a storage medium,

[1240] [A data processing device includes a transmission means for sending generated meeting documents to a user,

[1241] [The user's terminal has display means for displaying the received meeting documents,

[1242] A system that includes this.

[1243] (Claim 2)

[1244] The system according to claim 1, further comprising [configuration means for a user to set meeting details].

[1245] (Claim 3)

[1246] The system according to claim 1, further comprising [management means for a data processing device to manage meeting metadata].

[1247] "Application Example 1"

[1248] (Claim 1)

[1249] [Voice acquisition means for the user to perform voice input,

[1250] [A communication means for transmitting audio data to a server in real time,

[1251] [The server has speech recognition means for converting received audio data into text data,

[1252] [The server includes text classification means for classifying text data based on work instructions and progress status,

[1253] [The server provides document generation means for automatically generating progress reports and task lists from classified text data,

[1254] [The server has a means for saving the generated progress reports and task lists,

[1255] [The server provides a means for sending generated progress reports and task lists to the user,

[1256] A system that includes this.

[1257] (Claim 2)

[1258] [A system according to claim 1, which provides a setting means for the user to configure the work settings.

[1259] (Claim 3)

[1260] [A system according to claim 1, which has a management means for the server to manage metadata of the work.

[1261] "Example 2 of combining an emotion engine"

[1262] (Claim 1)

[1263] [Voice acquisition means for the user to perform voice input,

[1264] [A communication means for transmitting audio data to a server in real time,

[1265] [The server has speech recognition means for converting received audio data into text data,

[1266] [The server includes text classification means for classifying text data based on the content of the meeting,

[1267] [The server provides document generation means for automatically generating meeting documents from classified text data,

[1268] [The server has storage means for saving the generated meeting documents,

[1269] [The server provides a transmission means for sending the generated meeting document to the user,

[1270] [An emotion analysis means for the server to analyze emotions from voice data,

[1271] A system that includes this.

[1272] (Claim 2)

[1273] The system according to claim 1, further comprising [configuration means for a user to set up a meeting].

[1274] (Claim 3)

[1275] The system according to claim 1, further comprising [management means for the server to manage meeting metadata].

[1276] "Application example 2 of combining emotional engines"

[1277] (Claim 1)

[1278] [Voice acquisition means for the user to perform voice input,

[1279] [A communication means for transmitting audio data to a server in real time,

[1280] [The server has speech recognition means for converting received audio data into text data,

[1281] [The server includes text classification means for classifying text data based on the content of the meeting,

[1282] [The server provides document generation means for automatically generating meeting documents from classified text data,

[1283] [The server has storage means for saving the generated meeting documents,

[1284] [The server provides a transmission means for sending the generated meeting document to the user,

[1285] [An in-vehicle voice acquisition means for acquiring voice input inside an autonomous vehicle,

[1286] [An emotion analysis means for analyzing user emotions in real time,

[1287] [Display means for displaying the generated meeting documents on a vehicle display,

[1288] A system that includes this.

[1289] (Claim 2)

[1290] The system according to claim 1, further comprising [configuration means for a user to set up a meeting].

[1291] (Claim 3)

[1292] The system according to claim 1, further comprising [management means for the server to manage meeting metadata]. [Explanation of Symbols]

[1293] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. A means for acquiring voice input for the user, A communication means for transmitting audio data to a server in real time, The server includes speech recognition means for converting received audio data into text data, The server includes text classification means for classifying text data based on the content of the meeting, The server includes a document generation means for automatically generating meeting documents from classified text data, The server has storage means for saving the generated meeting documents, The server includes a transmission means for sending the generated meeting document to the user, A system that includes this.

2. The system according to claim 1, further comprising a setting means for a user to set up a meeting.

3. The system according to claim 1 or 2, further comprising management means for the server to manage meeting metadata.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A