System

The system addresses inefficiencies in recording and organizing meeting content by converting voice to text, analyzing for key points, and automatically formatting and sharing summaries, enhancing efficiency and accuracy.

JP2026037159APending Publication Date: 2026-03-06SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024140184
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-21
Publication Date
2026-03-06

AI Technical Summary

Technical Problem

Conventional methods for recording and organizing meeting or business negotiation contents are inefficient, requiring significant time and effort, especially for large meetings or long negotiations, and often result in inaccuracies and delays.

Method used

A system that converts voice data into text in real-time, analyzes the text to extract important points, formats the summary into a standard format, and automatically sends it to a specified destination, utilizing generative AI and automation tools.

Benefits of technology

This system automates the recording, summarization, and distribution of meeting minutes, significantly improving efficiency and accuracy by reducing manual effort and ensuring timely sharing of organized content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026037159000001_ABST
    Figure 2026037159000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system comprising: means for converting voice data into text in real time; means for analyzing the converted text data and generating a summary; means for converting the generated summary into a predetermined format; and means for automatically transmitting the generated formatted data to a designated destination.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] The task of effectively and quickly recording the contents of meetings and business negotiations and sharing them as organized minutes generally requires a lot of time and effort. Conventional methods require manually recording the contents of meetings and then later organizing and summarizing them, which is inefficient, especially for large meetings or long business negotiations. A system that solves this problem and reduces the burden on users is needed. [Means for solving the problem]

[0005] The present invention solves the above-mentioned problems by providing a system that includes the following means: First, it provides a means for converting voice data into text in real time. Second, it includes a means for analyzing the converted text data and summarizing the important points and content. Finally, it provides a means for converting the generated summary into a predetermined standard format and automatically sending the final formatted data to a specified destination. This makes it possible to quickly and effectively organize and share the contents of meetings and business negotiations.

[0006] "Voice data" refers to digital data that is an electronic recording of conversation or voice information.

[0007] "Real-time" refers to processing or conversion occurring immediately, without delay.

[0008] "Means for converting to text" refers to software or hardware for converting voice data into text information.

[0009] "Text data" is data stored as character information.

[0010] Analysis is the process of examining data in detail to extract and understand specific information.

[0011] A "summary generator" is software or an algorithm that extracts the main points or important details from long text data and shortens it.

[0012] A "predetermined format" refers to a document in which information is organized and laid out according to a specific format or template.

[0013] A "conversion means" is software or hardware that rearranges data into a different format or structure.

[0014] An "automated transmission method" is a mechanism for transferring data to a designated destination without human intervention.

[0015] "Destination" refers to the entity to which the data is ultimately sent, including, for example, an email address or a digital platform. [Brief explanation of the drawings]

[0016] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14]FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0017] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0018] First, the terms used in the following description will be explained.

[0019] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0020] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0021] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0022] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0023] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0024] [First embodiment]

[0025] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0026] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0027] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0028] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0029] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0030] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0031] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0032] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0033] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0034] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0035] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0036] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0037] This system provides a means to quickly record the contents of meetings and business negotiations and automatically share them as organized minutes. When a user holds a meeting or business negotiation, they first launch the voice-to-text tool. The voice data is converted into text data in real time. This converted text data is sent from the user's device to a server where it is analyzed. This analysis uses generative AI (e.g., GPT-4 (registered trademark)) to extract important points and content and generate a summary. The resulting summary data is then converted into a standard format. Finally, the formatted data is automatically sent to the specified destination.

[0038] A natural language description of the program's processing

[0039] Voice reading

[0040] User: When a meeting or business meeting begins, launch the voice-to-text tool and start recording.

[0041] On your device: A speech-to-text tool converts your conversation into text in real time, which is temporarily stored on your device.

[0042] Text Acquisition

[0043] Terminal: When the meeting or business negotiation ends, the saved text data is sent to the server.

[0044] Server: Receives the text data and begins analysis.

[0045] Text organization

[0046] Server: Using generative AI, the server analyzes the received text data and extracts key points and content, thereby summarizing the text.

[0047] Conversion to standard formats

[0048] Server: The extracted key points are embedded into a pre-defined format to complete the minutes template.

[0049] Automation workflow execution

[0050] Server: Using an automation tool such as MAKE, the generated minutes are automatically sent to the specified destination.

[0051] Specific examples

[0052] Examples of meeting content

[0053] Attendees: Yuko Sugimoto, Taro Tanaka, Ichiro Kimura

[0054] Subject: New product development project

[0055] Agenda: Development schedule, resource allocation, next meeting date

[0056] 1. The user launches the voice capture tool and begins recording the audio of the meeting.

[0057] 2. The device converts the conversation being recorded into text in real time and saves the text data.

[0058] 3. After the meeting ends, the device sends the saved text data to the server, and the server receives the text data.

[0059] 4. The server inputs the received text data into a generative AI (e.g., GPT-4) to summarize the meeting content. Specifically, it extracts important information related to the "development schedule," "resource allocation," and "next meeting date."

[0060] 5. The server embeds the extracted key points into a template to generate a standard formatted minutes. For example:

[0061] minutes

[0062] Meeting name: New product development project meeting

[0063] Attendees: Yuko Sugimoto, Taro Tanaka, Ichiro Kimura

[0064] Date and Time: [meeting date and time]

[0065] agenda:

[0066] 1. Development Schedule

[0067] [Summary]

[0068] 2. Resource Allocation

[0069] [Summary]

[0070] 3. Next meeting date

[0071] [Summary]

[0072] remarks:

[0073] There are no particular problems at the moment

[0074] 6. The server uses MAKE to execute a workflow to save the generated minutes in a Google® document or distribute them to all parties involved by email after the video conference ends.

[0075] 7. Users can see that the minutes are automatically shared immediately after the meeting ends and can make additional comments or corrections as needed.

[0076] This system automates a series of tasks, from recording meeting content to organizing and sharing it, greatly improving users' work efficiency.

[0077] The processing flow will be explained below.

[0078] Step 1:

[0079] When a user starts a meeting or business meeting, they launch Google Docs' voice reading tool and begin recording.

[0080] Step 2:

[0081] Your device will use Google Docs' speech-to-text tool to convert your conversation into text in real time.

[0082] Step 3:

[0083] The device temporarily stores the converted text data locally.

[0084] Step 4:

[0085] When the user finishes a meeting or business negotiation, the device automatically detects the end of the meeting and sends the saved text data to the server.

[0086] Step 5:

[0087] The server queues the received text data for analysis.

[0088] Step 6:

[0089] The server uses generative AI (e.g., GPT-4) to analyze the queued text data and extract key points and content.

[0090] Step 7:

[0091] The server summarizes the extracted data and summarizes the main points.

[0092] Step 8:

[0093] The server embeds the summary data into a standard format template and formats it.

[0094] Step 9:

[0095] The server uses an automation tool such as MAKE to set up a workflow in which the formatted minutes are sent to a specified destination.

[0096] Step 10:

[0097] The server executes the configured workflow and generates meeting minutes, which are then saved in a Google Doc or distributed to relevant parties via email or Slack.

[0098] Step 11:

[0099] The user reviews the shared minutes and makes corrections or additional comments as necessary.

[0100] Example 1

[0101] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0102] The traditional process of recording, organizing, and sharing the details of meetings and business negotiations is often manual, requiring time and effort, and there are problems with content being easily overlooked. Furthermore, the process of manually extracting key points and creating minutes is inaccurate and inefficient, often causing delays in business operations.

[0103] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0104] In this invention, the server includes a means for activating the voice reading tool and starting recording, a means for transmitting text data stored on the terminal to the server, a means for extracting important points and content from the text data using a generative AI model, and a means for transmitting data to a specified destination using an automation tool. This enables real-time conversion of voice data, automatic extraction of important content, standardized formatting of minutes, and automatic distribution of data.

[0105] "Voice data" refers to voice information collected during meetings and business negotiations.

[0106] A "real-time text conversion means" is a mechanism that collects audio data and simultaneously converts it into text form.

[0107] "Text data" refers to character information converted from audio data.

[0108] The "means for analyzing and generating summaries" is a mechanism for analyzing the converted text data, extracting important points and content, and summarizing them in a concise format.

[0109] The "means for converting into a predetermined format" is a mechanism for organizing the generated summary according to a particular template.

[0110] The "means for automatically sending to a specified destination" is a mechanism for automatically sending the generated standard format data to a pre-set recipient.

[0111] A "voice reading tool" is an application or device that collects audio from meetings or business negotiations and converts it into text data in real time.

[0112] A "terminal" is an electronic device used by a user, such as a computer, smartphone, or tablet.

[0113] A "server" is a computer system for performing tasks such as analyzing text data, generating summaries, formatting, and transmitting data.

[0114] A "generative AI model" is an artificial intelligence technology that uses natural language processing to extract important points and content from text data and generate summaries.

[0115] An "automation tool" is software or a platform for automating a series of tasks, such as sending and storing data.

[0116] A "designated destination" is a recipient of automatically delivered data, such as a specific email address, a messaging service account, or a cloud storage folder.

[0117] This invention is a system for quickly recording the contents of meetings and business negotiations and automatically sharing them as organized minutes. The main components include a user device (computer, smartphone, tablet, etc.), a voice reading tool, a server, a generative AI model, and an automation tool.

[0118] System configuration details

[0119] 1. Voice reading

[0120] User: When the meeting or business meeting begins, launch the voice reading tool on the device and begin recording.

[0121] On your device: A voice-to-text tool (such as speech recognition software) is used to convert your conversation into text in real time using speech recognition technology such as the Google Cloud Speech-to-Text API. This converted text is temporarily stored on your device's local storage.

[0122] 2. Sending and analyzing text data

[0123] Terminal: When the meeting or business meeting ends, the saved text data is sent to the server using the HTTPS protocol.

[0124] Server: Receives the text data and inputs it into a generative AI model (e.g., GPT-4) to begin analysis. The generative AI model uses natural language processing techniques to extract key points and content from the text data.

[0125] 3. Summarization and Formatting

[0126] Server: Summarizes the meeting content based on key points extracted by the generative AI model. The server then embeds the summary into a pre-defined minutes template, which includes information such as the meeting name, attendees, agenda, and key points.

[0127] 4. Automatic data distribution

[0128] Server: Use automation tools such as MAKE or Zapier to automatically send the generated minutes to a specified destination (e.g., Google Docs, email, messaging service).

[0129] Specific use cases

[0130] For example, in a meeting for a new product development project, the following steps are taken:

[0131] Attendees: Ichiro Tanaka, Jiro Suzuki, Saburo Sato

[0132] Subject: New product development project

[0133] Agenda: Product specification, development schedule, debugging plan

[0134] procedure

[0135] 1. The user launches the voice capture tool and begins recording the audio of the meeting.

[0136] 2. The device converts the conversation being recorded into text in real time and saves the text data.

[0137] 3. After the meeting ends, the device sends the saved text data to the server, and the server receives the text data.

[0138] 4. The server inputs the received text data into a generative AI model (e.g., GPT-4) to summarize the meeting content. Specifically, it extracts important information related to "product specification determination," "development schedule," and "debug plan."

[0139] 5. The server embeds the extracted key points into a template to generate a standard formatted minutes. For example:

[0140] markdown

[0141] minutes

[0142] Meeting name: New product development project meeting

[0143] Attendees: Ichiro Tanaka, Jiro Suzuki, Saburo Sato

[0144] Date and Time: [meeting date and time]

[0145] agenda:

[0146] 1. Deciding on product specifications

[0147] [Summary]

[0148] 2. Development Schedule

[0149] [Summary]

[0150] 3. Debugging Plan

[0151] [Summary]

[0152] remarks:

[0153] There are no particular problems at the moment

[0154] 6. The server uses MAKE to execute a workflow that saves the generated minutes in a Google Doc or distributes them to all parties involved by email after the meeting ends.

[0155] 7. Users can see that the minutes are automatically shared immediately after the meeting ends and can make additional comments or corrections as needed.

[0156] This system automates a series of tasks, from recording meeting content to organizing and sharing it, greatly improving users' work efficiency.

[0157] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0158] Step 1: Launch the voice reader and start recording

[0159] User: When a meeting or business meeting begins, the user launches the voice-reading tool on their device and starts recording. Specifically, the user clicks the tool icon and presses the record button.

[0160] Enter: Start of a meeting or business meeting.

[0161] Output: Recording begins.

[0162] Step 2: Real-time text conversion of audio data

[0163] On your device: The voice reader converts your conversation into text data in real time. It uses the Google Cloud Speech-to-Text API for speech recognition. The converted text data is temporarily stored in your device's local storage.

[0164] Input: Audio data from meetings and business negotiations.

[0165] Output: Text data generated in real time.

[0166] Specific operation: The speech recognition API converts the voice data into text, and the generated text is saved in the specified directory on the device.

[0167] Step 3: Send text data to the server

[0168] Terminal: When the meeting or business meeting ends, the saved text data is sent to the server. The data is transferred securely using the HTTPS protocol.

[0169] Input: Text data after the meeting ends.

[0170] Output: Text data stored on the server.

[0171] Specific operation: Sends a POST request to the API endpoint for sending text data, and receives a success message after sending.

[0172] Step 4: Parsing the incoming text data

[0173] Server: The received text data is input into the generative AI model (GPT-4) and analysis begins. The model extracts important points and keywords from the input text data.

[0174] Input: Received text data.

[0175] Output: Extracted key points and keywords.

[0176] How it works: The GPT-4 model breaks down text data into tokens, analyzes important context, and extracts key points.

[0177] Step 5: Generate a summary

[0178] Server: Summarizes the meeting content based on key points extracted by the generative AI model. It uses a summary generation algorithm to summarize the key points in a concise format.

[0179] Input: Extracted key points and keywords.

[0180] Output: The summarized text.

[0181] Specific operation: Reconstructs the main points into short sentences and performs text processing to generate a summary.

[0182] Step 6: Embedding in the meeting minutes template

[0183] Server: The summary is embedded into a pre-defined minutes template, which automatically includes information such as the meeting name, attendees, agenda, and key points.

[0184] Input: The summarized text.

[0185] Output: A completed meeting minutes template.

[0186] What it does: Places the summary text into the template placeholder and generates the final minutes document.

[0187] Step 7: Automatic data distribution

[0188] Server: Using automation tools such as MAKE or Zapier, the generated minutes are automatically sent to the specified recipients. The data is stored in a Google Doc or distributed via email or messaging services.

[0189] Input: Completed meeting minutes template.

[0190] Output: The minutes sent to the specified destination.

[0191] Specific behavior: Triggers the automation tool workflow and executes the data sending process, saving it to Google Docs and sending it by email.

[0192] (Application example 1)

[0193] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0194] In industrial sites, there is a need to reduce the time and effort required to record work instructions and generate reports, thereby achieving efficient management. There is also a need to improve the accuracy and speed of work by recording work instructions in real time and instantly sharing necessary information. Conventional methods require the tedious task of manually recording voice instructions and compiling them into reports, making automating this process a challenge.

[0195] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0196] In this invention, the server includes means for converting voice data into text in real time, means for analyzing the converted text data and generating a summary, and means for converting the generated summary into a predetermined format, thereby enabling industrial equipment to record voice instructions in real time and automatically generate work reports.

[0197] definition statement

[0198] "Voice data" refers to audio information such as conversations and instructions recorded in digital format.

[0199] "Text data" is character information generated by analyzing voice data.

[0200] "Real time" means that processing occurs at the exact moment an event occurs.

[0201] A "summary" is a shortened version of a large amount of information, extracting only the important points.

[0202] A "predetermined format" is a predetermined form or template.

[0203] "Destination" is the address or location of the recipient to which the generated data is sent.

[0204] "Industrial equipment" refers to all equipment such as machines and robots used in production sites such as factories.

[0205] A "job report" is a document that records the details of the work performed by industrial equipment.

[0206] "Analysis" is the process of examining data in detail and extracting hidden information and meaning.

[0207] "Generative AI" is an artificial intelligence system that mimics human intelligence and automatically generates and analyzes data.

[0208] "Natural language processing technology" is a technology that allows computers to process and understand human language.

[0209] "Email" is a means of sending and receiving messages electronically over the Internet.

[0210] "Messaging service" refers to a service that allows users to send and receive text messages over the Internet.

[0211] patent specification

[0212] An embodiment of the present invention relates to a system for issuing work instructions and reporting in meetings, business negotiations, and industrial sites. This system has the functions of converting voice data into text in real time, extracting and summarizing important information, converting the summary into a predetermined format, and automatically sending it to a specified destination.

[0213] Hardware and software used

[0214] Hardware: The system includes terminals with audio input devices (e.g., microphones), industrial equipment, and communication modules.

[0215] Software: We use the Google Speech-to-Text API for speech-to-text conversion, GPT-4 as the generative AI model, and MAKE as an automation tool for automatic data transmission and management.

[0216] Program processing

[0217] A user launches a voice-to-text tool and records a meeting, business negotiation, or work instructions at an industrial site. The device converts the recorded voice into text in real time and temporarily stores it. After the meeting or work is completed, the device sends the text data to a server. The server inputs the received text data into a generative AI (e.g., GPT-4) to extract key points. The resulting summary data is converted into a specified format and automatically generated as a work report or meeting minutes. Finally, the server uses MAKE to send the generated report or meeting minutes to the specified destination.

[0218] Specific examples

[0219] Examples of meeting content

[0220] Attendees: Mr. A, Mr. B, Mr. C

[0221] Subject: New product development project

[0222] Agenda: Development schedule, resource allocation, next meeting date

[0223] Example prompt sentence:

[0224] Instructions:

[0225] 1. Assembly Instructions

[0226] Connect part A to part B

[0227] Check the connection

[0228] 2. How to use the tools

[0229] Use a screwdriver and wrench

[0230] 3. Completion checkpoints

[0231] Operation check

[0232] In this example, project meetings can be recorded in real time, summarized, formatted, and finally shared automatically with all stakeholders. A similar process can be used for work instructions in industrial sites to automatically generate work reports that can be instantly sent to managers.

[0233] This system automates a series of tasks, from recording meetings, business negotiations, and work at industrial sites to generating and sharing reports, greatly improving work efficiency.

[0234] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0235] Processing flow

[0236] Program processing steps

[0237] Step 1:

[0238] A user launches a voice-to-speech tool to record a meeting, a business meeting, or work instructions on an industrial site.

[0239] Input: Audio data for meetings and work

[0240] Output: The voice reading tool retrieves the voice data.

[0241] Step 2:

[0242] The device converts the recorded audio into text in real time and stores it temporarily.

[0243] Input: Audio data

[0244] Data processing: Converting voice data into text data using the Google Speech-to-Text API

[0245] Output: Text data is saved to the device

[0246] Step 3:

[0247] After the meeting ends or when work is completed, the device sends the saved text data to the server.

[0248] Input: Text data

[0249] Data calculation: Upload text data to the server via the network

[0250] Output: Text data is sent to the server

[0251] Step 4:

[0252] The server inputs the received text data into a generative AI (e.g., GPT-4) to extract key points.

[0253] Input: Text data

[0254] Data Computing: Analyzing text data using GPT-4 to extract key points

[0255] Output: Extracted key points (summary data)

[0256] Step 5:

[0257] The server converts the generated summary data into a predetermined format.

[0258] Input: Summary data

[0259] Data processing: Embedding summary data into a defined template

[0260] Output: Formatted minutes and working papers

[0261] Step 6:

[0262] The server uses MAKE to send the generated formatted data to the specified destination.

[0263] Input: Formatted minutes and work reports

[0264] Data Calculation: Use MAKE to send data to email or a document management system

[0265] Output: Meeting minutes and work reports are sent to the specified destination.

[0266] Specific examples of operation

[0267] Specific actions of the meeting content

[0268] Meeting scenario: A user launches a voice-to-text tool to record the meeting audio.

[0269] Speech-to-text: Recorded audio is converted to text in real time using the Google Speech-to-Text API.

[0270] Data transmission: After the meeting ends, the device uploads the text data to the server.

[0271] Gist extraction: The server uses GPT-4 to analyze the text data and summarize the main topics.

[0272] Format conversion: Embed the summarized content into a predefined format.

[0273] Data transmission: Using MAKE, formatted minutes are emailed to all parties involved.

[0274] These processing steps enable efficient and automated work recording and reporting in meetings, business negotiations, and industrial sites, significantly improving business efficiency.

[0275] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0276] This invention is a system that quickly and efficiently records the contents of meetings and business negotiations and automatically shares them as organized minutes. It also has the ability to recognize user emotions and reflect that information in the minutes. This system includes a means for converting voice data into text in real time, analyzing it to generate summaries, and converting them into a standard format for sharing. It also incorporates an emotion engine that recognizes user emotions from voice data.

[0277] A natural language description of the program's processing

[0278] Voice reading

[0279] User: At the start of a meeting or business meeting, launch the Google Docs voice-to-text tool and start recording.

[0280] On your device: The speech-to-text tool converts your conversation into text in real time and temporarily stores this text locally.

[0281] Text Acquisition

[0282] Terminal: When the meeting or business negotiation ends, the saved text data is sent to the server.

[0283] Server: Receives text data and queues it for analysis.

[0284] Text organization and summary generation

[0285] Server: Uses generative AI (e.g., GPT-4) to analyze queued text data, extract key points and content, and generate summaries.

[0286] Server: Adds user emotion information recognized by the emotion engine to the summary data.

[0287] Conversion to standard formats

[0288] Server: Embeds and formats the extracted and summarized data and user emotional information into a predetermined standard format.

[0289] Automated workflow execution and sharing

[0290] Server: Uses an automation tool such as MAKE to set up and execute a workflow that sends formatted minutes to the specified destination.

[0291] Server: Save the generated minutes in a Google Doc or distribute them to relevant parties via email or Slack.

[0292] User: Review the shared minutes and make corrections or additional comments as needed.

[0293] Specific examples

[0294] Examples of meeting content

[0295] Attendees: Yuko Sugimoto, Taro Tanaka, Ichiro Kimura

[0296] Subject: New product development project

[0297] Agenda: Development schedule, resource allocation, next meeting date

[0298] 1. The user launches the voice capture tool and begins recording the audio of the meeting.

[0299] 2. The device converts the conversation being recorded into text in real time and saves the text data.

[0300] 3. After the meeting ends, the device sends the saved text data to the server, and the server receives the text data.

[0301] 4. The server inputs the received text data into a generative AI (e.g., GPT-4) to summarize the meeting content. Specifically, it extracts important information related to the "development schedule," "resource allocation," and "next meeting date."

[0302] 5. The server uses an emotion engine to analyze the user's emotions during the conversation. For example, it recognizes the speaker's emotions (excitement, relief, anxiety, etc.) separately from the content of the speech.

[0303] 6. The server embeds the extracted key points and sentiment information into a template to generate a standardized minutes of the meeting. For example:

[0304] minutes

[0305] Meeting name: New product development project meeting

[0306] Attendees: Yuko Sugimoto, Taro Tanaka, Ichiro Kimura

[0307] Date and Time: [meeting date and time]

[0308] agenda:

[0309] 1. Development Schedule

[0310] [Summary]

[0311] [Emotional information]

[0312] 2. Resource Allocation

[0313] [Summary]

[0314] [Emotional information]

[0315] 3. Next meeting date

[0316] [Summary]

[0317] [Emotional information]

[0318] remarks:

[0319] There are no particular problems at the moment

[0320] 7. The server uses MAKE to execute a workflow that saves the generated minutes in a Google Doc or distributes them to all parties involved by email after the meeting ends.

[0321] 8. Users can see that the minutes are automatically shared immediately after the meeting ends and can make additional comments or corrections as needed.

[0322] This system automates a series of tasks, from recording and organizing meeting content, to analyzing emotions and sharing, greatly improving users' work efficiency.

[0323] The processing flow will be explained below.

[0324] Step 1:

[0325] When a user starts a meeting or business meeting, they launch Google Docs' voice reading tool and begin recording.

[0326] Step 2:

[0327] Your device will use Google Docs' speech-to-text tool to convert your conversation into text in real time, which will then be temporarily stored locally.

[0328] Step 3:

[0329] The terminal transmits the saved text data together with the audio file to the server.

[0330] Step 4:

[0331] The server receives the text data and audio files and queues them for analysis.

[0332] Step 5:

[0333] The server uses generative AI (e.g., GPT-4) to analyze the queued text data and extract key points and content.

[0334] Step 6:

[0335] After the server extracts the main points, it uses an emotion engine to recognize the speaker's emotions from the audio file and analyzes the emotion information.

[0336] Step 7:

[0337] The server adds the user's emotion information recognized by the emotion engine to the generated summary data.

[0338] Step 8:

[0339] The server embeds the summary data and emotional information into a predetermined standard format (minutes template) and formats the formatted data.

[0340] Step 9:

[0341] The server uses MAKE's automation tools to set up a workflow that sends formatted minutes to a specified destination.

[0342] Step 10:

[0343] The server executes the configured workflow and generates meeting minutes, which are then saved in a Google Doc or distributed to relevant parties via email or Slack.

[0344] Step 11:

[0345] The user reviews the shared minutes and makes corrections or additional comments as necessary.

[0346] Example 2

[0347] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0348] Conventional meeting recording systems not only convert audio data into text, but also require manual data organization, summary generation, and emotional analysis, resulting in cumbersome post-meeting tasks that require time and effort. Furthermore, they are unable to reflect the emotional information of meeting participants, limiting the content of meeting minutes. This leads to issues such as a decline in the quality of meeting minutes and a decrease in communication efficiency.

[0349] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0350] In this invention, the server includes means for converting voice data into text in real time, means for analyzing the converted text data and generating a summary, means for adding user emotion information recognized by an emotion engine to the generated summary data, and means for automatically sending the generated formatted data to a specified destination. This automates data organization and summary generation after a meeting, and enables the generation and rapid sharing of detailed minutes including emotion information.

[0351] "Voice data" is data that formalizes voice signals collected during meetings, business negotiations, etc.

[0352] "Means for converting voice data into text in real time" refers to a technology or device that instantly analyzes voice signals and converts them into text information.

[0353] "Text data" is information obtained by converting voice data into text format.

[0354] "Analysis" is the process of examining data in detail and extracting specific patterns and important information.

[0355] A "summary" is a short summary of the main points extracted from a long piece of text or a large amount of information.

[0356] An "emotion engine" is a technology or software that analyzes voice or text data and recognizes the speaker's emotions.

[0357] "Format" refers to arranging data into a specific form or structure.

[0358] "Specified destination" means a predetermined recipient or location to which generated data is sent.

[0359] "Generative AI" is an artificial intelligence technology that learns patterns from various input data and generates new outputs.

[0360] An "automation tool" is software or a service that performs a specific process or task without human intervention.

[0361] MODE FOR CARRYING OUT THE INVENTION

[0362] The present invention is a system that quickly and efficiently records the contents of meetings and business negotiations and automatically shares them as organized minutes. It also has the ability to recognize user emotions and reflect that information in the minutes. The system uses a speech recognition tool to convert voice data into text in real time, a generative AI model to analyze and summarize the text data, an emotion engine for emotion analysis, and an automated tool for formatting and sharing. A specific embodiment of this system is described below.

[0363] Voice reading

[0364] When a user starts a meeting or business meeting, they activate a voice recognition tool (e.g., Google Docs' "voice input" function) and start recording. The voice recognition tool converts the spoken audio into text in real time and temporarily stores this text data locally.

[0365] Text Acquisition

[0366] When a meeting or business meeting ends, the device sends the text data stored locally to the server, which receives the text data and queues it for analysis. This queue manages the data waiting to be processed, enabling efficient analysis of the text data.

[0367] Text organization and summary generation

[0368] The server uses a generative AI model (e.g., GPT-4) to analyze the queued text data, extract key points and content, and generate a summary. Furthermore, it adds user emotion information recognized by an emotion engine (e.g., IBM Watson® emotion analysis function) to the generated summary data.

[0369] Conversion to standard formats

[0370] The server embeds the main points and emotional information into a predetermined format and formats it into minutes. This formatted data is stored as a temporary file.

[0371] Automated workflow execution and sharing

[0372] The server uses an automation tool such as MAKE to set up and execute a workflow that sends the minutes converted into a standard format to a specified destination. This workflow is configured to save the generated minutes in a specific document management system (e.g., Google Docs) and distribute them to relevant parties via email or a messaging service (e.g., Slack).

[0373] Specific examples

[0374] Examples of meeting content:

[0375] Attendees: A, B, C

[0376] Topic: New project progress

[0377] Agenda: Schedule confirmation, resource allocation, next meeting date

[0378] At the start of a meeting, a user activates Google Docs' voice input and begins recording the conversation. The recorded audio is converted into text in real time and saved to a local file. After the meeting ends, the text data is sent from the device to a server and queued for analysis. The server uses a generative AI model and emotion engine to analyze the text data, extract key points and emotional information, and generate meeting minutes.

[0379] The minutes are then formatted and saved in a Google Doc or distributed to relevant parties via email or Slack using MAKE's automation tools. Users can review the shared minutes and make additional comments or corrections as needed.

[0380] Example prompt sentence:

[0381] "I want to extract the key points from a meeting about the progress of a new project and generate a summary. I also want to analyze the emotions of the attendees and reflect them in the minutes."

[0382] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0383] Step 1:

[0384] When starting a meeting or business meeting, a user activates a voice recognition tool (e.g., Google Docs' "voice input" function) and starts recording.

[0385] Input: Meeting audio data

[0386] Output: Converted text data

[0387] Specific behavior:

[0388] The user opens Google Docs and selects "Voice Typing" from the Tools menu.

[0389] Once the meeting starts, click the "Speech to Text" button to start dictating. The speech recognition engine will capture the voice data and convert it into text in real time.

[0390] Step 2:

[0391] The device converts the recorded conversation into text in real time and temporarily stores this text data locally.

[0392] Input: Meeting audio data

[0393] Output: Locally saved text data

[0394] Specific behavior:

[0395] The speech recognition engine (Google's speech recognition API) converts the speech data into text information.

[0396] The converted text data is saved in a temporary directory in the terminal's local file system.

[0397] Step 3:

[0398] When the meeting or business negotiation ends, the terminal transmits the saved text data to the server.

[0399] Input: Locally stored text data

[0400] Output: Text data sent to the server

[0401] Specific behavior:

[0402] The user clicks the "Stop" button when the conference is over.

[0403] The device retrieves the text file from the temporary directory and uploads it to the server via an HTTP request.

[0404] Step 4:

[0405] The server receives the text data and queues it for analysis.

[0406] Input: Text data sent from the terminal

[0407] Output: Text data queued for analysis

[0408] Specific behavior:

[0409] The server stores the received data in a temporary storage directory.

[0410] The text data is registered as a task in the analysis queue.

[0411] Step 5:

[0412] The server uses a generative AI model (e.g., GPT-4) to analyze the queued text data, extract key points and content, and generate a summary.

[0413] Input: Text data queued for analysis

[0414] Output: Generated summary data

[0415] Specific behavior:

[0416] The server retrieves the text data from the queue.

[0417] A prompt sentence is input to the generative AI model and a summary result is received from the generative AI model.

[0418] Step 6:

[0419] The server adds user emotion information recognized by an emotion engine (e.g., IBM Watson's emotion analysis function) to the summary data.

[0420] Input: Generated summary data

[0421] Output: Summarized data with added sentiment information

[0422] Specific behavior:

[0423] The server sends the voice or text data to the sentiment analysis engine.

[0424] The received emotion information is incorporated into the summary data.

[0425] Step 7:

[0426] The server embeds the extracted main points and emotional information into a predetermined standard format and formats it.

[0427] Input: Summarized data with added sentiment information

[0428] Output: Minutes in a standard format

[0429] Specific behavior:

[0430] The server reads a template file with a fixed format and replaces the summary data and emotion information with the corresponding placeholders.

[0431] Save the formatted minutes as a temporary file.

[0432] Step 8:

[0433] The server uses an automation tool such as MAKE to set up and execute a workflow that sends the formatted minutes to the specified destination.

[0434] Input: Minutes in a standard format

[0435] Output: The minutes sent to the specified destination

[0436] Specific behavior:

[0437] The server sends an API request to an automation tool (e.g., MAKE) to trigger the workflow.

[0438] The workflow saves meeting minutes in a Google Doc and distributes them to relevant parties via email or messaging services.

[0439] Step 9:

[0440] Users can review the shared minutes and make corrections or additional comments as needed.

[0441] Input: Shared minutes

[0442] Output: Meeting minutes with corrections and additional comments

[0443] Specific behavior:

[0444] The user opens a Google Doc or an email they received and checks the minutes.

[0445] If any corrections or comments are required, please use the editing function of Google Docs to make additions or corrections.

[0446] (Application example 2)

[0447] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0448] Current meeting recording systems and work instruction recording systems simply convert audio data into text, without detailed analysis or summarization, including emotional information. As a result, the content of meetings and work instructions may not be accurately understood, leading to miscommunication and lack of awareness. Furthermore, methods for automatically sharing formatted data are limited, making efficient information sharing difficult. A system that solves these issues is needed.

[0449] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for converting voice data to text in real time, means for analyzing the converted text data and generating a summary, means for recognizing emotions from the voice data and reflecting the emotional information in the summary, means for converting the generated summary into a predetermined format, and means for automatically sending the generated formatted data to a specified destination. This enables accurate text conversion and analysis of voice data, addition of emotional information, and efficient information sharing.

[0450] "Audio data" refers to data that has been recorded and stored in digital form, such as voice or spoken words.

[0451] "Real-time text conversion means" refers to technology or equipment for instantly converting voice data into text form.

[0452] "Text data" refers to digital data expressed as character information.

[0453] "Analysis" refers to the process of analyzing and interpreting input data to understand its meaning and structure.

[0454] "Means for generating summaries" refers to technologies and devices that extract important information from large amounts of text data and summarize it concisely.

[0455] "Means for recognizing emotions and incorporating that emotional information into summaries" refers to technology and devices for identifying emotions from audio data and incorporating that emotional information into summaries.

[0456] The "means for converting into a predetermined format" refers to a technique or device for converting the generated summary into a predetermined format.

[0457] "Formatted data" refers to digital data that has been arranged in a particular format.

[0458] "Means for automatically sending to a specified destination" refers to technology or devices for automatically sending generated data to a specific recipient.

[0459] "Designated destination" refers to a recipient or location that has been pre-configured to receive data.

[0460] "Improved natural language processing technology" refers to technology for analyzing and processing natural language with improved accuracy and efficiency compared to conventional technology.

[0461] "Document Management System" means software or a platform for storing, managing, and accessing documents and data.

[0462] "Email or messaging service" means a service for sending and receiving messages or information over the Internet.

[0463] "Automation tool" refers to software that automatically performs designated tasks or processes.

[0464] To implement this invention, it is necessary to build a system that converts voice data into text in real time, analyzes the text data to generate a summary, converts it into a standard format after reflecting emotional information, and automatically sends it to a specified destination. This system aims to efficiently record work instructions in a factory and analyze the content of those instructions.

[0465] composition

[0466] The system components are as follows:

[0467] 1. Voice data collection devices: Microphones and voice-reading devices placed on the factory floor to record work instructions and conversations.

[0468] 2. Speech Recognition Software: Software that converts voice data into text in real time. Specifically, we use the Google Cloud Speech-to-Text API.

[0469] 3. Text analysis and summary generation software: Uses generative AI models to analyze converted text data and generate summaries. OpenAI's GPT-4 is a prime example.

[0470] 4. Emotion Recognition Software: Software to recognize emotions from audio data and incorporate that emotional information into summaries. Uses IBM Watson Tone Analyzer.

[0471] 5. Data conversion and sharing software: Software that converts the generated summaries into a standard format and automatically sends them to designated recipients. This includes automation tools such as Zapier.

[0472] Operation flow

[0473] 1. Voice reading

[0474] Users use a voice data capture device to record work instructions and conversations, and speech recognition software converts the recorded voice data into text in real time.

[0475] 2. Text Analysis

[0476] Text analysis and summary generation software analyzes the converted text data, extracts key points, and generates summaries using a generative AI model (e.g., GPT-4).

[0477] 3. Emotion recognition

[0478] Emotion recognition software (IBM Watson Tone Analyzer) is used to recognize emotions from audio data. The analyzed emotional information is reflected in the summary.

[0479] 4. Data Sharing

[0480] The generated data, including summaries and sentiment information, is converted into a predetermined format and sent to a designated destination via an automation tool (such as Zapier), allowing relevant parties to receive important information in real time.

[0481] Specific examples

[0482] If the factory manager instructs the robot to "run a test run of a new product on the next line," the system operates as follows:

[0483] An audio data collection device records the instructions.

[0484] Speech recognition software converts this audio data into text.

[0485] The text analysis and summary generation software summarizes it as "Test run of new product on the production line."

[0486] Emotion recognition software analyzes the "high level of attention" and "sense of security" of the person giving the instructions.

[0487] The generated summary and emotion information are converted into a predetermined format and sent to a specified destination.

[0488] Example prompt sentence:

[0489] "Please summarize the following text:

[0490] On the next line, I would like you to conduct a test run of a new product.

[0491] High level of focus and peace of mind

[0492] The above is a detailed embodiment for carrying out the invention. This system enables efficient and accurate recording and transmission of work instructions in a factory.

[0493] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0494] Step 1:

[0495] Voice reading

[0496] The user uses a voice data collection device to record work instructions and conversations.

[0497] Input: Audio data

[0498] Specific operation: Microphones and voice reading devices installed on the factory floor collect voices from factory managers and work instructors in real time.

[0499] Output: Real-time audio data

[0500] Step 2:

[0501] Real-time text transcription

[0502] The device converts the collected voice data into text in real time using voice recognition software (Google Cloud Speech-to-Text API).

[0503] Input: Real-time audio data

[0504] How it works: Speech recognition software analyzes speech waveforms and converts what is being said into text.

[0505] Output: Text data

[0506] Step 3:

[0507] Text data analysis

[0508] The device uses a generative AI model (OpenAI GPT-4) to analyze the converted text data, extract key points, and generate a summary.

[0509] Input: Text data

[0510] How it works: The generative AI model analyzes the input text data, uses prompts to identify key information, and summarizes it succinctly.

[0511] Output: Summary text

[0512] Step 4:

[0513] emotion recognition

[0514] The device uses emotion recognition software (IBM Watson Tone Analyzer) to recognize emotions from the analyzed voice data and reflects that emotional information in the summary.

[0515] Input: Text data

[0516] What it does: Emotion recognition software analyzes keywords and context within text data to identify the speaker's emotions (e.g., excitement, relief, anxiety, etc.).

[0517] Output: Summarized text with emotional information

[0518] Step 5:

[0519] Conversion to standard formats

[0520] The terminal converts the generated summary text into a predetermined format.

[0521] Input: Summarized text with sentiment information

[0522] What it does: The formatting function embeds summary and sentiment information into templates to generate standardized work instructions and meeting minutes.

[0523] Output: formatted data

[0524] Step 6:

[0525] Automatic Data Transmission and Sharing

[0526] The server uses an automation tool (such as Zapier) to automatically send the generated formatted data to the specified destination.

[0527] Input: Formatted data

[0528] What it does: The automation tool distributes data to a pre-specified list of recipients or sharers via email, messaging services, or document management systems.

[0529] Output: Shared data

[0530] Through these steps, a series of automated processes are implemented, from collecting voice data to analyzing it, recognizing emotions, converting formats, and sharing the information, enabling accurate and prompt recording and transmission of work instructions in factories.

[0531] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0532] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (registered trademark) (Internet search engine).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0533] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0534] [Second embodiment]

[0535] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0536] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0537] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0538] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0539] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0540] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0541] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0542] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0543] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0544] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0545] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0546] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0547] This system provides a means to quickly record the contents of meetings and business negotiations and automatically share them as organized minutes. When a user holds a meeting or business negotiation, they first launch the voice-to-text tool. The voice data is converted into text data in real time. This converted text data is sent from the user's device to a server where it is analyzed. This analysis uses generative AI (e.g., GPT-4) to extract important points and content and generate a summary. The resulting summary data is then converted into a standard format. Finally, the formatted data is automatically sent to the specified destination.

[0548] A natural language description of the program's processing

[0549] Voice reading

[0550] User: When a meeting or business meeting begins, launch the voice-to-text tool and start recording.

[0551] On your device: A speech-to-text tool converts your conversation into text in real time, which is temporarily stored on your device.

[0552] Text Acquisition

[0553] Terminal: When the meeting or business negotiation ends, the saved text data is sent to the server.

[0554] Server: Receives the text data and begins analysis.

[0555] Text organization

[0556] Server: Using generative AI, the server analyzes the received text data and extracts key points and content, thereby summarizing the text.

[0557] Conversion to standard formats

[0558] Server: The extracted key points are embedded into a pre-defined format to complete the minutes template.

[0559] Automation workflow execution

[0560] Server: Using an automation tool such as MAKE, the generated minutes are automatically sent to the specified destination.

[0561] Specific examples

[0562] Examples of meeting content

[0563] Attendees: Yuko Sugimoto, Taro Tanaka, Ichiro Kimura

[0564] Subject: New product development project

[0565] Agenda: Development schedule, resource allocation, next meeting date

[0566] 1. The user launches the voice capture tool and begins recording the audio of the meeting.

[0567] 2. The device converts the conversation being recorded into text in real time and saves the text data.

[0568] 3. After the meeting ends, the device sends the saved text data to the server, and the server receives the text data.

[0569] 4. The server inputs the received text data into a generative AI (e.g., GPT-4) to summarize the meeting content. Specifically, it extracts important information related to the "development schedule," "resource allocation," and "next meeting date."

[0570] 5. The server embeds the extracted key points into a template to generate a standard formatted minutes. For example:

[0571] minutes

[0572] Meeting name: New product development project meeting

[0573] Attendees: Yuko Sugimoto, Taro Tanaka, Ichiro Kimura

[0574] Date and Time: [meeting date and time]

[0575] agenda:

[0576] 1. Development Schedule

[0577] [Summary]

[0578] 2. Resource Allocation

[0579] [Summary]

[0580] 3. Next meeting date

[0581] [Summary]

[0582] remarks:

[0583] There are no particular problems at the moment

[0584] 6. The server uses MAKE to execute a workflow that saves the generated minutes in a Google Doc or distributes them to all parties involved by email after the video conference ends.

[0585] 7. Users can see that the minutes are automatically shared immediately after the meeting ends and can make additional comments or corrections as needed.

[0586] This system automates a series of tasks, from recording meeting content to organizing and sharing it, greatly improving users' work efficiency.

[0587] The processing flow will be explained below.

[0588] Step 1:

[0589] When a user starts a meeting or business meeting, they launch Google Docs' voice reading tool and begin recording.

[0590] Step 2:

[0591] Your device will use Google Docs' speech-to-text tool to convert your conversation into text in real time.

[0592] Step 3:

[0593] The device temporarily stores the converted text data locally.

[0594] Step 4:

[0595] When the user finishes a meeting or business negotiation, the device automatically detects the end of the meeting and sends the saved text data to the server.

[0596] Step 5:

[0597] The server queues the received text data for analysis.

[0598] Step 6:

[0599] The server uses generative AI (e.g., GPT-4) to analyze the queued text data and extract key points and content.

[0600] Step 7:

[0601] The server summarizes the extracted data and summarizes the main points.

[0602] Step 8:

[0603] The server embeds the summary data into a standard format template and formats it.

[0604] Step 9:

[0605] The server uses an automation tool such as MAKE to set up a workflow in which the formatted minutes are sent to a specified destination.

[0606] Step 10:

[0607] The server executes the configured workflow and generates meeting minutes, which are then saved in a Google Doc or distributed to relevant parties via email or Slack.

[0608] Step 11:

[0609] The user reviews the shared minutes and makes corrections or additional comments as necessary.

[0610] Example 1

[0611] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0612] The traditional process of recording, organizing, and sharing the details of meetings and business negotiations is often manual, requiring time and effort, and there are problems with content being easily overlooked. Furthermore, the process of manually extracting key points and creating minutes is inaccurate and inefficient, often causing delays in business operations.

[0613] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0614] In this invention, the server includes a means for activating the voice reading tool and starting recording, a means for transmitting text data stored on the terminal to the server, a means for extracting important points and content from the text data using a generative AI model, and a means for transmitting data to a specified destination using an automation tool. This enables real-time conversion of voice data, automatic extraction of important content, standardized formatting of minutes, and automatic distribution of data.

[0615] "Voice data" refers to voice information collected during meetings and business negotiations.

[0616] A "real-time text conversion means" is a mechanism that collects audio data and simultaneously converts it into text form.

[0617] "Text data" refers to character information converted from audio data.

[0618] The "means for analyzing and generating summaries" is a mechanism for analyzing the converted text data, extracting important points and content, and summarizing them in a concise format.

[0619] The "means for converting into a predetermined format" is a mechanism for organizing the generated summary according to a particular template.

[0620] The "means for automatically sending to a specified destination" is a mechanism for automatically sending the generated standard format data to a pre-set recipient.

[0621] A "voice reading tool" is an application or device that collects audio from meetings or business negotiations and converts it into text data in real time.

[0622] A "terminal" is an electronic device used by a user, such as a computer, smartphone, or tablet.

[0623] A "server" is a computer system for performing tasks such as analyzing text data, generating summaries, formatting, and transmitting data.

[0624] A "generative AI model" is an artificial intelligence technology that uses natural language processing to extract important points and content from text data and generate summaries.

[0625] An "automation tool" is software or a platform for automating a series of tasks, such as sending and storing data.

[0626] A "designated destination" is a recipient of automatically delivered data, such as a specific email address, a messaging service account, or a cloud storage folder.

[0627] This invention is a system for quickly recording the contents of meetings and business negotiations and automatically sharing them as organized minutes. The main components include a user device (computer, smartphone, tablet, etc.), a voice reading tool, a server, a generative AI model, and an automation tool.

[0628] System configuration details

[0629] 1. Voice reading

[0630] User: When the meeting or business meeting begins, launch the voice reading tool on the device and begin recording.

[0631] On your device: A voice-to-text tool (such as speech recognition software) is used to convert your conversation into text in real time using speech recognition technology such as the Google Cloud Speech-to-Text API. This converted text is temporarily stored on your device's local storage.

[0632] 2. Sending and analyzing text data

[0633] Terminal: When the meeting or business meeting ends, the saved text data is sent to the server using the HTTPS protocol.

[0634] Server: Receives the text data and inputs it into a generative AI model (e.g., GPT-4) to begin analysis. The generative AI model uses natural language processing techniques to extract key points and content from the text data.

[0635] 3. Summarization and Formatting

[0636] Server: Summarizes the meeting content based on key points extracted by the generative AI model. The server then embeds the summary into a pre-defined minutes template, which includes information such as the meeting name, attendees, agenda, and key points.

[0637] 4. Automatic data distribution

[0638] Server: Use automation tools such as MAKE or Zapier to automatically send the generated minutes to a specified destination (e.g., Google Docs, email, messaging service).

[0639] Specific use cases

[0640] For example, in a meeting for a new product development project, the following steps are taken:

[0641] Attendees: Ichiro Tanaka, Jiro Suzuki, Saburo Sato

[0642] Subject: New product development project

[0643] Agenda: Product specification, development schedule, debugging plan

[0644] procedure

[0645] 1. The user launches the voice capture tool and begins recording the audio of the meeting.

[0646] 2. The device converts the conversation being recorded into text in real time and saves the text data.

[0647] 3. After the meeting ends, the device sends the saved text data to the server, and the server receives the text data.

[0648] 4. The server inputs the received text data into a generative AI model (e.g., GPT-4) to summarize the meeting content. Specifically, it extracts important information related to "product specification determination," "development schedule," and "debug plan."

[0649] 5. The server embeds the extracted key points into a template to generate a standard formatted minutes. For example:

[0650] markdown

[0651] minutes

[0652] Meeting name: New product development project meeting

[0653] Attendees: Ichiro Tanaka, Jiro Suzuki, Saburo Sato

[0654] Date and Time: [meeting date and time]

[0655] agenda:

[0656] 1. Deciding on product specifications

[0657] [Summary]

[0658] 2. Development Schedule

[0659] [Summary]

[0660] 3. Debugging Plan

[0661] [Summary]

[0662] remarks:

[0663] There are no particular problems at the moment

[0664] 6. The server uses MAKE to execute a workflow that saves the generated minutes in a Google Doc or distributes them to all parties involved by email after the meeting ends.

[0665] 7. Users can see that the minutes are automatically shared immediately after the meeting ends and can make additional comments or corrections as needed.

[0666] This system automates a series of tasks, from recording meeting content to organizing and sharing it, greatly improving users' work efficiency.

[0667] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0668] Step 1: Launch the voice reader and start recording

[0669] User: When a meeting or business meeting begins, the user launches the voice-reading tool on their device and starts recording. Specifically, the user clicks the tool icon and presses the record button.

[0670] Enter: Start of a meeting or business meeting.

[0671] Output: Recording begins.

[0672] Step 2: Real-time text conversion of audio data

[0673] On your device: The voice reader converts your conversation into text data in real time. It uses the Google Cloud Speech-to-Text API for speech recognition. The converted text data is temporarily stored in your device's local storage.

[0674] Input: Audio data from meetings and business negotiations.

[0675] Output: Text data generated in real time.

[0676] Specific operation: The speech recognition API converts the voice data into text, and the generated text is saved in the specified directory on the device.

[0677] Step 3: Send text data to the server

[0678] Terminal: When the meeting or business meeting ends, the saved text data is sent to the server. The data is transferred securely using the HTTPS protocol.

[0679] Input: Text data after the meeting ends.

[0680] Output: Text data stored on the server.

[0681] Specific operation: Sends a POST request to the API endpoint for sending text data, and receives a success message after sending.

[0682] Step 4: Parsing the incoming text data

[0683] Server: The received text data is input into the generative AI model (GPT-4) and analysis begins. The model extracts important points and keywords from the input text data.

[0684] Input: Received text data.

[0685] Output: Extracted key points and keywords.

[0686] How it works: The GPT-4 model breaks down text data into tokens, analyzes important context, and extracts key points.

[0687] Step 5: Generate a summary

[0688] Server: Summarizes the meeting content based on key points extracted by the generative AI model. It uses a summary generation algorithm to summarize the key points in a concise format.

[0689] Input: Extracted key points and keywords.

[0690] Output: The summarized text.

[0691] Specific operation: Reconstructs the main points into short sentences and performs text processing to generate a summary.

[0692] Step 6: Embedding in the meeting minutes template

[0693] Server: The summary is embedded into a pre-defined minutes template, which automatically includes information such as the meeting name, attendees, agenda, and key points.

[0694] Input: The summarized text.

[0695] Output: A completed meeting minutes template.

[0696] What it does: Places the summary text into the template placeholder and generates the final minutes document.

[0697] Step 7: Automatic data distribution

[0698] Server: Using automation tools such as MAKE or Zapier, the generated minutes are automatically sent to the specified recipients. The data is stored in a Google Doc or distributed via email or messaging services.

[0699] Input: Completed meeting minutes template.

[0700] Output: The minutes sent to the specified destination.

[0701] Specific behavior: Triggers the automation tool workflow and executes the data sending process, saving it to Google Docs and sending it by email.

[0702] (Application example 1)

[0703] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0704] In industrial sites, there is a need to reduce the time and effort required to record work instructions and generate reports, thereby achieving efficient management. There is also a need to improve the accuracy and speed of work by recording work instructions in real time and instantly sharing necessary information. Conventional methods require the tedious task of manually recording voice instructions and compiling them into reports, making automating this process a challenge.

[0705] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0706] In this invention, the server includes means for converting voice data into text in real time, means for analyzing the converted text data and generating a summary, and means for converting the generated summary into a predetermined format, thereby enabling industrial equipment to record voice instructions in real time and automatically generate work reports.

[0707] definition statement

[0708] "Voice data" refers to audio information such as conversations and instructions recorded in digital format.

[0709] "Text data" is character information generated by analyzing voice data.

[0710] "Real time" means that processing occurs at the exact moment an event occurs.

[0711] A "summary" is a shortened version of a large amount of information, extracting only the important points.

[0712] A "predetermined format" is a predetermined form or template.

[0713] "Destination" is the address or location of the recipient to which the generated data is sent.

[0714] "Industrial equipment" refers to all equipment such as machines and robots used in production sites such as factories.

[0715] A "job report" is a document that records the details of the work performed by industrial equipment.

[0716] "Analysis" is the process of examining data in detail and extracting hidden information and meaning.

[0717] "Generative AI" is an artificial intelligence system that mimics human intelligence and automatically generates and analyzes data.

[0718] "Natural language processing technology" is a technology that allows computers to process and understand human language.

[0719] "Email" is a means of sending and receiving messages electronically over the Internet.

[0720] "Messaging service" refers to a service that allows users to send and receive text messages over the Internet.

[0721] patent specification

[0722] An embodiment of the present invention relates to a system for issuing work instructions and reporting in meetings, business negotiations, and industrial sites. This system has the functions of converting voice data into text in real time, extracting and summarizing important information, converting the summary into a predetermined format, and automatically sending it to a specified destination.

[0723] Hardware and software used

[0724] Hardware: The system includes terminals with audio input devices (e.g., microphones), industrial equipment, and communication modules.

[0725] Software: We use the Google Speech-to-Text API for speech-to-text conversion, GPT-4 as the generative AI model, and MAKE as an automation tool for automatic data transmission and management.

[0726] Program processing

[0727] A user launches a voice-to-text tool and records a meeting, business negotiation, or work instructions at an industrial site. The device converts the recorded voice into text in real time and temporarily stores it. After the meeting or work is completed, the device sends the text data to a server. The server inputs the received text data into a generative AI (e.g., GPT-4) to extract key points. The resulting summary data is converted into a specified format and automatically generated as a work report or meeting minutes. Finally, the server uses MAKE to send the generated report or meeting minutes to the specified destination.

[0728] Specific examples

[0729] Examples of meeting content

[0730] Attendees: Mr. A, Mr. B, Mr. C

[0731] Subject: New product development project

[0732] Agenda: Development schedule, resource allocation, next meeting date

[0733] Example prompt sentence:

[0734] Instructions:

[0735] 1. Assembly Instructions

[0736] Connect part A to part B

[0737] Check the connection

[0738] 2. How to use the tools

[0739] Use a screwdriver and wrench

[0740] 3. Completion checkpoints

[0741] Operation check

[0742] In this example, project meetings can be recorded in real time, summarized, formatted, and finally shared automatically with all stakeholders. A similar process can be used for work instructions in industrial sites to automatically generate work reports that can be instantly sent to managers.

[0743] This system automates a series of tasks, from recording meetings, business negotiations, and work at industrial sites to generating and sharing reports, greatly improving work efficiency.

[0744] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0745] Processing flow

[0746] Program processing steps

[0747] Step 1:

[0748] A user launches a voice-to-speech tool to record a meeting, a business meeting, or work instructions on an industrial site.

[0749] Input: Audio data for meetings and work

[0750] Output: The voice reading tool retrieves the voice data.

[0751] Step 2:

[0752] The device converts the recorded audio into text in real time and stores it temporarily.

[0753] Input: Audio data

[0754] Data processing: Converting voice data into text data using the Google Speech-to-Text API

[0755] Output: Text data is saved to the device

[0756] Step 3:

[0757] After the meeting ends or when work is completed, the device sends the saved text data to the server.

[0758] Input: Text data

[0759] Data calculation: Upload text data to the server via the network

[0760] Output: Text data is sent to the server

[0761] Step 4:

[0762] The server inputs the received text data into a generative AI (e.g., GPT-4) to extract key points.

[0763] Input: Text data

[0764] Data Computing: Analyzing text data using GPT-4 to extract key points

[0765] Output: Extracted key points (summary data)

[0766] Step 5:

[0767] The server converts the generated summary data into a predetermined format.

[0768] Input: Summary data

[0769] Data processing: Embedding summary data into a defined template

[0770] Output: Formatted minutes and working papers

[0771] Step 6:

[0772] The server uses MAKE to send the generated formatted data to the specified destination.

[0773] Input: Formatted minutes and work reports

[0774] Data Calculation: Use MAKE to send data to email or a document management system

[0775] Output: Meeting minutes and work reports are sent to the specified destination.

[0776] Specific examples of operation

[0777] Specific actions of the meeting content

[0778] Meeting scenario: A user launches a voice-to-text tool to record the meeting audio.

[0779] Speech-to-text: Recorded audio is converted to text in real time using the Google Speech-to-Text API.

[0780] Data transmission: After the meeting ends, the device uploads the text data to the server.

[0781] Gist extraction: The server uses GPT-4 to analyze the text data and summarize the main topics.

[0782] Format conversion: Embed the summarized content into a predefined format.

[0783] Data transmission: Using MAKE, formatted minutes are emailed to all parties involved.

[0784] These processing steps enable efficient and automated work recording and reporting in meetings, business negotiations, and industrial sites, significantly improving business efficiency.

[0785] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0786] This invention is a system that quickly and efficiently records the contents of meetings and business negotiations and automatically shares them as organized minutes. It also has the ability to recognize user emotions and reflect that information in the minutes. This system includes a means for converting voice data into text in real time, analyzing it to generate summaries, and converting them into a standard format for sharing. It also incorporates an emotion engine that recognizes user emotions from voice data.

[0787] A natural language description of the program's processing

[0788] Voice reading

[0789] User: At the start of a meeting or business meeting, launch the Google Docs voice-to-text tool and start recording.

[0790] On your device: The speech-to-text tool converts your conversation into text in real time and temporarily stores this text locally.

[0791] Text Acquisition

[0792] Terminal: When the meeting or business negotiation ends, the saved text data is sent to the server.

[0793] Server: Receives text data and queues it for analysis.

[0794] Text organization and summary generation

[0795] Server: Uses generative AI (e.g., GPT-4) to analyze queued text data, extract key points and content, and generate summaries.

[0796] Server: Adds user emotion information recognized by the emotion engine to the summary data.

[0797] Conversion to standard formats

[0798] Server: Embeds and formats the extracted and summarized data and user emotional information into a predetermined standard format.

[0799] Automated workflow execution and sharing

[0800] Server: Uses an automation tool such as MAKE to set up and execute a workflow that sends formatted minutes to the specified destination.

[0801] Server: Save the generated minutes in a Google Doc or distribute them to relevant parties via email or Slack.

[0802] User: Review the shared minutes and make corrections or additional comments as needed.

[0803] Specific examples

[0804] Examples of meeting content

[0805] Attendees: Yuko Sugimoto, Taro Tanaka, Ichiro Kimura

[0806] Subject: New product development project

[0807] Agenda: Development schedule, resource allocation, next meeting date

[0808] 1. The user launches the voice capture tool and begins recording the audio of the meeting.

[0809] 2. The device converts the conversation being recorded into text in real time and saves the text data.

[0810] 3. After the meeting ends, the device sends the saved text data to the server, and the server receives the text data.

[0811] 4. The server inputs the received text data into a generative AI (e.g., GPT-4) to summarize the meeting content. Specifically, it extracts important information related to the "development schedule," "resource allocation," and "next meeting date."

[0812] 5. The server uses an emotion engine to analyze the user's emotions during the conversation. For example, it recognizes the speaker's emotions (excitement, relief, anxiety, etc.) separately from the content of the speech.

[0813] 6. The server embeds the extracted key points and sentiment information into a template to generate a standardized minutes of the meeting. For example:

[0814] minutes

[0815] Meeting name: New product development project meeting

[0816] Attendees: Yuko Sugimoto, Taro Tanaka, Ichiro Kimura

[0817] Date and Time: [meeting date and time]

[0818] agenda:

[0819] 1. Development Schedule

[0820] [Summary]

[0821] [Emotional information]

[0822] 2. Resource Allocation

[0823] [Summary]

[0824] [Emotional information]

[0825] 3. Next meeting date

[0826] [Summary]

[0827] [Emotional information]

[0828] remarks:

[0829] There are no particular problems at the moment

[0830] 7. The server uses MAKE to execute a workflow that saves the generated minutes in a Google Doc or distributes them to all parties involved by email after the meeting ends.

[0831] 8. Users can see that the minutes are automatically shared immediately after the meeting ends and can make additional comments or corrections as needed.

[0832] This system automates a series of tasks, from recording and organizing meeting content, to analyzing emotions and sharing, greatly improving users' work efficiency.

[0833] The processing flow will be explained below.

[0834] Step 1:

[0835] When a user starts a meeting or business meeting, they launch Google Docs' voice reading tool and begin recording.

[0836] Step 2:

[0837] Your device will use Google Docs' speech-to-text tool to convert your conversation into text in real time, which will then be temporarily stored locally.

[0838] Step 3:

[0839] The terminal transmits the saved text data together with the audio file to the server.

[0840] Step 4:

[0841] The server receives the text data and audio files and queues them for analysis.

[0842] Step 5:

[0843] The server uses generative AI (e.g., GPT-4) to analyze the queued text data and extract key points and content.

[0844] Step 6:

[0845] After the server extracts the main points, it uses an emotion engine to recognize the speaker's emotions from the audio file and analyzes the emotion information.

[0846] Step 7:

[0847] The server adds the user's emotion information recognized by the emotion engine to the generated summary data.

[0848] Step 8:

[0849] The server embeds the summary data and emotional information into a predetermined standard format (minutes template) and formats the formatted data.

[0850] Step 9:

[0851] The server uses MAKE's automation tools to set up a workflow that sends formatted minutes to a specified destination.

[0852] Step 10:

[0853] The server executes the configured workflow and generates meeting minutes, which are then saved in a Google Doc or distributed to relevant parties via email or Slack.

[0854] Step 11:

[0855] The user reviews the shared minutes and makes corrections or additional comments as necessary.

[0856] Example 2

[0857] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0858] Conventional meeting recording systems not only convert audio data into text, but also require manual data organization, summary generation, and emotional analysis, resulting in cumbersome post-meeting tasks that require time and effort. Furthermore, they are unable to reflect the emotional information of meeting participants, limiting the content of meeting minutes. This leads to issues such as a decline in the quality of meeting minutes and a decrease in communication efficiency.

[0859] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0860] In this invention, the server includes means for converting voice data into text in real time, means for analyzing the converted text data and generating a summary, means for adding user emotion information recognized by an emotion engine to the generated summary data, and means for automatically sending the generated formatted data to a specified destination. This automates data organization and summary generation after a meeting, and enables the generation and rapid sharing of detailed minutes including emotion information.

[0861] "Voice data" is data that formalizes voice signals collected during meetings, business negotiations, etc.

[0862] "Means for converting voice data into text in real time" refers to a technology or device that instantly analyzes voice signals and converts them into text information.

[0863] "Text data" is information obtained by converting voice data into text format.

[0864] "Analysis" is the process of examining data in detail and extracting specific patterns and important information.

[0865] A "summary" is a short summary of the main points extracted from a long piece of text or a large amount of information.

[0866] An "emotion engine" is a technology or software that analyzes voice or text data and recognizes the speaker's emotions.

[0867] "Format" refers to arranging data into a specific form or structure.

[0868] "Specified destination" means a predetermined recipient or location to which generated data is sent.

[0869] "Generative AI" is an artificial intelligence technology that learns patterns from various input data and generates new outputs.

[0870] An "automation tool" is software or a service that performs a specific process or task without human intervention.

[0871] MODE FOR CARRYING OUT THE INVENTION

[0872] The present invention is a system that quickly and efficiently records the contents of meetings and business negotiations and automatically shares them as organized minutes. It also has the ability to recognize user emotions and reflect that information in the minutes. The system uses a speech recognition tool to convert voice data into text in real time, a generative AI model to analyze and summarize the text data, an emotion engine for emotion analysis, and an automated tool for formatting and sharing. A specific embodiment of this system is described below.

[0873] Voice reading

[0874] When a user starts a meeting or business meeting, they activate a voice recognition tool (e.g., Google Docs' "voice input" function) and start recording. The voice recognition tool converts the spoken audio into text in real time and temporarily stores this text data locally.

[0875] Text Acquisition

[0876] When a meeting or business meeting ends, the device sends the text data stored locally to the server, which receives the text data and queues it for analysis. This queue manages the data waiting to be processed, enabling efficient analysis of the text data.

[0877] Text organization and summary generation

[0878] The server uses a generative AI model (e.g., GPT-4) to analyze the queued text data, extract key points and content, and generate a summary. Furthermore, it adds user sentiment information recognized by an emotion engine (e.g., IBM Watson's sentiment analysis function) to the generated summary data.

[0879] Conversion to standard formats

[0880] The server embeds the main points and emotional information into a predetermined format and formats it into minutes. This formatted data is stored as a temporary file.

[0881] Automated workflow execution and sharing

[0882] The server uses an automation tool such as MAKE to set up and execute a workflow that sends the minutes converted into a standard format to a specified destination. This workflow is configured to save the generated minutes in a specific document management system (e.g., Google Docs) and distribute them to relevant parties via email or a messaging service (e.g., Slack).

[0883] Specific examples

[0884] Examples of meeting content:

[0885] Attendees: A, B, C

[0886] Topic: New project progress

[0887] Agenda: Schedule confirmation, resource allocation, next meeting date

[0888] At the start of a meeting, a user activates Google Docs' voice input and begins recording the conversation. The recorded audio is converted into text in real time and saved to a local file. After the meeting ends, the text data is sent from the device to a server and queued for analysis. The server uses a generative AI model and emotion engine to analyze the text data, extract key points and emotional information, and generate meeting minutes.

[0889] The minutes are then formatted and saved in a Google Doc or distributed to relevant parties via email or Slack using MAKE's automation tools. Users can review the shared minutes and make additional comments or corrections as needed.

[0890] Example prompt sentence:

[0891] "I want to extract the key points from a meeting about the progress of a new project and generate a summary. I also want to analyze the emotions of the attendees and reflect them in the minutes."

[0892] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0893] Step 1:

[0894] When starting a meeting or business meeting, a user activates a voice recognition tool (e.g., Google Docs' "voice input" function) and starts recording.

[0895] Input: Meeting audio data

[0896] Output: Converted text data

[0897] Specific behavior:

[0898] The user opens Google Docs and selects "Voice Typing" from the Tools menu.

[0899] Once the meeting starts, click the "Speech to Text" button to start dictating. The speech recognition engine will capture the voice data and convert it into text in real time.

[0900] Step 2:

[0901] The device converts the recorded conversation into text in real time and temporarily stores this text data locally.

[0902] Input: Meeting audio data

[0903] Output: Locally saved text data

[0904] Specific behavior:

[0905] The speech recognition engine (Google's speech recognition API) converts the speech data into text information.

[0906] The converted text data is saved in a temporary directory in the terminal's local file system.

[0907] Step 3:

[0908] When the meeting or business negotiation ends, the terminal transmits the saved text data to the server.

[0909] Input: Locally stored text data

[0910] Output: Text data sent to the server

[0911] Specific behavior:

[0912] The user clicks the "Stop" button when the conference is over.

[0913] The device retrieves the text file from the temporary directory and uploads it to the server via an HTTP request.

[0914] Step 4:

[0915] The server receives the text data and queues it for analysis.

[0916] Input: Text data sent from the terminal

[0917] Output: Text data queued for analysis

[0918] Specific behavior:

[0919] The server stores the received data in a temporary storage directory.

[0920] The text data is registered as a task in the analysis queue.

[0921] Step 5:

[0922] The server uses a generative AI model (e.g., GPT-4) to analyze the queued text data, extract key points and content, and generate a summary.

[0923] Input: Text data queued for analysis

[0924] Output: Generated summary data

[0925] Specific behavior:

[0926] The server retrieves the text data from the queue.

[0927] A prompt sentence is input to the generative AI model and a summary result is received from the generative AI model.

[0928] Step 6:

[0929] The server adds user emotion information recognized by an emotion engine (e.g., IBM Watson's emotion analysis function) to the summary data.

[0930] Input: Generated summary data

[0931] Output: Summarized data with added sentiment information

[0932] Specific behavior:

[0933] The server sends the voice or text data to the sentiment analysis engine.

[0934] The received emotion information is incorporated into the summary data.

[0935] Step 7:

[0936] The server embeds the extracted main points and emotional information into a predetermined standard format and formats it.

[0937] Input: Summarized data with added sentiment information

[0938] Output: Minutes in a standard format

[0939] Specific behavior:

[0940] The server reads a template file with a fixed format and replaces the summary data and emotion information with the corresponding placeholders.

[0941] Save the formatted minutes as a temporary file.

[0942] Step 8:

[0943] The server uses an automation tool such as MAKE to set up and execute a workflow that sends the formatted minutes to the specified destination.

[0944] Input: Minutes in a standard format

[0945] Output: The minutes sent to the specified destination

[0946] Specific behavior:

[0947] The server sends an API request to an automation tool (e.g., MAKE) to trigger the workflow.

[0948] The workflow saves meeting minutes in a Google Doc and distributes them to relevant parties via email or messaging services.

[0949] Step 9:

[0950] Users can review the shared minutes and make corrections or additional comments as needed.

[0951] Input: Shared minutes

[0952] Output: Meeting minutes with corrections and additional comments

[0953] Specific behavior:

[0954] The user opens a Google Doc or an email they received and checks the minutes.

[0955] If any corrections or comments are required, please use the editing function of Google Docs to make additions or corrections.

[0956] (Application example 2)

[0957] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0958] Current meeting recording systems and work instruction recording systems simply convert audio data into text, without detailed analysis or summarization, including emotional information. As a result, the content of meetings and work instructions may not be accurately understood, leading to miscommunication and lack of awareness. Furthermore, methods for automatically sharing formatted data are limited, making efficient information sharing difficult. A system that solves these issues is needed.

[0959] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for converting voice data to text in real time, means for analyzing the converted text data and generating a summary, means for recognizing emotions from the voice data and reflecting the emotional information in the summary, means for converting the generated summary into a predetermined format, and means for automatically sending the generated formatted data to a specified destination. This enables accurate text conversion and analysis of voice data, addition of emotional information, and efficient information sharing.

[0960] "Audio data" refers to data that has been recorded and stored in digital form, such as voice or spoken words.

[0961] "Real-time text conversion means" refers to technology or equipment for instantly converting voice data into text form.

[0962] "Text data" refers to digital data expressed as character information.

[0963] "Analysis" refers to the process of analyzing and interpreting input data to understand its meaning and structure.

[0964] "Means for generating summaries" refers to technologies and devices that extract important information from large amounts of text data and summarize it concisely.

[0965] "Means for recognizing emotions and incorporating that emotional information into summaries" refers to technology and devices for identifying emotions from audio data and incorporating that emotional information into summaries.

[0966] The "means for converting into a predetermined format" refers to a technique or device for converting the generated summary into a predetermined format.

[0967] "Formatted data" refers to digital data that has been arranged in a particular format.

[0968] "Means for automatically sending to a specified destination" refers to technology or devices for automatically sending generated data to a specific recipient.

[0969] "Designated destination" refers to a recipient or location that has been pre-configured to receive data.

[0970] "Improved natural language processing technology" refers to technology for analyzing and processing natural language with improved accuracy and efficiency compared to conventional technology.

[0971] "Document Management System" means software or a platform for storing, managing, and accessing documents and data.

[0972] "Email or messaging service" means a service for sending and receiving messages or information over the Internet.

[0973] "Automation tool" refers to software that automatically performs designated tasks or processes.

[0974] To implement this invention, it is necessary to build a system that converts voice data into text in real time, analyzes the text data to generate a summary, converts it into a standard format after reflecting emotional information, and automatically sends it to a specified destination. This system aims to efficiently record work instructions in a factory and analyze the content of those instructions.

[0975] composition

[0976] The system components are as follows:

[0977] 1. Voice data collection devices: Microphones and voice-reading devices placed on the factory floor to record work instructions and conversations.

[0978] 2. Speech Recognition Software: Software that converts voice data into text in real time. Specifically, we use the Google Cloud Speech-to-Text API.

[0979] 3. Text analysis and summary generation software: Uses generative AI models to analyze converted text data and generate summaries. OpenAI's GPT-4 is a prime example.

[0980] 4. Emotion Recognition Software: Software to recognize emotions from audio data and incorporate that emotional information into summaries. Uses IBM Watson Tone Analyzer.

[0981] 5. Data conversion and sharing software: Software that converts the generated summaries into a standard format and automatically sends them to designated recipients. This includes automation tools such as Zapier.

[0982] Operation flow

[0983] 1. Voice reading

[0984] Users use a voice data capture device to record work instructions and conversations, and speech recognition software converts the recorded voice data into text in real time.

[0985] 2. Text Analysis

[0986] Text analysis and summary generation software analyzes the converted text data, extracts key points, and generates summaries using a generative AI model (e.g., GPT-4).

[0987] 3. Emotion recognition

[0988] Emotion recognition software (IBM Watson Tone Analyzer) is used to recognize emotions from audio data. The analyzed emotional information is reflected in the summary.

[0989] 4. Data Sharing

[0990] The generated data, including summaries and sentiment information, is converted into a predetermined format and sent to a designated destination via an automation tool (such as Zapier), allowing relevant parties to receive important information in real time.

[0991] Specific examples

[0992] If the factory manager instructs the robot to "run a test run of a new product on the next line," the system operates as follows:

[0993] An audio data collection device records the instructions.

[0994] Speech recognition software converts this audio data into text.

[0995] The text analysis and summary generation software summarizes it as "Test run of new product on the production line."

[0996] Emotion recognition software analyzes the "high level of attention" and "sense of security" of the person giving the instructions.

[0997] The generated summary and emotion information are converted into a predetermined format and sent to a specified destination.

[0998] Example prompt sentence:

[0999] "Please summarize the following text:

[1000] On the next line, I would like you to conduct a test run of a new product.

[1001] High level of focus and peace of mind

[1002] The above is a detailed embodiment for carrying out the invention. This system enables efficient and accurate recording and transmission of work instructions in a factory.

[1003] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1004] Step 1:

[1005] Voice reading

[1006] The user uses a voice data collection device to record work instructions and conversations.

[1007] Input: Audio data

[1008] Specific operation: Microphones and voice reading devices installed on the factory floor collect voices from factory managers and work instructors in real time.

[1009] Output: Real-time audio data

[1010] Step 2:

[1011] Real-time text transcription

[1012] The device converts the collected voice data into text in real time using voice recognition software (Google Cloud Speech-to-Text API).

[1013] Input: Real-time audio data

[1014] How it works: Speech recognition software analyzes speech waveforms and converts what is being said into text.

[1015] Output: Text data

[1016] Step 3:

[1017] Text data analysis

[1018] The device uses a generative AI model (OpenAI GPT-4) to analyze the converted text data, extract key points, and generate a summary.

[1019] Input: Text data

[1020] How it works: The generative AI model analyzes the input text data, uses prompts to identify key information, and summarizes it succinctly.

[1021] Output: Summary text

[1022] Step 4:

[1023] emotion recognition

[1024] The device uses emotion recognition software (IBM Watson Tone Analyzer) to recognize emotions from the analyzed voice data and reflects that emotional information in the summary.

[1025] Input: Text data

[1026] What it does: Emotion recognition software analyzes keywords and context within text data to identify the speaker's emotions (e.g., excitement, relief, anxiety, etc.).

[1027] Output: Summarized text with emotional information

[1028] Step 5:

[1029] Conversion to standard formats

[1030] The terminal converts the generated summary text into a predetermined format.

[1031] Input: Summarized text with sentiment information

[1032] What it does: The formatting function embeds summary and sentiment information into templates to generate standardized work instructions and meeting minutes.

[1033] Output: formatted data

[1034] Step 6:

[1035] Automatic Data Transmission and Sharing

[1036] The server uses an automation tool (such as Zapier) to automatically send the generated formatted data to the specified destination.

[1037] Input: Formatted data

[1038] What it does: The automation tool distributes data to a pre-specified list of recipients or sharers via email, messaging services, or document management systems.

[1039] Output: Shared data

[1040] Through these steps, a series of automated processes are implemented, from collecting voice data to analyzing it, recognizing emotions, converting formats, and sharing the information, enabling accurate and prompt recording and transmission of work instructions in factories.

[1041] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1042] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1043] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[1044] [Third embodiment]

[1045] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[1046] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[1047] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1048] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[1049] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1050] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1051] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1052] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1053] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1054] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1055] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1056] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[1057] This system provides a means to quickly record the contents of meetings and business negotiations and automatically share them as organized minutes. When a user holds a meeting or business negotiation, they first launch the voice-to-text tool. The voice data is converted into text data in real time. This converted text data is sent from the user's device to a server where it is analyzed. This analysis uses generative AI (e.g., GPT-4) to extract important points and content and generate a summary. The resulting summary data is then converted into a standard format. Finally, the formatted data is automatically sent to the specified destination.

[1058] A natural language description of the program's processing

[1059] Voice reading

[1060] User: When a meeting or business meeting begins, launch the voice-to-text tool and start recording.

[1061] On your device: A speech-to-text tool converts your conversation into text in real time, which is temporarily stored on your device.

[1062] Text Acquisition

[1063] Terminal: When the meeting or business negotiation ends, the saved text data is sent to the server.

[1064] Server: Receives the text data and begins analysis.

[1065] Text organization

[1066] Server: Using generative AI, the server analyzes the received text data and extracts key points and content, thereby summarizing the text.

[1067] Conversion to standard formats

[1068] Server: The extracted key points are embedded into a pre-defined format to complete the minutes template.

[1069] Automation workflow execution

[1070] Server: Using an automation tool such as MAKE, the generated minutes are automatically sent to the specified destination.

[1071] Specific examples

[1072] Examples of meeting content

[1073] Attendees: Yuko Sugimoto, Taro Tanaka, Ichiro Kimura

[1074] Subject: New product development project

[1075] Agenda: Development schedule, resource allocation, next meeting date

[1076] 1. The user launches the voice capture tool and begins recording the audio of the meeting.

[1077] 2. The device converts the conversation being recorded into text in real time and saves the text data.

[1078] 3. After the meeting ends, the device sends the saved text data to the server, and the server receives the text data.

[1079] 4. The server inputs the received text data into a generative AI (e.g., GPT-4) to summarize the meeting content. Specifically, it extracts important information related to the "development schedule," "resource allocation," and "next meeting date."

[1080] 5. The server embeds the extracted key points into a template to generate a standard formatted minutes. For example:

[1081] minutes

[1082] Meeting name: New product development project meeting

[1083] Attendees: Yuko Sugimoto, Taro Tanaka, Ichiro Kimura

[1084] Date and Time: [meeting date and time]

[1085] agenda:

[1086] 1. Development Schedule

[1087] [Summary]

[1088] 2. Resource Allocation

[1089] [Summary]

[1090] 3. Next meeting date

[1091] [Summary]

[1092] remarks:

[1093] There are no particular problems at the moment

[1094] 6. The server uses MAKE to execute a workflow that saves the generated minutes in a Google Doc or distributes them to all parties involved by email after the video conference ends.

[1095] 7. Users can see that the minutes are automatically shared immediately after the meeting ends and can make additional comments or corrections as needed.

[1096] This system automates a series of tasks, from recording meeting content to organizing and sharing it, greatly improving users' work efficiency.

[1097] The processing flow will be explained below.

[1098] Step 1:

[1099] When a user starts a meeting or business meeting, they launch Google Docs' voice reading tool and begin recording.

[1100] Step 2:

[1101] Your device will use Google Docs' speech-to-text tool to convert your conversation into text in real time.

[1102] Step 3:

[1103] The device temporarily stores the converted text data locally.

[1104] Step 4:

[1105] When the user finishes a meeting or business negotiation, the device automatically detects the end of the meeting and sends the saved text data to the server.

[1106] Step 5:

[1107] The server queues the received text data for analysis.

[1108] Step 6:

[1109] The server uses generative AI (e.g., GPT-4) to analyze the queued text data and extract key points and content.

[1110] Step 7:

[1111] The server summarizes the extracted data and summarizes the main points.

[1112] Step 8:

[1113] The server embeds the summary data into a standard format template and formats it.

[1114] Step 9:

[1115] The server uses an automation tool such as MAKE to set up a workflow in which the formatted minutes are sent to a specified destination.

[1116] Step 10:

[1117] The server executes the configured workflow and generates meeting minutes, which are then saved in a Google Doc or distributed to relevant parties via email or Slack.

[1118] Step 11:

[1119] The user reviews the shared minutes and makes corrections or additional comments as necessary.

[1120] Example 1

[1121] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1122] The traditional process of recording, organizing, and sharing the details of meetings and business negotiations is often manual, requiring time and effort, and there are problems with content being easily overlooked. Furthermore, the process of manually extracting key points and creating minutes is inaccurate and inefficient, often causing delays in business operations.

[1123] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1124] In this invention, the server includes a means for activating the voice reading tool and starting recording, a means for transmitting text data stored on the terminal to the server, a means for extracting important points and content from the text data using a generative AI model, and a means for transmitting data to a specified destination using an automation tool. This enables real-time conversion of voice data, automatic extraction of important content, standardized formatting of minutes, and automatic distribution of data.

[1125] "Voice data" refers to voice information collected during meetings and business negotiations.

[1126] A "real-time text conversion means" is a mechanism that collects audio data and simultaneously converts it into text form.

[1127] "Text data" refers to character information converted from audio data.

[1128] The "means for analyzing and generating summaries" is a mechanism for analyzing the converted text data, extracting important points and content, and summarizing them in a concise format.

[1129] The "means for converting into a predetermined format" is a mechanism for organizing the generated summary according to a particular template.

[1130] The "means for automatically sending to a specified destination" is a mechanism for automatically sending the generated standard format data to a pre-set recipient.

[1131] A "voice reading tool" is an application or device that collects audio from meetings or business negotiations and converts it into text data in real time.

[1132] A "terminal" is an electronic device used by a user, such as a computer, smartphone, or tablet.

[1133] A "server" is a computer system for performing tasks such as analyzing text data, generating summaries, formatting, and transmitting data.

[1134] A "generative AI model" is an artificial intelligence technology that uses natural language processing to extract important points and content from text data and generate summaries.

[1135] An "automation tool" is software or a platform for automating a series of tasks, such as sending and storing data.

[1136] A "designated destination" is a recipient of automatically delivered data, such as a specific email address, a messaging service account, or a cloud storage folder.

[1137] This invention is a system for quickly recording the contents of meetings and business negotiations and automatically sharing them as organized minutes. The main components include a user device (computer, smartphone, tablet, etc.), a voice reading tool, a server, a generative AI model, and an automation tool.

[1138] System configuration details

[1139] 1. Voice reading

[1140] User: When the meeting or business meeting begins, launch the voice reading tool on the device and begin recording.

[1141] On your device: A voice-to-text tool (such as speech recognition software) is used to convert your conversation into text in real time using speech recognition technology such as the Google Cloud Speech-to-Text API. This converted text is temporarily stored on your device's local storage.

[1142] 2. Sending and analyzing text data

[1143] Terminal: When the meeting or business meeting ends, the saved text data is sent to the server using the HTTPS protocol.

[1144] Server: Receives the text data and inputs it into a generative AI model (e.g., GPT-4) to begin analysis. The generative AI model uses natural language processing techniques to extract key points and content from the text data.

[1145] 3. Summarization and Formatting

[1146] Server: Summarizes the meeting content based on key points extracted by the generative AI model. The server then embeds the summary into a pre-defined minutes template, which includes information such as the meeting name, attendees, agenda, and key points.

[1147] 4. Automatic data distribution

[1148] Server: Use automation tools such as MAKE or Zapier to automatically send the generated minutes to a specified destination (e.g., Google Docs, email, messaging service).

[1149] Specific use cases

[1150] For example, in a meeting for a new product development project, the following steps are taken:

[1151] Attendees: Ichiro Tanaka, Jiro Suzuki, Saburo Sato

[1152] Subject: New product development project

[1153] Agenda: Product specification, development schedule, debugging plan

[1154] procedure

[1155] 1. The user launches the voice capture tool and begins recording the audio of the meeting.

[1156] 2. The device converts the conversation being recorded into text in real time and saves the text data.

[1157] 3. After the meeting ends, the device sends the saved text data to the server, and the server receives the text data.

[1158] 4. The server inputs the received text data into a generative AI model (e.g., GPT-4) to summarize the meeting content. Specifically, it extracts important information related to "product specification determination," "development schedule," and "debug plan."

[1159] 5. The server embeds the extracted key points into a template to generate a standard formatted minutes. For example:

[1160] markdown

[1161] minutes

[1162] Meeting name: New product development project meeting

[1163] Attendees: Ichiro Tanaka, Jiro Suzuki, Saburo Sato

[1164] Date and Time: [meeting date and time]

[1165] agenda:

[1166] 1. Deciding on product specifications

[1167] [Summary]

[1168] 2. Development Schedule

[1169] [Summary]

[1170] 3. Debugging Plan

[1171] [Summary]

[1172] remarks:

[1173] There are no particular problems at the moment

[1174] 6. The server uses MAKE to execute a workflow that saves the generated minutes in a Google Doc or distributes them to all parties involved by email after the meeting ends.

[1175] 7. Users can see that the minutes are automatically shared immediately after the meeting ends and can make additional comments or corrections as needed.

[1176] This system automates a series of tasks, from recording meeting content to organizing and sharing it, greatly improving users' work efficiency.

[1177] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1178] Step 1: Launch the voice reader and start recording

[1179] User: When a meeting or business meeting begins, the user launches the voice-reading tool on their device and starts recording. Specifically, the user clicks the tool icon and presses the record button.

[1180] Enter: Start of a meeting or business meeting.

[1181] Output: Recording begins.

[1182] Step 2: Real-time text conversion of audio data

[1183] On your device: The voice reader converts your conversation into text data in real time. It uses the Google Cloud Speech-to-Text API for speech recognition. The converted text data is temporarily stored in your device's local storage.

[1184] Input: Audio data from meetings and business negotiations.

[1185] Output: Text data generated in real time.

[1186] Specific operation: The speech recognition API converts the voice data into text, and the generated text is saved in the specified directory on the device.

[1187] Step 3: Send text data to the server

[1188] Terminal: When the meeting or business meeting ends, the saved text data is sent to the server. The data is transferred securely using the HTTPS protocol.

[1189] Input: Text data after the meeting ends.

[1190] Output: Text data stored on the server.

[1191] Specific operation: Sends a POST request to the API endpoint for sending text data, and receives a success message after sending.

[1192] Step 4: Parsing the incoming text data

[1193] Server: The received text data is input into the generative AI model (GPT-4) and analysis begins. The model extracts important points and keywords from the input text data.

[1194] Input: Received text data.

[1195] Output: Extracted key points and keywords.

[1196] How it works: The GPT-4 model breaks down text data into tokens, analyzes important context, and extracts key points.

[1197] Step 5: Generate a summary

[1198] Server: Summarizes the meeting content based on key points extracted by the generative AI model. It uses a summary generation algorithm to summarize the key points in a concise format.

[1199] Input: Extracted key points and keywords.

[1200] Output: The summarized text.

[1201] Specific operation: Reconstructs the main points into short sentences and performs text processing to generate a summary.

[1202] Step 6: Embedding in the meeting minutes template

[1203] Server: The summary is embedded into a pre-defined minutes template, which automatically includes information such as the meeting name, attendees, agenda, and key points.

[1204] Input: The summarized text.

[1205] Output: A completed meeting minutes template.

[1206] What it does: Places the summary text into the template placeholder and generates the final minutes document.

[1207] Step 7: Automatic data distribution

[1208] Server: Using automation tools such as MAKE or Zapier, the generated minutes are automatically sent to the specified recipients. The data is stored in a Google Doc or distributed via email or messaging services.

[1209] Input: Completed meeting minutes template.

[1210] Output: The minutes sent to the specified destination.

[1211] Specific behavior: Triggers the automation tool workflow and executes the data sending process, saving it to Google Docs and sending it by email.

[1212] (Application example 1)

[1213] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1214] In industrial sites, there is a need to reduce the time and effort required to record work instructions and generate reports, thereby achieving efficient management. There is also a need to improve the accuracy and speed of work by recording work instructions in real time and instantly sharing necessary information. Conventional methods require the tedious task of manually recording voice instructions and compiling them into reports, making automating this process a challenge.

[1215] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1216] In this invention, the server includes means for converting voice data into text in real time, means for analyzing the converted text data and generating a summary, and means for converting the generated summary into a predetermined format, thereby enabling industrial equipment to record voice instructions in real time and automatically generate work reports.

[1217] definition statement

[1218] "Voice data" refers to audio information such as conversations and instructions recorded in digital format.

[1219] "Text data" is character information generated by analyzing voice data.

[1220] "Real time" means that processing occurs at the exact moment an event occurs.

[1221] A "summary" is a shortened version of a large amount of information, extracting only the important points.

[1222] A "predetermined format" is a predetermined form or template.

[1223] "Destination" is the address or location of the recipient to which the generated data is sent.

[1224] "Industrial equipment" refers to all equipment such as machines and robots used in production sites such as factories.

[1225] A "job report" is a document that records the details of the work performed by industrial equipment.

[1226] "Analysis" is the process of examining data in detail and extracting hidden information and meaning.

[1227] "Generative AI" is an artificial intelligence system that mimics human intelligence and automatically generates and analyzes data.

[1228] "Natural language processing technology" is a technology that allows computers to process and understand human language.

[1229] "Email" is a means of sending and receiving messages electronically over the Internet.

[1230] "Messaging service" refers to a service that allows users to send and receive text messages over the Internet.

[1231] patent specification

[1232] An embodiment of the present invention relates to a system for issuing work instructions and reporting in meetings, business negotiations, and industrial sites. This system has the functions of converting voice data into text in real time, extracting and summarizing important information, converting the summary into a predetermined format, and automatically sending it to a specified destination.

[1233] Hardware and software used

[1234] Hardware: The system includes terminals with audio input devices (e.g., microphones), industrial equipment, and communication modules.

[1235] Software: We use the Google Speech-to-Text API for speech-to-text conversion, GPT-4 as the generative AI model, and MAKE as an automation tool for automatic data transmission and management.

[1236] Program processing

[1237] A user launches a voice-to-text tool and records a meeting, business negotiation, or work instructions at an industrial site. The device converts the recorded voice into text in real time and temporarily stores it. After the meeting or work is completed, the device sends the text data to a server. The server inputs the received text data into a generative AI (e.g., GPT-4) to extract key points. The resulting summary data is converted into a specified format and automatically generated as a work report or meeting minutes. Finally, the server uses MAKE to send the generated report or meeting minutes to the specified destination.

[1238] Specific examples

[1239] Examples of meeting content

[1240] Attendees: Mr. A, Mr. B, Mr. C

[1241] Subject: New product development project

[1242] Agenda: Development schedule, resource allocation, next meeting date

[1243] Example prompt sentence:

[1244] Instructions:

[1245] 1. Assembly Instructions

[1246] Connect part A to part B

[1247] Check the connection

[1248] 2. How to use the tools

[1249] Use a screwdriver and wrench

[1250] 3. Completion checkpoints

[1251] Operation check

[1252] In this example, project meetings can be recorded in real time, summarized, formatted, and finally shared automatically with all stakeholders. A similar process can be used for work instructions in industrial sites to automatically generate work reports that can be instantly sent to managers.

[1253] This system automates a series of tasks, from recording meetings, business negotiations, and work at industrial sites to generating and sharing reports, greatly improving work efficiency.

[1254] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1255] Processing flow

[1256] Program processing steps

[1257] Step 1:

[1258] A user launches a voice-to-speech tool to record a meeting, a business meeting, or work instructions on an industrial site.

[1259] Input: Audio data for meetings and work

[1260] Output: The voice reading tool retrieves the voice data.

[1261] Step 2:

[1262] The device converts the recorded audio into text in real time and stores it temporarily.

[1263] Input: Audio data

[1264] Data processing: Converting voice data into text data using the Google Speech-to-Text API

[1265] Output: Text data is saved to the device

[1266] Step 3:

[1267] After the meeting ends or when work is completed, the device sends the saved text data to the server.

[1268] Input: Text data

[1269] Data calculation: Upload text data to the server via the network

[1270] Output: Text data is sent to the server

[1271] Step 4:

[1272] The server inputs the received text data into a generative AI (e.g., GPT-4) to extract key points.

[1273] Input: Text data

[1274] Data Computing: Analyzing text data using GPT-4 to extract key points

[1275] Output: Extracted key points (summary data)

[1276] Step 5:

[1277] The server converts the generated summary data into a predetermined format.

[1278] Input: Summary data

[1279] Data processing: Embedding summary data into a defined template

[1280] Output: Formatted minutes and working papers

[1281] Step 6:

[1282] The server uses MAKE to send the generated formatted data to the specified destination.

[1283] Input: Formatted minutes and work reports

[1284] Data Calculation: Use MAKE to send data to email or a document management system

[1285] Output: Meeting minutes and work reports are sent to the specified destination.

[1286] Specific examples of operation

[1287] Specific actions of the meeting content

[1288] Meeting scenario: A user launches a voice-to-text tool to record the meeting audio.

[1289] Speech-to-text: Recorded audio is converted to text in real time using the Google Speech-to-Text API.

[1290] Data transmission: After the meeting ends, the device uploads the text data to the server.

[1291] Gist extraction: The server uses GPT-4 to analyze the text data and summarize the main topics.

[1292] Format conversion: Embed the summarized content into a predefined format.

[1293] Data transmission: Using MAKE, formatted minutes are emailed to all parties involved.

[1294] These processing steps enable efficient and automated work recording and reporting in meetings, business negotiations, and industrial sites, significantly improving business efficiency.

[1295] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1296] This invention is a system that quickly and efficiently records the contents of meetings and business negotiations and automatically shares them as organized minutes. It also has the ability to recognize user emotions and reflect that information in the minutes. This system includes a means for converting voice data into text in real time, analyzing it to generate summaries, and converting them into a standard format for sharing. It also incorporates an emotion engine that recognizes user emotions from voice data.

[1297] A natural language description of the program's processing

[1298] Voice reading

[1299] User: At the start of a meeting or business meeting, launch the Google Docs voice-to-text tool and start recording.

[1300] On your device: The speech-to-text tool converts your conversation into text in real time and temporarily stores this text locally.

[1301] Text Acquisition

[1302] Terminal: When the meeting or business negotiation ends, the saved text data is sent to the server.

[1303] Server: Receives text data and queues it for analysis.

[1304] Text organization and summary generation

[1305] Server: Uses generative AI (e.g., GPT-4) to analyze queued text data, extract key points and content, and generate summaries.

[1306] Server: Adds user emotion information recognized by the emotion engine to the summary data.

[1307] Conversion to standard formats

[1308] Server: Embeds and formats the extracted and summarized data and user emotional information into a predetermined standard format.

[1309] Automated workflow execution and sharing

[1310] Server: Uses an automation tool such as MAKE to set up and execute a workflow that sends formatted minutes to the specified destination.

[1311] Server: Save the generated minutes in a Google Doc or distribute them to relevant parties via email or Slack.

[1312] User: Review the shared minutes and make corrections or additional comments as needed.

[1313] Specific examples

[1314] Examples of meeting content

[1315] Attendees: Yuko Sugimoto, Taro Tanaka, Ichiro Kimura

[1316] Subject: New product development project

[1317] Agenda: Development schedule, resource allocation, next meeting date

[1318] 1. The user launches the voice capture tool and begins recording the audio of the meeting.

[1319] 2. The device converts the conversation being recorded into text in real time and saves the text data.

[1320] 3. After the meeting ends, the device sends the saved text data to the server, and the server receives the text data.

[1321] 4. The server inputs the received text data into a generative AI (e.g., GPT-4) to summarize the meeting content. Specifically, it extracts important information related to the "development schedule," "resource allocation," and "next meeting date."

[1322] 5. The server uses an emotion engine to analyze the user's emotions during the conversation. For example, it recognizes the speaker's emotions (excitement, relief, anxiety, etc.) separately from the content of the speech.

[1323] 6. The server embeds the extracted key points and sentiment information into a template to generate a standardized minutes of the meeting. For example:

[1324] minutes

[1325] Meeting name: New product development project meeting

[1326] Attendees: Yuko Sugimoto, Taro Tanaka, Ichiro Kimura

[1327] Date and Time: [meeting date and time]

[1328] agenda:

[1329] 1. Development Schedule

[1330] [Summary]

[1331] [Emotional information]

[1332] 2. Resource Allocation

[1333] [Summary]

[1334] [Emotional information]

[1335] 3. Next meeting date

[1336] [Summary]

[1337] [Emotional information]

[1338] remarks:

[1339] There are no particular problems at the moment

[1340] 7. The server uses MAKE to execute a workflow that saves the generated minutes in a Google Doc or distributes them to all parties involved by email after the meeting ends.

[1341] 8. Users can see that the minutes are automatically shared immediately after the meeting ends and can make additional comments or corrections as needed.

[1342] This system automates a series of tasks, from recording and organizing meeting content, to analyzing emotions and sharing, greatly improving users' work efficiency.

[1343] The processing flow will be explained below.

[1344] Step 1:

[1345] When a user starts a meeting or business meeting, they launch Google Docs' voice reading tool and begin recording.

[1346] Step 2:

[1347] Your device will use Google Docs' speech-to-text tool to convert your conversation into text in real time, which will then be temporarily stored locally.

[1348] Step 3:

[1349] The terminal transmits the saved text data together with the audio file to the server.

[1350] Step 4:

[1351] The server receives the text data and audio files and queues them for analysis.

[1352] Step 5:

[1353] The server uses generative AI (e.g., GPT-4) to analyze the queued text data and extract key points and content.

[1354] Step 6:

[1355] After the server extracts the main points, it uses an emotion engine to recognize the speaker's emotions from the audio file and analyzes the emotion information.

[1356] Step 7:

[1357] The server adds the user's emotion information recognized by the emotion engine to the generated summary data.

[1358] Step 8:

[1359] The server embeds the summary data and emotional information into a predetermined standard format (minutes template) and formats the formatted data.

[1360] Step 9:

[1361] The server uses MAKE's automation tools to set up a workflow that sends formatted minutes to a specified destination.

[1362] Step 10:

[1363] The server executes the configured workflow and generates meeting minutes, which are then saved in a Google Doc or distributed to relevant parties via email or Slack.

[1364] Step 11:

[1365] The user reviews the shared minutes and makes corrections or additional comments as necessary.

[1366] Example 2

[1367] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1368] Conventional meeting recording systems not only convert audio data into text, but also require manual data organization, summary generation, and emotional analysis, resulting in cumbersome post-meeting tasks that require time and effort. Furthermore, they are unable to reflect the emotional information of meeting participants, limiting the content of meeting minutes. This leads to issues such as a decline in the quality of meeting minutes and a decrease in communication efficiency.

[1369] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1370] In this invention, the server includes means for converting voice data into text in real time, means for analyzing the converted text data and generating a summary, means for adding user emotion information recognized by an emotion engine to the generated summary data, and means for automatically sending the generated formatted data to a specified destination. This automates data organization and summary generation after a meeting, and enables the generation and rapid sharing of detailed minutes including emotion information.

[1371] "Voice data" is data that formalizes voice signals collected during meetings, business negotiations, etc.

[1372] "Means for converting voice data into text in real time" refers to a technology or device that instantly analyzes voice signals and converts them into text information.

[1373] "Text data" is information obtained by converting voice data into text format.

[1374] "Analysis" is the process of examining data in detail and extracting specific patterns and important information.

[1375] A "summary" is a short summary of the main points extracted from a long piece of text or a large amount of information.

[1376] An "emotion engine" is a technology or software that analyzes voice or text data and recognizes the speaker's emotions.

[1377] "Format" refers to arranging data into a specific form or structure.

[1378] "Specified destination" means a predetermined recipient or location to which generated data is sent.

[1379] "Generative AI" is an artificial intelligence technology that learns patterns from various input data and generates new outputs.

[1380] An "automation tool" is software or a service that performs a specific process or task without human intervention.

[1381] MODE FOR CARRYING OUT THE INVENTION

[1382] The present invention is a system that quickly and efficiently records the contents of meetings and business negotiations and automatically shares them as organized minutes. It also has the ability to recognize user emotions and reflect that information in the minutes. The system uses a speech recognition tool to convert voice data into text in real time, a generative AI model to analyze and summarize the text data, an emotion engine for emotion analysis, and an automated tool for formatting and sharing. A specific embodiment of this system is described below.

[1383] Voice reading

[1384] When a user starts a meeting or business meeting, they activate a voice recognition tool (e.g., Google Docs' "voice input" function) and start recording. The voice recognition tool converts the spoken audio into text in real time and temporarily stores this text data locally.

[1385] Text Acquisition

[1386] When a meeting or business meeting ends, the device sends the text data stored locally to the server, which receives the text data and queues it for analysis. This queue manages the data waiting to be processed, enabling efficient analysis of the text data.

[1387] Text organization and summary generation

[1388] The server uses a generative AI model (e.g., GPT-4) to analyze the queued text data, extract key points and content, and generate a summary. Furthermore, it adds user sentiment information recognized by an emotion engine (e.g., IBM Watson's sentiment analysis function) to the generated summary data.

[1389] Conversion to standard formats

[1390] The server embeds the main points and emotional information into a predetermined format and formats it into minutes. This formatted data is stored as a temporary file.

[1391] Automated workflow execution and sharing

[1392] The server uses an automation tool such as MAKE to set up and execute a workflow that sends the minutes converted into a standard format to a specified destination. This workflow is configured to save the generated minutes in a specific document management system (e.g., Google Docs) and distribute them to relevant parties via email or a messaging service (e.g., Slack).

[1393] Specific examples

[1394] Examples of meeting content:

[1395] Attendees: A, B, C

[1396] Topic: New project progress

[1397] Agenda: Schedule confirmation, resource allocation, next meeting date

[1398] At the start of a meeting, a user activates Google Docs' voice input and begins recording the conversation. The recorded audio is converted into text in real time and saved to a local file. After the meeting ends, the text data is sent from the device to a server and queued for analysis. The server uses a generative AI model and emotion engine to analyze the text data, extract key points and emotional information, and generate meeting minutes.

[1399] The minutes are then formatted and saved in a Google Doc or distributed to relevant parties via email or Slack using MAKE's automation tools. Users can review the shared minutes and make additional comments or corrections as needed.

[1400] Example prompt sentence:

[1401] "I want to extract the key points from a meeting about the progress of a new project and generate a summary. I also want to analyze the emotions of the attendees and reflect them in the minutes."

[1402] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1403] Step 1:

[1404] When starting a meeting or business meeting, a user activates a voice recognition tool (e.g., Google Docs' "voice input" function) and starts recording.

[1405] Input: Meeting audio data

[1406] Output: Converted text data

[1407] Specific behavior:

[1408] The user opens Google Docs and selects "Voice Typing" from the Tools menu.

[1409] Once the meeting starts, click the "Speech to Text" button to start dictating. The speech recognition engine will capture the voice data and convert it into text in real time.

[1410] Step 2:

[1411] The device converts the recorded conversation into text in real time and temporarily stores this text data locally.

[1412] Input: Meeting audio data

[1413] Output: Locally saved text data

[1414] Specific behavior:

[1415] The speech recognition engine (Google's speech recognition API) converts the speech data into text information.

[1416] The converted text data is saved in a temporary directory in the terminal's local file system.

[1417] Step 3:

[1418] When the meeting or business negotiation ends, the terminal transmits the saved text data to the server.

[1419] Input: Locally stored text data

[1420] Output: Text data sent to the server

[1421] Specific behavior:

[1422] The user clicks the "Stop" button when the conference is over.

[1423] The device retrieves the text file from the temporary directory and uploads it to the server via an HTTP request.

[1424] Step 4:

[1425] The server receives the text data and queues it for analysis.

[1426] Input: Text data sent from the terminal

[1427] Output: Text data queued for analysis

[1428] Specific behavior:

[1429] The server stores the received data in a temporary storage directory.

[1430] The text data is registered as a task in the analysis queue.

[1431] Step 5:

[1432] The server uses a generative AI model (e.g., GPT-4) to analyze the queued text data, extract key points and content, and generate a summary.

[1433] Input: Text data queued for analysis

[1434] Output: Generated summary data

[1435] Specific behavior:

[1436] The server retrieves the text data from the queue.

[1437] A prompt sentence is input to the generative AI model and a summary result is received from the generative AI model.

[1438] Step 6:

[1439] The server adds user emotion information recognized by an emotion engine (e.g., IBM Watson's emotion analysis function) to the summary data.

[1440] Input: Generated summary data

[1441] Output: Summarized data with added sentiment information

[1442] Specific behavior:

[1443] The server sends the voice or text data to the sentiment analysis engine.

[1444] The received emotion information is incorporated into the summary data.

[1445] Step 7:

[1446] The server embeds the extracted main points and emotional information into a predetermined standard format and formats it.

[1447] Input: Summarized data with added sentiment information

[1448] Output: Minutes in a standard format

[1449] Specific behavior:

[1450] The server reads a template file with a fixed format and replaces the summary data and emotion information with the corresponding placeholders.

[1451] Save the formatted minutes as a temporary file.

[1452] Step 8:

[1453] The server uses an automation tool such as MAKE to set up and execute a workflow that sends the formatted minutes to the specified destination.

[1454] Input: Minutes in a standard format

[1455] Output: The minutes sent to the specified destination

[1456] Specific behavior:

[1457] The server sends an API request to an automation tool (e.g., MAKE) to trigger the workflow.

[1458] The workflow saves meeting minutes in a Google Doc and distributes them to relevant parties via email or messaging services.

[1459] Step 9:

[1460] Users can review the shared minutes and make corrections or additional comments as needed.

[1461] Input: Shared minutes

[1462] Output: Meeting minutes with corrections and additional comments

[1463] Specific behavior:

[1464] The user opens a Google Doc or an email they received and checks the minutes.

[1465] If any corrections or comments are required, please use the editing function of Google Docs to make additions or corrections.

[1466] (Application example 2)

[1467] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1468] Current meeting recording systems and work instruction recording systems simply convert audio data into text, without detailed analysis or summarization, including emotional information. As a result, the content of meetings and work instructions may not be accurately understood, leading to miscommunication and lack of awareness. Furthermore, methods for automatically sharing formatted data are limited, making efficient information sharing difficult. A system that solves these issues is needed.

[1469] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for converting voice data to text in real time, means for analyzing the converted text data and generating a summary, means for recognizing emotions from the voice data and reflecting the emotional information in the summary, means for converting the generated summary into a predetermined format, and means for automatically sending the generated formatted data to a specified destination. This enables accurate text conversion and analysis of voice data, addition of emotional information, and efficient information sharing.

[1470] "Audio data" refers to data that has been recorded and stored in digital form, such as voice or spoken words.

[1471] "Real-time text conversion means" refers to technology or equipment for instantly converting voice data into text form.

[1472] "Text data" refers to digital data expressed as character information.

[1473] "Analysis" refers to the process of analyzing and interpreting input data to understand its meaning and structure.

[1474] "Means for generating summaries" refers to technologies and devices that extract important information from large amounts of text data and summarize it concisely.

[1475] "Means for recognizing emotions and incorporating that emotional information into summaries" refers to technology and devices for identifying emotions from audio data and incorporating that emotional information into summaries.

[1476] The "means for converting into a predetermined format" refers to a technique or device for converting the generated summary into a predetermined format.

[1477] "Formatted data" refers to digital data that has been arranged in a particular format.

[1478] "Means for automatically sending to a specified destination" refers to technology or devices for automatically sending generated data to a specific recipient.

[1479] "Designated destination" refers to a recipient or location that has been pre-configured to receive data.

[1480] "Improved natural language processing technology" refers to technology for analyzing and processing natural language with improved accuracy and efficiency compared to conventional technology.

[1481] "Document Management System" means software or a platform for storing, managing, and accessing documents and data.

[1482] "Email or messaging service" means a service for sending and receiving messages or information over the Internet.

[1483] "Automation tool" refers to software that automatically performs designated tasks or processes.

[1484] To implement this invention, it is necessary to build a system that converts voice data into text in real time, analyzes the text data to generate a summary, converts it into a standard format after reflecting emotional information, and automatically sends it to a specified destination. This system aims to efficiently record work instructions in a factory and analyze the content of those instructions.

[1485] composition

[1486] The system components are as follows:

[1487] 1. Voice data collection devices: Microphones and voice-reading devices placed on the factory floor to record work instructions and conversations.

[1488] 2. Speech Recognition Software: Software that converts voice data into text in real time. Specifically, we use the Google Cloud Speech-to-Text API.

[1489] 3. Text analysis and summary generation software: Uses generative AI models to analyze converted text data and generate summaries. OpenAI's GPT-4 is a prime example.

[1490] 4. Emotion Recognition Software: Software to recognize emotions from audio data and incorporate that emotional information into summaries. Uses IBM Watson Tone Analyzer.

[1491] 5. Data conversion and sharing software: Software that converts the generated summaries into a standard format and automatically sends them to designated recipients. This includes automation tools such as Zapier.

[1492] Operation flow

[1493] 1. Voice reading

[1494] Users use a voice data capture device to record work instructions and conversations, and speech recognition software converts the recorded voice data into text in real time.

[1495] 2. Text Analysis

[1496] Text analysis and summary generation software analyzes the converted text data, extracts key points, and generates summaries using a generative AI model (e.g., GPT-4).

[1497] 3. Emotion recognition

[1498] Emotion recognition software (IBM Watson Tone Analyzer) is used to recognize emotions from audio data. The analyzed emotional information is reflected in the summary.

[1499] 4. Data Sharing

[1500] The generated data, including summaries and sentiment information, is converted into a predetermined format and sent to a designated destination via an automation tool (such as Zapier), allowing relevant parties to receive important information in real time.

[1501] Specific examples

[1502] If the factory manager instructs the robot to "run a test run of a new product on the next line," the system operates as follows:

[1503] An audio data collection device records the instructions.

[1504] Speech recognition software converts this audio data into text.

[1505] The text analysis and summary generation software summarizes it as "Test run of new product on the production line."

[1506] Emotion recognition software analyzes the "high level of attention" and "sense of security" of the person giving the instructions.

[1507] The generated summary and emotion information are converted into a predetermined format and sent to a specified destination.

[1508] Example prompt sentence:

[1509] "Please summarize the following text:

[1510] On the next line, I would like you to conduct a test run of a new product.

[1511] High level of focus and peace of mind

[1512] The above is a detailed embodiment for carrying out the invention. This system enables efficient and accurate recording and transmission of work instructions in a factory.

[1513] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1514] Step 1:

[1515] Voice reading

[1516] The user uses a voice data collection device to record work instructions and conversations.

[1517] Input: Audio data

[1518] Specific operation: Microphones and voice reading devices installed on the factory floor collect voices from factory managers and work instructors in real time.

[1519] Output: Real-time audio data

[1520] Step 2:

[1521] Real-time text transcription

[1522] The device converts the collected voice data into text in real time using voice recognition software (Google Cloud Speech-to-Text API).

[1523] Input: Real-time audio data

[1524] How it works: Speech recognition software analyzes speech waveforms and converts what is being said into text.

[1525] Output: Text data

[1526] Step 3:

[1527] Text data analysis

[1528] The device uses a generative AI model (OpenAI GPT-4) to analyze the converted text data, extract key points, and generate a summary.

[1529] Input: Text data

[1530] How it works: The generative AI model analyzes the input text data, uses prompts to identify key information, and summarizes it succinctly.

[1531] Output: Summary text

[1532] Step 4:

[1533] emotion recognition

[1534] The device uses emotion recognition software (IBM Watson Tone Analyzer) to recognize emotions from the analyzed voice data and reflects that emotional information in the summary.

[1535] Input: Text data

[1536] What it does: Emotion recognition software analyzes keywords and context within text data to identify the speaker's emotions (e.g., excitement, relief, anxiety, etc.).

[1537] Output: Summarized text with emotional information

[1538] Step 5:

[1539] Conversion to standard formats

[1540] The terminal converts the generated summary text into a predetermined format.

[1541] Input: Summarized text with sentiment information

[1542] What it does: The formatting function embeds summary and sentiment information into templates to generate standardized work instructions and meeting minutes.

[1543] Output: formatted data

[1544] Step 6:

[1545] Automatic Data Transmission and Sharing

[1546] The server uses an automation tool (such as Zapier) to automatically send the generated formatted data to the specified destination.

[1547] Input: Formatted data

[1548] What it does: The automation tool distributes data to a pre-specified list of recipients or sharers via email, messaging services, or document management systems.

[1549] Output: Shared data

[1550] Through these steps, a series of automated processes are implemented, from collecting voice data to analyzing it, recognizing emotions, converting formats, and sharing the information, enabling accurate and prompt recording and transmission of work instructions in factories.

[1551] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1552] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1553] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1554] [Fourth embodiment]

[1555] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1556] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1557] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1558] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1559] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1560] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1561] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1562] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1563] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1564] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1565] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1566] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1567] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1568] This system provides a means to quickly record the contents of meetings and business negotiations and automatically share them as organized minutes. When a user holds a meeting or business negotiation, they first launch the voice-to-text tool. The voice data is converted into text data in real time. This converted text data is sent from the user's device to a server where it is analyzed. This analysis uses generative AI (e.g., GPT-4) to extract important points and content and generate a summary. The resulting summary data is then converted into a standard format. Finally, the formatted data is automatically sent to the specified destination.

[1569] A natural language description of the program's processing

[1570] Voice reading

[1571] User: When a meeting or business meeting begins, launch the voice-to-text tool and start recording.

[1572] On your device: A speech-to-text tool converts your conversation into text in real time, which is temporarily stored on your device.

[1573] Text Acquisition

[1574] Terminal: When the meeting or business negotiation ends, the saved text data is sent to the server.

[1575] Server: Receives the text data and begins analysis.

[1576] Text organization

[1577] Server: Using generative AI, the server analyzes the received text data and extracts key points and content, thereby summarizing the text.

[1578] Conversion to standard formats

[1579] Server: The extracted key points are embedded into a pre-defined format to complete the minutes template.

[1580] Automation workflow execution

[1581] Server: Using an automation tool such as MAKE, the generated minutes are automatically sent to the specified destination.

[1582] Specific examples

[1583] Examples of meeting content

[1584] Attendees: Yuko Sugimoto, Taro Tanaka, Ichiro Kimura

[1585] Subject: New product development project

[1586] Agenda: Development schedule, resource allocation, next meeting date

[1587] 1. The user launches the voice capture tool and begins recording the audio of the meeting.

[1588] 2. The device converts the conversation being recorded into text in real time and saves the text data.

[1589] 3. After the meeting ends, the device sends the saved text data to the server, and the server receives the text data.

[1590] 4. The server inputs the received text data into a generative AI (e.g., GPT-4) to summarize the meeting content. Specifically, it extracts important information related to the "development schedule," "resource allocation," and "next meeting date."

[1591] 5. The server embeds the extracted key points into a template to generate a standard formatted minutes. For example:

[1592] minutes

[1593] Meeting name: New product development project meeting

[1594] Attendees: Yuko Sugimoto, Taro Tanaka, Ichiro Kimura

[1595] Date and Time: [meeting date and time]

[1596] agenda:

[1597] 1. Development Schedule

[1598] [Summary]

[1599] 2. Resource Allocation

[1600] [Summary]

[1601] 3. Next meeting date

[1602] [Summary]

[1603] remarks:

[1604] There are no particular problems at the moment

[1605] 6. The server uses MAKE to execute a workflow that saves the generated minutes in a Google Doc or distributes them to all parties involved by email after the video conference ends.

[1606] 7. Users can see that the minutes are automatically shared immediately after the meeting ends and can make additional comments or corrections as needed.

[1607] This system automates a series of tasks, from recording meeting content to organizing and sharing it, greatly improving users' work efficiency.

[1608] The processing flow will be explained below.

[1609] Step 1:

[1610] When a user starts a meeting or business meeting, they launch Google Docs' voice reading tool and begin recording.

[1611] Step 2:

[1612] Your device will use Google Docs' speech-to-text tool to convert your conversation into text in real time.

[1613] Step 3:

[1614] The device temporarily stores the converted text data locally.

[1615] Step 4:

[1616] When the user finishes a meeting or business negotiation, the device automatically detects the end of the meeting and sends the saved text data to the server.

[1617] Step 5:

[1618] The server queues the received text data for analysis.

[1619] Step 6:

[1620] The server uses generative AI (e.g., GPT-4) to analyze the queued text data and extract key points and content.

[1621] Step 7:

[1622] The server summarizes the extracted data and summarizes the main points.

[1623] Step 8:

[1624] The server embeds the summary data into a standard format template and formats it.

[1625] Step 9:

[1626] The server uses an automation tool such as MAKE to set up a workflow in which the formatted minutes are sent to a specified destination.

[1627] Step 10:

[1628] The server executes the configured workflow and generates meeting minutes, which are then saved in a Google Doc or distributed to relevant parties via email or Slack.

[1629] Step 11:

[1630] The user reviews the shared minutes and makes corrections or additional comments as necessary.

[1631] Example 1

[1632] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1633] The traditional process of recording, organizing, and sharing the details of meetings and business negotiations is often manual, requiring time and effort, and there are problems with content being easily overlooked. Furthermore, the process of manually extracting key points and creating minutes is inaccurate and inefficient, often causing delays in business operations.

[1634] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1635] In this invention, the server includes a means for activating the voice reading tool and starting recording, a means for transmitting text data stored on the terminal to the server, a means for extracting important points and content from the text data using a generative AI model, and a means for transmitting data to a specified destination using an automation tool. This enables real-time conversion of voice data, automatic extraction of important content, standardized formatting of minutes, and automatic distribution of data.

[1636] "Voice data" refers to voice information collected during meetings and business negotiations.

[1637] A "real-time text conversion means" is a mechanism that collects audio data and simultaneously converts it into text form.

[1638] "Text data" refers to character information converted from audio data.

[1639] The "means for analyzing and generating summaries" is a mechanism for analyzing the converted text data, extracting important points and content, and summarizing them in a concise format.

[1640] The "means for converting into a predetermined format" is a mechanism for organizing the generated summary according to a particular template.

[1641] The "means for automatically sending to a specified destination" is a mechanism for automatically sending the generated standard format data to a pre-set recipient.

[1642] A "voice reading tool" is an application or device that collects audio from meetings or business negotiations and converts it into text data in real time.

[1643] A "terminal" is an electronic device used by a user, such as a computer, smartphone, or tablet.

[1644] A "server" is a computer system for performing tasks such as analyzing text data, generating summaries, formatting, and transmitting data.

[1645] A "generative AI model" is an artificial intelligence technology that uses natural language processing to extract important points and content from text data and generate summaries.

[1646] An "automation tool" is software or a platform for automating a series of tasks, such as sending and storing data.

[1647] A "designated destination" is a recipient of automatically delivered data, such as a specific email address, a messaging service account, or a cloud storage folder.

[1648] This invention is a system for quickly recording the contents of meetings and business negotiations and automatically sharing them as organized minutes. The main components include a user device (computer, smartphone, tablet, etc.), a voice reading tool, a server, a generative AI model, and an automation tool.

[1649] System configuration details

[1650] 1. Voice reading

[1651] User: When the meeting or business meeting begins, launch the voice reading tool on the device and begin recording.

[1652] On your device: A voice-to-text tool (such as speech recognition software) is used to convert your conversation into text in real time using speech recognition technology such as the Google Cloud Speech-to-Text API. This converted text is temporarily stored on your device's local storage.

[1653] 2. Sending and analyzing text data

[1654] Terminal: When the meeting or business meeting ends, the saved text data is sent to the server using the HTTPS protocol.

[1655] Server: Receives the text data and inputs it into a generative AI model (e.g., GPT-4) to begin analysis. The generative AI model uses natural language processing techniques to extract key points and content from the text data.

[1656] 3. Summarization and Formatting

[1657] Server: Summarizes the meeting content based on key points extracted by the generative AI model. The server then embeds the summary into a pre-defined minutes template, which includes information such as the meeting name, attendees, agenda, and key points.

[1658] 4. Automatic data distribution

[1659] Server: Use automation tools such as MAKE or Zapier to automatically send the generated minutes to a specified destination (e.g., Google Docs, email, messaging service).

[1660] Specific use cases

[1661] For example, in a meeting for a new product development project, the following steps are taken:

[1662] Attendees: Ichiro Tanaka, Jiro Suzuki, Saburo Sato

[1663] Subject: New product development project

[1664] Agenda: Product specification, development schedule, debugging plan

[1665] procedure

[1666] 1. The user launches the voice capture tool and begins recording the audio of the meeting.

[1667] 2. The device converts the conversation being recorded into text in real time and saves the text data.

[1668] 3. After the meeting ends, the device sends the saved text data to the server, and the server receives the text data.

[1669] 4. The server inputs the received text data into a generative AI model (e.g., GPT-4) to summarize the meeting content. Specifically, it extracts important information related to "product specification determination," "development schedule," and "debug plan."

[1670] 5. The server embeds the extracted key points into a template to generate a standard formatted minutes. For example:

[1671] markdown

[1672] minutes

[1673] Meeting name: New product development project meeting

[1674] Attendees: Ichiro Tanaka, Jiro Suzuki, Saburo Sato

[1675] Date and Time: [meeting date and time]

[1676] agenda:

[1677] 1. Deciding on product specifications

[1678] [Summary]

[1679] 2. Development Schedule

[1680] [Summary]

[1681] 3. Debugging Plan

[1682] [Summary]

[1683] remarks:

[1684] There are no particular problems at the moment

[1685] 6. The server uses MAKE to execute a workflow that saves the generated minutes in a Google Doc or distributes them to all parties involved by email after the meeting ends.

[1686] 7. Users can see that the minutes are automatically shared immediately after the meeting ends and can make additional comments or corrections as needed.

[1687] This system automates a series of tasks, from recording meeting content to organizing and sharing it, greatly improving users' work efficiency.

[1688] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1689] Step 1: Launch the voice reader and start recording

[1690] User: When a meeting or business meeting begins, the user launches the voice-reading tool on their device and starts recording. Specifically, the user clicks the tool icon and presses the record button.

[1691] Enter: Start of a meeting or business meeting.

[1692] Output: Recording begins.

[1693] Step 2: Real-time text conversion of audio data

[1694] On your device: The voice reader converts your conversation into text data in real time. It uses the Google Cloud Speech-to-Text API for speech recognition. The converted text data is temporarily stored in your device's local storage.

[1695] Input: Audio data from meetings and business negotiations.

[1696] Output: Text data generated in real time.

[1697] Specific operation: The speech recognition API converts the voice data into text, and the generated text is saved in the specified directory on the device.

[1698] Step 3: Send text data to the server

[1699] Terminal: When the meeting or business meeting ends, the saved text data is sent to the server. The data is transferred securely using the HTTPS protocol.

[1700] Input: Text data after the meeting ends.

[1701] Output: Text data stored on the server.

[1702] Specific operation: Sends a POST request to the API endpoint for sending text data, and receives a success message after sending.

[1703] Step 4: Parsing the incoming text data

[1704] Server: The received text data is input into the generative AI model (GPT-4) and analysis begins. The model extracts important points and keywords from the input text data.

[1705] Input: Received text data.

[1706] Output: Extracted key points and keywords.

[1707] How it works: The GPT-4 model breaks down text data into tokens, analyzes important context, and extracts key points.

[1708] Step 5: Generate a summary

[1709] Server: Summarizes the meeting content based on key points extracted by the generative AI model. It uses a summary generation algorithm to summarize the key points in a concise format.

[1710] Input: Extracted key points and keywords.

[1711] Output: The summarized text.

[1712] Specific operation: Reconstructs the main points into short sentences and performs text processing to generate a summary.

[1713] Step 6: Embedding in the meeting minutes template

[1714] Server: The summary is embedded into a pre-defined minutes template, which automatically includes information such as the meeting name, attendees, agenda, and key points.

[1715] Input: The summarized text.

[1716] Output: A completed meeting minutes template.

[1717] What it does: Places the summary text into the template placeholder and generates the final minutes document.

[1718] Step 7: Automatic data distribution

[1719] Server: Using automation tools such as MAKE or Zapier, the generated minutes are automatically sent to the specified recipients. The data is stored in a Google Doc or distributed via email or messaging services.

[1720] Input: Completed meeting minutes template.

[1721] Output: The minutes sent to the specified destination.

[1722] Specific behavior: Triggers the automation tool workflow and executes the data sending process, saving it to Google Docs and sending it by email.

[1723] (Application example 1)

[1724] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1725] In industrial sites, there is a need to reduce the time and effort required to record work instructions and generate reports, thereby achieving efficient management. There is also a need to improve the accuracy and speed of work by recording work instructions in real time and instantly sharing necessary information. Conventional methods require the tedious task of manually recording voice instructions and compiling them into reports, making automating this process a challenge.

[1726] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1727] In this invention, the server includes means for converting voice data into text in real time, means for analyzing the converted text data and generating a summary, and means for converting the generated summary into a predetermined format, thereby enabling industrial equipment to record voice instructions in real time and automatically generate work reports.

[1728] definition statement

[1729] "Voice data" refers to audio information such as conversations and instructions recorded in digital format.

[1730] "Text data" is character information generated by analyzing voice data.

[1731] "Real time" means that processing occurs at the exact moment an event occurs.

[1732] A "summary" is a shortened version of a large amount of information, extracting only the important points.

[1733] A "predetermined format" is a predetermined form or template.

[1734] "Destination" is the address or location of the recipient to which the generated data is sent.

[1735] "Industrial equipment" refers to all equipment such as machines and robots used in production sites such as factories.

[1736] A "job report" is a document that records the details of the work performed by industrial equipment.

[1737] "Analysis" is the process of examining data in detail and extracting hidden information and meaning.

[1738] "Generative AI" is an artificial intelligence system that mimics human intelligence and automatically generates and analyzes data.

[1739] "Natural language processing technology" is a technology that allows computers to process and understand human language.

[1740] "Email" is a means of sending and receiving messages electronically over the Internet.

[1741] "Messaging service" refers to a service that allows users to send and receive text messages over the Internet.

[1742] patent specification

[1743] An embodiment of the present invention relates to a system for issuing work instructions and reporting in meetings, business negotiations, and industrial sites. This system has the functions of converting voice data into text in real time, extracting and summarizing important information, converting the summary into a predetermined format, and automatically sending it to a specified destination.

[1744] Hardware and software used

[1745] Hardware: The system includes terminals with audio input devices (e.g., microphones), industrial equipment, and communication modules.

[1746] Software: We use the Google Speech-to-Text API for speech-to-text conversion, GPT-4 as the generative AI model, and MAKE as an automation tool for automatic data transmission and management.

[1747] Program processing

[1748] A user launches a voice-to-text tool and records a meeting, business negotiation, or work instructions at an industrial site. The device converts the recorded voice into text in real time and temporarily stores it. After the meeting or work is completed, the device sends the text data to a server. The server inputs the received text data into a generative AI (e.g., GPT-4) to extract key points. The resulting summary data is converted into a specified format and automatically generated as a work report or meeting minutes. Finally, the server uses MAKE to send the generated report or meeting minutes to the specified destination.

[1749] Specific examples

[1750] Examples of meeting content

[1751] Attendees: Mr. A, Mr. B, Mr. C

[1752] Subject: New product development project

[1753] Agenda: Development schedule, resource allocation, next meeting date

[1754] Example prompt sentence:

[1755] Instructions:

[1756] 1. Assembly Instructions

[1757] Connect part A to part B

[1758] Check the connection

[1759] 2. How to use the tools

[1760] Use a screwdriver and wrench

[1761] 3. Completion checkpoints

[1762] Operation check

[1763] In this example, project meetings can be recorded in real time, summarized, formatted, and finally shared automatically with all stakeholders. A similar process can be used for work instructions in industrial sites to automatically generate work reports that can be instantly sent to managers.

[1764] This system automates a series of tasks, from recording meetings, business negotiations, and work at industrial sites to generating and sharing reports, greatly improving work efficiency.

[1765] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1766] Processing flow

[1767] Program processing steps

[1768] Step 1:

[1769] A user launches a voice-to-speech tool to record a meeting, a business meeting, or work instructions on an industrial site.

[1770] Input: Audio data for meetings and work

[1771] Output: The voice reading tool retrieves the voice data.

[1772] Step 2:

[1773] The device converts the recorded audio into text in real time and stores it temporarily.

[1774] Input: Audio data

[1775] Data processing: Converting voice data into text data using the Google Speech-to-Text API

[1776] Output: Text data is saved to the device

[1777] Step 3:

[1778] After the meeting ends or when work is completed, the device sends the saved text data to the server.

[1779] Input: Text data

[1780] Data calculation: Upload text data to the server via the network

[1781] Output: Text data is sent to the server

[1782] Step 4:

[1783] The server inputs the received text data into a generative AI (e.g., GPT-4) to extract key points.

[1784] Input: Text data

[1785] Data Computing: Analyzing text data using GPT-4 to extract key points

[1786] Output: Extracted key points (summary data)

[1787] Step 5:

[1788] The server converts the generated summary data into a predetermined format.

[1789] Input: Summary data

[1790] Data processing: Embedding summary data into a defined template

[1791] Output: Formatted minutes and working papers

[1792] Step 6:

[1793] The server uses MAKE to send the generated formatted data to the specified destination.

[1794] Input: Formatted minutes and work reports

[1795] Data Calculation: Use MAKE to send data to email or a document management system

[1796] Output: Meeting minutes and work reports are sent to the specified destination.

[1797] Specific examples of operation

[1798] Specific actions of the meeting content

[1799] Meeting scenario: A user launches a voice-to-text tool to record the meeting audio.

[1800] Speech-to-text: Recorded audio is converted to text in real time using the Google Speech-to-Text API.

[1801] Data transmission: After the meeting ends, the device uploads the text data to the server.

[1802] Gist extraction: The server uses GPT-4 to analyze the text data and summarize the main topics.

[1803] Format conversion: Embed the summarized content into a predefined format.

[1804] Data transmission: Using MAKE, formatted minutes are emailed to all parties involved.

[1805] These processing steps enable efficient and automated work recording and reporting in meetings, business negotiations, and industrial sites, significantly improving business efficiency.

[1806] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1807] This invention is a system that quickly and efficiently records the contents of meetings and business negotiations and automatically shares them as organized minutes. It also has the ability to recognize user emotions and reflect that information in the minutes. This system includes a means for converting voice data into text in real time, analyzing it to generate summaries, and converting them into a standard format for sharing. It also incorporates an emotion engine that recognizes user emotions from voice data.

[1808] A natural language description of the program's processing

[1809] Voice reading

[1810] User: At the start of a meeting or business meeting, launch the Google Docs voice-to-text tool and start recording.

[1811] On your device: The speech-to-text tool converts your conversation into text in real time and temporarily stores this text locally.

[1812] Text Acquisition

[1813] Terminal: When the meeting or business negotiation ends, the saved text data is sent to the server.

[1814] Server: Receives text data and queues it for analysis.

[1815] Text organization and summary generation

[1816] Server: Uses generative AI (e.g., GPT-4) to analyze queued text data, extract key points and content, and generate summaries.

[1817] Server: Adds user emotion information recognized by the emotion engine to the summary data.

[1818] Conversion to standard formats

[1819] Server: Embeds and formats the extracted and summarized data and user emotional information into a predetermined standard format.

[1820] Automated workflow execution and sharing

[1821] Server: Uses an automation tool such as MAKE to set up and execute a workflow that sends formatted minutes to the specified destination.

[1822] Server: Save the generated minutes in a Google Doc or distribute them to relevant parties via email or Slack.

[1823] User: Review the shared minutes and make corrections or additional comments as needed.

[1824] Specific examples

[1825] Examples of meeting content

[1826] Attendees: Yuko Sugimoto, Taro Tanaka, Ichiro Kimura

[1827] Subject: New product development project

[1828] Agenda: Development schedule, resource allocation, next meeting date

[1829] 1. The user launches the voice capture tool and begins recording the audio of the meeting.

[1830] 2. The device converts the conversation being recorded into text in real time and saves the text data.

[1831] 3. After the meeting ends, the device sends the saved text data to the server, and the server receives the text data.

[1832] 4. The server inputs the received text data into a generative AI (e.g., GPT-4) to summarize the meeting content. Specifically, it extracts important information related to the "development schedule," "resource allocation," and "next meeting date."

[1833] 5. The server uses an emotion engine to analyze the user's emotions during the conversation. For example, it recognizes the speaker's emotions (excitement, relief, anxiety, etc.) separately from the content of the speech.

[1834] 6. The server embeds the extracted key points and sentiment information into a template to generate a standardized minutes of the meeting. For example:

[1835] minutes

[1836] Meeting name: New product development project meeting

[1837] Attendees: Yuko Sugimoto, Taro Tanaka, Ichiro Kimura

[1838] Date and Time: [meeting date and time]

[1839] agenda:

[1840] 1. Development Schedule

[1841] [Summary]

[1842] [Emotional information]

[1843] 2. Resource Allocation

[1844] [Summary]

[1845] [Emotional information]

[1846] 3. Next meeting date

[1847] [Summary]

[1848] [Emotional information]

[1849] remarks:

[1850] There are no particular problems at the moment

[1851] 7. The server uses MAKE to execute a workflow that saves the generated minutes in a Google Doc or distributes them to all parties involved by email after the meeting ends.

[1852] 8. Users can see that the minutes are automatically shared immediately after the meeting ends and can make additional comments or corrections as needed.

[1853] This system automates a series of tasks, from recording and organizing meeting content, to analyzing emotions and sharing, greatly improving users' work efficiency.

[1854] The processing flow will be explained below.

[1855] Step 1:

[1856] When a user starts a meeting or business meeting, they launch Google Docs' voice reading tool and begin recording.

[1857] Step 2:

[1858] Your device will use Google Docs' speech-to-text tool to convert your conversation into text in real time, which will then be temporarily stored locally.

[1859] Step 3:

[1860] The terminal transmits the saved text data together with the audio file to the server.

[1861] Step 4:

[1862] The server receives the text data and audio files and queues them for analysis.

[1863] Step 5:

[1864] The server uses generative AI (e.g., GPT-4) to analyze the queued text data and extract key points and content.

[1865] Step 6:

[1866] After the server extracts the main points, it uses an emotion engine to recognize the speaker's emotions from the audio file and analyzes the emotion information.

[1867] Step 7:

[1868] The server adds the user's emotion information recognized by the emotion engine to the generated summary data.

[1869] Step 8:

[1870] The server embeds the summary data and emotional information into a predetermined standard format (minutes template) and formats the formatted data.

[1871] Step 9:

[1872] The server uses MAKE's automation tools to set up a workflow that sends formatted minutes to a specified destination.

[1873] Step 10:

[1874] The server executes the configured workflow and generates meeting minutes, which are then saved in a Google Doc or distributed to relevant parties via email or Slack.

[1875] Step 11:

[1876] The user reviews the shared minutes and makes corrections or additional comments as necessary.

[1877] Example 2

[1878] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1879] Conventional meeting recording systems not only convert audio data into text, but also require manual data organization, summary generation, and emotional analysis, resulting in cumbersome post-meeting tasks that require time and effort. Furthermore, they are unable to reflect the emotional information of meeting participants, limiting the content of meeting minutes. This leads to issues such as a decline in the quality of meeting minutes and a decrease in communication efficiency.

[1880] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1881] In this invention, the server includes means for converting voice data into text in real time, means for analyzing the converted text data and generating a summary, means for adding user emotion information recognized by an emotion engine to the generated summary data, and means for automatically sending the generated formatted data to a specified destination. This automates data organization and summary generation after a meeting, and enables the generation and rapid sharing of detailed minutes including emotion information.

[1882] "Voice data" is data that formalizes voice signals collected during meetings, business negotiations, etc.

[1883] "Means for converting voice data into text in real time" refers to a technology or device that instantly analyzes voice signals and converts them into text information.

[1884] "Text data" is information obtained by converting voice data into text format.

[1885] "Analysis" is the process of examining data in detail and extracting specific patterns and important information.

[1886] A "summary" is a short summary of the main points extracted from a long piece of text or a large amount of information.

[1887] An "emotion engine" is a technology or software that analyzes voice or text data and recognizes the speaker's emotions.

[1888] "Format" refers to arranging data into a specific form or structure.

[1889] "Specified destination" means a predetermined recipient or location to which generated data is sent.

[1890] "Generative AI" is an artificial intelligence technology that learns patterns from various input data and generates new outputs.

[1891] An "automation tool" is software or a service that performs a specific process or task without human intervention.

[1892] MODE FOR CARRYING OUT THE INVENTION

[1893] The present invention is a system that quickly and efficiently records the contents of meetings and business negotiations and automatically shares them as organized minutes. It also has the ability to recognize user emotions and reflect that information in the minutes. The system uses a speech recognition tool to convert voice data into text in real time, a generative AI model to analyze and summarize the text data, an emotion engine for emotion analysis, and an automated tool for formatting and sharing. A specific embodiment of this system is described below.

[1894] Voice reading

[1895] When a user starts a meeting or business meeting, they activate a voice recognition tool (e.g., Google Docs' "voice input" function) and start recording. The voice recognition tool converts the spoken audio into text in real time and temporarily stores this text data locally.

[1896] Text Acquisition

[1897] When a meeting or business meeting ends, the device sends the text data stored locally to the server, which receives the text data and queues it for analysis. This queue manages the data waiting to be processed, enabling efficient analysis of the text data.

[1898] Text organization and summary generation

[1899] The server uses a generative AI model (e.g., GPT-4) to analyze the queued text data, extract key points and content, and generate a summary. Furthermore, it adds user sentiment information recognized by an emotion engine (e.g., IBM Watson's sentiment analysis function) to the generated summary data.

[1900] Conversion to standard formats

[1901] The server embeds the main points and emotional information into a predetermined format and formats it into minutes. This formatted data is stored as a temporary file.

[1902] Automated workflow execution and sharing

[1903] The server uses an automation tool such as MAKE to set up and execute a workflow that sends the minutes converted into a standard format to a specified destination. This workflow is configured to save the generated minutes in a specific document management system (e.g., Google Docs) and distribute them to relevant parties via email or a messaging service (e.g., Slack).

[1904] Specific examples

[1905] Examples of meeting content:

[1906] Attendees: A, B, C

[1907] Topic: New project progress

[1908] Agenda: Schedule confirmation, resource allocation, next meeting date

[1909] At the start of a meeting, a user activates Google Docs' voice input and begins recording the conversation. The recorded audio is converted into text in real time and saved to a local file. After the meeting ends, the text data is sent from the device to a server and queued for analysis. The server uses a generative AI model and emotion engine to analyze the text data, extract key points and emotional information, and generate meeting minutes.

[1910] The minutes are then formatted and saved in a Google Doc or distributed to relevant parties via email or Slack using MAKE's automation tools. Users can review the shared minutes and make additional comments or corrections as needed.

[1911] Example prompt sentence:

[1912] "I want to extract the key points from a meeting about the progress of a new project and generate a summary. I also want to analyze the emotions of the attendees and reflect them in the minutes."

[1913] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1914] Step 1:

[1915] When starting a meeting or business meeting, a user activates a voice recognition tool (e.g., Google Docs' "voice input" function) and starts recording.

[1916] Input: Meeting audio data

[1917] Output: Converted text data

[1918] Specific behavior:

[1919] The user opens Google Docs and selects "Voice Typing" from the Tools menu.

[1920] Once the meeting starts, click the "Speech to Text" button to start dictating. The speech recognition engine will capture the voice data and convert it into text in real time.

[1921] Step 2:

[1922] The device converts the recorded conversation into text in real time and temporarily stores this text data locally.

[1923] Input: Meeting audio data

[1924] Output: Locally saved text data

[1925] Specific behavior:

[1926] The speech recognition engine (Google's speech recognition API) converts the speech data into text information.

[1927] The converted text data is saved in a temporary directory in the terminal's local file system.

[1928] Step 3:

[1929] When the meeting or business negotiation ends, the terminal transmits the saved text data to the server.

[1930] Input: Locally stored text data

[1931] Output: Text data sent to the server

[1932] Specific behavior:

[1933] The user clicks the "Stop" button when the conference is over.

[1934] The device retrieves the text file from the temporary directory and uploads it to the server via an HTTP request.

[1935] Step 4:

[1936] The server receives the text data and queues it for analysis.

[1937] Input: Text data sent from the terminal

[1938] Output: Text data queued for analysis

[1939] Specific behavior:

[1940] The server stores the received data in a temporary storage directory.

[1941] The text data is registered as a task in the analysis queue.

[1942] Step 5:

[1943] The server uses a generative AI model (e.g., GPT-4) to analyze the queued text data, extract key points and content, and generate a summary.

[1944] Input: Text data queued for analysis

[1945] Output: Generated summary data

[1946] Specific behavior:

[1947] The server retrieves the text data from the queue.

[1948] A prompt sentence is input to the generative AI model and a summary result is received from the generative AI model.

[1949] Step 6:

[1950] The server adds user emotion information recognized by an emotion engine (e.g., IBM Watson's emotion analysis function) to the summary data.

[1951] Input: Generated summary data

[1952] Output: Summarized data with added sentiment information

[1953] Specific behavior:

[1954] The server sends the voice or text data to the sentiment analysis engine.

[1955] The received emotion information is incorporated into the summary data.

[1956] Step 7:

[1957] The server embeds the extracted main points and emotional information into a predetermined standard format and formats it.

[1958] Input: Summarized data with added sentiment information

[1959] Output: Minutes in a standard format

[1960] Specific behavior:

[1961] The server reads a template file with a fixed format and replaces the summary data and emotion information with the corresponding placeholders.

[1962] Save the formatted minutes as a temporary file.

[1963] Step 8:

[1964] The server uses an automation tool such as MAKE to set up and execute a workflow that sends the formatted minutes to the specified destination.

[1965] Input: Minutes in a standard format

[1966] Output: The minutes sent to the specified destination

[1967] Specific behavior:

[1968] The server sends an API request to an automation tool (e.g., MAKE) to trigger the workflow.

[1969] The workflow saves meeting minutes in a Google Doc and distributes them to relevant parties via email or messaging services.

[1970] Step 9:

[1971] Users can review the shared minutes and make corrections or additional comments as needed.

[1972] Input: Shared minutes

[1973] Output: Meeting minutes with corrections and additional comments

[1974] Specific behavior:

[1975] The user opens a Google Doc or an email they received and checks the minutes.

[1976] If any corrections or comments are required, please use the editing function of Google Docs to make additions or corrections.

[1977] (Application example 2)

[1978] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1979] Current meeting recording systems and work instruction recording systems simply convert audio data into text, without detailed analysis or summarization, including emotional information. As a result, the content of meetings and work instructions may not be accurately understood, leading to miscommunication and lack of awareness. Furthermore, methods for automatically sharing formatted data are limited, making efficient information sharing difficult. A system that solves these issues is needed.

[1980] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for converting voice data to text in real time, means for analyzing the converted text data and generating a summary, means for recognizing emotions from the voice data and reflecting the emotional information in the summary, means for converting the generated summary into a predetermined format, and means for automatically sending the generated formatted data to a specified destination. This enables accurate text conversion and analysis of voice data, addition of emotional information, and efficient information sharing.

[1981] "Audio data" refers to data that has been recorded and stored in digital form, such as voice or spoken words.

[1982] "Real-time text conversion means" refers to technology or equipment for instantly converting voice data into text form.

[1983] "Text data" refers to digital data expressed as character information.

[1984] "Analysis" refers to the process of analyzing and interpreting input data to understand its meaning and structure.

[1985] "Means for generating summaries" refers to technologies and devices that extract important information from large amounts of text data and summarize it concisely.

[1986] "Means for recognizing emotions and incorporating that emotional information into summaries" refers to technology and devices for identifying emotions from audio data and incorporating that emotional information into summaries.

[1987] The "means for converting into a predetermined format" refers to a technique or device for converting the generated summary into a predetermined format.

[1988] "Formatted data" refers to digital data that has been arranged in a particular format.

[1989] "Means for automatically sending to a specified destination" refers to technology or devices for automatically sending generated data to a specific recipient.

[1990] "Designated destination" refers to a recipient or location that has been pre-configured to receive data.

[1991] "Improved natural language processing technology" refers to technology for analyzing and processing natural language with improved accuracy and efficiency compared to conventional technology.

[1992] "Document Management System" means software or a platform for storing, managing, and accessing documents and data.

[1993] "Email or messaging service" means a service for sending and receiving messages or information over the Internet.

[1994] "Automation tool" refers to software that automatically performs designated tasks or processes.

[1995] To implement this invention, it is necessary to build a system that converts voice data into text in real time, analyzes the text data to generate a summary, converts it into a standard format after reflecting emotional information, and automatically sends it to a specified destination. This system aims to efficiently record work instructions in a factory and analyze the content of those instructions.

[1996] composition

[1997] The system components are as follows:

[1998] 1. Voice data collection devices: Microphones and voice-reading devices placed on the factory floor to record work instructions and conversations.

[1999] 2. Speech Recognition Software: Software that converts voice data into text in real time. Specifically, we use the Google Cloud Speech-to-Text API.

[2000] 3. Text analysis and summary generation software: Uses generative AI models to analyze converted text data and generate summaries. OpenAI's GPT-4 is a prime example.

[2001] 4. Emotion Recognition Software: Software to recognize emotions from audio data and incorporate that emotional information into summaries. Uses IBM Watson Tone Analyzer.

[2002] 5. Data conversion and sharing software: Software that converts the generated summaries into a standard format and automatically sends them to designated recipients. This includes automation tools such as Zapier.

[2003] Operation flow

[2004] 1. Voice reading

[2005] Users use a voice data capture device to record work instructions and conversations, and speech recognition software converts the recorded voice data into text in real time.

[2006] 2. Text Analysis

[2007] Text analysis and summary generation software analyzes the converted text data, extracts key points, and generates summaries using a generative AI model (e.g., GPT-4).

[2008] 3. Emotion recognition

[2009] Emotion recognition software (IBM Watson Tone Analyzer) is used to recognize emotions from audio data. The analyzed emotional information is reflected in the summary.

[2010] 4. Data Sharing

[2011] The generated data, including summaries and sentiment information, is converted into a predetermined format and sent to a designated destination via an automation tool (such as Zapier), allowing relevant parties to receive important information in real time.

[2012] Specific examples

[2013] If the factory manager instructs the robot to "run a test run of a new product on the next line," the system operates as follows:

[2014] An audio data collection device records the instructions.

[2015] Speech recognition software converts this audio data into text.

[2016] The text analysis and summary generation software summarizes it as "Test run of new product on the production line."

[2017] Emotion recognition software analyzes the "high level of attention" and "sense of security" of the person giving the instructions.

[2018] The generated summary and emotion information are converted into a predetermined format and sent to a specified destination.

[2019] Example prompt sentence:

[2020] "Please summarize the following text:

[2021] On the next line, I would like you to conduct a test run of a new product.

[2022] High level of focus and peace of mind

[2023] The above is a detailed embodiment for carrying out the invention. This system enables efficient and accurate recording and transmission of work instructions in a factory.

[2024] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[2025] Step 1:

[2026] Voice reading

[2027] The user uses a voice data collection device to record work instructions and conversations.

[2028] Input: Audio data

[2029] Specific operation: Microphones and voice reading devices installed on the factory floor collect voices from factory managers and work instructors in real time.

[2030] Output: Real-time audio data

[2031] Step 2:

[2032] Real-time text transcription

[2033] The device converts the collected voice data into text in real time using voice recognition software (Google Cloud Speech-to-Text API).

[2034] Input: Real-time audio data

[2035] How it works: Speech recognition software analyzes speech waveforms and converts what is being said into text.

[2036] Output: Text data

[2037] Step 3:

[2038] Text data analysis

[2039] The device uses a generative AI model (OpenAI GPT-4) to analyze the converted text data, extract key points, and generate a summary.

[2040] Input: Text data

[2041] How it works: The generative AI model analyzes the input text data, uses prompts to identify key information, and summarizes it succinctly.

[2042] Output: Summary text

[2043] Step 4:

[2044] emotion recognition

[2045] The device uses emotion recognition software (IBM Watson Tone Analyzer) to recognize emotions from the analyzed voice data and reflects that emotional information in the summary.

[2046] Input: Text data

[2047] What it does: Emotion recognition software analyzes keywords and context within text data to identify the speaker's emotions (e.g., excitement, relief, anxiety, etc.).

[2048] Output: Summarized text with emotional information

[2049] Step 5:

[2050] Conversion to standard formats

[2051] The terminal converts the generated summary text into a predetermined format.

[2052] Input: Summarized text with sentiment information

[2053] What it does: The formatting function embeds summary and sentiment information into templates to generate standardized work instructions and meeting minutes.

[2054] Output: formatted data

[2055] Step 6:

[2056] Automatic Data Transmission and Sharing

[2057] The server uses an automation tool (such as Zapier) to automatically send the generated formatted data to the specified destination.

[2058] Input: Formatted data

[2059] What it does: The automation tool distributes data to a pre-specified list of recipients or sharers via email, messaging services, or document management systems.

[2060] Output: Shared data

[2061] Through these steps, a series of automated processes are implemented, from collecting voice data to analyzing it, recognizing emotions, converting formats, and sharing the information, enabling accurate and prompt recording and transmission of work instructions in factories.

[2062] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[2063] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[2064] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[2065] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[2066] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[2067] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[2068] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[2069] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[2070] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[2071] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[2072] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[2073] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[2074] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[2075] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[2076] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[2077] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[2078] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[2079] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[2080] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[2081] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[2082] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[2083] The following is further disclosed regarding the above embodiment.

[2084] (Claim 1)

[2085] a means for converting voice data into text in real time;

[2086] means for analyzing the converted text data and generating a summary;

[2087] means for converting the generated summary into a predetermined format;

[2088] means for automatically transmitting the generated formatted data to a specified destination;

[2089] A system including:

[2090] (Claim 2)

[2091] 10. The system of claim 1, wherein the converted text data is analyzed and the generated summary is generated using improved natural language processing techniques.

[2092] (Claim 3)

[2093] 10. The system of claim 1, configured to store the formatted data in a specific document management system or send it via email or messaging service.

[2094] "Example 1"

[2095] (Claim 1)

[2096] a means for converting voice data into text in real time;

[2097] means for analyzing the converted text data and generating a summary;

[2098] means for converting the generated summary into a predetermined format;

[2099] means for automatically transmitting the generated formatted data to a specified destination;

[2100] a means for launching the voice reading tool and starting recording;

[2101] means for transmitting text data stored in the terminal to a server;

[2102] A means of extracting key points and content from text data using generative AI models,

[2103] means for transmitting the data to a specified destination using an automated tool;

[2104] A system including:

[2105] (Claim 2)

[2106] 2. The system of claim 1, configured to analyze the converted text data, extract key points and content using a generative AI model, and embed the generated key points into a minutes template.

[2107] (Claim 3)

[2108] 10. The system of claim 1, configured to store the formatted data in a specific document management system or send it via email or messaging service.

[2109] "Application Example 1"

[2110] (Claim 1)

[2111] a means for converting voice data into text in real time;

[2112] means for analyzing the converted text data and generating a summary;

[2113] means for converting the generated summary into a predetermined format;

[2114] means for automatically transmitting the generated formatted data to a specified destination;

[2115] means for the industrial device to record voice instructions in real time and automatically generate a work report;

[2116] A system including:

[2117] (Claim 2)

[2118] 10. The system of claim 1, wherein the converted text data is analyzed and the generated summary is generated using improved natural language processing techniques.

[2119] (Claim 3)

[2120] 10. The system of claim 1, configured to store the formatted data in a specific document management system or send it via email or messaging service.

[2121] "Example 2: Combining Emotion Engines"

[2122] (Claim 1)

[2123] a means for converting voice data into text in real time;

[2124] means for analyzing the converted text data and generating a summary;

[2125] means for adding user emotion information recognized by the emotion engine to the generated summary data;

[2126] means for automatically transmitting the generated formatted data to a specified destination;

[2127] A system including:

[2128] (Claim 2)

[2129] 10. The system of claim 1, wherein the converted text data is analyzed and the generated summary is generated using improved natural language processing and emotion recognition techniques.

[2130] (Claim 3)

[2131] 10. The system of claim 1, configured to transmit the formatted data via a specific document management system, email, or messaging service.

[2132] "Application example 2 when combining emotion engines"

[2133] (Claim 1)

[2134] a means for converting voice data into text in real time;

[2135] means for analyzing the converted text data and generating a summary;

[2136] A means for recognizing emotions from speech data and reflecting the emotional information in a summary;

[2137] means for converting the generated summary into a predetermined format;

[2138] means for automatically transmitting the generated formatted data to a specified destination;

[2139] A system including:

[2140] (Claim 2)

[2141] 10. The system of claim 1, further comprising: analyzing the converted text data; applying improved natural language processing techniques to the generated summary; and adding emotional information.

[2142] (Claim 3)

[2143] 10. The system of claim 1, configured to store the formatted data in a specific document management system, or send it via email or messaging services, and share it using any automation tools necessary. [Explanation of symbols]

[2144] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. a means for converting voice data into text in real time; means for analyzing the converted text data and generating a summary; means for converting the generated summary into a predetermined format; means for automatically transmitting the generated formatted data to a specified destination; A system including:

2. 10. The system of claim 1, wherein the converted text data is analyzed and the generated summary is generated using improved natural language processing techniques.

3. 10. The system of claim 1, configured to store the formatted data in a particular document management system or send it via email or messaging service.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A