System

The system addresses inefficiencies in traditional meetings by automating the recording, conversion, and summarization of meeting content, enhancing productivity through real-time access to past knowledge and solutions.

JP2026018060APending Publication Date: 2026-02-05SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024119121
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-24
Publication Date
2026-02-05

AI Technical Summary

Technical Problem

Traditional meetings are inefficient due to time-consuming manual recording of minutes and lack of systems to utilize meeting content effectively, leading to reduced productivity and ineffective use of past knowledge and solutions.

Method used

A system that records meeting content, converts audio to text, automatically creates summaries, stores them in a database, and searches for and proposes related information, enabling real-time utilization of past knowledge and solutions.

Benefits of technology

Automates the recording and management of meetings, allowing for quick search and sharing of relevant information, thereby improving efficiency and productivity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026018060000001_ABST
    Figure 2026018060000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system comprising: means for recording contents of a meeting; means for converting voice data into text data; means for automatically creating a summary of minutes from the generated text data; means for storing the summary of minutes in a database; means for searching for and proposing related information from the database based on a problem proposed during the meeting; and means for displaying the proposed related information and providing means for contacting other meeting participants.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] In traditional meetings, not only was it time-consuming and labor-intensive to create minutes, but many of the ideas and issues proposed during the meeting often ended up getting lost. Furthermore, there was a lack of a system for effectively utilizing the content discussed in the meeting, meaning that past knowledge and solutions were not fully utilized. This resulted in problems that reduced the productivity and effectiveness of meetings. Therefore, there is a need for a system that can efficiently record meeting content and utilize that information in real time. [Means for solving the problem]

[0005] The present invention is a system including a means for recording the contents of a meeting, a means for converting the audio data into text data, a means for automatically creating a summary of the minutes from the generated text data, a means for saving the summary of the minutes in a database, a means for searching for and proposing related information from the database based on issues proposed during the meeting, and a means for displaying the proposed related information and providing a means for contacting other meeting participants. This makes it possible to instantly refer to and utilize past knowledge and solutions for issues proposed during the meeting, and to automate and streamline the process of creating minutes.

[0006] "Meeting Content" refers to all matters discussed, statements, proposals, decisions, and action items during the meeting.

[0007] "Recording means" refers to a device or software for recording sound as digital data.

[0008] "Audio data" refers to data that represents recorded audio signals in digital form.

[0009] "Text data" refers to data containing character information converted from audio data.

[0010] "Generated text data" refers to data resulting from converting voice data into text.

[0011] A "minutes summary" refers to a document that summarizes the main points and important matters from the meeting.

[0012] "Automatic creation means" refers to a program that has the ability to generate or construct specific data through a system without human intervention.

[0013] "Database" refers to a system for managing an organized collection of stored digital information.

[0014] "Means of storage" refers to storage devices and systems for long-term retention and management of data.

[0015] "Proposed issues" refers to problems or matters to be resolved that were discussed or proposed during the meeting.

[0016] "Means for searching a database" refers to a program or function for searching for information in a database based on specified criteria.

[0017] "Related information" refers to data on past meetings and solutions that are deemed useful for the proposed problem.

[0018] "Means for suggesting" refers to a system or program for presenting related information to the user.

[0019] "Means for displaying" refers to a monitor or display that allows the user to view relevant information.

[0020] "Means of contact" refers to communication facilities for contacting participants of different conferences. [Brief explanation of the drawings]

[0021] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6]FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0022] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0023] First, the terms used in the following description will be explained.

[0024] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0025] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0026] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0027] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0028] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0029] [First embodiment]

[0030] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0031] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0032] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0033] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0034] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0035] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0036] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0037] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0038] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0039] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0040] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0041] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0042] The system of the present invention records the contents of a meeting, converts the recording data into text data, automatically creates a summary of the minutes, and stores the summary in a database. Furthermore, it can search for and propose related information from the database based on issues raised during the meeting. This system allows for efficient recording of meeting contents and makes it possible to utilize past knowledge and solutions in real time.

[0043] Specifically, the system operates in the following steps:

[0044] Program processing

[0045] 1. Recording of meeting content

[0046] User: Start the recording device as soon as the meeting starts and record the conversation using a dedicated recording app.

[0047] 2. Sending recording data

[0048] Device: After the meeting ends, the recording data is sent to the server. The recording app automatically sends the data.

[0049] 3. Converting Audio Data to Text

[0050] Server: Receives the transmitted voice data and converts it into text data using a speech recognition API. The converted text data is temporarily saved.

[0051] 4. Automatic summary creation

[0052] Server: Based on the text data converted from the voice data, the generative AI extracts important points and decisions, and automatically creates a summary of the minutes.

[0053] 5. Save summary

[0054] Server: Saves the summary of the minutes to a database. The database stores not only the summary but also the original text data.

[0055] 6. Search and suggest related information

[0056] Server: Based on the issues recorded by users during the meeting, the server searches the database to find relevant past minutes and solutions.

[0057] Server: Compiles relevant information into suggestions and sends them to the user's device.

[0058] 7. Information Display and Contact Functions

[0059] Terminal: Providing suggested relevant information to the user, the user has the option to review the required information and contact other meeting participants.

[0060] User: Uses the suggested information to send a message to relevant meeting participants and schedule a meeting if necessary.

[0061] Specific examples

[0062] For example, when a new product development team holds a meeting, they launch a dedicated recording app at the start of the meeting. After the meeting ends, the recording data is automatically sent to a server. The server receives the data and converts it into text. Generative AI extracts key points from the text data, creates a summary, and stores it in a database.

[0063] Next, the system searches the database for past discussion topics related to new product features proposed during the meeting. If relevant past meeting minutes or ideas are found, they are suggested to the user, who can use this information to develop new ideas.

[0064] Thus, the present invention can significantly improve the efficiency and productivity of meetings.

[0065] The processing flow will be explained below.

[0066] Step 1:

[0067] When a user starts a meeting, the recording device is activated and the conversation is recorded using a dedicated recording app.

[0068] Step 2:

[0069] When the meeting ends, the device automatically sends the recording data to the server, which is done automatically by the recording app.

[0070] Step 3:

[0071] The server receives the transmitted audio data and stores it temporarily.

[0072] Step 4:

[0073] The server uses a speech recognition API to convert the voice data into text data, which is then temporarily saved.

[0074] Step 5:

[0075] The server inputs the text data into the AI ​​generator, which automatically creates a summary of the minutes, extracting key points and decisions.

[0076] Step 6:

[0077] The server saves the created summary in a database, along with the original text data.

[0078] Step 7:

[0079] Based on the issues users record during meetings, the server searches the database to find minutes of past meetings and solutions.

[0080] Step 8:

[0081] The server compiles relevant information into suggestions and sends them to the terminal, where they are displayed to the user.

[0082] Step 9:

[0083] The user reviews the suggested information and selects the information they require. The user has the option to contact other meeting participants.

[0084] Step 10:

[0085] The user can then use the suggested information to send messages to relevant meeting participants and schedule meetings as needed.

[0086] Step 11:

[0087] The server feeds back the utilized information and results to the database, enabling more accurate proposals to be made in future meetings.

[0088] Example 1

[0089] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0090] Conventional meeting recording systems often require manual recording and management of meeting content, which requires time and effort. It is also difficult to quickly search past minutes and related information, making it difficult to receive immediate and effective feedback on issues raised during meetings. Furthermore, they lack the ability to share relevant information with other meeting participants, potentially reducing the efficiency and productivity of meetings.

[0091] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0092] In this invention, the server includes means for recording the contents of the meeting, means for converting the audio data into text data, means for automatically creating a summary of the minutes from the generated text data, means for saving the summary of the minutes in a database, means for searching and proposing related information from the database based on issues proposed during the meeting, means for displaying the proposed related information and providing a means for contacting other meeting participants, means for extracting important points from the text data using a generation AI, and means for transmitting the recorded data to the server. This automates the recording and management of meetings and enables past minutes and related information to be quickly searched and shared, thereby improving the efficiency and productivity of meetings.

[0093] "Means for recording the contents of a meeting" refers to a device or application for recording statements and conversations made during a meeting as audio data.

[0094] "Means for converting voice data into text data" refers to voice recognition technology or software for analyzing recorded voice data and converting it into text data.

[0095] The "means for automatically creating a summary of meeting minutes from generated text data" is a generative AI system that analyzes text data to extract and summarize important points and decisions.

[0096] The "means for storing the summary of the minutes in a database" refers to a device or system for storing the automatically created summary of the minutes in a relational database, cloud storage, or the like.

[0097] "A means for searching and proposing relevant information from a database based on issues proposed during a meeting" is an engine that searches a database for past minutes and solutions related to the issues entered by the user, extracts appropriate information, and presents it.

[0098] "Means for displaying suggested related information and providing a means for contacting other conference participants" refers to functionality or software that displays related information on a user interface and enables users to send messages to other conference participants or schedule meetings.

[0099] "Means for extracting key points from text data using generative AI" refers to technology that uses a generative AI model (e.g., a natural language processing model) to mechanically extract key points and decisions from text data.

[0100] The "means for transmitting recorded data to a server" refers to a communication function or protocol for transmitting recorded voice data to a server via the Internet.

[0101] This invention is a system that records the contents of a meeting, converts the recorded data into text data, automatically creates a summary of the minutes, stores it in a database, and searches for and suggests related information.

[0102] First, users turn on the recording device at the start of a meeting and record the conversation using a dedicated recording app, which can be software installed on a smartphone, tablet, or PC, such as Otter.ai or a similar application.

[0103] Then, after the meeting ends, the device sends the recording data to the server. The recording app does this automatically, and the data is sent using an encryption protocol (e.g., TLS). The data transmission is completed without any special operation by the user.

[0104] The server sends the received voice data to the Google Cloud Speech-to-Text API, which converts the voice data into text data. The converted text data is temporarily stored on the server. After conversion is complete, the server uses a generative AI (e.g., OpenAI GPT-3) to extract important points and decisions from the text data and create a summary of the meeting minutes.

[0105] The server stores the generated summary in a database (e.g., MySQL, PostgreSQL). The database stores not only the summary of the minutes but also the original text data, ensuring the integrity of the information.

[0106] Furthermore, based on the issues proposed during the meeting, the server searches the database and provides related past meeting minutes and solutions. This search uses an SQL query such as "SELECT FROM summaries WHERE text LIKE '%keywords related to the issue%'". The server compiles the relevant information found as suggestions and sends them to the user's device. This is sent in JSON format, and a dedicated suggestion display app receives and displays it.

[0107] The device displays the suggested related information to the user. The user can check the displayed information and, if necessary, send a message to other meeting participants or schedule a new meeting. By clicking the "Contact" button in the suggestion display app, a message sending screen will appear, allowing the user to send a message to the relevant participants.

[0108] Specific examples

[0109] When a new product development team holds a meeting, the user launches a dedicated recording app at the start of the meeting to record the conversation. After the meeting ends, the recorded data is automatically sent from the device to the server. The server receives this data and converts the speech to text using the Google Cloud Speech-to-Text API. OpenAI GPT-3 extracts key points from the converted text data and generates a summary of the meeting minutes. The generated summary is stored in a MySQL database.

[0110] Next, for new product features proposed during a meeting, the server searches the database to see if similar topics have been discussed in the past. If relevant minutes or ideas from past meetings are found, the server compiles them into a proposal and sends them in JSON format to the user's device. The user can then view the information through the proposal display app, send messages to participants in related meetings, and schedule new meetings.

[0111] Prompt Sentence Examples

[0112] "Please extract the key points and decisions from the following text data and create a summary of the meeting minutes.

[0113] Text data:

[0114] Proposing new product features at a meeting

[0115] Supervisor approves

[0116] We plan to work out the details at the next meeting.

[0117] ...

[0118] "

[0119] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0120] Step 1:

[0121] When the meeting starts, the user activates the recording device and records the conversation using a dedicated recording app. The recording app is installed on a smartphone, tablet, or PC, and an example is "Otter.ai." The user starts recording by pressing the "Start Recording" button. The input is an audio signal, and the output is voice data.

[0122] Step 2:

[0123] After the meeting ends, the device sends the recorded data to the server. When the recording app detects the end of the meeting, it automatically compresses the recorded data and sends it to the server using an encryption protocol (e.g., TLS). The input is the audio data, and the output is the data sent to the server.

[0124] Step 3:

[0125] The server stores the received voice data and sends it to the Google Cloud Speech-to-Text API to convert it into text. This API analyzes the voice data and outputs it as a string. The input is voice data and the output is text data.

[0126] Step 4:

[0127] The server temporarily stores the text data received from the Google Cloud Speech-to-Text API. It then supplies the text data to a generative AI (e.g., OpenAI GPT-3) with a prompt: "Please extract the key points and decisions from the following text data and create a summary of the minutes." The generative AI extracts the key points and generates a summary of the minutes. The input is the text data and the prompt, and the output is a summary of the minutes.

[0128] Step 5:

[0129] The server saves the generated summary of the minutes to a database. The server first establishes a database connection and executes the SQL query "INSERT INTO summaries (text, summary) VALUES (?, ?)". The original text data is also saved to the same database. The input is the summary of the minutes and the text data, and the output is the information saved in the database.

[0130] Step 6:

[0131] The server searches the database based on the issues recorded by the user during the meeting. For example, it runs the SQL query "SELECT FROM summaries WHERE text LIKE '%keywords related to the issue%'". The server filters the relevant information found and extracts the relevant information. The input is the issue keywords and the output is the relevant information.

[0132] Step 7:

[0133] The server compiles related information as suggestions and sends them to the user's device. The suggestions are sent in JSON format, and a dedicated suggestion display app receives them and displays them on the user interface. The input is the related information, and the output is the data sent to the device.

[0134] Step 8:

[0135] The device displays the suggested related information to the user. The suggestion display app parses the received JSON data and presents it to the user in a list format. The user can click on an interesting suggestion from the list to view details. The input is JSON data, and the output is the information displayed to the user.

[0136] Step 9:

[0137] Based on the proposed information, the user can send a message to the relevant meeting participants and set up a new meeting. For example, clicking the "Contact" button in the proposal display app will display the message sending screen, allowing the user to send a message to the relevant participants. The input is the user's action and the proposed information, and the output is the sent message.

[0138] (Application example 1)

[0139] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0140] Currently, in workplaces that use industrial machinery, it is difficult to efficiently create meeting records and utilize past related information in real time. Furthermore, there is a lack of a method to efficiently search and propose solutions to issues raised during meetings, which leads to inefficiency and reduced productivity. Furthermore, in industrial machinery workplaces, there is a need for a method to directly reflect the information in meeting minutes in machine operation.

[0141] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0142] In this invention, the server includes: means for recording the contents of a meeting; means for converting voice data into text data; means for automatically creating a summary of the minutes from the generated text data; means for saving the summary of the minutes in a database; means for searching and proposing related information from the database based on issues proposed during the meeting; means for displaying the proposed related information and providing a means for contacting other meeting participants; means for automatically generating meeting minutes and presenting search results for related information on the display device of the industrial machine in cooperation with the industrial machine control device; means for generating the proposed related information using a generative AI model; and means for generating a summary of the minutes using prompt sentences. This enables efficient recording of meeting contents and real-time utilization of past meeting information, thereby improving work efficiency and productivity at industrial machine sites.

[0143] A "means for recording the contents of a meeting" is a device or software that can record the audio content spoken during a meeting.

[0144] "Means for converting voice data into text data" refers to algorithms or tools for analyzing recorded voice data and converting it into text information.

[0145] The "means for automatically creating a summary of minutes from generated text data" is a means for extracting important points and decisions based on text data converted from speech and generating a summary.

[0146] The "means for saving the summary of the minutes in a database" refers to a function or device for writing and saving the generated summary in a database.

[0147] "A means for searching and proposing relevant information from a database based on issues proposed during a meeting" is a means for inputting the agenda items or issues raised during a meeting, searching a database for past information related to them, and presenting it.

[0148] "Means for displaying suggested related information and providing means for contacting other conference participants" refers to means for displaying searched related information in a user interface and providing options for communicating with other conference participants.

[0149] "Means for automatically generating meeting minutes and presenting search results for related information on the display device of the industrial machine in cooperation with the control device of the industrial machine" refers to means for exchanging data with the control device of the industrial machine and visually displaying the meeting minutes and related information on the display device of the machine.

[0150] A "means for generating suggested relevant information using a generative AI model" is a means for using an artificial intelligence model to automatically generate highly relevant information or solutions based on input data.

[0151] The "means for generating a summary of meeting minutes using prompt sentences" is a means for efficiently summarizing meeting minutes using pre-set questions or instructions (prompt sentences).

[0152] The system of the present invention realizes efficient creation of meeting records and real-time search for related information in workplaces where industrial machinery is heavily used. This system records the contents of meetings, converts the recording data into text data, automatically creates a summary of the minutes, and stores the summary in a database. Furthermore, it can search for and propose related information from the database based on issues raised during the meeting. This system allows efficient recording of meeting contents and utilization of past knowledge and solutions in real time.

[0153] Program processing

[0154] At the start of the meeting, the user launches a dedicated recording app on the industrial machine's tablet device and records the conversation. After the meeting ends, the recording data is automatically sent to the server. The server receives it and converts the audio data into text data using the Google Cloud Speech-to-Text API. A summary is then automatically created from the text data using a generative AI model such as OpenAI GPT-4. This generative AI model generates a summary using a prompt sentence. For example, the following prompt sentence can be used:

[0155] Example prompt: "Meeting minutes text: Today's meeting discussed new product development plans. Please write a summary."

[0156] The generated summary is stored in a MySQL database. Furthermore, based on issues recorded during the meeting, the database is searched to find relevant past meeting records and solutions. This involves using a generative AI model to generate relevant information and provide responses in the form of prompts based on the user's request. This relevant information is then presented on the display of the industrial machine, allowing the user to confirm the required information and contact other meeting participants.

[0157] Hardware and Software

[0158] The hardware used includes tablet devices and control devices for industrial machines, and the software used includes a meeting recording application, Google Cloud Speech-to-Text API, OpenAI GPT-4 generative AI models, and a MySQL database.

[0159] Specific examples

[0160] For example, the procedure for a new product development meeting is as follows: At the start of the meeting, the user launches a recording app on their tablet device and records the conversation. After the meeting ends, the recording app sends the audio data to a server, which then converts it into text data using the Google Cloud Speech-to-Text API. A generative AI model (OpenAI GPT-4) extracts key points from this text data and generates a summary of the meeting minutes. This summary is then stored in a MySQL database. Past related information about the new features proposed in the meeting is then searched for and presented on a display device. This information promotes the development of new ideas.

[0161] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0162] Step 1:

[0163] The user starts a dedicated recording app on the tablet device of the industrial machine and records the contents of the meeting. Remarks and discussions during the meeting are recorded as audio data in real time.

[0164] Input: Voice spoken during a meeting

[0165] Output: Recorded audio data file

[0166] Step 2:

[0167] After the conference ends, the device automatically sends the recorded data to the server. This data transfer occurs when the recording application is terminated.

[0168] Input: Recorded audio data file

[0169] Output: Audio data sent to the server

[0170] Step 3:

[0171] The server converts the received voice data into text data using the Google Cloud Speech-to-Text API. The voice data is analyzed and saved as text information.

[0172] Input: Audio data sent to the server

[0173] Output: Data converted to text

[0174] Step 4:

[0175] The server uses a generative AI model (OpenAI GPT-4) to automatically create a summary of the minutes from the converted text data. It uses prompts to extract key points and decisions and generate a summary.

[0176] Input: Text data

[0177] Output: A summary created by the generative AI model

[0178] Example prompt: "Meeting minutes text: Today's meeting discussed new product development plans. Please write a summary."

[0179] Step 5:

[0180] The server stores the generated summaries in a MySQL database, where they can be later stored in a searchable format, along with the original text data.

[0181] Input: Generated summary

[0182] Output: Summary and original text data stored in a database

[0183] Step 6:

[0184] The server searches a database for past meeting records and solutions based on the issues proposed during the meeting, generates relevant information using a generative AI model, and provides responses in the form of prompts tailored to the user's request.

[0185] Input: Issues proposed during the meeting

[0186] Output: Relevant information retrieved from the database

[0187] Step 7:

[0188] The terminal displays the suggested related information on the display device of the industrial machine, allowing the user to check the required information and contact other conference participants.

[0189] Input: Relevant information retrieved from the database

[0190] Output: relevant information presented on a display device

[0191] Processing flow example

[0192] 1. When the meeting starts, the user launches the recording app on the tablet device and records the conversation (step 1).

[0193] 2. After the meeting ends, the app automatically sends the recording data to the server (step 2).

[0194] 3. The server converts the audio data into text using the Google Cloud Speech-to-Text API (step 3).

[0195] 4. The server uses a generative AI model such as OpenAI GPT-4 to automatically create a summary from the generated text data, using the prompt sentence (Step 4).

[0196] 5. The generated summaries and the original text data are stored in a MySQL database (Step 5).

[0197] 6. Based on the issues proposed during the meeting, the server searches the database for past meeting records and solutions, and generates relevant information (step 6).

[0198] 7. The generated related information is presented on the display device of the industrial machine, allowing the user to check the necessary information and contact other conference participants (step 7).

[0199] Through the above steps, the system of the present invention realizes efficient recording of meeting contents and real-time utilization of past knowledge, thereby significantly improving work efficiency at industrial machinery sites.

[0200] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0201] The system according to the present invention records the contents of a meeting, converts the recording data into text data, automatically creates a summary of the minutes, and stores the summary in a database. Furthermore, the system can search for and suggest related information from the database based on the topics proposed during the meeting. Furthermore, by combining it with an emotion engine that recognizes the user's emotions, it can generate minutes with emotional information added and suggest information according to the emotions.

[0202] Specifically, the system operates in the following steps:

[0203] Program processing

[0204] 1. Recording of meeting content

[0205] User: When the meeting starts, the recording device is activated and the conversation is recorded. The user uses a dedicated recording app.

[0206] 2. Sending recording data

[0207] On your device: After the meeting ends, the recording data is automatically sent to the server. The recording app does this automatically.

[0208] 3. Converting Audio Data to Text

[0209] Server: Receives the transmitted audio data. The audio data is temporarily stored.

[0210] Server: Using a speech recognition API, the voice data is converted into text data, which is then temporarily saved.

[0211] 4. Acquiring emotional information

[0212] Server: Analyzes the emotions of meeting participants using an emotion engine from the recorded data. The analyzed emotional information is also temporarily saved.

[0213] 5. Automatic summary creation

[0214] Server: Based on text data and emotional information, the generative AI automatically creates a summary of the minutes, including key points and decisions, as well as analyzed emotional information.

[0215] 6. Save Summary

[0216] Server: The summary of the minutes is saved in a database. Not only the summary but also the original text data and emotional information are saved.

[0217] 7. Search and suggest related information

[0218] Server: Based on the issues recorded by users during meetings, the server searches the database to find relevant past meeting minutes and solutions, and also references emotional information to improve the accuracy of the search.

[0219] Server: The server compiles the searched related information into suggestions and sends them to the device. The suggested information is then displayed to the user.

[0220] 8. Information Display and Contact Functions

[0221] Terminal: Providing suggested relevant information to the user, the user has the option to review the required information and contact other meeting participants.

[0222] User: Uses the suggested information to send a message to relevant meeting participants and schedule a meeting if necessary.

[0223] Specific examples

[0224] For example, when a new product development team holds a meeting, they launch a recording app at the beginning of the meeting. After the meeting ends, the recording is automatically sent to the server. The server receives the recording and converts it into text using a speech recognition API. It then uses an emotion engine to analyze the emotions of the meeting participants.

[0225] The generative AI extracts key points and decisions based on text data and emotional information, and creates a summary. This summary is saved in a database. Next, it searches the database for past discussion topics similar to the new product features proposed during the meeting. By also referring to emotional information, more accurate related information is found. If relevant minutes or ideas from past meetings are found, they are suggested to the user.

[0226] The user can use this information to develop new ideas and contact relevant meeting participants as needed, thus significantly improving the efficiency and productivity of meetings.

[0227] The processing flow will be explained below.

[0228] Step 1:

[0229] When a user starts a meeting, the recording device is activated and the conversation is recorded using a dedicated recording app.

[0230] Step 2:

[0231] When the meeting ends, the device automatically sends the recording data to the server, which is done automatically by the recording app.

[0232] Step 3:

[0233] The server receives the transmitted audio data and stores it temporarily.

[0234] Step 4:

[0235] The server uses a speech recognition API to convert the voice data into text data, which is then temporarily saved.

[0236] Step 5:

[0237] The server uses an emotion engine to analyze the emotions of the conference participants from the received voice data, and the analyzed emotion information is also temporarily saved.

[0238] Step 6:

[0239] The server uses text data and emotional information to automatically generate a summary of the minutes using a generation AI, which extracts emotional information in addition to important points and decisions.

[0240] Step 7:

[0241] The server stores the generated summary in a database, which includes the original text data and sentiment information.

[0242] Step 8:

[0243] The server searches the database based on the issues recorded by the user during the meeting, finding minutes of past meetings and solutions, and also using emotional information to improve the accuracy of the search.

[0244] Step 9:

[0245] The server compiles the retrieved related information into suggestions and sends them to the terminal, where they are displayed to the user.

[0246] Step 10:

[0247] The user reviews the suggested information and selects the information they require. The user has the option to contact other meeting participants.

[0248] Step 11:

[0249] The user can then use the suggested information to send messages to relevant meeting participants and schedule meetings as needed.

[0250] Step 12:

[0251] The server feeds back the utilized information and results to the database, enabling more accurate proposals to be made in future meetings.

[0252] Example 2

[0253] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0254] Conventional meeting recording systems only record meeting content and convert the audio data into text to create minutes, but they do not offer advanced features such as creating minutes that take into account the emotional information of participants during the meeting, or searching and suggesting related information based on proposed issues. This makes it difficult to maximize the efficiency and effectiveness of meetings.

[0255] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for recording the contents of the meeting, means for converting voice data into text data, means for automatically creating a summary of the minutes from the generated text data, means for saving the summary of the minutes in a database, means for searching for and proposing related information from the database based on issues proposed during the meeting, means for displaying the proposed related information and contacting other meeting participants, means for acquiring emotional information and creating a summary of the minutes based on the emotions of the meeting participants, and means for searching for and proposing related information based on the emotional information. This makes it possible to create minutes that take into account not only the content of the meeting but also the emotions of the participants, thereby maximizing the efficiency and effectiveness of the meeting.

[0256] A "meeting recording device" is a device or software that captures and digitally stores audio from a meeting.

[0257] A "means for converting audio data to text data" is software or algorithms for analyzing recorded audio and converting it into a corresponding text format.

[0258] The "means for automatically creating a summary of meeting minutes from generated text data" is a system that analyzes text data, extracts important points and decisions from the meeting, and compiles them into a short summary.

[0259] The "means for storing the summary of the minutes in a database" refers to a data storage system for storing the generated summary of the minutes.

[0260] "A means for searching and proposing relevant information from a database based on issues proposed during a meeting" is a system for searching a database to find relevant information based on issues or questions raised during a meeting.

[0261] The "means for displaying suggested related information and contacting other conference participants" is a system for displaying suggested information to the user as a search result and contacting other conference participants as needed.

[0262] "Means for acquiring emotional information and creating a summary of meeting minutes based on the emotions of meeting participants" is a system that analyzes the emotions of participants from audio and text data during a meeting and creates a summary of meeting minutes based on that information.

[0263] "Means for searching and suggesting related information based on emotional information" is a system that searches for related information based on emotional analysis and makes more appropriate and useful suggestions.

[0264] The system according to the present invention records the contents of a meeting, converts the recording data into text data, automatically creates a summary of the minutes, and stores the summary in a database. It also has a function to search the database for relevant information based on the topics proposed during the meeting and suggest it. It also uses an emotion engine to analyze the user's emotions, and can create minutes and suggest relevant information based on the emotional information.

[0265] Hardware and software used

[0266] Hardware

[0267] Recording device: Smartphone or dedicated recorder

[0268] Server: A high-performance computer for data processing

[0269] Devices: Laptops, tablets, smartphones

[0270] software

[0271] Recording app: An application that records the contents of a meeting and sends the data to a server.

[0272] Speech Recognition API: Google Cloud Speech-to-Text

[0273] Emotion engine: IBM Watson Tone Analyzer

[0274] Generative AI model: OpenAI GPT

[0275] Database system: A relational database management system (RDBMS) for storing data.

[0276] Program Overview

[0277] 1. The user launches a recording app at the start of a meeting and records the contents of the meeting.

[0278] 2. After the meeting ends, the device automatically sends the recording data to the server.

[0279] 3. The server converts the received voice data into text using the Google Cloud Speech-to-Text API.

[0280] 4. The server runs IBM Watson Tone Analyzer on the text and audio data to obtain emotion information.

[0281] 5. The server uses OpenAI GPT to automatically create a summary of the minutes based on the text data and emotional information.

[0282] 6. The server stores the generated summary of the minutes, the original text data, and the emotion information in a database.

[0283] 7. The server searches the database based on the issues described during the meeting and suggests relevant information. The suggestions also refer to emotional information.

[0284] 8. The terminal displays the suggested relevant information to the user, providing the user with the option to review the required information and contact other conference participants as needed.

[0285] Specific operation example

[0286] Working example 1:

[0287] When a new product development team holds a meeting, they follow these steps:

[0288] 1. The user launches the recording app at the start of the meeting.

[0289] 2. After the meeting ends, the recording data is automatically sent to the server.

[0290] 3. The server receives the recording and converts it into text using Google Cloud Speech-to-Text.

[0291] 4. Additionally, IBM Watson Tone Analyzer is used to analyze the emotions of meeting participants.

[0292] 5. Create a summary of meeting minutes from text data and sentiment information using OpenAI GPT.

[0293] 6. The summary, original text data, and emotion information are stored in a database.

[0294] 7. The server searches the database based on the topic at hand and suggests relevant information.

[0295] 8. The terminal displays the information to the user and also provides options for contacting other conference participants.

[0296] Prompt Sentence Examples

[0297] A new product development meeting is being held. The recording app is launched to record the conversation, and after the meeting ends, the recording is automatically sent to the server. The server converts the text using Google Cloud Speech-to-Text and performs sentiment analysis using IBM Watson Tone Analyzer. A summary is created using OpenAI GPT and stored in a database. Related information is also searched and suggested. For example, the system searches past meeting minutes and related information to suggest new product features discussed during the meeting.

[0298] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0299] Step 1:

[0300] User: At the start of a meeting, the user launches a recording app. Specifically, the user taps the "Start Recording" button on the recording app installed on their smartphone or tablet. The input is the audio from the meeting, and the output is an audio file saved as recording data in the internal storage.

[0301] Step 2:

[0302] Device: After the meeting ends, the recording data is automatically sent to the server. When the user taps the "End Meeting" button, the app uploads the recording data to the server via the Internet. The input is the audio file stored in the device's internal storage, and the output is the audio file transferred to the server. Specifically, the file is uploaded using an HTTPS request.

[0303] Step 3:

[0304] Server: Receives and temporarily stores recorded data. Specifically, the server stores uploaded audio files in a storage directory. The input is the audio file sent from the device, and the output is the audio data stored in the server's internal storage.

[0305] Step 4:

[0306] Server: Uses a speech recognition API (Google Cloud Speech-to-Text) to convert voice data into text data. The server passes the voice data to the API and receives the converted text data. The input is an audio file as recorded data, and the output is conversation data in text format. Specifically, it sends requests to the API and receives the results.

[0307] Step 5:

[0308] Server: Analyzes the emotions of meeting participants from the recording data using an emotion engine (IBM Watson Tone Analyzer). The server sends text data to the API and receives the emotion information resulting from the analysis. The input is text data and the output is emotion information. Specifically, it sends requests to the emotion engine API and receives the results.

[0309] Step 6:

[0310] Server: Automatically creates a summary of meeting minutes based on text data and emotional information using a generative AI model (OpenAI GPT). The input is text data and emotional information, and the output is the generated summary. Specifically, the input data is passed to the generative AI model to generate the summary.

[0311] Step 7:

[0312] Server: The created summary of the minutes is saved in a database. The original text data and emotion information are also saved in the database. The input is the generated summary, text data, and emotion information, and the output is each data entry saved in the database. Specifically, data is added to the database using SQL statements.

[0313] Step 8:

[0314] Server: Searches the database based on the issues recorded by the user during the meeting and suggests related information. The server receives the user's issue information as input and searches for related entries in the database. The input is the issue information, and the output is related information as a search result. Specifically, it generates and executes the database query.

[0315] Step 9:

[0316] Terminal: Displays suggested relevant information to the user. The user can check the information through the app interface and contact other meeting participants if necessary. The input is the relevant information sent from the server, and the output is the information displayed on the user interface. Specifically, the terminal receives data from the API and displays it on the UI.

[0317] Step 10:

[0318] User: Based on the suggested information, send a message to relevant meeting participants and schedule a meeting if necessary. This is executed when the user enters a message in the app and taps the "Send" button. The input is the suggested related information and the message entered by the user, and the output is the sent message. Specifically, the message is sent using the message sending API.

[0319] (Application example 2)

[0320] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0321] Conventional meeting recording and automatic minutes creation systems have difficulty efficiently searching and suggesting emotional information and related past meeting information that can be used to solve problems in factories and other workplaces. Furthermore, the generated minutes summaries lack emotional information and provide insufficient detailed information that reflects the atmosphere of the discussion and the emotions of the meeting participants. This hinders rapid problem solving and efficient information sharing, especially in the workplace.

[0322] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes: means for recording the contents of the meeting; means for converting audio data into text data; means for automatically creating a summary of the minutes from the generated text data; means for saving the summary of the minutes in a database; means for searching and proposing related information from the database based on issues proposed during the meeting; means for displaying the proposed related information and contacting other meeting participants; means for recognizing emotional information and analyzing the emotions of the meeting participants; means for adding the analyzed emotional information to the summary; and means for providing meeting information for problem solving to machines installed in a factory. This enables accurate recording of the contents of the meeting and generation of a detailed summary including emotional information. Furthermore, by searching for and proposing related information based on issues proposed during the meeting, rapid problem solving and efficient information sharing can be achieved on-site.

[0323] "Means for recording the contents of a conference" refers to a device or method for recording the audio during a conference as digital data.

[0324] "Means for converting voice data into text data" refers to voice recognition technology or software for converting recorded voice data into text information.

[0325] The "means for automatically creating a summary of minutes from the generated text data" refers to an algorithm or program that summarizes the text data, extracts important information, and generates an outline of the minutes.

[0326] The "means for storing the summary of the minutes in a database" refers to a method or device for recording the generated summary of the minutes in an information storage system called a database.

[0327] "Means for searching and suggesting relevant information from a database based on an issue raised during a meeting" refers to a method or device for automatically searching a database for information related to an issue raised during a meeting and presenting that information to a user.

[0328] "Means for displaying suggested related information and contacting other conference participants" refers to an interface or method for displaying the searched information on a display or the like and for contacting other conference participants as necessary.

[0329] "Means for recognizing emotional information and analyzing the emotions of meeting participants" refers to technology and software for estimating and analyzing the emotional state of participants from the content of statements made during a meeting and their tone of voice.

[0330] The "means for adding analyzed emotional information to a summary" refers to a method or program for adding the results of analyzing the emotions of meeting participants to the summary of the minutes.

[0331] The "means for providing meeting information for problem solving to machines installed in the factory" refers to an interface or method for linking meeting minutes and related information to machines and management systems in the factory.

[0332] The program for implementing this invention records the audio of meetings and problem-solving meetings within a factory, processes the data to generate a summary of the minutes, and then performs emotion analysis to support efficient information sharing and problem solving. The specific processing and the hardware and software used are described below.

[0333] Hardware and Software Configuration

[0334] The server uses the following software:

[0335] Speech recognition API: A technology for converting voice data into text data. A typical example is Google's speech recognition API.

[0336] Emotion recognition engine: A technology for analyzing the emotions of meeting participants. The EmotionRecognition library is one example.

[0337] Generative AI model: A technology for generating a summary of meeting minutes from text data. A generative AI model using the Transformer library falls into this category.

[0338] Database: A system for storing generated minutes summaries and related information. It uses an SQLite database.

[0339] System processing overview

[0340] The server receives the audio data of the meeting recorded by the user and converts it into text data using a speech recognition API. The converted text data is then temporarily stored. Next, an emotion recognition engine is used to analyze emotional information from the text data, which is also temporarily stored. A generative AI model is then used to generate a summary of the minutes based on the text data and emotional information. This summary includes key points, decisions, and emotional information. The generated summary is then stored in a database.

[0341] Based on the topic proposed during the meeting, the system searches for related past information from the database. The server also refers to emotional information to improve the accuracy of the search. The retrieved related information is compiled as suggestions and sent to the user's device. The suggested information is displayed to the user, allowing them to contact other meeting participants.

[0342] This information is provided to machines and management systems within the factory, helping to quickly resolve problems on site.

[0343] Specific example explanation

[0344] For example, consider a meeting to improve the production efficiency of a new product. During the meeting, a recording device is running and the audio data is automatically sent to a server. The server converts the audio data into text data and performs sentiment analysis. A generative AI model uses the text data and sentiment information to generate a summary of the meeting minutes, including key points and decisions. This summary is stored in a database, and can be quickly searched and suggested when related information is needed in the future.

[0345] The minutes of past discussions regarding the "new production setup" problem proposed during this meeting can be referenced, allowing users to find effective solutions and contact other participants if necessary.

[0346] Prompt Sentence Examples

[0347] Here are some example prompts to input to a generative AI model:

[0348] "Based on the minutes of previous meetings, please tell us how discussions regarding bottlenecks on production lines have gone in the past."

[0349] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0350] Step 1:

[0351] At the start of a conference, a user activates a recording device to record the contents of the conference. The recording device records the conversation as digital audio data and temporarily stores the data. The input is the conversation, and the output is digital audio data.

[0352] Step 2:

[0353] After the meeting ends, the device automatically sends the recording data to the server. Communication between the recording device and the server is performed using a secure protocol. The input here is digital audio data, and the output is data transmission to the server.

[0354] Step 3:

[0355] The server temporarily stores the received voice data. Then, it converts the voice data into text data using a speech recognition API (for example, Google's speech recognition service). The input here is the received voice data, and the output is the converted text data.

[0356] Step 4:

[0357] The server uses an emotion recognition engine (e.g., the EmotionRecognition library) based on the text data to analyze the emotional information of the conference participants. This analysis result is also temporarily saved. The input is text data, and the output is the emotion analysis result.

[0358] Step 5:

[0359] The server uses a generative AI model (e.g., the Transformer library) to automatically generate a summary of the minutes from the text data and emotional information. The generated summary includes emotional information as well as key points and decisions. This summary is stored in a database. The input is text data and emotional information, and the output is a summary of the minutes.

[0360] Step 6:

[0361] The server searches the database for relevant past information based on the topic proposed during the meeting. Emotional information is also referenced to improve the accuracy of the search. The input is the proposed topic and emotional information, and the output is relevant past information.

[0362] Step 7:

[0363] The server summarizes the retrieved related information as suggestions and sends them to the terminal. The suggested information is displayed on the user's terminal. The input is related past information, and the output is the suggested information displayed on the user's terminal.

[0364] Step 8:

[0365] The user reviews the proposed information and contacts other meeting participants as needed, allowing them to send messages to related meeting participants and set up new meetings. The input is the proposed information, and the output is contact with participants and meeting settings.

[0366] The above is the specific processing flow of the present invention.

[0367] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0368] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0369] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0370] [Second embodiment]

[0371] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0372] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0373] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0374] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0375] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0376] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0377] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0378] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0379] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0380] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0381] In the smart glasses 214, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0382] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0383] The system of the present invention records the contents of a meeting, converts the recording data into text data, automatically creates a summary of the minutes, and stores the summary in a database. Furthermore, it can search for and propose related information from the database based on issues raised during the meeting. This system allows for efficient recording of meeting contents and makes it possible to utilize past knowledge and solutions in real time.

[0384] Specifically, the system operates in the following steps:

[0385] Program processing

[0386] 1. Recording of meeting content

[0387] User: Start the recording device as soon as the meeting starts and record the conversation using a dedicated recording app.

[0388] 2. Sending recording data

[0389] Device: After the meeting ends, the recording data is sent to the server. The recording app automatically sends the data.

[0390] 3. Converting Audio Data to Text

[0391] Server: Receives the transmitted voice data and converts it into text data using a speech recognition API. The converted text data is temporarily saved.

[0392] 4. Automatic summary creation

[0393] Server: Based on the text data converted from the voice data, the generative AI extracts important points and decisions, and automatically creates a summary of the minutes.

[0394] 5. Save summary

[0395] Server: Saves the summary of the minutes to a database. The database stores not only the summary but also the original text data.

[0396] 6. Search and suggest related information

[0397] Server: Based on the issues recorded by users during the meeting, the server searches the database to find relevant past minutes and solutions.

[0398] Server: Compiles relevant information into suggestions and sends them to the user's device.

[0399] 7. Information Display and Contact Functions

[0400] Terminal: Providing suggested relevant information to the user, the user has the option to review the required information and contact other meeting participants.

[0401] User: Uses the suggested information to send a message to relevant meeting participants and schedule a meeting if necessary.

[0402] Specific examples

[0403] For example, when a new product development team holds a meeting, they launch a dedicated recording app at the start of the meeting. After the meeting ends, the recording data is automatically sent to a server. The server receives the data and converts it into text. Generative AI extracts key points from the text data, creates a summary, and stores it in a database.

[0404] Next, the system searches the database for past discussion topics related to new product features proposed during the meeting. If relevant past meeting minutes or ideas are found, they are suggested to the user, who can use this information to develop new ideas.

[0405] Thus, the present invention can significantly improve the efficiency and productivity of meetings.

[0406] The processing flow will be explained below.

[0407] Step 1:

[0408] When a user starts a meeting, the recording device is activated and the conversation is recorded using a dedicated recording app.

[0409] Step 2:

[0410] When the meeting ends, the device automatically sends the recording data to the server, which is done automatically by the recording app.

[0411] Step 3:

[0412] The server receives the transmitted audio data and stores it temporarily.

[0413] Step 4:

[0414] The server uses a speech recognition API to convert the voice data into text data, which is then temporarily saved.

[0415] Step 5:

[0416] The server inputs the text data into the AI ​​generator, which automatically creates a summary of the minutes, extracting key points and decisions.

[0417] Step 6:

[0418] The server saves the created summary in a database, along with the original text data.

[0419] Step 7:

[0420] Based on the issues users record during meetings, the server searches the database to find minutes of past meetings and solutions.

[0421] Step 8:

[0422] The server compiles relevant information into suggestions and sends them to the terminal, where they are displayed to the user.

[0423] Step 9:

[0424] The user reviews the suggested information and selects the information they require. The user has the option to contact other meeting participants.

[0425] Step 10:

[0426] The user can then use the suggested information to send messages to relevant meeting participants and schedule meetings as needed.

[0427] Step 11:

[0428] The server feeds back the utilized information and results to the database, enabling more accurate proposals to be made in future meetings.

[0429] Example 1

[0430] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0431] Conventional meeting recording systems often require manual recording and management of meeting content, which requires time and effort. It is also difficult to quickly search past minutes and related information, making it difficult to receive immediate and effective feedback on issues raised during meetings. Furthermore, they lack the ability to share relevant information with other meeting participants, potentially reducing the efficiency and productivity of meetings.

[0432] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0433] In this invention, the server includes means for recording the contents of the meeting, means for converting the audio data into text data, means for automatically creating a summary of the minutes from the generated text data, means for saving the summary of the minutes in a database, means for searching and proposing related information from the database based on issues proposed during the meeting, means for displaying the proposed related information and providing a means for contacting other meeting participants, means for extracting important points from the text data using a generation AI, and means for transmitting the recorded data to the server. This automates the recording and management of meetings and enables past minutes and related information to be quickly searched and shared, thereby improving the efficiency and productivity of meetings.

[0434] "Means for recording the contents of a meeting" refers to a device or application for recording statements and conversations made during a meeting as audio data.

[0435] "Means for converting voice data into text data" refers to voice recognition technology or software for analyzing recorded voice data and converting it into text data.

[0436] The "means for automatically creating a summary of meeting minutes from generated text data" is a generative AI system that analyzes text data to extract and summarize important points and decisions.

[0437] The "means for storing the summary of the minutes in a database" refers to a device or system for storing the automatically created summary of the minutes in a relational database, cloud storage, or the like.

[0438] "A means for searching and proposing relevant information from a database based on issues proposed during a meeting" is an engine that searches a database for past minutes and solutions related to the issues entered by the user, extracts appropriate information, and presents it.

[0439] "Means for displaying suggested related information and providing a means for contacting other conference participants" refers to functionality or software that displays related information on a user interface and enables users to send messages to other conference participants or schedule meetings.

[0440] "Means for extracting key points from text data using generative AI" refers to technology that uses a generative AI model (e.g., a natural language processing model) to mechanically extract key points and decisions from text data.

[0441] The "means for transmitting recorded data to a server" refers to a communication function or protocol for transmitting recorded voice data to a server via the Internet.

[0442] This invention is a system that records the contents of a meeting, converts the recorded data into text data, automatically creates a summary of the minutes, stores it in a database, and searches for and suggests related information.

[0443] First, users turn on the recording device at the start of a meeting and record the conversation using a dedicated recording app, which can be software installed on a smartphone, tablet, or PC, such as Otter.ai or a similar application.

[0444] Then, after the meeting ends, the device sends the recording data to the server. The recording app does this automatically, and the data is sent using an encryption protocol (e.g., TLS). The data transmission is completed without any special operation by the user.

[0445] The server sends the received voice data to the Google Cloud Speech-to-Text API, which converts the voice data into text data. The converted text data is temporarily stored on the server. After conversion is complete, the server uses a generative AI (e.g., OpenAI GPT-3) to extract important points and decisions from the text data and create a summary of the meeting minutes.

[0446] The server stores the generated summary in a database (e.g., MySQL, PostgreSQL). The database stores not only the summary of the minutes but also the original text data, ensuring the integrity of the information.

[0447] Furthermore, based on the issues proposed during the meeting, the server searches the database and provides related past meeting minutes and solutions. This search uses an SQL query such as "SELECT FROM summaries WHERE text LIKE '%keywords related to the issue%'". The server compiles the relevant information found as suggestions and sends them to the user's device. This is sent in JSON format, and a dedicated suggestion display app receives and displays it.

[0448] The device displays the suggested related information to the user. The user can check the displayed information and, if necessary, send a message to other meeting participants or schedule a new meeting. By clicking the "Contact" button in the suggestion display app, a message sending screen will appear, allowing the user to send a message to the relevant participants.

[0449] Specific examples

[0450] When a new product development team holds a meeting, the user launches a dedicated recording app at the start of the meeting to record the conversation. After the meeting ends, the recorded data is automatically sent from the device to the server. The server receives this data and converts the speech to text using the Google Cloud Speech-to-Text API. OpenAI GPT-3 extracts key points from the converted text data and generates a summary of the meeting minutes. The generated summary is stored in a MySQL database.

[0451] Next, for new product features proposed during a meeting, the server searches the database to see if similar topics have been discussed in the past. If relevant minutes or ideas from past meetings are found, the server compiles them into a proposal and sends them in JSON format to the user's device. The user can then view the information through the proposal display app, send messages to participants in related meetings, and schedule new meetings.

[0452] Prompt Sentence Examples

[0453] "Please extract the key points and decisions from the following text data and create a summary of the meeting minutes.

[0454] Text data:

[0455] Proposing new product features at a meeting

[0456] Supervisor approves

[0457] We plan to work out the details at the next meeting.

[0458] ...

[0459] "

[0460] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0461] Step 1:

[0462] When the meeting starts, the user activates the recording device and records the conversation using a dedicated recording app. The recording app is installed on a smartphone, tablet, or PC, and an example is "Otter.ai." The user starts recording by pressing the "Start Recording" button. The input is an audio signal, and the output is voice data.

[0463] Step 2:

[0464] After the meeting ends, the device sends the recorded data to the server. When the recording app detects the end of the meeting, it automatically compresses the recorded data and sends it to the server using an encryption protocol (e.g., TLS). The input is the audio data, and the output is the data sent to the server.

[0465] Step 3:

[0466] The server stores the received voice data and sends it to the Google Cloud Speech-to-Text API to convert it into text. This API analyzes the voice data and outputs it as a string. The input is voice data and the output is text data.

[0467] Step 4:

[0468] The server temporarily stores the text data received from the Google Cloud Speech-to-Text API. It then supplies the text data to a generative AI (e.g., OpenAI GPT-3) with a prompt: "Please extract the key points and decisions from the following text data and create a summary of the minutes." The generative AI extracts the key points and generates a summary of the minutes. The input is the text data and the prompt, and the output is a summary of the minutes.

[0469] Step 5:

[0470] The server saves the generated summary of the minutes to a database. The server first establishes a database connection and executes the SQL query "INSERT INTO summaries (text, summary) VALUES (?, ?)". The original text data is also saved to the same database. The input is the summary of the minutes and the text data, and the output is the information saved in the database.

[0471] Step 6:

[0472] The server searches the database based on the issues recorded by the user during the meeting. For example, it runs the SQL query "SELECT FROM summaries WHERE text LIKE '%keywords related to the issue%'". The server filters the relevant information found and extracts the relevant information. The input is the issue keywords and the output is the relevant information.

[0473] Step 7:

[0474] The server compiles related information as suggestions and sends them to the user's device. The suggestions are sent in JSON format, and a dedicated suggestion display app receives them and displays them on the user interface. The input is the related information, and the output is the data sent to the device.

[0475] Step 8:

[0476] The device displays the suggested related information to the user. The suggestion display app parses the received JSON data and presents it to the user in a list format. The user can click on an interesting suggestion from the list to view details. The input is JSON data, and the output is the information displayed to the user.

[0477] Step 9:

[0478] Based on the proposed information, the user can send a message to the relevant meeting participants and set up a new meeting. For example, clicking the "Contact" button in the proposal display app will display the message sending screen, allowing the user to send a message to the relevant participants. The input is the user's action and the proposed information, and the output is the sent message.

[0479] (Application example 1)

[0480] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0481] Currently, in workplaces that use industrial machinery, it is difficult to efficiently create meeting records and utilize past related information in real time. Furthermore, there is a lack of a method to efficiently search and propose solutions to issues raised during meetings, which leads to inefficiency and reduced productivity. Furthermore, in industrial machinery workplaces, there is a need for a method to directly reflect the information in meeting minutes in machine operation.

[0482] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0483] In this invention, the server includes: means for recording the contents of a meeting; means for converting voice data into text data; means for automatically creating a summary of the minutes from the generated text data; means for saving the summary of the minutes in a database; means for searching and proposing related information from the database based on issues proposed during the meeting; means for displaying the proposed related information and providing a means for contacting other meeting participants; means for automatically generating meeting minutes and presenting search results for related information on the display device of the industrial machine in cooperation with the industrial machine control device; means for generating the proposed related information using a generative AI model; and means for generating a summary of the minutes using prompt sentences. This enables efficient recording of meeting contents and real-time utilization of past meeting information, thereby improving work efficiency and productivity at industrial machine sites.

[0484] A "means for recording the contents of a meeting" is a device or software that can record the audio content spoken during a meeting.

[0485] "Means for converting voice data into text data" refers to algorithms or tools for analyzing recorded voice data and converting it into text information.

[0486] The "means for automatically creating a summary of minutes from generated text data" is a means for extracting important points and decisions based on text data converted from speech and generating a summary.

[0487] The "means for saving the summary of the minutes in a database" refers to a function or device for writing and saving the generated summary in a database.

[0488] "A means for searching and proposing relevant information from a database based on issues proposed during a meeting" is a means for inputting the agenda items or issues raised during a meeting, searching a database for past information related to them, and presenting it.

[0489] "Means for displaying suggested related information and providing means for contacting other conference participants" refers to means for displaying searched related information in a user interface and providing options for communicating with other conference participants.

[0490] "Means for automatically generating meeting minutes and presenting search results for related information on the display device of the industrial machine in cooperation with the control device of the industrial machine" refers to means for exchanging data with the control device of the industrial machine and visually displaying the meeting minutes and related information on the display device of the machine.

[0491] A "means for generating suggested relevant information using a generative AI model" is a means for using an artificial intelligence model to automatically generate highly relevant information or solutions based on input data.

[0492] The "means for generating a summary of meeting minutes using prompt sentences" is a means for efficiently summarizing meeting minutes using pre-set questions or instructions (prompt sentences).

[0493] The system of the present invention realizes efficient creation of meeting records and real-time search for related information in workplaces where industrial machinery is heavily used. This system records the contents of meetings, converts the recording data into text data, automatically creates a summary of the minutes, and stores the summary in a database. Furthermore, it can search for and propose related information from the database based on issues raised during the meeting. This system allows efficient recording of meeting contents and utilization of past knowledge and solutions in real time.

[0494] Program processing

[0495] At the start of the meeting, the user launches a dedicated recording app on the industrial machine's tablet device and records the conversation. After the meeting ends, the recording data is automatically sent to the server. The server receives it and converts the audio data into text data using the Google Cloud Speech-to-Text API. A summary is then automatically created from the text data using a generative AI model such as OpenAI GPT-4. This generative AI model generates a summary using a prompt sentence. For example, the following prompt sentence can be used:

[0496] Example prompt: "Meeting minutes text: Today's meeting discussed new product development plans. Please write a summary."

[0497] The generated summary is stored in a MySQL database. Furthermore, based on issues recorded during the meeting, the database is searched to find relevant past meeting records and solutions. This involves using a generative AI model to generate relevant information and provide responses in the form of prompts based on the user's request. This relevant information is then presented on the display of the industrial machine, allowing the user to confirm the required information and contact other meeting participants.

[0498] Hardware and Software

[0499] The hardware used includes tablet devices and control devices for industrial machines, and the software used includes a meeting recording application, Google Cloud Speech-to-Text API, OpenAI GPT-4 generative AI models, and a MySQL database.

[0500] Specific examples

[0501] For example, the procedure for a new product development meeting is as follows: At the start of the meeting, the user launches a recording app on their tablet device and records the conversation. After the meeting ends, the recording app sends the audio data to a server, which then converts it into text data using the Google Cloud Speech-to-Text API. A generative AI model (OpenAI GPT-4) extracts key points from this text data and generates a summary of the meeting minutes. This summary is then stored in a MySQL database. Past related information about the new features proposed in the meeting is then searched for and presented on a display device. This information promotes the development of new ideas.

[0502] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0503] Step 1:

[0504] The user starts a dedicated recording app on the tablet device of the industrial machine and records the contents of the meeting. Remarks and discussions during the meeting are recorded as audio data in real time.

[0505] Input: Voice spoken during a meeting

[0506] Output: Recorded audio data file

[0507] Step 2:

[0508] After the conference ends, the device automatically sends the recorded data to the server. This data transfer occurs when the recording application is terminated.

[0509] Input: Recorded audio data file

[0510] Output: Audio data sent to the server

[0511] Step 3:

[0512] The server converts the received voice data into text data using the Google Cloud Speech-to-Text API. The voice data is analyzed and saved as text information.

[0513] Input: Audio data sent to the server

[0514] Output: Data converted to text

[0515] Step 4:

[0516] The server uses a generative AI model (OpenAI GPT-4) to automatically create a summary of the minutes from the converted text data. It uses prompts to extract key points and decisions and generate a summary.

[0517] Input: Text data

[0518] Output: A summary created by the generative AI model

[0519] Example prompt: "Meeting minutes text: Today's meeting discussed new product development plans. Please write a summary."

[0520] Step 5:

[0521] The server stores the generated summaries in a MySQL database, where they can be later stored in a searchable format, along with the original text data.

[0522] Input: Generated summary

[0523] Output: Summary and original text data stored in a database

[0524] Step 6:

[0525] The server searches a database for past meeting records and solutions based on the issues proposed during the meeting, generates relevant information using a generative AI model, and provides responses in the form of prompts tailored to the user's request.

[0526] Input: Issues proposed during the meeting

[0527] Output: Relevant information retrieved from the database

[0528] Step 7:

[0529] The terminal displays the suggested related information on the display device of the industrial machine, allowing the user to check the required information and contact other conference participants.

[0530] Input: Relevant information retrieved from the database

[0531] Output: relevant information presented on a display device

[0532] Processing flow example

[0533] 1. When the meeting starts, the user launches the recording app on the tablet device and records the conversation (step 1).

[0534] 2. After the meeting ends, the app automatically sends the recording data to the server (step 2).

[0535] 3. The server converts the audio data into text using the Google Cloud Speech-to-Text API (step 3).

[0536] 4. The server uses a generative AI model such as OpenAI GPT-4 to automatically create a summary from the generated text data, using the prompt sentence (Step 4).

[0537] 5. The generated summaries and the original text data are stored in a MySQL database (Step 5).

[0538] 6. Based on the issues proposed during the meeting, the server searches the database for past meeting records and solutions, and generates relevant information (step 6).

[0539] 7. The generated related information is presented on the display device of the industrial machine, allowing the user to check the necessary information and contact other conference participants (step 7).

[0540] Through the above steps, the system of the present invention realizes efficient recording of meeting contents and real-time utilization of past knowledge, thereby significantly improving work efficiency at industrial machinery sites.

[0541] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0542] The system according to the present invention records the contents of a meeting, converts the recording data into text data, automatically creates a summary of the minutes, and stores the summary in a database. Furthermore, the system can search for and suggest related information from the database based on the topics proposed during the meeting. Furthermore, by combining it with an emotion engine that recognizes the user's emotions, it can generate minutes with emotional information added and suggest information according to the emotions.

[0543] Specifically, the system operates in the following steps:

[0544] Program processing

[0545] 1. Recording of meeting content

[0546] User: When the meeting starts, the recording device is activated and the conversation is recorded. The user uses a dedicated recording app.

[0547] 2. Sending recording data

[0548] On your device: After the meeting ends, the recording data is automatically sent to the server. The recording app does this automatically.

[0549] 3. Converting Audio Data to Text

[0550] Server: Receives the transmitted audio data. The audio data is temporarily stored.

[0551] Server: Using a speech recognition API, the voice data is converted into text data, which is then temporarily saved.

[0552] 4. Acquiring emotional information

[0553] Server: Analyzes the emotions of meeting participants using an emotion engine from the recorded data. The analyzed emotional information is also temporarily saved.

[0554] 5. Automatic summary creation

[0555] Server: Based on text data and emotional information, the generative AI automatically creates a summary of the minutes, including key points and decisions, as well as analyzed emotional information.

[0556] 6. Save Summary

[0557] Server: The summary of the minutes is saved in a database. Not only the summary but also the original text data and emotional information are saved.

[0558] 7. Search and suggest related information

[0559] Server: Based on the issues recorded by users during meetings, the server searches the database to find relevant past meeting minutes and solutions, and also references emotional information to improve the accuracy of the search.

[0560] Server: The server compiles the searched related information into suggestions and sends them to the device. The suggested information is then displayed to the user.

[0561] 8. Information Display and Contact Functions

[0562] Terminal: Providing suggested relevant information to the user, the user has the option to review the required information and contact other meeting participants.

[0563] User: Uses the suggested information to send a message to relevant meeting participants and schedule a meeting if necessary.

[0564] Specific examples

[0565] For example, when a new product development team holds a meeting, they launch a recording app at the beginning of the meeting. After the meeting ends, the recording is automatically sent to the server. The server receives the recording and converts it into text using a speech recognition API. It then uses an emotion engine to analyze the emotions of the meeting participants.

[0566] The generative AI extracts key points and decisions based on text data and emotional information, and creates a summary. This summary is saved in a database. Next, it searches the database for past discussion topics similar to the new product features proposed during the meeting. By also referring to emotional information, more accurate related information is found. If relevant minutes or ideas from past meetings are found, they are suggested to the user.

[0567] The user can use this information to develop new ideas and contact relevant meeting participants as needed, thus significantly improving the efficiency and productivity of meetings.

[0568] The processing flow will be explained below.

[0569] Step 1:

[0570] When a user starts a meeting, the recording device is activated and the conversation is recorded using a dedicated recording app.

[0571] Step 2:

[0572] When the meeting ends, the device automatically sends the recording data to the server, which is done automatically by the recording app.

[0573] Step 3:

[0574] The server receives the transmitted audio data and stores it temporarily.

[0575] Step 4:

[0576] The server uses a speech recognition API to convert the voice data into text data, which is then temporarily saved.

[0577] Step 5:

[0578] The server uses an emotion engine to analyze the emotions of the conference participants from the received voice data, and the analyzed emotion information is also temporarily saved.

[0579] Step 6:

[0580] The server uses text data and emotional information to automatically generate a summary of the minutes using a generation AI, which extracts emotional information in addition to important points and decisions.

[0581] Step 7:

[0582] The server stores the generated summary in a database, which includes the original text data and sentiment information.

[0583] Step 8:

[0584] The server searches the database based on the issues recorded by the user during the meeting, finding minutes of past meetings and solutions, and also using emotional information to improve the accuracy of the search.

[0585] Step 9:

[0586] The server compiles the retrieved related information into suggestions and sends them to the terminal, where they are displayed to the user.

[0587] Step 10:

[0588] The user reviews the suggested information and selects the information they require. The user has the option to contact other meeting participants.

[0589] Step 11:

[0590] The user can then use the suggested information to send messages to relevant meeting participants and schedule meetings as needed.

[0591] Step 12:

[0592] The server feeds back the utilized information and results to the database, enabling more accurate proposals to be made in future meetings.

[0593] Example 2

[0594] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0595] Conventional meeting recording systems only record meeting content and convert the audio data into text to create minutes, but they do not offer advanced features such as creating minutes that take into account the emotional information of participants during the meeting, or searching and suggesting related information based on proposed issues. This makes it difficult to maximize the efficiency and effectiveness of meetings.

[0596] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for recording the contents of the meeting, means for converting voice data into text data, means for automatically creating a summary of the minutes from the generated text data, means for saving the summary of the minutes in a database, means for searching for and proposing related information from the database based on issues proposed during the meeting, means for displaying the proposed related information and contacting other meeting participants, means for acquiring emotional information and creating a summary of the minutes based on the emotions of the meeting participants, and means for searching for and proposing related information based on the emotional information. This makes it possible to create minutes that take into account not only the content of the meeting but also the emotions of the participants, thereby maximizing the efficiency and effectiveness of the meeting.

[0597] A "meeting recording device" is a device or software that captures and digitally stores audio from a meeting.

[0598] A "means for converting audio data to text data" is software or algorithms for analyzing recorded audio and converting it into a corresponding text format.

[0599] The "means for automatically creating a summary of meeting minutes from generated text data" is a system that analyzes text data, extracts important points and decisions from the meeting, and compiles them into a short summary.

[0600] The "means for storing the summary of the minutes in a database" refers to a data storage system for storing the generated summary of the minutes.

[0601] "A means for searching and proposing relevant information from a database based on issues proposed during a meeting" is a system for searching a database to find relevant information based on issues or questions raised during a meeting.

[0602] The "means for displaying suggested related information and contacting other conference participants" is a system for displaying suggested information to the user as a search result and contacting other conference participants as needed.

[0603] "Means for acquiring emotional information and creating a summary of meeting minutes based on the emotions of meeting participants" is a system that analyzes the emotions of participants from audio and text data during a meeting and creates a summary of meeting minutes based on that information.

[0604] "Means for searching and suggesting related information based on emotional information" is a system that searches for related information based on emotional analysis and makes more appropriate and useful suggestions.

[0605] The system according to the present invention records the contents of a meeting, converts the recording data into text data, automatically creates a summary of the minutes, and stores the summary in a database. It also has a function to search the database for relevant information based on the topics proposed during the meeting and suggest it. It also uses an emotion engine to analyze the user's emotions, and can create minutes and suggest relevant information based on the emotional information.

[0606] Hardware and software used

[0607] Hardware

[0608] Recording device: Smartphone or dedicated recorder

[0609] Server: A high-performance computer for data processing

[0610] Devices: Laptops, tablets, smartphones

[0611] software

[0612] Recording app: An application that records the contents of a meeting and sends the data to a server.

[0613] Speech Recognition API: Google Cloud Speech-to-Text

[0614] Emotion engine: IBM Watson Tone Analyzer

[0615] Generative AI model: OpenAI GPT

[0616] Database system: A relational database management system (RDBMS) for storing data.

[0617] Program Overview

[0618] 1. The user launches a recording app at the start of a meeting and records the contents of the meeting.

[0619] 2. After the meeting ends, the device automatically sends the recording data to the server.

[0620] 3. The server converts the received voice data into text using the Google Cloud Speech-to-Text API.

[0621] 4. The server runs IBM Watson Tone Analyzer on the text and audio data to obtain emotion information.

[0622] 5. The server uses OpenAI GPT to automatically create a summary of the minutes based on the text data and emotional information.

[0623] 6. The server stores the generated summary of the minutes, the original text data, and the emotion information in a database.

[0624] 7. The server searches the database based on the issues described during the meeting and suggests relevant information. The suggestions also refer to emotional information.

[0625] 8. The terminal displays the suggested relevant information to the user, providing the user with the option to review the required information and contact other conference participants as needed.

[0626] Specific operation example

[0627] Working example 1:

[0628] When a new product development team holds a meeting, they follow these steps:

[0629] 1. The user launches the recording app at the start of the meeting.

[0630] 2. After the meeting ends, the recording data is automatically sent to the server.

[0631] 3. The server receives the recording and converts it into text using Google Cloud Speech-to-Text.

[0632] 4. Additionally, IBM Watson Tone Analyzer is used to analyze the emotions of meeting participants.

[0633] 5. Create a summary of meeting minutes from text data and sentiment information using OpenAI GPT.

[0634] 6. The summary, original text data, and emotion information are stored in a database.

[0635] 7. The server searches the database based on the topic at hand and suggests relevant information.

[0636] 8. The terminal displays the information to the user and also provides options for contacting other conference participants.

[0637] Prompt Sentence Examples

[0638] A new product development meeting is being held. The recording app is launched to record the conversation, and after the meeting ends, the recording is automatically sent to the server. The server converts the text using Google Cloud Speech-to-Text and performs sentiment analysis using IBM Watson Tone Analyzer. A summary is created using OpenAI GPT and stored in a database. Related information is also searched and suggested. For example, the system searches past meeting minutes and related information to suggest new product features discussed during the meeting.

[0639] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0640] Step 1:

[0641] User: At the start of a meeting, the user launches a recording app. Specifically, the user taps the "Start Recording" button on the recording app installed on their smartphone or tablet. The input is the audio from the meeting, and the output is an audio file saved as recording data in the internal storage.

[0642] Step 2:

[0643] Device: After the meeting ends, the recording data is automatically sent to the server. When the user taps the "End Meeting" button, the app uploads the recording data to the server via the Internet. The input is the audio file stored in the device's internal storage, and the output is the audio file transferred to the server. Specifically, the file is uploaded using an HTTPS request.

[0644] Step 3:

[0645] Server: Receives and temporarily stores recorded data. Specifically, the server stores uploaded audio files in a storage directory. The input is the audio file sent from the device, and the output is the audio data stored in the server's internal storage.

[0646] Step 4:

[0647] Server: Uses a speech recognition API (Google Cloud Speech-to-Text) to convert voice data into text data. The server passes the voice data to the API and receives the converted text data. The input is an audio file as recorded data, and the output is conversation data in text format. Specifically, it sends requests to the API and receives the results.

[0648] Step 5:

[0649] Server: Analyzes the emotions of meeting participants from the recording data using an emotion engine (IBM Watson Tone Analyzer). The server sends text data to the API and receives the emotion information resulting from the analysis. The input is text data and the output is emotion information. Specifically, it sends requests to the emotion engine API and receives the results.

[0650] Step 6:

[0651] Server: Automatically creates a summary of meeting minutes based on text data and emotional information using a generative AI model (OpenAI GPT). The input is text data and emotional information, and the output is the generated summary. Specifically, the input data is passed to the generative AI model to generate the summary.

[0652] Step 7:

[0653] Server: The created summary of the minutes is saved in a database. The original text data and emotion information are also saved in the database. The input is the generated summary, text data, and emotion information, and the output is each data entry saved in the database. Specifically, data is added to the database using SQL statements.

[0654] Step 8:

[0655] Server: Searches the database based on the issues recorded by the user during the meeting and suggests related information. The server receives the user's issue information as input and searches for related entries in the database. The input is the issue information, and the output is related information as a search result. Specifically, it generates and executes the database query.

[0656] Step 9:

[0657] Terminal: Displays suggested relevant information to the user. The user can check the information through the app interface and contact other meeting participants if necessary. The input is the relevant information sent from the server, and the output is the information displayed on the user interface. Specifically, the terminal receives data from the API and displays it on the UI.

[0658] Step 10:

[0659] User: Based on the suggested information, send a message to relevant meeting participants and schedule a meeting if necessary. This is executed when the user enters a message in the app and taps the "Send" button. The input is the suggested related information and the message entered by the user, and the output is the sent message. Specifically, the message is sent using the message sending API.

[0660] (Application example 2)

[0661] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0662] Conventional meeting recording and automatic minutes creation systems have difficulty efficiently searching and suggesting emotional information and related past meeting information that can be used to solve problems in factories and other workplaces. Furthermore, the generated minutes summaries lack emotional information and provide insufficient detailed information that reflects the atmosphere of the discussion and the emotions of the meeting participants. This hinders rapid problem solving and efficient information sharing, especially in the workplace.

[0663] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes: means for recording the contents of the meeting; means for converting audio data into text data; means for automatically creating a summary of the minutes from the generated text data; means for saving the summary of the minutes in a database; means for searching and proposing related information from the database based on issues proposed during the meeting; means for displaying the proposed related information and contacting other meeting participants; means for recognizing emotional information and analyzing the emotions of the meeting participants; means for adding the analyzed emotional information to the summary; and means for providing meeting information for problem solving to machines installed in a factory. This enables accurate recording of the contents of the meeting and generation of a detailed summary including emotional information. Furthermore, by searching for and proposing related information based on issues proposed during the meeting, rapid problem solving and efficient information sharing can be achieved on-site.

[0664] "Means for recording the contents of a conference" refers to a device or method for recording the audio during a conference as digital data.

[0665] "Means for converting voice data into text data" refers to voice recognition technology or software for converting recorded voice data into text information.

[0666] The "means for automatically creating a summary of minutes from the generated text data" refers to an algorithm or program that summarizes the text data, extracts important information, and generates an outline of the minutes.

[0667] The "means for storing the summary of the minutes in a database" refers to a method or device for recording the generated summary of the minutes in an information storage system called a database.

[0668] "Means for searching and suggesting relevant information from a database based on an issue raised during a meeting" refers to a method or device for automatically searching a database for information related to an issue raised during a meeting and presenting that information to a user.

[0669] "Means for displaying suggested related information and contacting other conference participants" refers to an interface or method for displaying the searched information on a display or the like and for contacting other conference participants as necessary.

[0670] "Means for recognizing emotional information and analyzing the emotions of meeting participants" refers to technology and software for estimating and analyzing the emotional state of participants from the content of statements made during a meeting and their tone of voice.

[0671] The "means for adding analyzed emotional information to a summary" refers to a method or program for adding the results of analyzing the emotions of meeting participants to the summary of the minutes.

[0672] The "means for providing meeting information for problem solving to machines installed in the factory" refers to an interface or method for linking meeting minutes and related information to machines and management systems in the factory.

[0673] The program for implementing this invention records the audio of meetings and problem-solving meetings within a factory, processes the data to generate a summary of the minutes, and then performs emotion analysis to support efficient information sharing and problem solving. The specific processing and the hardware and software used are described below.

[0674] Hardware and Software Configuration

[0675] The server uses the following software:

[0676] Speech recognition API: A technology for converting voice data into text data. A typical example is Google's speech recognition API.

[0677] Emotion recognition engine: A technology for analyzing the emotions of meeting participants. The EmotionRecognition library is one example.

[0678] Generative AI model: A technology for generating a summary of meeting minutes from text data. A generative AI model using the Transformer library falls into this category.

[0679] Database: A system for storing generated minutes summaries and related information. It uses an SQLite database.

[0680] System processing overview

[0681] The server receives the audio data of the meeting recorded by the user and converts it into text data using a speech recognition API. The converted text data is then temporarily stored. Next, an emotion recognition engine is used to analyze emotional information from the text data, which is also temporarily stored. A generative AI model is then used to generate a summary of the minutes based on the text data and emotional information. This summary includes key points, decisions, and emotional information. The generated summary is then stored in a database.

[0682] Based on the topic proposed during the meeting, the system searches for related past information from the database. The server also refers to emotional information to improve the accuracy of the search. The retrieved related information is compiled as suggestions and sent to the user's device. The suggested information is displayed to the user, allowing them to contact other meeting participants.

[0683] This information is provided to machines and management systems within the factory, helping to quickly resolve problems on site.

[0684] Specific example explanation

[0685] For example, consider a meeting to improve the production efficiency of a new product. During the meeting, a recording device is running and the audio data is automatically sent to a server. The server converts the audio data into text data and performs sentiment analysis. A generative AI model uses the text data and sentiment information to generate a summary of the meeting minutes, including key points and decisions. This summary is stored in a database, and can be quickly searched and suggested when related information is needed in the future.

[0686] The minutes of past discussions regarding the "new production setup" problem proposed during this meeting can be referenced, allowing users to find effective solutions and contact other participants if necessary.

[0687] Prompt Sentence Examples

[0688] Here are some example prompts to input to a generative AI model:

[0689] "Based on the minutes of previous meetings, please tell us how discussions regarding bottlenecks on production lines have gone in the past."

[0690] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0691] Step 1:

[0692] At the start of a conference, a user activates a recording device to record the contents of the conference. The recording device records the conversation as digital audio data and temporarily stores the data. The input is the conversation, and the output is digital audio data.

[0693] Step 2:

[0694] After the meeting ends, the device automatically sends the recording data to the server. Communication between the recording device and the server is performed using a secure protocol. The input here is digital audio data, and the output is data transmission to the server.

[0695] Step 3:

[0696] The server temporarily stores the received voice data. Then, it converts the voice data into text data using a speech recognition API (for example, Google's speech recognition service). The input here is the received voice data, and the output is the converted text data.

[0697] Step 4:

[0698] The server uses an emotion recognition engine (e.g., the EmotionRecognition library) based on the text data to analyze the emotional information of the conference participants. This analysis result is also temporarily saved. The input is text data, and the output is the emotion analysis result.

[0699] Step 5:

[0700] The server uses a generative AI model (e.g., the Transformer library) to automatically generate a summary of the minutes from the text data and emotional information. The generated summary includes emotional information as well as key points and decisions. This summary is stored in a database. The input is text data and emotional information, and the output is a summary of the minutes.

[0701] Step 6:

[0702] The server searches the database for relevant past information based on the topic proposed during the meeting. Emotional information is also referenced to improve the accuracy of the search. The input is the proposed topic and emotional information, and the output is relevant past information.

[0703] Step 7:

[0704] The server summarizes the retrieved related information as suggestions and sends them to the terminal. The suggested information is displayed on the user's terminal. The input is related past information, and the output is the suggested information displayed on the user's terminal.

[0705] Step 8:

[0706] The user reviews the proposed information and contacts other meeting participants as needed, allowing them to send messages to related meeting participants and set up new meetings. The input is the proposed information, and the output is contact with participants and meeting settings.

[0707] The above is the specific processing flow of the present invention.

[0708] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0709] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0710] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0711] [Third embodiment]

[0712] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0713] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0714] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0715] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0716] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0717] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0718] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0719] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0720] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0721] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0722] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0723] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0724] The system of the present invention records the contents of a meeting, converts the recording data into text data, automatically creates a summary of the minutes, and stores the summary in a database. Furthermore, it can search for and propose related information from the database based on issues raised during the meeting. This system allows for efficient recording of meeting contents and makes it possible to utilize past knowledge and solutions in real time.

[0725] Specifically, the system operates in the following steps:

[0726] Program processing

[0727] 1. Recording of meeting content

[0728] User: Start the recording device as soon as the meeting starts and record the conversation using a dedicated recording app.

[0729] 2. Sending recording data

[0730] Device: After the meeting ends, the recording data is sent to the server. The recording app automatically sends the data.

[0731] 3. Converting Audio Data to Text

[0732] Server: Receives the transmitted voice data and converts it into text data using a speech recognition API. The converted text data is temporarily saved.

[0733] 4. Automatic summary creation

[0734] Server: Based on the text data converted from the voice data, the generative AI extracts important points and decisions, and automatically creates a summary of the minutes.

[0735] 5. Save summary

[0736] Server: Saves the summary of the minutes to a database. The database stores not only the summary but also the original text data.

[0737] 6. Search and suggest related information

[0738] Server: Based on the issues recorded by users during the meeting, the server searches the database to find relevant past minutes and solutions.

[0739] Server: Compiles relevant information into suggestions and sends them to the user's device.

[0740] 7. Information Display and Contact Functions

[0741] Terminal: Providing suggested relevant information to the user, the user has the option to review the required information and contact other meeting participants.

[0742] User: Uses the suggested information to send a message to relevant meeting participants and schedule a meeting if necessary.

[0743] Specific examples

[0744] For example, when a new product development team holds a meeting, they launch a dedicated recording app at the start of the meeting. After the meeting ends, the recording data is automatically sent to a server. The server receives the data and converts it into text. Generative AI extracts key points from the text data, creates a summary, and stores it in a database.

[0745] Next, the system searches the database for past discussion topics related to new product features proposed during the meeting. If relevant past meeting minutes or ideas are found, they are suggested to the user, who can use this information to develop new ideas.

[0746] Thus, the present invention can significantly improve the efficiency and productivity of meetings.

[0747] The processing flow will be explained below.

[0748] Step 1:

[0749] When a user starts a meeting, the recording device is activated and the conversation is recorded using a dedicated recording app.

[0750] Step 2:

[0751] When the meeting ends, the device automatically sends the recording data to the server, which is done automatically by the recording app.

[0752] Step 3:

[0753] The server receives the transmitted audio data and stores it temporarily.

[0754] Step 4:

[0755] The server uses a speech recognition API to convert the voice data into text data, which is then temporarily saved.

[0756] Step 5:

[0757] The server inputs the text data into the AI ​​generator, which automatically creates a summary of the minutes, extracting key points and decisions.

[0758] Step 6:

[0759] The server saves the created summary in a database, along with the original text data.

[0760] Step 7:

[0761] Based on the issues users record during meetings, the server searches the database to find minutes of past meetings and solutions.

[0762] Step 8:

[0763] The server compiles relevant information into suggestions and sends them to the terminal, where they are displayed to the user.

[0764] Step 9:

[0765] The user reviews the suggested information and selects the information they require. The user has the option to contact other meeting participants.

[0766] Step 10:

[0767] The user can then use the suggested information to send messages to relevant meeting participants and schedule meetings as needed.

[0768] Step 11:

[0769] The server feeds back the utilized information and results to the database, enabling more accurate proposals to be made in future meetings.

[0770] Example 1

[0771] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0772] Conventional meeting recording systems often require manual recording and management of meeting content, which requires time and effort. It is also difficult to quickly search past minutes and related information, making it difficult to receive immediate and effective feedback on issues raised during meetings. Furthermore, they lack the ability to share relevant information with other meeting participants, potentially reducing the efficiency and productivity of meetings.

[0773] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0774] In this invention, the server includes means for recording the contents of the meeting, means for converting the audio data into text data, means for automatically creating a summary of the minutes from the generated text data, means for saving the summary of the minutes in a database, means for searching and proposing related information from the database based on issues proposed during the meeting, means for displaying the proposed related information and providing a means for contacting other meeting participants, means for extracting important points from the text data using a generation AI, and means for transmitting the recorded data to the server. This automates the recording and management of meetings and enables past minutes and related information to be quickly searched and shared, thereby improving the efficiency and productivity of meetings.

[0775] "Means for recording the contents of a meeting" refers to a device or application for recording statements and conversations made during a meeting as audio data.

[0776] "Means for converting voice data into text data" refers to voice recognition technology or software for analyzing recorded voice data and converting it into text data.

[0777] The "means for automatically creating a summary of meeting minutes from generated text data" is a generative AI system that analyzes text data to extract and summarize important points and decisions.

[0778] The "means for storing the summary of the minutes in a database" refers to a device or system for storing the automatically created summary of the minutes in a relational database, cloud storage, or the like.

[0779] "A means for searching and proposing relevant information from a database based on issues proposed during a meeting" is an engine that searches a database for past minutes and solutions related to the issues entered by the user, extracts appropriate information, and presents it.

[0780] "Means for displaying suggested related information and providing a means for contacting other conference participants" refers to functionality or software that displays related information on a user interface and enables users to send messages to other conference participants or schedule meetings.

[0781] "Means for extracting key points from text data using generative AI" refers to technology that uses a generative AI model (e.g., a natural language processing model) to mechanically extract key points and decisions from text data.

[0782] The "means for transmitting recorded data to a server" refers to a communication function or protocol for transmitting recorded voice data to a server via the Internet.

[0783] This invention is a system that records the contents of a meeting, converts the recorded data into text data, automatically creates a summary of the minutes, stores it in a database, and searches for and suggests related information.

[0784] First, users turn on the recording device at the start of a meeting and record the conversation using a dedicated recording app, which can be software installed on a smartphone, tablet, or PC, such as Otter.ai or a similar application.

[0785] Then, after the meeting ends, the device sends the recording data to the server. The recording app does this automatically, and the data is sent using an encryption protocol (e.g., TLS). The data transmission is completed without any special operation by the user.

[0786] The server sends the received voice data to the Google Cloud Speech-to-Text API, which converts the voice data into text data. The converted text data is temporarily stored on the server. After conversion is complete, the server uses a generative AI (e.g., OpenAI GPT-3) to extract important points and decisions from the text data and create a summary of the meeting minutes.

[0787] The server stores the generated summary in a database (e.g., MySQL, PostgreSQL). The database stores not only the summary of the minutes but also the original text data, ensuring the integrity of the information.

[0788] Furthermore, based on the issues proposed during the meeting, the server searches the database and provides related past meeting minutes and solutions. This search uses an SQL query such as "SELECT FROM summaries WHERE text LIKE '%keywords related to the issue%'". The server compiles the relevant information found as suggestions and sends them to the user's device. This is sent in JSON format, and a dedicated suggestion display app receives and displays it.

[0789] The device displays the suggested related information to the user. The user can check the displayed information and, if necessary, send a message to other meeting participants or schedule a new meeting. By clicking the "Contact" button in the suggestion display app, a message sending screen will appear, allowing the user to send a message to the relevant participants.

[0790] Specific examples

[0791] When a new product development team holds a meeting, the user launches a dedicated recording app at the start of the meeting to record the conversation. After the meeting ends, the recorded data is automatically sent from the device to the server. The server receives this data and converts the speech to text using the Google Cloud Speech-to-Text API. OpenAI GPT-3 extracts key points from the converted text data and generates a summary of the meeting minutes. The generated summary is stored in a MySQL database.

[0792] Next, for new product features proposed during a meeting, the server searches the database to see if similar topics have been discussed in the past. If relevant minutes or ideas from past meetings are found, the server compiles them into a proposal and sends them in JSON format to the user's device. The user can then view the information through the proposal display app, send messages to participants in related meetings, and schedule new meetings.

[0793] Prompt Sentence Examples

[0794] "Please extract the key points and decisions from the following text data and create a summary of the meeting minutes.

[0795] Text data:

[0796] Proposing new product features at a meeting

[0797] Supervisor approves

[0798] We plan to work out the details at the next meeting.

[0799] ...

[0800] "

[0801] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0802] Step 1:

[0803] When the meeting starts, the user activates the recording device and records the conversation using a dedicated recording app. The recording app is installed on a smartphone, tablet, or PC, and an example is "Otter.ai." The user starts recording by pressing the "Start Recording" button. The input is an audio signal, and the output is voice data.

[0804] Step 2:

[0805] After the meeting ends, the device sends the recorded data to the server. When the recording app detects the end of the meeting, it automatically compresses the recorded data and sends it to the server using an encryption protocol (e.g., TLS). The input is the audio data, and the output is the data sent to the server.

[0806] Step 3:

[0807] The server stores the received voice data and sends it to the Google Cloud Speech-to-Text API to convert it into text. This API analyzes the voice data and outputs it as a string. The input is voice data and the output is text data.

[0808] Step 4:

[0809] The server temporarily stores the text data received from the Google Cloud Speech-to-Text API. It then supplies the text data to a generative AI (e.g., OpenAI GPT-3) with a prompt: "Please extract the key points and decisions from the following text data and create a summary of the minutes." The generative AI extracts the key points and generates a summary of the minutes. The input is the text data and the prompt, and the output is a summary of the minutes.

[0810] Step 5:

[0811] The server saves the generated summary of the minutes to a database. The server first establishes a database connection and executes the SQL query "INSERT INTO summaries (text, summary) VALUES (?, ?)". The original text data is also saved to the same database. The input is the summary of the minutes and the text data, and the output is the information saved in the database.

[0812] Step 6:

[0813] The server searches the database based on the issues recorded by the user during the meeting. For example, it runs the SQL query "SELECT FROM summaries WHERE text LIKE '%keywords related to the issue%'". The server filters the relevant information found and extracts the relevant information. The input is the issue keywords and the output is the relevant information.

[0814] Step 7:

[0815] The server compiles related information as suggestions and sends them to the user's device. The suggestions are sent in JSON format, and a dedicated suggestion display app receives them and displays them on the user interface. The input is the related information, and the output is the data sent to the device.

[0816] Step 8:

[0817] The device displays the suggested related information to the user. The suggestion display app parses the received JSON data and presents it to the user in a list format. The user can click on an interesting suggestion from the list to view details. The input is JSON data, and the output is the information displayed to the user.

[0818] Step 9:

[0819] Based on the proposed information, the user can send a message to the relevant meeting participants and set up a new meeting. For example, clicking the "Contact" button in the proposal display app will display the message sending screen, allowing the user to send a message to the relevant participants. The input is the user's action and the proposed information, and the output is the sent message.

[0820] (Application example 1)

[0821] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0822] Currently, in workplaces that use industrial machinery, it is difficult to efficiently create meeting records and utilize past related information in real time. Furthermore, there is a lack of a method to efficiently search and propose solutions to issues raised during meetings, which leads to inefficiency and reduced productivity. Furthermore, in industrial machinery workplaces, there is a need for a method to directly reflect the information in meeting minutes in machine operation.

[0823] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0824] In this invention, the server includes: means for recording the contents of a meeting; means for converting voice data into text data; means for automatically creating a summary of the minutes from the generated text data; means for saving the summary of the minutes in a database; means for searching and proposing related information from the database based on issues proposed during the meeting; means for displaying the proposed related information and providing a means for contacting other meeting participants; means for automatically generating meeting minutes and presenting search results for related information on the display device of the industrial machine in cooperation with the industrial machine control device; means for generating the proposed related information using a generative AI model; and means for generating a summary of the minutes using prompt sentences. This enables efficient recording of meeting contents and real-time utilization of past meeting information, thereby improving work efficiency and productivity at industrial machine sites.

[0825] A "means for recording the contents of a meeting" is a device or software that can record the audio content spoken during a meeting.

[0826] "Means for converting voice data into text data" refers to algorithms or tools for analyzing recorded voice data and converting it into text information.

[0827] The "means for automatically creating a summary of minutes from generated text data" is a means for extracting important points and decisions based on text data converted from speech and generating a summary.

[0828] The "means for saving the summary of the minutes in a database" refers to a function or device for writing and saving the generated summary in a database.

[0829] "A means for searching and proposing relevant information from a database based on issues proposed during a meeting" is a means for inputting the agenda items or issues raised during a meeting, searching a database for past information related to them, and presenting it.

[0830] "Means for displaying suggested related information and providing means for contacting other conference participants" refers to means for displaying searched related information in a user interface and providing options for communicating with other conference participants.

[0831] "Means for automatically generating meeting minutes and presenting search results for related information on the display device of the industrial machine in cooperation with the control device of the industrial machine" refers to means for exchanging data with the control device of the industrial machine and visually displaying the meeting minutes and related information on the display device of the machine.

[0832] A "means for generating suggested relevant information using a generative AI model" is a means for using an artificial intelligence model to automatically generate highly relevant information or solutions based on input data.

[0833] The "means for generating a summary of meeting minutes using prompt sentences" is a means for efficiently summarizing meeting minutes using pre-set questions or instructions (prompt sentences).

[0834] The system of the present invention realizes efficient creation of meeting records and real-time search for related information in workplaces where industrial machinery is heavily used. This system records the contents of meetings, converts the recording data into text data, automatically creates a summary of the minutes, and stores the summary in a database. Furthermore, it can search for and propose related information from the database based on issues raised during the meeting. This system allows efficient recording of meeting contents and utilization of past knowledge and solutions in real time.

[0835] Program processing

[0836] At the start of the meeting, the user launches a dedicated recording app on the industrial machine's tablet device and records the conversation. After the meeting ends, the recording data is automatically sent to the server. The server receives it and converts the audio data into text data using the Google Cloud Speech-to-Text API. A summary is then automatically created from the text data using a generative AI model such as OpenAI GPT-4. This generative AI model generates a summary using a prompt sentence. For example, the following prompt sentence can be used:

[0837] Example prompt: "Meeting minutes text: Today's meeting discussed new product development plans. Please write a summary."

[0838] The generated summary is stored in a MySQL database. Furthermore, based on issues recorded during the meeting, the database is searched to find relevant past meeting records and solutions. This involves using a generative AI model to generate relevant information and provide responses in the form of prompts based on the user's request. This relevant information is then presented on the display of the industrial machine, allowing the user to confirm the required information and contact other meeting participants.

[0839] Hardware and Software

[0840] The hardware used includes tablet devices and control devices for industrial machines, and the software used includes a meeting recording application, Google Cloud Speech-to-Text API, OpenAI GPT-4 generative AI models, and a MySQL database.

[0841] Specific examples

[0842] For example, the procedure for a new product development meeting is as follows: At the start of the meeting, the user launches a recording app on their tablet device and records the conversation. After the meeting ends, the recording app sends the audio data to a server, which then converts it into text data using the Google Cloud Speech-to-Text API. A generative AI model (OpenAI GPT-4) extracts key points from this text data and generates a summary of the meeting minutes. This summary is then stored in a MySQL database. Past related information about the new features proposed in the meeting is then searched for and presented on a display device. This information promotes the development of new ideas.

[0843] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0844] Step 1:

[0845] The user starts a dedicated recording app on the tablet device of the industrial machine and records the contents of the meeting. Remarks and discussions during the meeting are recorded as audio data in real time.

[0846] Input: Voice spoken during a meeting

[0847] Output: Recorded audio data file

[0848] Step 2:

[0849] After the conference ends, the device automatically sends the recorded data to the server. This data transfer occurs when the recording application is terminated.

[0850] Input: Recorded audio data file

[0851] Output: Audio data sent to the server

[0852] Step 3:

[0853] The server converts the received voice data into text data using the Google Cloud Speech-to-Text API. The voice data is analyzed and saved as text information.

[0854] Input: Audio data sent to the server

[0855] Output: Data converted to text

[0856] Step 4:

[0857] The server uses a generative AI model (OpenAI GPT-4) to automatically create a summary of the minutes from the converted text data. It uses prompts to extract key points and decisions and generate a summary.

[0858] Input: Text data

[0859] Output: A summary created by the generative AI model

[0860] Example prompt: "Meeting minutes text: Today's meeting discussed new product development plans. Please write a summary."

[0861] Step 5:

[0862] The server stores the generated summaries in a MySQL database, where they can be later stored in a searchable format, along with the original text data.

[0863] Input: Generated summary

[0864] Output: Summary and original text data stored in a database

[0865] Step 6:

[0866] The server searches a database for past meeting records and solutions based on the issues proposed during the meeting, generates relevant information using a generative AI model, and provides responses in the form of prompts tailored to the user's request.

[0867] Input: Issues proposed during the meeting

[0868] Output: Relevant information retrieved from the database

[0869] Step 7:

[0870] The terminal displays the suggested related information on the display device of the industrial machine, allowing the user to check the required information and contact other conference participants.

[0871] Input: Relevant information retrieved from the database

[0872] Output: relevant information presented on a display device

[0873] Processing flow example

[0874] 1. When the meeting starts, the user launches the recording app on the tablet device and records the conversation (step 1).

[0875] 2. After the meeting ends, the app automatically sends the recording data to the server (step 2).

[0876] 3. The server converts the audio data into text using the Google Cloud Speech-to-Text API (step 3).

[0877] 4. The server uses a generative AI model such as OpenAI GPT-4 to automatically create a summary from the generated text data, using the prompt sentence (Step 4).

[0878] 5. The generated summaries and the original text data are stored in a MySQL database (Step 5).

[0879] 6. Based on the issues proposed during the meeting, the server searches the database for past meeting records and solutions, and generates relevant information (step 6).

[0880] 7. The generated related information is presented on the display device of the industrial machine, allowing the user to check the necessary information and contact other conference participants (step 7).

[0881] Through the above steps, the system of the present invention realizes efficient recording of meeting contents and real-time utilization of past knowledge, thereby significantly improving work efficiency at industrial machinery sites.

[0882] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0883] The system according to the present invention records the contents of a meeting, converts the recording data into text data, automatically creates a summary of the minutes, and stores the summary in a database. Furthermore, the system can search for and suggest related information from the database based on the topics proposed during the meeting. Furthermore, by combining it with an emotion engine that recognizes the user's emotions, it can generate minutes with emotional information added and suggest information according to the emotions.

[0884] Specifically, the system operates in the following steps:

[0885] Program processing

[0886] 1. Recording of meeting content

[0887] User: When the meeting starts, the recording device is activated and the conversation is recorded. The user uses a dedicated recording app.

[0888] 2. Sending recording data

[0889] On your device: After the meeting ends, the recording data is automatically sent to the server. The recording app does this automatically.

[0890] 3. Converting Audio Data to Text

[0891] Server: Receives the transmitted audio data. The audio data is temporarily stored.

[0892] Server: Using a speech recognition API, the voice data is converted into text data, which is then temporarily saved.

[0893] 4. Acquiring emotional information

[0894] Server: Analyzes the emotions of meeting participants using an emotion engine from the recorded data. The analyzed emotional information is also temporarily saved.

[0895] 5. Automatic summary creation

[0896] Server: Based on text data and emotional information, the generative AI automatically creates a summary of the minutes, including key points and decisions, as well as analyzed emotional information.

[0897] 6. Save Summary

[0898] Server: The summary of the minutes is saved in a database. Not only the summary but also the original text data and emotional information are saved.

[0899] 7. Search and suggest related information

[0900] Server: Based on the issues recorded by users during meetings, the server searches the database to find relevant past meeting minutes and solutions, and also references emotional information to improve the accuracy of the search.

[0901] Server: The server compiles the searched related information into suggestions and sends them to the device. The suggested information is then displayed to the user.

[0902] 8. Information Display and Contact Functions

[0903] Terminal: Providing suggested relevant information to the user, the user has the option to review the required information and contact other meeting participants.

[0904] User: Uses the suggested information to send a message to relevant meeting participants and schedule a meeting if necessary.

[0905] Specific examples

[0906] For example, when a new product development team holds a meeting, they launch a recording app at the beginning of the meeting. After the meeting ends, the recording is automatically sent to the server. The server receives the recording and converts it into text using a speech recognition API. It then uses an emotion engine to analyze the emotions of the meeting participants.

[0907] The generative AI extracts key points and decisions based on text data and emotional information, and creates a summary. This summary is saved in a database. Next, it searches the database for past discussion topics similar to the new product features proposed during the meeting. By also referring to emotional information, more accurate related information is found. If relevant minutes or ideas from past meetings are found, they are suggested to the user.

[0908] The user can use this information to develop new ideas and contact relevant meeting participants as needed, thus significantly improving the efficiency and productivity of meetings.

[0909] The processing flow will be explained below.

[0910] Step 1:

[0911] When a user starts a meeting, the recording device is activated and the conversation is recorded using a dedicated recording app.

[0912] Step 2:

[0913] When the meeting ends, the device automatically sends the recording data to the server, which is done automatically by the recording app.

[0914] Step 3:

[0915] The server receives the transmitted audio data and stores it temporarily.

[0916] Step 4:

[0917] The server uses a speech recognition API to convert the voice data into text data, which is then temporarily saved.

[0918] Step 5:

[0919] The server uses an emotion engine to analyze the emotions of the conference participants from the received voice data, and the analyzed emotion information is also temporarily saved.

[0920] Step 6:

[0921] The server uses text data and emotional information to automatically generate a summary of the minutes using a generation AI, which extracts emotional information in addition to important points and decisions.

[0922] Step 7:

[0923] The server stores the generated summary in a database, which includes the original text data and sentiment information.

[0924] Step 8:

[0925] The server searches the database based on the issues recorded by the user during the meeting, finding minutes of past meetings and solutions, and also using emotional information to improve the accuracy of the search.

[0926] Step 9:

[0927] The server compiles the retrieved related information into suggestions and sends them to the terminal, where they are displayed to the user.

[0928] Step 10:

[0929] The user reviews the suggested information and selects the information they require. The user has the option to contact other meeting participants.

[0930] Step 11:

[0931] The user can then use the suggested information to send messages to relevant meeting participants and schedule meetings as needed.

[0932] Step 12:

[0933] The server feeds back the utilized information and results to the database, enabling more accurate proposals to be made in future meetings.

[0934] Example 2

[0935] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0936] Conventional meeting recording systems only record meeting content and convert the audio data into text to create minutes, but they do not offer advanced features such as creating minutes that take into account the emotional information of participants during the meeting, or searching and suggesting related information based on proposed issues. This makes it difficult to maximize the efficiency and effectiveness of meetings.

[0937] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for recording the contents of the meeting, means for converting voice data into text data, means for automatically creating a summary of the minutes from the generated text data, means for saving the summary of the minutes in a database, means for searching for and proposing related information from the database based on issues proposed during the meeting, means for displaying the proposed related information and contacting other meeting participants, means for acquiring emotional information and creating a summary of the minutes based on the emotions of the meeting participants, and means for searching for and proposing related information based on the emotional information. This makes it possible to create minutes that take into account not only the content of the meeting but also the emotions of the participants, thereby maximizing the efficiency and effectiveness of the meeting.

[0938] A "meeting recording device" is a device or software that captures and digitally stores audio from a meeting.

[0939] A "means for converting audio data to text data" is software or algorithms for analyzing recorded audio and converting it into a corresponding text format.

[0940] The "means for automatically creating a summary of meeting minutes from generated text data" is a system that analyzes text data, extracts important points and decisions from the meeting, and compiles them into a short summary.

[0941] The "means for storing the summary of the minutes in a database" refers to a data storage system for storing the generated summary of the minutes.

[0942] "A means for searching and proposing relevant information from a database based on issues proposed during a meeting" is a system for searching a database to find relevant information based on issues or questions raised during a meeting.

[0943] The "means for displaying suggested related information and contacting other conference participants" is a system for displaying suggested information to the user as a search result and contacting other conference participants as needed.

[0944] "Means for acquiring emotional information and creating a summary of meeting minutes based on the emotions of meeting participants" is a system that analyzes the emotions of participants from audio and text data during a meeting and creates a summary of meeting minutes based on that information.

[0945] "Means for searching and suggesting related information based on emotional information" is a system that searches for related information based on emotional analysis and makes more appropriate and useful suggestions.

[0946] The system according to the present invention records the contents of a meeting, converts the recording data into text data, automatically creates a summary of the minutes, and stores the summary in a database. It also has a function to search the database for relevant information based on the topics proposed during the meeting and suggest it. It also uses an emotion engine to analyze the user's emotions, and can create minutes and suggest relevant information based on the emotional information.

[0947] Hardware and software used

[0948] Hardware

[0949] Recording device: Smartphone or dedicated recorder

[0950] Server: A high-performance computer for data processing

[0951] Devices: Laptops, tablets, smartphones

[0952] software

[0953] Recording app: An application that records the contents of a meeting and sends the data to a server.

[0954] Speech Recognition API: Google Cloud Speech-to-Text

[0955] Emotion engine: IBM Watson Tone Analyzer

[0956] Generative AI model: OpenAI GPT

[0957] Database system: A relational database management system (RDBMS) for storing data.

[0958] Program Overview

[0959] 1. The user launches a recording app at the start of a meeting and records the contents of the meeting.

[0960] 2. After the meeting ends, the device automatically sends the recording data to the server.

[0961] 3. The server converts the received voice data into text using the Google Cloud Speech-to-Text API.

[0962] 4. The server runs IBM Watson Tone Analyzer on the text and audio data to obtain emotion information.

[0963] 5. The server uses OpenAI GPT to automatically create a summary of the minutes based on the text data and emotional information.

[0964] 6. The server stores the generated summary of the minutes, the original text data, and the emotion information in a database.

[0965] 7. The server searches the database based on the issues described during the meeting and suggests relevant information. The suggestions also refer to emotional information.

[0966] 8. The terminal displays the suggested relevant information to the user, providing the user with the option to review the required information and contact other conference participants as needed.

[0967] Specific operation example

[0968] Working example 1:

[0969] When a new product development team holds a meeting, they follow these steps:

[0970] 1. The user launches the recording app at the start of the meeting.

[0971] 2. After the meeting ends, the recording data is automatically sent to the server.

[0972] 3. The server receives the recording and converts it into text using Google Cloud Speech-to-Text.

[0973] 4. Additionally, IBM Watson Tone Analyzer is used to analyze the emotions of meeting participants.

[0974] 5. Create a summary of meeting minutes from text data and sentiment information using OpenAI GPT.

[0975] 6. The summary, original text data, and emotion information are stored in a database.

[0976] 7. The server searches the database based on the topic at hand and suggests relevant information.

[0977] 8. The terminal displays the information to the user and also provides options for contacting other conference participants.

[0978] Prompt Sentence Examples

[0979] A new product development meeting is being held. The recording app is launched to record the conversation, and after the meeting ends, the recording is automatically sent to the server. The server converts the text using Google Cloud Speech-to-Text and performs sentiment analysis using IBM Watson Tone Analyzer. A summary is created using OpenAI GPT and stored in a database. Related information is also searched and suggested. For example, the system searches past meeting minutes and related information to suggest new product features discussed during the meeting.

[0980] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0981] Step 1:

[0982] User: At the start of a meeting, the user launches a recording app. Specifically, the user taps the "Start Recording" button on the recording app installed on their smartphone or tablet. The input is the audio from the meeting, and the output is an audio file saved as recording data in the internal storage.

[0983] Step 2:

[0984] Device: After the meeting ends, the recording data is automatically sent to the server. When the user taps the "End Meeting" button, the app uploads the recording data to the server via the Internet. The input is the audio file stored in the device's internal storage, and the output is the audio file transferred to the server. Specifically, the file is uploaded using an HTTPS request.

[0985] Step 3:

[0986] Server: Receives and temporarily stores recorded data. Specifically, the server stores uploaded audio files in a storage directory. The input is the audio file sent from the device, and the output is the audio data stored in the server's internal storage.

[0987] Step 4:

[0988] Server: Uses a speech recognition API (Google Cloud Speech-to-Text) to convert voice data into text data. The server passes the voice data to the API and receives the converted text data. The input is an audio file as recorded data, and the output is conversation data in text format. Specifically, it sends requests to the API and receives the results.

[0989] Step 5:

[0990] Server: Analyzes the emotions of meeting participants from the recording data using an emotion engine (IBM Watson Tone Analyzer). The server sends text data to the API and receives the emotion information resulting from the analysis. The input is text data and the output is emotion information. Specifically, it sends requests to the emotion engine API and receives the results.

[0991] Step 6:

[0992] Server: Automatically creates a summary of meeting minutes based on text data and emotional information using a generative AI model (OpenAI GPT). The input is text data and emotional information, and the output is the generated summary. Specifically, the input data is passed to the generative AI model to generate the summary.

[0993] Step 7:

[0994] Server: The created summary of the minutes is saved in a database. The original text data and emotion information are also saved in the database. The input is the generated summary, text data, and emotion information, and the output is each data entry saved in the database. Specifically, data is added to the database using SQL statements.

[0995] Step 8:

[0996] Server: Searches the database based on the issues recorded by the user during the meeting and suggests related information. The server receives the user's issue information as input and searches for related entries in the database. The input is the issue information, and the output is related information as a search result. Specifically, it generates and executes the database query.

[0997] Step 9:

[0998] Terminal: Displays suggested relevant information to the user. The user can check the information through the app interface and contact other meeting participants if necessary. The input is the relevant information sent from the server, and the output is the information displayed on the user interface. Specifically, the terminal receives data from the API and displays it on the UI.

[0999] Step 10:

[1000] User: Based on the suggested information, send a message to relevant meeting participants and schedule a meeting if necessary. This is executed when the user enters a message in the app and taps the "Send" button. The input is the suggested related information and the message entered by the user, and the output is the sent message. Specifically, the message is sent using the message sending API.

[1001] (Application example 2)

[1002] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1003] Conventional meeting recording and automatic minutes creation systems have difficulty efficiently searching and suggesting emotional information and related past meeting information that can be used to solve problems in factories and other workplaces. Furthermore, the generated minutes summaries lack emotional information and provide insufficient detailed information that reflects the atmosphere of the discussion and the emotions of the meeting participants. This hinders rapid problem solving and efficient information sharing, especially in the workplace.

[1004] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes: means for recording the contents of the meeting; means for converting audio data into text data; means for automatically creating a summary of the minutes from the generated text data; means for saving the summary of the minutes in a database; means for searching and proposing related information from the database based on issues proposed during the meeting; means for displaying the proposed related information and contacting other meeting participants; means for recognizing emotional information and analyzing the emotions of the meeting participants; means for adding the analyzed emotional information to the summary; and means for providing meeting information for problem solving to machines installed in a factory. This enables accurate recording of the contents of the meeting and generation of a detailed summary including emotional information. Furthermore, by searching for and proposing related information based on issues proposed during the meeting, rapid problem solving and efficient information sharing can be achieved on-site.

[1005] "Means for recording the contents of a conference" refers to a device or method for recording the audio during a conference as digital data.

[1006] "Means for converting voice data into text data" refers to voice recognition technology or software for converting recorded voice data into text information.

[1007] The "means for automatically creating a summary of minutes from the generated text data" refers to an algorithm or program that summarizes the text data, extracts important information, and generates an outline of the minutes.

[1008] The "means for storing the summary of the minutes in a database" refers to a method or device for recording the generated summary of the minutes in an information storage system called a database.

[1009] "Means for searching and suggesting relevant information from a database based on an issue raised during a meeting" refers to a method or device for automatically searching a database for information related to an issue raised during a meeting and presenting that information to a user.

[1010] "Means for displaying suggested related information and contacting other conference participants" refers to an interface or method for displaying the searched information on a display or the like and for contacting other conference participants as necessary.

[1011] "Means for recognizing emotional information and analyzing the emotions of meeting participants" refers to technology and software for estimating and analyzing the emotional state of participants from the content of statements made during a meeting and their tone of voice.

[1012] The "means for adding analyzed emotional information to a summary" refers to a method or program for adding the results of analyzing the emotions of meeting participants to the summary of the minutes.

[1013] The "means for providing meeting information for problem solving to machines installed in the factory" refers to an interface or method for linking meeting minutes and related information to machines and management systems in the factory.

[1014] The program for implementing this invention records the audio of meetings and problem-solving meetings within a factory, processes the data to generate a summary of the minutes, and then performs emotion analysis to support efficient information sharing and problem solving. The specific processing and the hardware and software used are described below.

[1015] Hardware and Software Configuration

[1016] The server uses the following software:

[1017] Speech recognition API: A technology for converting voice data into text data. A typical example is Google's speech recognition API.

[1018] Emotion recognition engine: A technology for analyzing the emotions of meeting participants. The EmotionRecognition library is one example.

[1019] Generative AI model: A technology for generating a summary of meeting minutes from text data. A generative AI model using the Transformer library falls into this category.

[1020] Database: A system for storing generated minutes summaries and related information. It uses an SQLite database.

[1021] System processing overview

[1022] The server receives the audio data of the meeting recorded by the user and converts it into text data using a speech recognition API. The converted text data is then temporarily stored. Next, an emotion recognition engine is used to analyze emotional information from the text data, which is also temporarily stored. A generative AI model is then used to generate a summary of the minutes based on the text data and emotional information. This summary includes key points, decisions, and emotional information. The generated summary is then stored in a database.

[1023] Based on the topic proposed during the meeting, the system searches for related past information from the database. The server also refers to emotional information to improve the accuracy of the search. The retrieved related information is compiled as suggestions and sent to the user's device. The suggested information is displayed to the user, allowing them to contact other meeting participants.

[1024] This information is provided to machines and management systems within the factory, helping to quickly resolve problems on site.

[1025] Specific example explanation

[1026] For example, consider a meeting to improve the production efficiency of a new product. During the meeting, a recording device is running and the audio data is automatically sent to a server. The server converts the audio data into text data and performs sentiment analysis. A generative AI model uses the text data and sentiment information to generate a summary of the meeting minutes, including key points and decisions. This summary is stored in a database, and can be quickly searched and suggested when related information is needed in the future.

[1027] The minutes of past discussions regarding the "new production setup" problem proposed during this meeting can be referenced, allowing users to find effective solutions and contact other participants if necessary.

[1028] Prompt Sentence Examples

[1029] Here are some example prompts to input to a generative AI model:

[1030] "Based on the minutes of previous meetings, please tell us how discussions regarding bottlenecks on production lines have gone in the past."

[1031] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1032] Step 1:

[1033] At the start of a conference, a user activates a recording device to record the contents of the conference. The recording device records the conversation as digital audio data and temporarily stores the data. The input is the conversation, and the output is digital audio data.

[1034] Step 2:

[1035] After the meeting ends, the device automatically sends the recording data to the server. Communication between the recording device and the server is performed using a secure protocol. The input here is digital audio data, and the output is data transmission to the server.

[1036] Step 3:

[1037] The server temporarily stores the received voice data. Then, it converts the voice data into text data using a speech recognition API (for example, Google's speech recognition service). The input here is the received voice data, and the output is the converted text data.

[1038] Step 4:

[1039] The server uses an emotion recognition engine (e.g., the EmotionRecognition library) based on the text data to analyze the emotional information of the conference participants. This analysis result is also temporarily saved. The input is text data, and the output is the emotion analysis result.

[1040] Step 5:

[1041] The server uses a generative AI model (e.g., the Transformer library) to automatically generate a summary of the minutes from the text data and emotional information. The generated summary includes emotional information as well as key points and decisions. This summary is stored in a database. The input is text data and emotional information, and the output is a summary of the minutes.

[1042] Step 6:

[1043] The server searches the database for relevant past information based on the topic proposed during the meeting. Emotional information is also referenced to improve the accuracy of the search. The input is the proposed topic and emotional information, and the output is relevant past information.

[1044] Step 7:

[1045] The server summarizes the retrieved related information as suggestions and sends them to the terminal. The suggested information is displayed on the user's terminal. The input is related past information, and the output is the suggested information displayed on the user's terminal.

[1046] Step 8:

[1047] The user reviews the proposed information and contacts other meeting participants as needed, allowing them to send messages to related meeting participants and set up new meetings. The input is the proposed information, and the output is contact with participants and meeting settings.

[1048] The above is the specific processing flow of the present invention.

[1049] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1050] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1051] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1052] [Fourth embodiment]

[1053] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1054] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1055] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1056] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1057] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1058] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1059] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1060] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1061] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1062] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1063] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1064] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1065] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1066] The system of the present invention records the contents of a meeting, converts the recording data into text data, automatically creates a summary of the minutes, and stores the summary in a database. Furthermore, it can search for and propose related information from the database based on issues raised during the meeting. This system allows for efficient recording of meeting contents and makes it possible to utilize past knowledge and solutions in real time.

[1067] Specifically, the system operates in the following steps:

[1068] Program processing

[1069] 1. Recording of meeting content

[1070] User: Start the recording device as soon as the meeting starts and record the conversation using a dedicated recording app.

[1071] 2. Sending recording data

[1072] Device: After the meeting ends, the recording data is sent to the server. The recording app automatically sends the data.

[1073] 3. Converting Audio Data to Text

[1074] Server: Receives the transmitted voice data and converts it into text data using a speech recognition API. The converted text data is temporarily saved.

[1075] 4. Automatic summary creation

[1076] Server: Based on the text data converted from the voice data, the generative AI extracts important points and decisions, and automatically creates a summary of the minutes.

[1077] 5. Save summary

[1078] Server: Saves the summary of the minutes to a database. The database stores not only the summary but also the original text data.

[1079] 6. Search and suggest related information

[1080] Server: Based on the issues recorded by users during the meeting, the server searches the database to find relevant past minutes and solutions.

[1081] Server: Compiles relevant information into suggestions and sends them to the user's device.

[1082] 7. Information Display and Contact Functions

[1083] Terminal: Providing suggested relevant information to the user, the user has the option to review the required information and contact other meeting participants.

[1084] User: Uses the suggested information to send a message to relevant meeting participants and schedule a meeting if necessary.

[1085] Specific examples

[1086] For example, when a new product development team holds a meeting, they launch a dedicated recording app at the start of the meeting. After the meeting ends, the recording data is automatically sent to a server. The server receives the data and converts it into text. Generative AI extracts key points from the text data, creates a summary, and stores it in a database.

[1087] Next, the system searches the database for past discussion topics related to new product features proposed during the meeting. If relevant past meeting minutes or ideas are found, they are suggested to the user, who can use this information to develop new ideas.

[1088] Thus, the present invention can significantly improve the efficiency and productivity of meetings.

[1089] The processing flow will be explained below.

[1090] Step 1:

[1091] When a user starts a meeting, the recording device is activated and the conversation is recorded using a dedicated recording app.

[1092] Step 2:

[1093] When the meeting ends, the device automatically sends the recording data to the server, which is done automatically by the recording app.

[1094] Step 3:

[1095] The server receives the transmitted audio data and stores it temporarily.

[1096] Step 4:

[1097] The server uses a speech recognition API to convert the voice data into text data, which is then temporarily saved.

[1098] Step 5:

[1099] The server inputs the text data into the AI ​​generator, which automatically creates a summary of the minutes, extracting key points and decisions.

[1100] Step 6:

[1101] The server saves the created summary in a database, along with the original text data.

[1102] Step 7:

[1103] Based on the issues users record during meetings, the server searches the database to find minutes of past meetings and solutions.

[1104] Step 8:

[1105] The server compiles relevant information into suggestions and sends them to the terminal, where they are displayed to the user.

[1106] Step 9:

[1107] The user reviews the suggested information and selects the information they require. The user has the option to contact other meeting participants.

[1108] Step 10:

[1109] The user can then use the suggested information to send messages to relevant meeting participants and schedule meetings as needed.

[1110] Step 11:

[1111] The server feeds back the utilized information and results to the database, enabling more accurate proposals to be made in future meetings.

[1112] Example 1

[1113] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1114] Conventional meeting recording systems often require manual recording and management of meeting content, which requires time and effort. It is also difficult to quickly search past minutes and related information, making it difficult to receive immediate and effective feedback on issues raised during meetings. Furthermore, they lack the ability to share relevant information with other meeting participants, potentially reducing the efficiency and productivity of meetings.

[1115] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1116] In this invention, the server includes means for recording the contents of the meeting, means for converting the audio data into text data, means for automatically creating a summary of the minutes from the generated text data, means for saving the summary of the minutes in a database, means for searching and proposing related information from the database based on issues proposed during the meeting, means for displaying the proposed related information and providing a means for contacting other meeting participants, means for extracting important points from the text data using a generation AI, and means for transmitting the recorded data to the server. This automates the recording and management of meetings and enables past minutes and related information to be quickly searched and shared, thereby improving the efficiency and productivity of meetings.

[1117] "Means for recording the contents of a meeting" refers to a device or application for recording statements and conversations made during a meeting as audio data.

[1118] "Means for converting voice data into text data" refers to voice recognition technology or software for analyzing recorded voice data and converting it into text data.

[1119] The "means for automatically creating a summary of meeting minutes from generated text data" is a generative AI system that analyzes text data to extract and summarize important points and decisions.

[1120] The "means for storing the summary of the minutes in a database" refers to a device or system for storing the automatically created summary of the minutes in a relational database, cloud storage, or the like.

[1121] "A means for searching and proposing relevant information from a database based on issues proposed during a meeting" is an engine that searches a database for past minutes and solutions related to the issues entered by the user, extracts appropriate information, and presents it.

[1122] "Means for displaying suggested related information and providing a means for contacting other conference participants" refers to functionality or software that displays related information on a user interface and enables users to send messages to other conference participants or schedule meetings.

[1123] "Means for extracting key points from text data using generative AI" refers to technology that uses a generative AI model (e.g., a natural language processing model) to mechanically extract key points and decisions from text data.

[1124] The "means for transmitting recorded data to a server" refers to a communication function or protocol for transmitting recorded voice data to a server via the Internet.

[1125] This invention is a system that records the contents of a meeting, converts the recorded data into text data, automatically creates a summary of the minutes, stores it in a database, and searches for and suggests related information.

[1126] First, users turn on the recording device at the start of a meeting and record the conversation using a dedicated recording app, which can be software installed on a smartphone, tablet, or PC, such as Otter.ai or a similar application.

[1127] Then, after the meeting ends, the device sends the recording data to the server. The recording app does this automatically, and the data is sent using an encryption protocol (e.g., TLS). The data transmission is completed without any special operation by the user.

[1128] The server sends the received voice data to the Google Cloud Speech-to-Text API, which converts the voice data into text data. The converted text data is temporarily stored on the server. After conversion is complete, the server uses a generative AI (e.g., OpenAI GPT-3) to extract important points and decisions from the text data and create a summary of the meeting minutes.

[1129] The server stores the generated summary in a database (e.g., MySQL, PostgreSQL). The database stores not only the summary of the minutes but also the original text data, ensuring the integrity of the information.

[1130] Furthermore, based on the issues proposed during the meeting, the server searches the database and provides related past meeting minutes and solutions. This search uses an SQL query such as "SELECT FROM summaries WHERE text LIKE '%keywords related to the issue%'". The server compiles the relevant information found as suggestions and sends them to the user's device. This is sent in JSON format, and a dedicated suggestion display app receives and displays it.

[1131] The device displays the suggested related information to the user. The user can check the displayed information and, if necessary, send a message to other meeting participants or schedule a new meeting. By clicking the "Contact" button in the suggestion display app, a message sending screen will appear, allowing the user to send a message to the relevant participants.

[1132] Specific examples

[1133] When a new product development team holds a meeting, the user launches a dedicated recording app at the start of the meeting to record the conversation. After the meeting ends, the recorded data is automatically sent from the device to the server. The server receives this data and converts the speech to text using the Google Cloud Speech-to-Text API. OpenAI GPT-3 extracts key points from the converted text data and generates a summary of the meeting minutes. The generated summary is stored in a MySQL database.

[1134] Next, for new product features proposed during a meeting, the server searches the database to see if similar topics have been discussed in the past. If relevant minutes or ideas from past meetings are found, the server compiles them into a proposal and sends them in JSON format to the user's device. The user can then view the information through the proposal display app, send messages to participants in related meetings, and schedule new meetings.

[1135] Prompt Sentence Examples

[1136] "Please extract the key points and decisions from the following text data and create a summary of the meeting minutes.

[1137] Text data:

[1138] Proposing new product features at a meeting

[1139] Supervisor approves

[1140] We plan to work out the details at the next meeting.

[1141] ...

[1142] "

[1143] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1144] Step 1:

[1145] When the meeting starts, the user activates the recording device and records the conversation using a dedicated recording app. The recording app is installed on a smartphone, tablet, or PC, and an example is "Otter.ai." The user starts recording by pressing the "Start Recording" button. The input is an audio signal, and the output is voice data.

[1146] Step 2:

[1147] After the meeting ends, the device sends the recorded data to the server. When the recording app detects the end of the meeting, it automatically compresses the recorded data and sends it to the server using an encryption protocol (e.g., TLS). The input is the audio data, and the output is the data sent to the server.

[1148] Step 3:

[1149] The server stores the received voice data and sends it to the Google Cloud Speech-to-Text API to convert it into text. This API analyzes the voice data and outputs it as a string. The input is voice data and the output is text data.

[1150] Step 4:

[1151] The server temporarily stores the text data received from the Google Cloud Speech-to-Text API. It then supplies the text data to a generative AI (e.g., OpenAI GPT-3) with a prompt: "Please extract the key points and decisions from the following text data and create a summary of the minutes." The generative AI extracts the key points and generates a summary of the minutes. The input is the text data and the prompt, and the output is a summary of the minutes.

[1152] Step 5:

[1153] The server saves the generated summary of the minutes to a database. The server first establishes a database connection and executes the SQL query "INSERT INTO summaries (text, summary) VALUES (?, ?)". The original text data is also saved to the same database. The input is the summary of the minutes and the text data, and the output is the information saved in the database.

[1154] Step 6:

[1155] The server searches the database based on the issues recorded by the user during the meeting. For example, it runs the SQL query "SELECT FROM summaries WHERE text LIKE '%keywords related to the issue%'". The server filters the relevant information found and extracts the relevant information. The input is the issue keywords and the output is the relevant information.

[1156] Step 7:

[1157] The server compiles related information as suggestions and sends them to the user's device. The suggestions are sent in JSON format, and a dedicated suggestion display app receives them and displays them on the user interface. The input is the related information, and the output is the data sent to the device.

[1158] Step 8:

[1159] The device displays the suggested related information to the user. The suggestion display app parses the received JSON data and presents it to the user in a list format. The user can click on an interesting suggestion from the list to view details. The input is JSON data, and the output is the information displayed to the user.

[1160] Step 9:

[1161] Based on the proposed information, the user can send a message to the relevant meeting participants and set up a new meeting. For example, clicking the "Contact" button in the proposal display app will display the message sending screen, allowing the user to send a message to the relevant participants. The input is the user's action and the proposed information, and the output is the sent message.

[1162] (Application example 1)

[1163] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1164] Currently, in workplaces that use industrial machinery, it is difficult to efficiently create meeting records and utilize past related information in real time. Furthermore, there is a lack of a method to efficiently search and propose solutions to issues raised during meetings, which leads to inefficiency and reduced productivity. Furthermore, in industrial machinery workplaces, there is a need for a method to directly reflect the information in meeting minutes in machine operation.

[1165] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1166] In this invention, the server includes: means for recording the contents of a meeting; means for converting voice data into text data; means for automatically creating a summary of the minutes from the generated text data; means for saving the summary of the minutes in a database; means for searching and proposing related information from the database based on issues proposed during the meeting; means for displaying the proposed related information and providing a means for contacting other meeting participants; means for automatically generating meeting minutes and presenting search results for related information on the display device of the industrial machine in cooperation with the industrial machine control device; means for generating the proposed related information using a generative AI model; and means for generating a summary of the minutes using prompt sentences. This enables efficient recording of meeting contents and real-time utilization of past meeting information, thereby improving work efficiency and productivity at industrial machine sites.

[1167] A "means for recording the contents of a meeting" is a device or software that can record the audio content spoken during a meeting.

[1168] "Means for converting voice data into text data" refers to algorithms or tools for analyzing recorded voice data and converting it into text information.

[1169] The "means for automatically creating a summary of minutes from generated text data" is a means for extracting important points and decisions based on text data converted from speech and generating a summary.

[1170] The "means for saving the summary of the minutes in a database" refers to a function or device for writing and saving the generated summary in a database.

[1171] "A means for searching and proposing relevant information from a database based on issues proposed during a meeting" is a means for inputting the agenda items or issues raised during a meeting, searching a database for past information related to them, and presenting it.

[1172] "Means for displaying suggested related information and providing means for contacting other conference participants" refers to means for displaying searched related information in a user interface and providing options for communicating with other conference participants.

[1173] "Means for automatically generating meeting minutes and presenting search results for related information on the display device of the industrial machine in cooperation with the control device of the industrial machine" refers to means for exchanging data with the control device of the industrial machine and visually displaying the meeting minutes and related information on the display device of the machine.

[1174] A "means for generating suggested relevant information using a generative AI model" is a means for using an artificial intelligence model to automatically generate highly relevant information or solutions based on input data.

[1175] The "means for generating a summary of meeting minutes using prompt sentences" is a means for efficiently summarizing meeting minutes using pre-set questions or instructions (prompt sentences).

[1176] The system of the present invention realizes efficient creation of meeting records and real-time search for related information in workplaces where industrial machinery is heavily used. This system records the contents of meetings, converts the recording data into text data, automatically creates a summary of the minutes, and stores the summary in a database. Furthermore, it can search for and propose related information from the database based on issues raised during the meeting. This system allows efficient recording of meeting contents and utilization of past knowledge and solutions in real time.

[1177] Program processing

[1178] At the start of the meeting, the user launches a dedicated recording app on the industrial machine's tablet device and records the conversation. After the meeting ends, the recording data is automatically sent to the server. The server receives it and converts the audio data into text data using the Google Cloud Speech-to-Text API. A summary is then automatically created from the text data using a generative AI model such as OpenAI GPT-4. This generative AI model generates a summary using a prompt sentence. For example, the following prompt sentence can be used:

[1179] Example prompt: "Meeting minutes text: Today's meeting discussed new product development plans. Please write a summary."

[1180] The generated summary is stored in a MySQL database. Furthermore, based on issues recorded during the meeting, the database is searched to find relevant past meeting records and solutions. This involves using a generative AI model to generate relevant information and provide responses in the form of prompts based on the user's request. This relevant information is then presented on the display of the industrial machine, allowing the user to confirm the required information and contact other meeting participants.

[1181] Hardware and Software

[1182] The hardware used includes tablet devices and control devices for industrial machines, and the software used includes a meeting recording application, Google Cloud Speech-to-Text API, OpenAI GPT-4 generative AI models, and a MySQL database.

[1183] Specific examples

[1184] For example, the procedure for a new product development meeting is as follows: At the start of the meeting, the user launches a recording app on their tablet device and records the conversation. After the meeting ends, the recording app sends the audio data to a server, which then converts it into text data using the Google Cloud Speech-to-Text API. A generative AI model (OpenAI GPT-4) extracts key points from this text data and generates a summary of the meeting minutes. This summary is then stored in a MySQL database. Past related information about the new features proposed in the meeting is then searched for and presented on a display device. This information promotes the development of new ideas.

[1185] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1186] Step 1:

[1187] The user starts a dedicated recording app on the tablet device of the industrial machine and records the contents of the meeting. Remarks and discussions during the meeting are recorded as audio data in real time.

[1188] Input: Voice spoken during a meeting

[1189] Output: Recorded audio data file

[1190] Step 2:

[1191] After the conference ends, the device automatically sends the recorded data to the server. This data transfer occurs when the recording application is terminated.

[1192] Input: Recorded audio data file

[1193] Output: Audio data sent to the server

[1194] Step 3:

[1195] The server converts the received voice data into text data using the Google Cloud Speech-to-Text API. The voice data is analyzed and saved as text information.

[1196] Input: Audio data sent to the server

[1197] Output: Data converted to text

[1198] Step 4:

[1199] The server uses a generative AI model (OpenAI GPT-4) to automatically create a summary of the minutes from the converted text data. It uses prompts to extract key points and decisions and generate a summary.

[1200] Input: Text data

[1201] Output: A summary created by the generative AI model

[1202] Example prompt: "Meeting minutes text: Today's meeting discussed new product development plans. Please write a summary."

[1203] Step 5:

[1204] The server stores the generated summaries in a MySQL database, where they can be later stored in a searchable format, along with the original text data.

[1205] Input: Generated summary

[1206] Output: Summary and original text data stored in a database

[1207] Step 6:

[1208] The server searches a database for past meeting records and solutions based on the issues proposed during the meeting, generates relevant information using a generative AI model, and provides responses in the form of prompts tailored to the user's request.

[1209] Input: Issues proposed during the meeting

[1210] Output: Relevant information retrieved from the database

[1211] Step 7:

[1212] The terminal displays the suggested related information on the display device of the industrial machine, allowing the user to check the required information and contact other conference participants.

[1213] Input: Relevant information retrieved from the database

[1214] Output: relevant information presented on a display device

[1215] Processing flow example

[1216] 1. When the meeting starts, the user launches the recording app on the tablet device and records the conversation (step 1).

[1217] 2. After the meeting ends, the app automatically sends the recording data to the server (step 2).

[1218] 3. The server converts the audio data into text using the Google Cloud Speech-to-Text API (step 3).

[1219] 4. The server uses a generative AI model such as OpenAI GPT-4 to automatically create a summary from the generated text data, using the prompt sentence (Step 4).

[1220] 5. The generated summaries and the original text data are stored in a MySQL database (Step 5).

[1221] 6. Based on the issues proposed during the meeting, the server searches the database for past meeting records and solutions, and generates relevant information (step 6).

[1222] 7. The generated related information is presented on the display device of the industrial machine, allowing the user to check the necessary information and contact other conference participants (step 7).

[1223] Through the above steps, the system of the present invention realizes efficient recording of meeting contents and real-time utilization of past knowledge, thereby significantly improving work efficiency at industrial machinery sites.

[1224] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1225] The system according to the present invention records the contents of a meeting, converts the recording data into text data, automatically creates a summary of the minutes, and stores the summary in a database. Furthermore, the system can search for and suggest related information from the database based on the topics proposed during the meeting. Furthermore, by combining it with an emotion engine that recognizes the user's emotions, it can generate minutes with emotional information added and suggest information according to the emotions.

[1226] Specifically, the system operates in the following steps:

[1227] Program processing

[1228] 1. Recording of meeting content

[1229] User: When the meeting starts, the recording device is activated and the conversation is recorded. The user uses a dedicated recording app.

[1230] 2. Sending recording data

[1231] On your device: After the meeting ends, the recording data is automatically sent to the server. The recording app does this automatically.

[1232] 3. Converting Audio Data to Text

[1233] Server: Receives the transmitted audio data. The audio data is temporarily stored.

[1234] Server: Using a speech recognition API, the voice data is converted into text data, which is then temporarily saved.

[1235] 4. Acquiring emotional information

[1236] Server: Analyzes the emotions of meeting participants using an emotion engine from the recorded data. The analyzed emotional information is also temporarily saved.

[1237] 5. Automatic summary creation

[1238] Server: Based on text data and emotional information, the generative AI automatically creates a summary of the minutes, including key points and decisions, as well as analyzed emotional information.

[1239] 6. Save Summary

[1240] Server: The summary of the minutes is saved in a database. Not only the summary but also the original text data and emotional information are saved.

[1241] 7. Search and suggest related information

[1242] Server: Based on the issues recorded by users during meetings, the server searches the database to find relevant past meeting minutes and solutions, and also references emotional information to improve the accuracy of the search.

[1243] Server: The server compiles the searched related information into suggestions and sends them to the device. The suggested information is then displayed to the user.

[1244] 8. Information Display and Contact Functions

[1245] Terminal: Providing suggested relevant information to the user, the user has the option to review the required information and contact other meeting participants.

[1246] User: Uses the suggested information to send a message to relevant meeting participants and schedule a meeting if necessary.

[1247] Specific examples

[1248] For example, when a new product development team holds a meeting, they launch a recording app at the beginning of the meeting. After the meeting ends, the recording is automatically sent to the server. The server receives the recording and converts it into text using a speech recognition API. It then uses an emotion engine to analyze the emotions of the meeting participants.

[1249] The generative AI extracts key points and decisions based on text data and emotional information, and creates a summary. This summary is saved in a database. Next, it searches the database for past discussion topics similar to the new product features proposed during the meeting. By also referring to emotional information, more accurate related information is found. If relevant minutes or ideas from past meetings are found, they are suggested to the user.

[1250] The user can use this information to develop new ideas and contact relevant meeting participants as needed, thus significantly improving the efficiency and productivity of meetings.

[1251] The processing flow will be explained below.

[1252] Step 1:

[1253] When a user starts a meeting, the recording device is activated and the conversation is recorded using a dedicated recording app.

[1254] Step 2:

[1255] When the meeting ends, the device automatically sends the recording data to the server, which is done automatically by the recording app.

[1256] Step 3:

[1257] The server receives the transmitted audio data and stores it temporarily.

[1258] Step 4:

[1259] The server uses a speech recognition API to convert the voice data into text data, which is then temporarily saved.

[1260] Step 5:

[1261] The server uses an emotion engine to analyze the emotions of the conference participants from the received voice data, and the analyzed emotion information is also temporarily saved.

[1262] Step 6:

[1263] The server uses text data and emotional information to automatically generate a summary of the minutes using a generation AI, which extracts emotional information in addition to important points and decisions.

[1264] Step 7:

[1265] The server stores the generated summary in a database, which includes the original text data and sentiment information.

[1266] Step 8:

[1267] The server searches the database based on the issues recorded by the user during the meeting, finding minutes of past meetings and solutions, and also using emotional information to improve the accuracy of the search.

[1268] Step 9:

[1269] The server compiles the retrieved related information into suggestions and sends them to the terminal, where they are displayed to the user.

[1270] Step 10:

[1271] The user reviews the suggested information and selects the information they require. The user has the option to contact other meeting participants.

[1272] Step 11:

[1273] The user can then use the suggested information to send messages to relevant meeting participants and schedule meetings as needed.

[1274] Step 12:

[1275] The server feeds back the utilized information and results to the database, enabling more accurate proposals to be made in future meetings.

[1276] Example 2

[1277] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1278] Conventional meeting recording systems only record meeting content and convert the audio data into text to create minutes, but they do not offer advanced features such as creating minutes that take into account the emotional information of participants during the meeting, or searching and suggesting related information based on proposed issues. This makes it difficult to maximize the efficiency and effectiveness of meetings.

[1279] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for recording the contents of the meeting, means for converting voice data into text data, means for automatically creating a summary of the minutes from the generated text data, means for saving the summary of the minutes in a database, means for searching for and proposing related information from the database based on issues proposed during the meeting, means for displaying the proposed related information and contacting other meeting participants, means for acquiring emotional information and creating a summary of the minutes based on the emotions of the meeting participants, and means for searching for and proposing related information based on the emotional information. This makes it possible to create minutes that take into account not only the content of the meeting but also the emotions of the participants, thereby maximizing the efficiency and effectiveness of the meeting.

[1280] A "meeting recording device" is a device or software that captures and digitally stores audio from a meeting.

[1281] A "means for converting audio data to text data" is software or algorithms for analyzing recorded audio and converting it into a corresponding text format.

[1282] The "means for automatically creating a summary of meeting minutes from generated text data" is a system that analyzes text data, extracts important points and decisions from the meeting, and compiles them into a short summary.

[1283] The "means for storing the summary of the minutes in a database" refers to a data storage system for storing the generated summary of the minutes.

[1284] "A means for searching and proposing relevant information from a database based on issues proposed during a meeting" is a system for searching a database to find relevant information based on issues or questions raised during a meeting.

[1285] The "means for displaying suggested related information and contacting other conference participants" is a system for displaying suggested information to the user as a search result and contacting other conference participants as needed.

[1286] "Means for acquiring emotional information and creating a summary of meeting minutes based on the emotions of meeting participants" is a system that analyzes the emotions of participants from audio and text data during a meeting and creates a summary of meeting minutes based on that information.

[1287] "Means for searching and suggesting related information based on emotional information" is a system that searches for related information based on emotional analysis and makes more appropriate and useful suggestions.

[1288] The system according to the present invention records the contents of a meeting, converts the recording data into text data, automatically creates a summary of the minutes, and stores the summary in a database. It also has a function to search the database for relevant information based on the topics proposed during the meeting and suggest it. It also uses an emotion engine to analyze the user's emotions, and can create minutes and suggest relevant information based on the emotional information.

[1289] Hardware and software used

[1290] Hardware

[1291] Recording device: Smartphone or dedicated recorder

[1292] Server: A high-performance computer for data processing

[1293] Devices: Laptops, tablets, smartphones

[1294] software

[1295] Recording app: An application that records the contents of a meeting and sends the data to a server.

[1296] Speech Recognition API: Google Cloud Speech-to-Text

[1297] Emotion engine: IBM Watson Tone Analyzer

[1298] Generative AI model: OpenAI GPT

[1299] Database system: A relational database management system (RDBMS) for storing data.

[1300] Program Overview

[1301] 1. The user launches a recording app at the start of a meeting and records the contents of the meeting.

[1302] 2. After the meeting ends, the device automatically sends the recording data to the server.

[1303] 3. The server converts the received voice data into text using the Google Cloud Speech-to-Text API.

[1304] 4. The server runs IBM Watson Tone Analyzer on the text and audio data to obtain emotion information.

[1305] 5. The server uses OpenAI GPT to automatically create a summary of the minutes based on the text data and emotional information.

[1306] 6. The server stores the generated summary of the minutes, the original text data, and the emotion information in a database.

[1307] 7. The server searches the database based on the issues described during the meeting and suggests relevant information. The suggestions also refer to emotional information.

[1308] 8. The terminal displays the suggested relevant information to the user, providing the user with the option to review the required information and contact other conference participants as needed.

[1309] Specific operation example

[1310] Working example 1:

[1311] When a new product development team holds a meeting, they follow these steps:

[1312] 1. The user launches the recording app at the start of the meeting.

[1313] 2. After the meeting ends, the recording data is automatically sent to the server.

[1314] 3. The server receives the recording and converts it into text using Google Cloud Speech-to-Text.

[1315] 4. Additionally, IBM Watson Tone Analyzer is used to analyze the emotions of meeting participants.

[1316] 5. Create a summary of meeting minutes from text data and sentiment information using OpenAI GPT.

[1317] 6. The summary, original text data, and emotion information are stored in a database.

[1318] 7. The server searches the database based on the topic at hand and suggests relevant information.

[1319] 8. The terminal displays the information to the user and also provides options for contacting other conference participants.

[1320] Prompt Sentence Examples

[1321] A new product development meeting is being held. The recording app is launched to record the conversation, and after the meeting ends, the recording is automatically sent to the server. The server converts the text using Google Cloud Speech-to-Text and performs sentiment analysis using IBM Watson Tone Analyzer. A summary is created using OpenAI GPT and stored in a database. Related information is also searched and suggested. For example, the system searches past meeting minutes and related information to suggest new product features discussed during the meeting.

[1322] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1323] Step 1:

[1324] User: At the start of a meeting, the user launches a recording app. Specifically, the user taps the "Start Recording" button on the recording app installed on their smartphone or tablet. The input is the audio from the meeting, and the output is an audio file saved as recording data in the internal storage.

[1325] Step 2:

[1326] Device: After the meeting ends, the recording data is automatically sent to the server. When the user taps the "End Meeting" button, the app uploads the recording data to the server via the Internet. The input is the audio file stored in the device's internal storage, and the output is the audio file transferred to the server. Specifically, the file is uploaded using an HTTPS request.

[1327] Step 3:

[1328] Server: Receives and temporarily stores recorded data. Specifically, the server stores uploaded audio files in a storage directory. The input is the audio file sent from the device, and the output is the audio data stored in the server's internal storage.

[1329] Step 4:

[1330] Server: Uses a speech recognition API (Google Cloud Speech-to-Text) to convert voice data into text data. The server passes the voice data to the API and receives the converted text data. The input is an audio file as recorded data, and the output is conversation data in text format. Specifically, it sends requests to the API and receives the results.

[1331] Step 5:

[1332] Server: Analyzes the emotions of meeting participants from the recording data using an emotion engine (IBM Watson Tone Analyzer). The server sends text data to the API and receives the emotion information resulting from the analysis. The input is text data and the output is emotion information. Specifically, it sends requests to the emotion engine API and receives the results.

[1333] Step 6:

[1334] Server: Automatically creates a summary of meeting minutes based on text data and emotional information using a generative AI model (OpenAI GPT). The input is text data and emotional information, and the output is the generated summary. Specifically, the input data is passed to the generative AI model to generate the summary.

[1335] Step 7:

[1336] Server: The created summary of the minutes is saved in a database. The original text data and emotion information are also saved in the database. The input is the generated summary, text data, and emotion information, and the output is each data entry saved in the database. Specifically, data is added to the database using SQL statements.

[1337] Step 8:

[1338] Server: Searches the database based on the issues recorded by the user during the meeting and suggests related information. The server receives the user's issue information as input and searches for related entries in the database. The input is the issue information, and the output is related information as a search result. Specifically, it generates and executes the database query.

[1339] Step 9:

[1340] Terminal: Displays suggested relevant information to the user. The user can check the information through the app interface and contact other meeting participants if necessary. The input is the relevant information sent from the server, and the output is the information displayed on the user interface. Specifically, the terminal receives data from the API and displays it on the UI.

[1341] Step 10:

[1342] User: Based on the suggested information, send a message to relevant meeting participants and schedule a meeting if necessary. This is executed when the user enters a message in the app and taps the "Send" button. The input is the suggested related information and the message entered by the user, and the output is the sent message. Specifically, the message is sent using the message sending API.

[1343] (Application example 2)

[1344] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1345] Conventional meeting recording and automatic minutes creation systems have difficulty efficiently searching and suggesting emotional information and related past meeting information that can be used to solve problems in factories and other workplaces. Furthermore, the generated minutes summaries lack emotional information and provide insufficient detailed information that reflects the atmosphere of the discussion and the emotions of the meeting participants. This hinders rapid problem solving and efficient information sharing, especially in the workplace.

[1346] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes: means for recording the contents of the meeting; means for converting audio data into text data; means for automatically creating a summary of the minutes from the generated text data; means for saving the summary of the minutes in a database; means for searching and proposing related information from the database based on issues proposed during the meeting; means for displaying the proposed related information and contacting other meeting participants; means for recognizing emotional information and analyzing the emotions of the meeting participants; means for adding the analyzed emotional information to the summary; and means for providing meeting information for problem solving to machines installed in a factory. This enables accurate recording of the contents of the meeting and generation of a detailed summary including emotional information. Furthermore, by searching for and proposing related information based on issues proposed during the meeting, rapid problem solving and efficient information sharing can be achieved on-site.

[1347] "Means for recording the contents of a conference" refers to a device or method for recording the audio during a conference as digital data.

[1348] "Means for converting voice data into text data" refers to voice recognition technology or software for converting recorded voice data into text information.

[1349] The "means for automatically creating a summary of minutes from the generated text data" refers to an algorithm or program that summarizes the text data, extracts important information, and generates an outline of the minutes.

[1350] The "means for storing the summary of the minutes in a database" refers to a method or device for recording the generated summary of the minutes in an information storage system called a database.

[1351] "Means for searching and suggesting relevant information from a database based on an issue raised during a meeting" refers to a method or device for automatically searching a database for information related to an issue raised during a meeting and presenting that information to a user.

[1352] "Means for displaying suggested related information and contacting other conference participants" refers to an interface or method for displaying the searched information on a display or the like and for contacting other conference participants as necessary.

[1353] "Means for recognizing emotional information and analyzing the emotions of meeting participants" refers to technology and software for estimating and analyzing the emotional state of participants from the content of statements made during a meeting and their tone of voice.

[1354] The "means for adding analyzed emotional information to a summary" refers to a method or program for adding the results of analyzing the emotions of meeting participants to the summary of the minutes.

[1355] The "means for providing meeting information for problem solving to machines installed in the factory" refers to an interface or method for linking meeting minutes and related information to machines and management systems in the factory.

[1356] The program for implementing this invention records the audio of meetings and problem-solving meetings within a factory, processes the data to generate a summary of the minutes, and then performs emotion analysis to support efficient information sharing and problem solving. The specific processing and the hardware and software used are described below.

[1357] Hardware and Software Configuration

[1358] The server uses the following software:

[1359] Speech recognition API: A technology for converting voice data into text data. A typical example is Google's speech recognition API.

[1360] Emotion recognition engine: A technology for analyzing the emotions of meeting participants. The EmotionRecognition library is one example.

[1361] Generative AI model: A technology for generating a summary of meeting minutes from text data. A generative AI model using the Transformer library falls into this category.

[1362] Database: A system for storing generated minutes summaries and related information. It uses an SQLite database.

[1363] System processing overview

[1364] The server receives the audio data of the meeting recorded by the user and converts it into text data using a speech recognition API. The converted text data is then temporarily stored. Next, an emotion recognition engine is used to analyze emotional information from the text data, which is also temporarily stored. A generative AI model is then used to generate a summary of the minutes based on the text data and emotional information. This summary includes key points, decisions, and emotional information. The generated summary is then stored in a database.

[1365] Based on the topic proposed during the meeting, the system searches for related past information from the database. The server also refers to emotional information to improve the accuracy of the search. The retrieved related information is compiled as suggestions and sent to the user's device. The suggested information is displayed to the user, allowing them to contact other meeting participants.

[1366] This information is provided to machines and management systems within the factory, helping to quickly resolve problems on site.

[1367] Specific example explanation

[1368] For example, consider a meeting to improve the production efficiency of a new product. During the meeting, a recording device is running and the audio data is automatically sent to a server. The server converts the audio data into text data and performs sentiment analysis. A generative AI model uses the text data and sentiment information to generate a summary of the meeting minutes, including key points and decisions. This summary is stored in a database, and can be quickly searched and suggested when related information is needed in the future.

[1369] The minutes of past discussions regarding the "new production setup" problem proposed during this meeting can be referenced, allowing users to find effective solutions and contact other participants if necessary.

[1370] Prompt Sentence Examples

[1371] Here are some example prompts to input to a generative AI model:

[1372] "Based on the minutes of previous meetings, please tell us how discussions regarding bottlenecks on production lines have gone in the past."

[1373] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1374] Step 1:

[1375] At the start of a conference, a user activates a recording device to record the contents of the conference. The recording device records the conversation as digital audio data and temporarily stores the data. The input is the conversation, and the output is digital audio data.

[1376] Step 2:

[1377] After the meeting ends, the device automatically sends the recording data to the server. Communication between the recording device and the server is performed using a secure protocol. The input here is digital audio data, and the output is data transmission to the server.

[1378] Step 3:

[1379] The server temporarily stores the received voice data. Then, it converts the voice data into text data using a speech recognition API (for example, Google's speech recognition service). The input here is the received voice data, and the output is the converted text data.

[1380] Step 4:

[1381] The server uses an emotion recognition engine (e.g., the EmotionRecognition library) based on the text data to analyze the emotional information of the conference participants. This analysis result is also temporarily saved. The input is text data, and the output is the emotion analysis result.

[1382] Step 5:

[1383] The server uses a generative AI model (e.g., the Transformer library) to automatically generate a summary of the minutes from the text data and emotional information. The generated summary includes emotional information as well as key points and decisions. This summary is stored in a database. The input is text data and emotional information, and the output is a summary of the minutes.

[1384] Step 6:

[1385] The server searches the database for relevant past information based on the topic proposed during the meeting. Emotional information is also referenced to improve the accuracy of the search. The input is the proposed topic and emotional information, and the output is relevant past information.

[1386] Step 7:

[1387] The server summarizes the retrieved related information as suggestions and sends them to the terminal. The suggested information is displayed on the user's terminal. The input is related past information, and the output is the suggested information displayed on the user's terminal.

[1388] Step 8:

[1389] The user reviews the proposed information and contacts other meeting participants as needed, allowing them to send messages to related meeting participants and set up new meetings. The input is the proposed information, and the output is contact with participants and meeting settings.

[1390] The above is the specific processing flow of the present invention.

[1391] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1392] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1393] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1394] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1395] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1396] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1397] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1398] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1399] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1400] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1401] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1402] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1403] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1404] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1405] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1406] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1407] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1408] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1409] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1410] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1411] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1412] The following is further disclosed regarding the above embodiment.

[1413] (Claim 1)

[1414] A means of recording the meeting content;

[1415] means for converting voice data into text data;

[1416] A means for automatically creating a summary of the minutes from the generated text data;

[1417] means for storing a summary of the minutes in a database;

[1418] A means for searching and suggesting relevant information from a database based on the topic proposed during the meeting;

[1419] means for displaying suggested related information and providing a means for contacting other conference participants;

[1420] A system including:

[1421] (Claim 2)

[1422] 2. The system of claim 1, further comprising means for a user to review the relevant information retrieved from the database and select the information required.

[1423] (Claim 3)

[1424] 10. The system of claim 1, further comprising a means for a user to send messages to other conference participants and set up meetings based on the suggested related information.

[1425] "Example 1"

[1426] (Claim 1)

[1427] A means of recording the meeting content;

[1428] means for converting voice data into text data;

[1429] A means for automatically creating a summary of the minutes from the generated text data;

[1430] means for storing a summary of the minutes in a database;

[1431] A means for searching and suggesting relevant information from a database based on the topic proposed during the meeting;

[1432] means for displaying suggested related information and providing a means for contacting other conference participants;

[1433] A means of extracting key points from text data using generative AI;

[1434] means for transmitting the recording data to a server;

[1435] A system including:

[1436] (Claim 2)

[1437] 2. The system of claim 1, further comprising means for a user to review the relevant information retrieved from the database and select the information required.

[1438] (Claim 3)

[1439] 10. The system of claim 1, further comprising a means for a user to send messages to other conference participants and set up meetings based on the suggested related information.

[1440] "Application Example 1"

[1441] (Claim 1)

[1442] A means of recording the meeting content;

[1443] means for converting voice data into text data;

[1444] A means for automatically creating a summary of the minutes from the generated text data;

[1445] means for storing a summary of the minutes in a database;

[1446] A means for searching and suggesting relevant information from a database based on the topic proposed during the meeting;

[1447] means for displaying suggested related information and providing a means for contacting other conference participants;

[1448] a means for automatically generating minutes of a meeting and displaying search results for related information on a display device of the industrial machine in cooperation with the control device of the industrial machine;

[1449] means for generating suggested related information using a generative AI model;

[1450] a means for generating a summary of the proceedings using prompt statements;

[1451] A system including:

[1452] (Claim 2)

[1453] 2. The system of claim 1, further comprising means for a user to review the relevant information retrieved from the database and select the information required.

[1454] (Claim 3)

[1455] 10. The system of claim 1, further comprising a means for a user to send messages to other conference participants and set up meetings based on the suggested related information.

[1456] "Example 2: Combining Emotion Engines"

[1457] (Claim 1)

[1458] A means of recording the meeting content;

[1459] means for converting voice data into text data;

[1460] A means for automatically creating a summary of the minutes from the generated text data;

[1461] means for storing a summary of the minutes in a database;

[1462] A means for searching and suggesting relevant information from a database based on the topic proposed during the meeting;

[1463] a means of viewing suggested related information and contacting other meeting participants;

[1464] A means for acquiring emotion information and creating a summary of the minutes based on the emotions of the meeting participants;

[1465] A means for searching and suggesting related information based on emotion information;

[1466] A system including:

[1467] (Claim 2)

[1468] 2. The system of claim 1, further comprising means for a user to review the relevant information retrieved from the database and select the information required.

[1469] (Claim 3)

[1470] 10. The system of claim 1, further comprising a means for a user to send messages to other conference participants and set up meetings based on the suggested related information.

[1471] "Application example 2 when combining emotion engines"

[1472] (Claim 1)

[1473] A means of recording the meeting content;

[1474] means for converting voice data into text data;

[1475] A means for automatically creating a summary of the minutes from the generated text data;

[1476] means for storing a summary of the minutes in a database;

[1477] A means for searching and suggesting relevant information from a database based on the topic proposed during the meeting;

[1478] a means of viewing suggested related information and contacting other meeting participants;

[1479] means for recognizing emotional information and analyzing the emotions of conference participants;

[1480] means for adding the analyzed emotion information to the summary;

[1481] a means for providing problem-solving conference information to machines installed in the factory;

[1482] A system including:

[1483] (Claim 2)

[1484] a means for a user to review the relevant information retrieved from the database and select the information required;

[1485] 10. The system of claim 1, further comprising means for verifying information analyzing emotions of conference participants.

[1486] (Claim 3)

[1487] a means for the user to send a message to other conference participants and set up a meeting based on the suggested related information;

[1488] 10. The system of claim 1, further comprising a means for suggesting appropriate solutions based on sentiment analysis. [Explanation of symbols]

[1489] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. A means of recording the meeting content; means for converting voice data into text data; A means for automatically creating a summary of the minutes from the generated text data; means for storing a summary of the minutes in a database; A means for searching and suggesting relevant information from a database based on the topic proposed during the meeting; means for displaying suggested related information and providing a means for contacting other conference participants; A system including:

2. 2. The system according to claim 1, further comprising means for allowing a user to review the relevant information retrieved from the database and select the information required.

3. 10. The system of claim 1, further comprising a means for a user to send messages to other conference participants and set up meetings based on the suggested related information.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A