System
The system addresses inefficient information management by converting voice/data to text, extracting key points, generating organized minutes, and tracking decision history, enhancing project efficiency.
Patent Information
- Application Number
- JP2024124054
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-30
- Publication Date
- 2026-02-12
AI Technical Summary
In fast-paced business environments like finance and start-ups, managing rapidly updated information and accessing past meetings and decision-making histories is cumbersome, leading to inefficiencies for new members trying to catch up on projects.
A system that inputs voice or text data, converts it into text, extracts important points and decisions, generates and organizes minutes chronologically, creates a keyword map, and records decision-making history for easy access and search.
Facilitates quick understanding of past meetings and decision-making processes, improving project efficiency by enabling efficient information access and management.
Smart Images

Figure 2026022537000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] Describe the "problem the invention is trying to solve" and the "means for solving the problem."
[0005] In today's business environment, particularly in finance, research, and start-up companies, information management can be difficult in fields where information is updated rapidly and requires specialized knowledge. Accessing the contents of past meetings and decision-making histories is crucial for new members to quickly catch up on projects and work efficiently. However, if this information is managed in a cumbersome manner, it can be difficult for new members to properly acquire and understand the information, resulting in delays in project progress. The present invention aims to solve this problem by enabling new members to quickly access information and improve project efficiency. [Means for solving the problem]
[0006] The system of the present invention includes a means for inputting voice data or text data and converting the voice data into text data. It also includes a means for extracting important points and decisions from the converted text data and automatically generating minutes. The generated minutes are organized in chronological order and stored in a database. A means is also provided that allows users to access the chronologically organized minutes and search for minutes from a specific date. The system also includes a means for extracting technical terms from the minutes and creating a keyword map, and a means for recording and storing past decision-making history in a traceable format. This allows new participants to quickly understand the content of past meetings, technical terms, and decision-making processes, which is expected to improve project efficiency.
[0007] Understood. The following definitions are provided for key terms contained in the claims.
[0008] "Audio data" refers to digital data that records the contents of meetings or discussions in audio format.
[0009] "Text data" refers to character string data that has been converted from voice data or that has been directly input.
[0010] "Convert" refers to the operation of replacing audio data with text data.
[0011] "Extraction" refers to the process of extracting important information or keywords from text data.
[0012] A "minutes" is a document that summarizes and records the contents of a meeting or discussion.
[0013] "Organizing chronologically" means rearranging data in chronological order.
[0014] A "database" is a computer system or software for systematically storing and managing information.
[0015] "Access" refers to the operation of making specific information available for reading or use.
[0016] "Searching" is the act of manipulating a database to find the information you need.
[0017] "Jargon" refers to the specific words and expressions used in a particular field or industry.
[0018] A "keyword map" is a diagram or table that visually shows the relationships between extracted keywords.
[0019] "Decision-making history" refers to data that records the content of decisions made in an organization or project and any changes made to those decisions.
[0020] "Recording" is the act of storing information or data in a certain format.
[0021] "Tracing" refers to the act of examining a history to see what actions or changes have been taken in the past.
[0022] "Storage" refers to the act of saving and managing data or information. [Brief explanation of the drawings]
[0023] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4]FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0024] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0025] First, the terms used in the following description will be explained.
[0026] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0027] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0028] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0029] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0030] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0031] [First embodiment]
[0032] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0033] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0034] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0035] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0036] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0037] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0038] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0039] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0040] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0041] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0042] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0043] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0044] This invention is a system that can input voice data or text data, automatically generate minutes based on that data, and extract, store, and search important information. An embodiment of this system will be described below in natural language.
[0045] Data Input and Transformation
[0046] Users record audio data during a meeting and upload the audio file to their device after the meeting ends. The device then sends the uploaded audio file to the server. The server processes the received audio data by sending it to a generation AI, which converts it into text data. The generation AI then converts the audio into text and returns the text data to the server.
[0047] Automatic generation of meeting minutes
[0048] The server reprocesses the text data received from the generation AI and extracts important points and decisions. The generation AI automatically generates minutes based on this text data and returns them to the server. The server then stores the generated minutes in a database in chronological order.
[0049] Finding and Accessing Data
[0050] A user requests a search for a specific date to access the minutes of a past meeting. The server receives the request, retrieves the minutes corresponding to that date from the database, and sends them to the terminal. The terminal displays the minutes received from the server to the user.
[0051] Extraction and visualization of technical terms
[0052] After the minutes are generated, the server sends the data to the AI again to extract technical terms from the minutes. The AI analyzes the relationships between these technical terms and generates a keyword map. The generated keyword map is then stored in a database by the server.
[0053] Record and track decision history
[0054] Decisions made during meetings and their progress are recorded along with minutes and stored in a database by the server. When a user wants to check the decision-making history of a specific project, the server retrieves the relevant history from the database and sends it to the device. The device then displays the decision-making history to the user.
[0055] Specific examples
[0056] For example, if a user wants to search for a meeting record for a specific date, they can request, "Show me the meeting record for October 10, 2023." The server processes this request and retrieves the corresponding minutes from the database. The retrieved minutes are then sent to the terminal and displayed to the user.
[0057] Furthermore, if the user wants to check the decision-making history regarding the progress of the project, he or she can instruct "Show the decision-making history of project X." The server retrieves the decision-making history related to project X from the database, sends it to the terminal, and displays it to the user.
[0058] This allows users to quickly and efficiently understand the content of past and present meetings, the history of decision-making, and the relationships between technical terms used.
[0059] The processing flow will be explained below.
[0060] Understood. The program's processing will be explained in the following steps.
[0061] Step 1:
[0062] A user records audio data during a conference and uploads the audio file to a terminal after the conference ends.
[0063] Step 2:
[0064] The device sends the uploaded audio file to the server.
[0065] Step 3:
[0066] The server sends the received voice data to the generation AI, which converts the voice data into text data.
[0067] Step 4:
[0068] The generating AI converts the speech into text data and sends that text data back to the server.
[0069] Step 5:
[0070] The server reprocesses the text data received from the generation AI and sends it back to the generation AI to extract key points and decisions.
[0071] Step 6:
[0072] The generation AI extracts important points and decisions from the text data and automatically generates minutes, which are then sent back to the server.
[0073] Step 7:
[0074] The server stores the generated minutes in a database, assigning a timestamp to the minutes and organizing them in chronological order.
[0075] Step 8:
[0076] A user requests a search for a particular date to access minutes from a past meeting.
[0077] Step 9:
[0078] The server receives a request from the user, retrieves the minutes of the specified date from the database, and transmits them to the terminal.
[0079] Step 10:
[0080] The terminal displays the minutes received from the server to the user.
[0081] Step 11:
[0082] The server retrieves the minutes from the database, sends them to the generation AI, and extracts technical terms from the minutes.
[0083] Step 12:
[0084] The generative AI extracts technical terms from the minutes, analyzes their relationships, and generates a keyword map, which is then sent back to the server.
[0085] Step 13:
[0086] The server stores the generated keyword map in a database.
[0087] Step 14:
[0088] A user requests a search by specifying the project name to check the decision-making history of a particular project.
[0089] Step 15:
[0090] The server receives a request from a user, retrieves the decision-making history of the specified project from the database, and transmits it to the terminal.
[0091] Step 16:
[0092] The terminal displays the decision-making history received from the server to the user.
[0093] The above is the specific processing flow of this system.
[0094] Example 1
[0095] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0096] In today's business environment, it is important to efficiently manage meeting records and decision-making histories. However, manually creating meeting minutes and sorting through technical terms takes time and effort, and there is a high possibility of missing important information or delaying retrieval. There is also a lack of consistent record management to track past decision-making history. A system to solve these issues and improve meeting efficiency is needed.
[0097] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0098] In this invention, the server includes: means for inputting voice data or text data and converting the voice data to text data; means for extracting important points and decisions from the converted text data and automatically generating minutes; means for organizing the generated minutes in chronological order and saving them in a database; means for accessing the chronologically organized minutes and searching for minutes of a specific date; means for extracting technical terms from the minutes and creating a keyword map; means for recording and storing past decision-making history in a traceable form; means for converting voice data to text data using a generative AI model and extracting important information from the text data to automatically generate minutes; and means for searching for minutes of a specified date or project decision-making history and displaying them to the user. This allows for quick and accurate recording of meeting content, facilitating search and history tracking, and improving work efficiency.
[0099] "Audio Data" means human speech or other acoustic signals recorded during a meeting and stored in digital form.
[0100] "Text data" refers to character string information obtained as a result of analyzing voice data, or directly input character information.
[0101] "Conversion means" refers to technical equipment or software for converting audio data into text data, and may for example be a generative AI model.
[0102] Minutes are documents that organize and summarize the contents of a meeting and record the main topics and decisions made.
[0103] "Key points" or "decisions" refer to particularly important information or conclusions from a meeting and are the main contents that should be recorded in the minutes.
[0104] "Generating means" refers to technical devices or software for automatically creating minutes, including, for example, means for creating minutes from text data using a generative AI model.
[0105] "Database" refers to a system for efficiently storing, managing, and searching collected text data and generated minutes.
[0106] "Search means" refers to technical devices or software that search the minutes stored in the database based on specific criteria.
[0107] "Jargon" refers to specialized words and phrases frequently used in a particular field or industry.
[0108] A "keyword map" refers to a diagram or table that visually shows the relationships between technical terms extracted from minutes.
[0109] "Decision-making history" refers to historical information that records the decisions made during a meeting and the process behind them.
[0110] "Storage means" refers to technical devices and software that store decision-making history and generated minutes in a database in a traceable form.
[0111] A "generative AI model" refers to a machine learning model that uses artificial intelligence technology to analyze voice data and convert it into text data, or to generate minutes based on text data.
[0112] "User" refers to an individual or organizational member who operates the system to upload audio data, search meeting minutes, and check decision-making history.
[0113] This invention is a system that can input voice data or text data, automatically generate minutes based on that data, and extract, store, and search important information. This system aims to efficiently manage meeting records and includes the following components:
[0114] Data Input and Transformation
[0115] Users record audio data during a meeting and upload the audio file to their device after the meeting ends. The uploaded audio file is sent from the device to a server. The server then sends the received audio data to a generative AI model, which converts the audio data into text data. The generative AI model can use speech recognition technology such as the Google Cloud Speech-to-Text API. The generated text data is then returned to the server.
[0116] Automatic generation of meeting minutes
[0117] The server reprocesses the text data received from the generative AI model and uses natural language processing (NLP) techniques to extract key points and decisions. Based on the extracted information, the generative AI model automatically generates meeting minutes. The generated minutes are then stored in a database in chronological order by the server. For example, a database management system such as MySQL or MongoDB can be used.
[0118] Finding and Accessing Data
[0119] A user can search for minutes of past meetings by specifying a specific date. For example, a user can instruct "Show the meeting records for October 10, 2023." Upon receiving this request, the server searches the database for the corresponding minutes and sends them to the terminal. The terminal then displays the retrieved minutes to the user.
[0120] Extraction and visualization of technical terms
[0121] After the minutes are generated, the server sends the data to the generative AI model again to extract technical terms from the minutes. The generative AI model analyzes the relationships between these technical terms and generates a keyword map. The generated keyword map is stored in a database by the server and displayed to the user as needed.
[0122] Record and track decision history
[0123] The server records the decisions made during the meeting and their progress along with the minutes, and stores them in a database. If a user wants to check the decision-making history of a specific project, they can instruct the server to "show the decision-making history for project X." The server receives this request, retrieves the relevant history from the database, and sends it to the terminal. The terminal then displays the decision-making history to the user.
[0124] Specific examples
[0125] For example, if a user wants to search for meeting records for a specific date, they can request, "Show me the meeting records for October 10, 2023." The server processes this request and retrieves the corresponding minutes from the database. The retrieved minutes are then sent to the terminal and displayed to the user.
[0126] Furthermore, if the user wants to check the decision-making history regarding the progress of a project, he or she can instruct, for example, "Show the decision-making history of project X." The server retrieves the decision-making history related to project X from the database, sends it to the terminal, and displays it to the user.
[0127] This allows users to quickly and efficiently grasp the content of past and present meetings, the history of decision-making, and the relationships between technical terms used. Furthermore, the specific operation of this system improves work efficiency and the accuracy of information management.
[0128] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0129] Step 1:
[0130] A user records audio data during a meeting. Specifically, the user collects the audio data using a smartphone or a dedicated recording device. This audio data is stored in a digital format for subsequent processing.
[0131] Step 2:
[0132] After the meeting, the user uploads the recorded audio file to the device. The user accesses the system's web interface using a browser on their computer or smartphone, selects the audio file, and clicks the upload button. This saves the audio data to the device.
[0133] Step 3:
[0134] The device sends the uploaded audio file to the server. Specifically, it uses an HTTP request to send the audio data and user authentication information to the server. The server receives this.
[0135] Step 4:
[0136] The server receives the voice data and sends it to the generative AI model. The server converts the voice data into an appropriate format and sends a speech recognition request to the generative AI model (for example, Google Cloud Speech-to-Text API). The input data is an audio file, and the output data is text data.
[0137] Step 5:
[0138] The generative AI model converts the voice data into text data and returns the text data to the server. The generative AI model analyzes the voice data and outputs it as text data. Specifically, an algorithm is used to extract character string information from the voice. The server receives this text data and proceeds to the next step.
[0139] Step 6:
[0140] The server reprocesses the text data received from the generative AI model. The input data is text data, and the server uses natural language processing (NLP) techniques to extract key points and decisions. Specifically, it analyzes keywords in the text and identifies the necessary information.
[0141] Step 7:
[0142] The generative AI model automatically generates minutes based on the extracted information. The server sends the extracted information to the generative AI model, which then generates the minutes. The input data are key points and decisions, and the output data are the minutes. The generated minutes are well-formed documents suitable for future reference.
[0143] Step 8:
[0144] The server saves the generated minutes in a database in chronological order. The input data is the minutes, and a date tag is added when saving. For example, the minutes can be saved using a database management system such as MySQL or MongoDB. This makes it easy to search and organize the minutes.
[0145] Step 9:
[0146] A user searches for past meeting minutes by specifying a specific date. Specifically, the user enters "Show me the meeting minutes for October 10, 2023" into the system. This request is sent to the server.
[0147] Step 10:
[0148] The server receives a request from a user and retrieves the minutes corresponding to that date from the database. The input data is a search request, and the server searches for the corresponding minutes using an SQL query or similar. The output data is the retrieved minutes.
[0149] Step 11:
[0150] The server sends the acquired minutes to the terminal. Specifically, it returns the minutes data as an HTTP response. The terminal receives this.
[0151] Step 12:
[0152] The terminal displays the minutes received from the server to the user. The user can view the minutes via a browser or other device. The displayed minutes clearly show the contents of the meeting and the decisions made.
[0153] Step 13:
[0154] The server sends the generated minutes to the generative AI model again to extract technical terms. The input data is the minutes, and the generative AI model identifies technical terms and returns a list of technical terms as output data.
[0155] Step 14:
[0156] The generative AI model analyzes the relationships between technical terms and generates a keyword map. The input data is a list of technical terms, and the output data is a keyword map. Specifically, a diagram is generated that visually shows the relationships between technical terms.
[0157] Step 15:
[0158] The server saves the generated keyword map in a database. The input data is a keyword map, which is saved in a database in an appropriate format, allowing for later verification of terminology relationships.
[0159] Step 16:
[0160] A user submits a request to view the decision history of a particular project, for example, by asking the system, "Show me the decision history of project X." This request is sent to the server.
[0161] Step 17:
[0162] The server retrieves the relevant decision-making history from the database and sends it to the terminal. Input data includes project names and dates, and the server searches for the corresponding history and retrieves the decision-making history as output data.
[0163] Step 18:
[0164] The device displays the decision-making history to the user, who can then view it via a browser or other device. The displayed history clearly shows the progress of the project and important decisions.
[0165] (Application example 1)
[0166] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0167] In today's corporate environment, especially in factories and corporate meetings, there is a need to accurately record meeting content and refer to it later. However, manually creating meeting minutes is time-consuming and laborious, and it is difficult to grasp the relationships between important decisions and technical terms. In addition, there is a lack of systems that can process meeting audio data in real time and instantly generate meeting minutes. It is necessary to solve these issues and record and manage meeting content efficiently and accurately.
[0168] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0169] In this invention, the server includes means for recording voice data in real time and instantly converting it into text data, means for extracting important points and decisions from the converted text data in real time, and means for automatically saving the generated minutes and important points in a database. This makes it possible to efficiently record the contents of a meeting in real time and quickly extract and save important information.
[0170] "Audio data" refers to data in which the sounds of a meeting, conversation, etc. are recorded in digital format.
[0171] "Text data" is digital data that expresses voice data as character information.
[0172] The "means for converting voice data into text data" is a function that executes a process of converting voice data into text information using voice recognition technology.
[0173] "Means for extracting key points and decisions" is a process for automatically identifying and extracting the main agenda items and important decisions of a meeting from text data.
[0174] "Means for automatically generating minutes" is a function that creates a systematic record of the meeting contents based on extracted important points and decisions.
[0175] "Means of organizing in chronological order and storing in a database" refers to the process of organizing the generated minutes in chronological order and storing them in a digital database.
[0176] "A means to search for minutes of a specific date" is a function that searches for minutes saved in the past by specifying a date and displays the minutes you need.
[0177] "Means for extracting technical terms and creating and displaying a keyword map" refers to the process of automatically extracting related technical terms from text data, and creating and displaying a keyword map that visually shows the relevance of these terms.
[0178] "A means of recording decisions and storing the history in a traceable form" is a function that records the history of decisions made in meetings and stores them in a form that allows you to check later when and how the decisions were made.
[0179] The "means for recording in real time and instantly converting it into text data" is a process for recording audio data during a meeting on the spot and converting it into text data in real time without stopping.
[0180] "Means for automatically saving to a database" is a function that automatically saves the generated minutes and extracted important information to a database.
[0181] MODE FOR CARRYING OUT THE INVENTION
[0182] This invention is a system that converts voice data from meetings in factories or companies into text data in real time and automatically creates and saves minutes of the meetings. Specific embodiments of the system are described below.
[0183] System Overview
[0184] The system records audio data in real time and instantly converts it into text. It extracts important points and decisions from the generated text, automatically generates and saves meeting minutes, and extracts technical terms to create a keyword map and store past decision-making history in a traceable format.
[0185] Data Input and Transformation
[0186] During a meeting, users use a hardware device (such as a microphone or smartphone) to record audio data. The recorded audio data is automatically sent to a server. The server processes the received audio data by sending it to a generative AI model that uses speech recognition technology and converting it into text data. The generative AI model uses Hugging Face's pipeline ("automatic-speech-recognition").
[0187] Automatic generation of meeting minutes
[0188] The server reprocesses the text data obtained by speech recognition and extracts key points and decisions in real time using a generative AI summarization model (for example, Hugging Face's pipeline ("summarization")). The automatically generated minutes are organized in chronological order and stored in a database.
[0189] Finding and Accessing Data
[0190] A user can request a search for a specific date to access the minutes of past meetings. The server receives the request from the user, retrieves the minutes corresponding to that date from the database, and sends them to the terminal. The terminal displays the minutes received from the server to the user.
[0191] Extraction and visualization of technical terms
[0192] After the minutes are generated, the server uses the generative AI model again to extract technical terms from the minutes. The generative AI analyzes the relationships between these technical terms and generates a keyword map. The generated keyword map is then stored in a database by the server.
[0193] Record and track decision history
[0194] The decision-making process during the meeting is recorded along with the minutes, and the server stores them in a database. When a user wants to check the decision-making history for a specific project, the server retrieves the relevant history from the database, sends it to the user's device, and displays it to the user.
[0195] Specific examples
[0196] For example, if a user wants to search for a meeting record for a specific date, they can request, "Show me the meeting record for October 10, 2023." The server processes this request and retrieves the corresponding minutes from the database. The retrieved minutes are then sent to the terminal and displayed to the user.
[0197] Furthermore, if a user wants to check the decision-making history regarding the progress of a project, they can instruct "Display the decision-making history for project X." The server retrieves the decision-making history related to project X from the database, sends it to the terminal, and displays it to the user. This allows the user to quickly and efficiently understand the content of past and present meetings, the history of decision-making, and the relationships between the technical terms used.
[0198] Prompt Sentence Examples
[0199] Examples of input prompts for generative AI models include:
[0200] "View the meeting minutes for October 10, 2023"
[0201] "Show the decision history for Project X"
[0202] This allows users to use the system efficiently and easily access the information they need.
[0203] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0204] Step 1:
[0205] Users record audio data during a meeting. Audio data is collected using smart devices (smartphones and microphones). The input is in the form of audio data, which is sent to the next step.
[0206] Step 2:
[0207] The device uploads the recorded audio data to the server. The input is the audio data file, and the output is the transfer of the audio data to the server. Specifically, the device selects an audio file and sends it to the server via the Internet.
[0208] Step 3:
[0209] The server sends the received voice data to the generative AI model and converts it into text data. The input is voice data and the output is text data. Specifically, speech recognition is performed using Hugging Face's pipeline ("automatic-speech-recognition").
[0210] Step 4:
[0211] The server reprocesses the generated text data and extracts key points and decisions. The input is the text data, and the output is the extracted key points and decisions. The summary is generated using a generative AI summarization model (Hugging Face's pipeline("summarization")).
[0212] Step 5:
[0213] The server automatically generates minutes based on the extracted content and stores them in a database. The input is important points and decisions, and the output is the automatically generated minutes. Specifically, the minutes are organized in chronological order and stored in a database.
[0214] Step 6:
[0215] A user sends a search request to the server specifying a specific date to search for past meeting minutes. The input is a search query (e.g., "Show meeting minutes from October 10, 2023"), which the server receives.
[0216] Step 7:
[0217] The server processes the request from the user and retrieves the relevant minutes from the database. The input is the search query, and the output is the relevant minutes data. Specifically, the server queries the database based on the date and retrieves the results.
[0218] Step 8:
[0219] The server sends the acquired minutes data to the terminal. The input is the minutes data, and the output is the data transfer to the terminal. Specifically, the server sends the data to the terminal via the Internet.
[0220] Step 9:
[0221] The terminal receives the minutes from the server and displays them to the user. The input is the minutes data, and the output is a visual display to the user. The specific operation is that the minutes are displayed on the terminal's display.
[0222] Step 10:
[0223] The server extracts technical terms from the minutes and generates a keyword map. The input is the minutes data, and the output is a keyword map. A generative AI model is used to analyze and visualize the relationships between technical terms.
[0224] Step 11:
[0225] The server saves the keyword map in a database. The input is the keyword map and the output is the saved data. As a specific operation, the generated keyword map is stored in the database.
[0226] Step 12:
[0227] When a user wants to check the decision-making history of a particular project, the server retrieves the relevant history from the database, sends it to the terminal, and displays it to the user. The input is a database query for the decision-making history, and the output is the decision-making history data and its display. Specifically, the server executes the query, obtains the results, and sends them to the terminal, which then displays the received data to the user.
[0228] This series of processes allows for efficient recording of meeting content in real time, allowing important information to be instantly extracted, stored, and searched.
[0229] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0230] This invention is a system that can input voice data or text data, automatically generate minutes based on that data, and extract, store, and search important information, and it also combines an emotion engine that recognizes the user's emotions. Below, an embodiment of this system is explained in natural language.
[0231] Data Input and Transformation
[0232] Users record audio data during a conference and upload the audio file to their terminals after the conference ends. The terminals then transmit the uploaded audio file to a server.
[0233] Speech-to-text conversion
[0234] The server sends the received voice data to the generation AI, which converts the voice data into text data. The generation AI converts the voice into text and returns the text data to the server.
[0235] Emotion recognition
[0236] The server sends the converted text and voice data to the emotion engine, which then extracts emotion data from the voice and text and returns the emotion data to the server.
[0237] Automatic generation of meeting minutes
[0238] The server reprocesses the text data and emotion data received from the generation AI to extract important points and decisions. The generation AI automatically generates minutes based on this data and returns them to the server. The server then stores the generated minutes in a database in chronological order along with the emotion data.
[0239] Finding and Accessing Data
[0240] To access the minutes of past meetings, a user requests a search by specifying a specific date. The server receives the request from the user, retrieves the minutes and emotion data for the corresponding date from the database, and sends them to the device. The device then displays the minutes and emotion data received from the server to the user.
[0241] Extraction and visualization of technical terms
[0242] After the minutes are generated, the server sends the data to the generation AI again, which extracts technical terms from the minutes. The generation AI analyzes the relationships between these technical terms and generates a keyword map. The generated keyword map is then stored in a database by the server. Emotional data is also reflected in the keyword map, visualizing changes in emotions.
[0243] Record and track decision history
[0244] Decisions made during meetings and their progress are recorded along with minutes and stored in a database by the server. When a user wants to check the decision-making history of a specific project, the server retrieves the relevant history from the database and sends it along with emotion data to the device, which then displays it to the user.
[0245] Specific examples
[0246] For example, if a user wants to search for a meeting record for a specific date, they can request, "Show me the meeting record for October 10, 2023." The server processes this request and retrieves the corresponding meeting minutes and emotion data from the database. The retrieved meeting minutes are sent to the device and displayed to the user. Based on the emotion data, the server also displays the overall emotional changes during the meeting.
[0247] Furthermore, if the user wants to check the decision-making history regarding the progress of the project, they can instruct "Show the decision-making history for project X." The server retrieves the decision-making history and emotion data related to project X from the database, sends them to the terminal, and displays them to the user. By also displaying the emotion data, it is possible to grasp the trend of emotions as the project progresses.
[0248] This allows users to quickly and efficiently grasp the content of past and present meetings, the history of decision-making, the relationships between technical terms used, and changes in emotions.
[0249] The processing flow will be explained below.
[0250] Understood. The specific operation will be explained below, divided into processing steps.
[0251] Step 1:
[0252] A user records audio data during a conference and uploads the audio file to a terminal after the conference ends.
[0253] Step 2:
[0254] The device sends the uploaded audio file to the server.
[0255] Step 3:
[0256] The server sends the received voice data to the emotion engine and generation AI, which converts the voice data into text data and recognizes emotions.
[0257] Step 4:
[0258] The generating AI converts the speech into text data and sends that text data back to the server.
[0259] Step 5:
[0260] The emotion engine recognizes the user's emotion from the voice data and returns the emotion data to the server.
[0261] Step 6:
[0262] The server reprocesses the text data received from the generation AI and the emotion data received from the emotion engine, and sends it back to the generation AI to extract important points and decisions.
[0263] Step 7:
[0264] The generation AI extracts important points and decisions from the text data and automatically generates minutes, which are then sent back to the server.
[0265] Step 8:
[0266] The server stores the generated minutes and emotion data in a database. The minutes are time-stamped and organized in chronological order.
[0267] Step 9:
[0268] A user submits a request specifying a specific date to search for minutes of past meetings.
[0269] Step 10:
[0270] The server receives a request from the user, retrieves the minutes of the specified date and the corresponding emotion data from the database, and transmits them to the terminal.
[0271] Step 11:
[0272] The device receives the minutes and emotion data from the server and displays them to the user, allowing the user to check the emotional trends during the meeting as well as the content of the meeting.
[0273] Step 12:
[0274] The server retrieves the minutes from the database, sends them to the generation AI, and extracts technical terms from the minutes.
[0275] Step 13:
[0276] The generative AI extracts technical terms from the minutes, analyzes their relationships, and generates a keyword map, which is then sent back to the server.
[0277] Step 14:
[0278] The server stores the generated keyword map in a database, and also reflects the emotion data in this map so that it can be displayed to the user.
[0279] Step 15:
[0280] A user submits a request specifying the project name to check the decision history of a specific project.
[0281] Step 16:
[0282] The server receives a request from a user, retrieves the decision-making history and emotion data for the specified project from the database, and sends them to the terminal.
[0283] Step 17:
[0284] The device receives the decision-making history and emotional data from the server and displays it to the user, allowing the user to understand the decision-making process as the project progresses and the emotional trends that occurred during that process.
[0285] The above is the specific processing flow of this system that combines an emotion engine.
[0286] Example 2
[0287] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0288] In today's business environment, efficient meeting progress and recording are crucial. However, accurately recording important discussions and decisions made during meetings, as well as participants' emotions, and storing and retrieving these in a format that can be used later, can be time-consuming. Furthermore, extracting relevant technical terms and emotional changes from minutes and visualizing them can be difficult. For these reasons, there is a need for a way for users to quickly grasp the details of past meeting records and decisions.
[0289] The specific processing by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for inputting voice data or text data and converting the voice data into text data, means for extracting emotion data from the converted text data, means for extracting important points and decisions from the text data and emotion data and automatically generating minutes, means for organizing the generated minutes and emotion data in chronological order and storing them in a database, means for accessing the chronologically organized minutes and emotion data and searching for minutes of a specific date, means for extracting technical terms from the minutes, creating a keyword map, and visualizing changes in emotion, and means for recording past decision-making history and storing it in a traceable form. This allows users to efficiently manage and quickly access important information from meetings.
[0290] "Audio data" refers to digital data of audio recorded during a meeting or conversation.
[0291] "Text data" is digital data of character information converted from audio data.
[0292] "Means of conversion" refers to the technology or device that analyzes voice data and converts it into text data as character information.
[0293] "Emotion data" refers to data that indicates emotional states and changes extracted from text and speech analysis.
[0294] "Extraction means" refers to technology or equipment that extracts key points, decisions, or emotional data from text data or audio data.
[0295] "Means for automatically generating minutes" refers to technology or devices that automatically generate documents summarizing meeting content and decisions based on text data and emotion data.
[0296] A "database" is a digital system or software for systematically storing and managing minutes and emotional data.
[0297] "Searchable means" refers to technology or devices that search for required data based on specific dates or conditions from information stored in a database.
[0298] "Terminology" refers to specialized terms and expressions used in a particular field or industry.
[0299] A "keyword map" is a diagram or chart that visually represents the relationships and interactions of extracted technical terms.
[0300] "Visualization means" refers to technology or devices that visually display extracted data.
[0301] "Decision-making history" is data that records the details and process of decisions made in meetings and projects.
[0302] "Means for storing data in a traceable form" refers to technology or devices that store decision-making history in a database for future reference.
[0303] This invention provides a system for processing audio data and text data from a meeting, automatically generating minutes, and efficiently managing important information. This system includes functions for converting audio data, recognizing emotions, automatically generating minutes, and searching and visualizing data. Specific embodiments of the system are described below.
[0304] Entering data
[0305] Users upload audio data recorded during a meeting to their device, which saves the audio in a standard digital audio format (e.g., MP3 or WAV files), and the device sends the uploaded audio file to the server.
[0306] Speech-to-text conversion
[0307] The server sends the received voice data to a speech recognition service (such as Google Cloud Speech-to-Text API), which is a generating AI, and converts the voice data into text data. The generating AI analyzes the voice and converts the content into text data as character information. The converted text data is then sent back to the server.
[0308] Emotion recognition
[0309] The server then sends the converted text and voice data to an emotion recognition engine (e.g., IBM Watson Tone Analyzer) to recognize the user's emotions. The emotion engine extracts emotion data (e.g., joy, anger, sadness, etc.) from the voice tone and text content. This emotion data is then sent back to the server.
[0310] Automatic generation of meeting minutes
[0311] The server uses a generative AI model (e.g., OpenAI's GPT-3) to extract key points and decisions based on the generated text data and emotion data, and automatically generates meeting minutes. The generated minutes include important discussions and decisions made during the meeting. The minutes and emotion data are organized chronologically by the server and stored in a database.
[0312] Finding and Accessing Data
[0313] A user can submit a search request to access meeting minutes for a specific date, for example, using a specific prompt such as "Show me the meeting minutes for October 10, 2023." The server receives this request and retrieves the corresponding meeting minutes and emotion data from the database. The retrieved data is then sent to the device and displayed to the user.
[0314] Extracting technical terms and generating keyword maps
[0315] The server then sends the generated minutes to a generation AI (e.g., BERT) to extract technical terms from the minutes. The generation AI then analyzes the relationships between the extracted technical terms and generates a keyword map. This keyword map is also stored in a database and displayed visually along with sentiment data.
[0316] Record and track decision history
[0317] The server records the decisions made during the meeting and their process, and stores them in a database along with the minutes. If a user wants to check the decision-making history of a specific project, they can use a specific prompt, such as "Show the decision-making history of project X." The server retrieves the relevant data from the database, sends it to the terminal, and displays it to the user.
[0318] This allows users to quickly and efficiently manage and understand past and present meeting details, decision details, terminology relationships, and emotional changes.
[0319] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0320] Step 1:
[0321] Users upload audio data recorded during a meeting to their device. The audio data is saved in a digital audio file format such as MP3 or WAV. The device then sends the uploaded audio file to the server. The input is the audio data, and the output is the audio file sent to the server.
[0322] Step 2:
[0323] The server sends the received voice data to a generation AI (for example, Google Cloud Speech-to-Text API), which is a voice recognition service. The generation AI analyzes the voice data and generates text data as character information. The input is voice data, and the output from the generation AI is text data. The server receives this text data and proceeds to the next step of processing.
[0324] Step 3:
[0325] The server sends the converted text data and the original voice data to an emotion recognition engine (e.g., IBM Watson Tone Analyzer). The emotion recognition engine extracts emotion data from these data. The input is text data and voice data, and the output is emotion data. The server receives the emotion data and proceeds to the next step of processing.
[0326] Step 4:
[0327] The server uses a generative AI model (e.g., OpenAI's GPT-3) based on the text data and emotion data to extract key points and decisions, and automatically generate meeting minutes. The input is text data and emotion data, and the output is automatically generated meeting minutes. The generated minutes are used for further processing.
[0328] Step 5:
[0329] The server organizes the generated minutes and emotion data in chronological order and stores them in a database (e.g., MySQL or PostgreSQL). The input is the minutes and emotion data, and the output is visualized data stored in the database.
[0330] Step 6:
[0331] To access minutes for a specific date, a user sends a search request through their terminal. For example, they request "Show the meeting records for October 10, 2023." The input is the search request, and the server receives this request and retrieves the corresponding minutes and emotion data from the database. The output is the retrieved minutes and emotion data.
[0332] Step 7:
[0333] The terminal displays the minutes and emotion data obtained from the server to the user through a user interface (e.g., a web browser or a dedicated app). The input is the obtained minutes and emotion data, and the output is the meeting record displayed to the user.
[0334] Step 8:
[0335] The server then sends the generated transcripts to a generation AI (e.g., BERT) to extract technical terms from the transcripts and generate a keyword map. The input is the transcripts, and the output is the generated keyword map. The generated keyword map is also stored in a database and can be visualized along with the sentiment data.
[0336] Step 9:
[0337] The server stores the decision-making process and its details in a database along with the meeting minutes. If a user wants to check the decision-making history of a specific project, for example, they can request, "Show the decision-making history of project X." The server retrieves the relevant data from the database, sends it to the terminal, and displays it to the user. The input is the search request and the decision-making history, and the output is the project's decision-making history displayed to the user.
[0338] (Application example 2)
[0339] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0340] Conventional meeting and online event minutes generation systems require manual minutes creation, which is time-consuming and labor-intensive. Furthermore, it is difficult to accurately capture important information and changes in user emotions, making it difficult to quickly grasp decision-making and analyze emotional trends. This creates the risk that some meeting content may be overlooked, and it is difficult to accurately track the decision-making process.
[0341] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for inputting voice data or text data and converting the voice data into text data, means for extracting important points and decisions from the converted text data and automatically generating minutes, means for organizing the generated minutes in chronological order and saving them in a database, means for accessing the chronologically organized minutes and searching for minutes of a specific date, means for extracting technical terms from the minutes and creating a keyword map, means for recording past decision-making history and storing it in a traceable form, means for visualizing user emotional trends using the generated minutes and emotion data, and means for processing the voice data based on prompts and searching for and displaying specific events or meeting records. This enables the content of meetings and online events to be quickly and efficiently recorded, searched, and analyzed.
[0342] "Audio data" refers to data in which audio is recorded in digital format.
[0343] "Text data" refers to data in which character information is recorded in digital format.
[0344] "Means for converting" refers to an algorithm or machine for converting audio data into text data.
[0345] "Key points" are information or matters that should be given special importance in the content of a meeting or event.
[0346] "Decisions" are matters that were discussed at meetings or events and ultimately decided upon.
[0347] Minutes are documents that record the contents of a meeting or event in chronological order.
[0348] "Chronology" refers to the arrangement of events or data in chronological order.
[0349] A "database" is a storage system that systematically stores data and allows it to be searched and updated.
[0350] "Searchable means" refers to methods and technologies for extracting and viewing data based on specific criteria.
[0351] "Terminology" is a term with a specific meaning used in a particular field or industry.
[0352] A "keyword map" is a graphical tool for visually displaying relevant keywords within minutes or other documents.
[0353] "Decision-making history" is a record of decisions and the processes involved in past meetings and events.
[0354] "Emotion data" is data that represents the user's emotional state and is the result of emotion analysis.
[0355] A "means for visualizing emotion trends" is a method for visually displaying changes in a user's emotions over time.
[0356] A "prompt sentence" is a specific instruction sentence that is input to a generative AI model.
[0357] "Processing means" refers to the technology and devices used to analyze, convert and display audio and text data.
[0358] "Means for searching and displaying events and meeting records" means methods or technologies for searching events and meeting records based on specific criteria and displaying the results.
[0359] This invention is a system that converts audio data from meetings and online events into text data in real time, extracts important points and decisions from the text data, and automatically generates minutes.The system also has the function of visualizing users' emotional trends using the generated minutes and emotion data, and searching and displaying specific event or meeting records based on prompt sentences.
[0360] The system utilizes the following hardware and software:
[0361] Hardware used:
[0362] Microphone: For voice input
[0363] Server: storing and processing audio files
[0364] User device: Device used to display results (smartphone, tablet, PC, etc.)
[0365] Software used:
[0366] speech_recognition library: convert speech to text
[0367] transformers library: sentiment analysis
[0368] Custom database module: storing meeting minutes data
[0369] pipeline (natural language processing): processing emotional data
[0370] Process flow:
[0371] During a meeting, users record audio data using a microphone. This audio data is uploaded to the server via the terminal after the meeting ends. The server converts the uploaded audio file into text data using the speech_recognition library.
[0372] The server then sends the converted text data to the transformers library for sentiment analysis. The resulting sentiment data is then sent back to the server. The server then extracts key points and decisions from the text and sentiment data and automatically generates meeting minutes. The generated minutes are then organized and stored in chronological order using a custom database module.
[0373] Users can access meeting minutes by specifying a specific date. The server can also search for past meeting records and emotion data, and display related events and meeting records on the device based on a specific prompt. For example, by entering the prompt "View the meeting records for the live event on October 10, 2023," the corresponding meeting minutes and emotion data will be displayed.
[0374] Furthermore, the server visualizes the user's emotional trends using the generated minutes and emotional data, allowing users to visually grasp changes in emotions during meetings and events.
[0375] In this way, the content of meetings and online events can be recorded, searched and analyzed quickly and efficiently.
[0376] Examples of prompts:
[0377] "View the recording of the live event on October 10, 2023"
[0378] The prompt allows the user to easily search and view meeting records for a particular date.
[0379] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0380] Step 1:
[0381] During a meeting, a user uses a microphone to record audio. The recorded audio data is saved on the device. The input is audio data, and the output is an audio file.
[0382] Step 2:
[0383] After the conference, the user uploads the recorded audio data to the server using the terminal. In this process, the audio file is the input, and the audio file sent to the server is the output.
[0384] Step 3:
[0385] The server receives the uploaded audio file and converts it into text using the speech_recognition library. The input is the audio file, and the output is the generated text. The server then processes the sound waves to convert them into text.
[0386] Step 4:
[0387] The converted text data is sent by the server to the sentiment analysis function of the transformers library. The input is text data and the output is emotion data. The server recognizes the emotion of the text content and generates emotion data.
[0388] Step 5:
[0389] The server automatically generates minutes by extracting important points and decisions based on the generated text data and emotion data. The input is text data and emotion data, and the output is the automatically generated minutes. The server uses natural language processing to extract important information.
[0390] Step 6:
[0391] The generated minutes are organized in chronological order by the server and saved in a custom database module. The input is the minutes, and the output is the minutes data organized in chronological order. The server stores the data in the database.
[0392] Step 7:
[0393] A user searches for meeting minutes for a specific date on their device by entering the prompt "Show meeting notes for the live event on October 10, 2023." The input is the search prompt, and the output is the meeting minutes for the specified date and emotion data.
[0394] Step 8:
[0395] The server searches the database based on the prompt sentence, obtains the relevant minutes and emotion data, and sends them to the terminal. The input is the prompt sentence and the database query, and the output is the search results: the minutes and emotion data.
[0396] Step 9:
[0397] The user checks the minutes and emotion data displayed on the device. The server generates graphs and dashboards to visually display emotion trends based on the emotion data. The input is emotion data, and the output is visualized emotion trends.
[0398] This series of processing steps enables the content of meetings and online events to be recorded, searched, and analyzed quickly and efficiently.
[0399] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0400] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0401] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0402] [Second embodiment]
[0403] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0404] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0405] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0406] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0407] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0408] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0409] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0410] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0411] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0412] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0413] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0414] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0415] This invention is a system that can input voice data or text data, automatically generate minutes based on that data, and extract, store, and search important information. An embodiment of this system will be described below in natural language.
[0416] Data Input and Transformation
[0417] Users record audio data during a meeting and upload the audio file to their device after the meeting ends. The device then sends the uploaded audio file to the server. The server processes the received audio data by sending it to a generation AI, which converts it into text data. The generation AI then converts the audio into text and returns the text data to the server.
[0418] Automatic generation of meeting minutes
[0419] The server reprocesses the text data received from the generation AI and extracts important points and decisions. The generation AI automatically generates minutes based on this text data and returns them to the server. The server then stores the generated minutes in a database in chronological order.
[0420] Finding and Accessing Data
[0421] A user requests a search for a specific date to access the minutes of a past meeting. The server receives the request, retrieves the minutes corresponding to that date from the database, and sends them to the terminal. The terminal displays the minutes received from the server to the user.
[0422] Extraction and visualization of technical terms
[0423] After the minutes are generated, the server sends the data to the AI again to extract technical terms from the minutes. The AI analyzes the relationships between these technical terms and generates a keyword map. The generated keyword map is then stored in a database by the server.
[0424] Record and track decision history
[0425] Decisions made during meetings and their progress are recorded along with minutes and stored in a database by the server. When a user wants to check the decision-making history of a particular project, the server retrieves the relevant history from the database and sends it to the device. The device then displays the decision-making history to the user.
[0426] Specific examples
[0427] For example, if a user wants to search for a meeting record for a specific date, they can request, "Show me the meeting record for October 10, 2023." The server processes this request and retrieves the corresponding minutes from the database. The retrieved minutes are then sent to the terminal and displayed to the user.
[0428] Furthermore, if the user wants to check the decision-making history regarding the progress of the project, he or she can instruct "Show the decision-making history of project X." The server retrieves the decision-making history related to project X from the database, sends it to the terminal, and displays it to the user.
[0429] This allows users to quickly and efficiently understand the content of past and present meetings, the history of decision-making, and the relationships between technical terms used.
[0430] The processing flow will be explained below.
[0431] Understood. The program's processing will be explained in the following steps.
[0432] Step 1:
[0433] A user records audio data during a conference and uploads the audio file to a terminal after the conference ends.
[0434] Step 2:
[0435] The device sends the uploaded audio file to the server.
[0436] Step 3:
[0437] The server sends the received voice data to the generation AI, which converts the voice data into text data.
[0438] Step 4:
[0439] The generating AI converts the speech into text data and sends that text data back to the server.
[0440] Step 5:
[0441] The server reprocesses the text data received from the generation AI and sends it back to the generation AI to extract key points and decisions.
[0442] Step 6:
[0443] The generation AI extracts important points and decisions from the text data and automatically generates minutes, which are then sent back to the server.
[0444] Step 7:
[0445] The server stores the generated minutes in a database, assigning a timestamp to the minutes and organizing them in chronological order.
[0446] Step 8:
[0447] A user requests a search for a particular date to access minutes from a past meeting.
[0448] Step 9:
[0449] The server receives a request from the user, retrieves the minutes of the specified date from the database, and transmits them to the terminal.
[0450] Step 10:
[0451] The terminal displays the minutes received from the server to the user.
[0452] Step 11:
[0453] The server retrieves the minutes from the database, sends them to the generation AI, and extracts technical terms from the minutes.
[0454] Step 12:
[0455] The generative AI extracts technical terms from the minutes, analyzes their relationships, and generates a keyword map, which is then sent back to the server.
[0456] Step 13:
[0457] The server stores the generated keyword map in a database.
[0458] Step 14:
[0459] A user requests a search by specifying the project name to check the decision-making history of a particular project.
[0460] Step 15:
[0461] The server receives a request from a user, retrieves the decision-making history of the specified project from the database, and transmits it to the terminal.
[0462] Step 16:
[0463] The terminal displays the decision-making history received from the server to the user.
[0464] The above is the specific processing flow of this system.
[0465] Example 1
[0466] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0467] In today's business environment, it is important to efficiently manage meeting records and decision-making histories. However, manually creating meeting minutes and sorting through technical terms takes time and effort, and there is a high possibility of missing important information or delaying retrieval. There is also a lack of consistent record management to track past decision-making history. A system to solve these issues and improve meeting efficiency is needed.
[0468] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0469] In this invention, the server includes: means for inputting voice data or text data and converting the voice data to text data; means for extracting important points and decisions from the converted text data and automatically generating minutes; means for organizing the generated minutes in chronological order and saving them in a database; means for accessing the chronologically organized minutes and searching for minutes of a specific date; means for extracting technical terms from the minutes and creating a keyword map; means for recording and storing past decision-making history in a traceable form; means for converting voice data to text data using a generative AI model and extracting important information from the text data to automatically generate minutes; and means for searching for minutes of a specified date or project decision-making history and displaying them to the user. This allows for quick and accurate recording of meeting content, facilitating search and history tracking, and improving work efficiency.
[0470] "Audio Data" means human speech or other acoustic signals recorded during a meeting and stored in digital form.
[0471] "Text data" refers to character string information obtained as a result of analyzing voice data, or directly input character information.
[0472] "Conversion means" refers to technical equipment or software for converting audio data into text data, and may for example be a generative AI model.
[0473] Minutes are documents that organize and summarize the contents of a meeting and record the main topics and decisions made.
[0474] "Key points" or "decisions" refer to particularly important information or conclusions from a meeting and are the main contents that should be recorded in the minutes.
[0475] "Generating means" refers to technical devices or software for automatically creating minutes, including, for example, means for creating minutes from text data using a generative AI model.
[0476] "Database" refers to a system for efficiently storing, managing, and searching collected text data and generated minutes.
[0477] "Search means" refers to technical devices or software that search the minutes stored in the database based on specific criteria.
[0478] "Jargon" refers to specialized words and phrases frequently used in a particular field or industry.
[0479] A "keyword map" refers to a diagram or table that visually shows the relationships between technical terms extracted from minutes.
[0480] "Decision-making history" refers to historical information that records the decisions made during a meeting and the process behind them.
[0481] "Storage means" refers to technical devices and software that store decision-making history and generated minutes in a database in a traceable form.
[0482] A "generative AI model" refers to a machine learning model that uses artificial intelligence technology to analyze voice data and convert it into text data, or to generate minutes based on text data.
[0483] "User" refers to an individual or organizational member who operates the system to upload audio data, search meeting minutes, and check decision-making history.
[0484] This invention is a system that can input voice data or text data, automatically generate minutes based on that data, and extract, store, and search important information. This system aims to efficiently manage meeting records and includes the following components:
[0485] Data Input and Transformation
[0486] Users record audio data during a meeting and upload the audio file to their device after the meeting ends. The uploaded audio file is sent from the device to a server. The server then sends the received audio data to a generative AI model, which converts the audio data into text data. The generative AI model can use speech recognition technology such as the Google Cloud Speech-to-Text API. The generated text data is then returned to the server.
[0487] Automatic generation of meeting minutes
[0488] The server reprocesses the text data received from the generative AI model and uses natural language processing (NLP) techniques to extract key points and decisions. Based on the extracted information, the generative AI model automatically generates meeting minutes. The generated minutes are then stored in a database in chronological order by the server. For example, a database management system such as MySQL or MongoDB can be used.
[0489] Finding and Accessing Data
[0490] A user can search for minutes of past meetings by specifying a specific date. For example, a user can instruct "Show the meeting records for October 10, 2023." Upon receiving this request, the server searches the database for the corresponding minutes and sends them to the terminal. The terminal then displays the retrieved minutes to the user.
[0491] Extraction and visualization of technical terms
[0492] After the minutes are generated, the server sends the data to the generative AI model again to extract technical terms from the minutes. The generative AI model analyzes the relationships between these technical terms and generates a keyword map. The generated keyword map is stored in a database by the server and displayed to the user as needed.
[0493] Record and track decision history
[0494] The server records the decisions made during the meeting and their progress along with the minutes, and stores them in a database. If a user wants to check the decision-making history of a specific project, they can instruct the server to "show the decision-making history for project X." The server receives this request, retrieves the relevant history from the database, and sends it to the terminal. The terminal then displays the decision-making history to the user.
[0495] Specific examples
[0496] For example, if a user wants to search for meeting records for a specific date, they can request, "Show me the meeting records for October 10, 2023." The server processes this request and retrieves the corresponding minutes from the database. The retrieved minutes are then sent to the terminal and displayed to the user.
[0497] Furthermore, if the user wants to check the decision-making history regarding the progress of a project, he or she can instruct, for example, "Show the decision-making history of project X." The server retrieves the decision-making history related to project X from the database, sends it to the terminal, and displays it to the user.
[0498] This allows users to quickly and efficiently grasp the content of past and present meetings, the history of decision-making, and the relationships between technical terms used. Furthermore, the specific operation of this system improves work efficiency and the accuracy of information management.
[0499] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0500] Step 1:
[0501] A user records audio data during a meeting. Specifically, the user collects the audio data using a smartphone or a dedicated recording device. This audio data is stored in a digital format for subsequent processing.
[0502] Step 2:
[0503] After the meeting, the user uploads the recorded audio file to the device. The user accesses the system's web interface using a browser on their computer or smartphone, selects the audio file, and clicks the upload button. This saves the audio data to the device.
[0504] Step 3:
[0505] The device sends the uploaded audio file to the server. Specifically, it uses an HTTP request to send the audio data and user authentication information to the server. The server receives this.
[0506] Step 4:
[0507] The server receives the voice data and sends it to the generative AI model. The server converts the voice data into an appropriate format and sends a speech recognition request to the generative AI model (for example, Google Cloud Speech-to-Text API). The input data is an audio file, and the output data is text data.
[0508] Step 5:
[0509] The generative AI model converts the voice data into text data and returns the text data to the server. The generative AI model analyzes the voice data and outputs it as text data. Specifically, an algorithm is used to extract character string information from the voice. The server receives this text data and proceeds to the next step.
[0510] Step 6:
[0511] The server reprocesses the text data received from the generative AI model. The input data is text data, and the server uses natural language processing (NLP) techniques to extract key points and decisions. Specifically, it analyzes keywords in the text and identifies the necessary information.
[0512] Step 7:
[0513] The generative AI model automatically generates minutes based on the extracted information. The server sends the extracted information to the generative AI model, which then generates the minutes. The input data are key points and decisions, and the output data are the minutes. The generated minutes are well-formed documents suitable for future reference.
[0514] Step 8:
[0515] The server saves the generated minutes in a database in chronological order. The input data is the minutes, and a date tag is added when saving. For example, the minutes can be saved using a database management system such as MySQL or MongoDB. This makes it easy to search and organize the minutes.
[0516] Step 9:
[0517] A user searches for past meeting minutes by specifying a specific date. Specifically, the user enters "Show me the meeting minutes for October 10, 2023" into the system. This request is sent to the server.
[0518] Step 10:
[0519] The server receives a request from a user and retrieves the minutes corresponding to that date from the database. The input data is a search request, and the server searches for the corresponding minutes using an SQL query or similar. The output data is the retrieved minutes.
[0520] Step 11:
[0521] The server sends the acquired minutes to the terminal. Specifically, it returns the minutes data as an HTTP response. The terminal receives this.
[0522] Step 12:
[0523] The terminal displays the minutes received from the server to the user. The user can view the minutes via a browser or other device. The displayed minutes clearly show the contents of the meeting and the decisions made.
[0524] Step 13:
[0525] The server sends the generated minutes to the generative AI model again to extract technical terms. The input data is the minutes, and the generative AI model identifies technical terms and returns a list of technical terms as output data.
[0526] Step 14:
[0527] The generative AI model analyzes the relationships between technical terms and generates a keyword map. The input data is a list of technical terms, and the output data is a keyword map. Specifically, a diagram is generated that visually shows the relationships between technical terms.
[0528] Step 15:
[0529] The server saves the generated keyword map in a database. The input data is a keyword map, which is saved in a database in an appropriate format, allowing for later verification of terminology relationships.
[0530] Step 16:
[0531] A user submits a request to view the decision history of a particular project, for example, by asking the system, "Show me the decision history of project X." This request is sent to the server.
[0532] Step 17:
[0533] The server retrieves the relevant decision-making history from the database and sends it to the terminal. Input data includes project names and dates, and the server searches for the corresponding history and retrieves the decision-making history as output data.
[0534] Step 18:
[0535] The device displays the decision-making history to the user, who can then view it via a browser or other device. The displayed history clearly shows the progress of the project and important decisions.
[0536] (Application example 1)
[0537] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0538] In today's corporate environment, especially in factories and corporate meetings, there is a need to accurately record meeting content and refer to it later. However, manually creating meeting minutes is time-consuming and laborious, and it is difficult to grasp the relationships between important decisions and technical terms. In addition, there is a lack of systems that can process meeting audio data in real time and instantly generate meeting minutes. It is necessary to solve these issues and record and manage meeting content efficiently and accurately.
[0539] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0540] In this invention, the server includes means for recording voice data in real time and instantly converting it into text data, means for extracting important points and decisions from the converted text data in real time, and means for automatically saving the generated minutes and important points in a database. This makes it possible to efficiently record the contents of a meeting in real time and quickly extract and save important information.
[0541] "Audio data" refers to data in which the sounds of a meeting, conversation, etc. are recorded in digital format.
[0542] "Text data" is digital data that expresses voice data as character information.
[0543] The "means for converting voice data into text data" is a function that executes a process of converting voice data into text information using voice recognition technology.
[0544] "Means for extracting key points and decisions" is a process for automatically identifying and extracting the main agenda items and important decisions of a meeting from text data.
[0545] "Means for automatically generating minutes" is a function that creates a systematic record of the meeting contents based on extracted important points and decisions.
[0546] "Means of organizing in chronological order and storing in a database" refers to the process of organizing the generated minutes in chronological order and storing them in a digital database.
[0547] "A means to search for minutes of a specific date" is a function that searches for minutes saved in the past by specifying a date and displays the minutes you need.
[0548] "Means for extracting technical terms and creating and displaying a keyword map" refers to the process of automatically extracting related technical terms from text data, and creating and displaying a keyword map that visually shows the relevance of these terms.
[0549] "A means of recording decisions and storing the history in a traceable form" is a function that records the history of decisions made in meetings and stores them in a form that allows you to check later when and how the decisions were made.
[0550] The "means for recording in real time and instantly converting it into text data" is a process for recording audio data during a meeting on the spot and converting it into text data in real time without stopping.
[0551] "Means for automatically saving to a database" is a function that automatically saves the generated minutes and extracted important information to a database.
[0552] MODE FOR CARRYING OUT THE INVENTION
[0553] This invention is a system that converts voice data from meetings in factories or companies into text data in real time and automatically creates and saves minutes of the meetings. Specific embodiments of the system are described below.
[0554] System Overview
[0555] The system records audio data in real time and instantly converts it into text. It extracts important points and decisions from the generated text, automatically generates and saves meeting minutes, and extracts technical terms to create a keyword map and store past decision-making history in a traceable format.
[0556] Data Input and Transformation
[0557] During a meeting, users use a hardware device (such as a microphone or smartphone) to record audio data. The recorded audio data is automatically sent to a server. The server processes the received audio data by sending it to a generative AI model that uses speech recognition technology and converting it into text data. The generative AI model uses Hugging Face's pipeline ("automatic-speech-recognition").
[0558] Automatic generation of meeting minutes
[0559] The server reprocesses the text data obtained by speech recognition and extracts key points and decisions in real time using a generative AI summarization model (for example, Hugging Face's pipeline ("summarization")). The automatically generated minutes are organized in chronological order and stored in a database.
[0560] Finding and Accessing Data
[0561] A user can request a search for a specific date to access the minutes of past meetings. The server receives the request from the user, retrieves the minutes corresponding to that date from the database, and sends them to the terminal. The terminal displays the minutes received from the server to the user.
[0562] Extraction and visualization of technical terms
[0563] After the minutes are generated, the server uses the generative AI model again to extract technical terms from the minutes. The generative AI analyzes the relationships between these technical terms and generates a keyword map. The generated keyword map is then stored in a database by the server.
[0564] Record and track decision history
[0565] The decision-making process during the meeting is recorded along with the minutes, and the server stores them in a database. When a user wants to check the decision-making history for a specific project, the server retrieves the relevant history from the database, sends it to the user's device, and displays it to the user.
[0566] Specific examples
[0567] For example, if a user wants to search for a meeting record for a specific date, they can request, "Show me the meeting record for October 10, 2023." The server processes this request and retrieves the corresponding minutes from the database. The retrieved minutes are then sent to the terminal and displayed to the user.
[0568] Furthermore, if a user wants to check the decision-making history regarding the progress of a project, they can instruct "Display the decision-making history for project X." The server retrieves the decision-making history related to project X from the database, sends it to the terminal, and displays it to the user. This allows the user to quickly and efficiently understand the content of past and present meetings, the history of decision-making, and the relationships between the technical terms used.
[0569] Prompt Sentence Examples
[0570] Examples of input prompts for generative AI models include:
[0571] "View the meeting minutes for October 10, 2023"
[0572] "Show the decision history for Project X"
[0573] This allows users to use the system efficiently and easily access the information they need.
[0574] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0575] Step 1:
[0576] Users record audio data during a meeting. Audio data is collected using smart devices (smartphones and microphones). The input is in the form of audio data, which is sent to the next step.
[0577] Step 2:
[0578] The device uploads the recorded audio data to the server. The input is the audio data file, and the output is the transfer of the audio data to the server. Specifically, the device selects an audio file and sends it to the server via the Internet.
[0579] Step 3:
[0580] The server sends the received voice data to the generative AI model and converts it into text data. The input is voice data and the output is text data. Specifically, speech recognition is performed using Hugging Face's pipeline ("automatic-speech-recognition").
[0581] Step 4:
[0582] The server reprocesses the generated text data and extracts key points and decisions. The input is the text data, and the output is the extracted key points and decisions. The summary is generated using a generative AI summarization model (Hugging Face's pipeline("summarization")).
[0583] Step 5:
[0584] The server automatically generates minutes based on the extracted content and stores them in a database. The input is important points and decisions, and the output is the automatically generated minutes. Specifically, the minutes are organized in chronological order and stored in a database.
[0585] Step 6:
[0586] A user sends a search request to the server specifying a specific date to search for past meeting minutes. The input is a search query (e.g., "Show meeting minutes from October 10, 2023"), which the server receives.
[0587] Step 7:
[0588] The server processes the request from the user and retrieves the relevant minutes from the database. The input is the search query, and the output is the relevant minutes data. Specifically, the server queries the database based on the date and retrieves the results.
[0589] Step 8:
[0590] The server sends the acquired minutes data to the terminal. The input is the minutes data, and the output is the data transfer to the terminal. Specifically, the server sends the data to the terminal via the Internet.
[0591] Step 9:
[0592] The terminal receives the minutes from the server and displays them to the user. The input is the minutes data, and the output is a visual display to the user. The specific operation is that the minutes are displayed on the terminal's display.
[0593] Step 10:
[0594] The server extracts technical terms from the minutes and generates a keyword map. The input is the minutes data, and the output is a keyword map. A generative AI model is used to analyze and visualize the relationships between technical terms.
[0595] Step 11:
[0596] The server saves the keyword map in a database. The input is the keyword map and the output is the saved data. As a specific operation, the generated keyword map is stored in the database.
[0597] Step 12:
[0598] When a user wants to check the decision-making history of a particular project, the server retrieves the relevant history from the database, sends it to the terminal, and displays it to the user. The input is a database query for the decision-making history, and the output is the decision-making history data and its display. Specifically, the server executes the query, obtains the results, and sends them to the terminal, which then displays the received data to the user.
[0599] This series of processes allows for efficient recording of meeting content in real time, allowing important information to be instantly extracted, stored, and searched.
[0600] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0601] This invention is a system that can input voice data or text data, automatically generate minutes based on that data, and extract, store, and search important information, and it also combines an emotion engine that recognizes the user's emotions. Below, an embodiment of this system is explained in natural language.
[0602] Data Input and Transformation
[0603] Users record audio data during a conference and upload the audio file to their terminals after the conference ends. The terminals then transmit the uploaded audio file to a server.
[0604] Speech-to-text conversion
[0605] The server sends the received voice data to the generation AI, which converts the voice data into text data. The generation AI converts the voice into text and returns the text data to the server.
[0606] Emotion recognition
[0607] The server sends the converted text and voice data to the emotion engine, which then extracts emotion data from the voice and text and returns the emotion data to the server.
[0608] Automatic generation of meeting minutes
[0609] The server reprocesses the text data and emotion data received from the generation AI to extract important points and decisions. The generation AI automatically generates minutes based on this data and returns them to the server. The server then stores the generated minutes in a database in chronological order along with the emotion data.
[0610] Finding and Accessing Data
[0611] To access the minutes of past meetings, a user requests a search by specifying a specific date. The server receives the request from the user, retrieves the minutes and emotion data for the corresponding date from the database, and sends them to the device. The device then displays the minutes and emotion data received from the server to the user.
[0612] Extraction and visualization of technical terms
[0613] After the minutes are generated, the server sends the data to the generation AI again, which extracts technical terms from the minutes. The generation AI analyzes the relationships between these technical terms and generates a keyword map. The generated keyword map is then stored in a database by the server. Emotional data is also reflected in the keyword map, visualizing changes in emotions.
[0614] Record and track decision history
[0615] Decisions made during meetings and their progress are recorded along with minutes and stored in a database by the server. When a user wants to check the decision-making history of a specific project, the server retrieves the relevant history from the database and sends it along with emotion data to the device, which then displays it to the user.
[0616] Specific examples
[0617] For example, if a user wants to search for a meeting record for a specific date, they can request, "Show me the meeting record for October 10, 2023." The server processes this request and retrieves the corresponding meeting minutes and emotion data from the database. The retrieved meeting minutes are sent to the device and displayed to the user. Based on the emotion data, the server also displays the overall emotional changes during the meeting.
[0618] Furthermore, if the user wants to check the decision-making history regarding the progress of the project, they can instruct "Show the decision-making history for project X." The server retrieves the decision-making history and emotion data related to project X from the database, sends them to the terminal, and displays them to the user. By also displaying the emotion data, it is possible to grasp the trend of emotions as the project progresses.
[0619] This allows users to quickly and efficiently grasp the content of past and present meetings, the history of decision-making, the relationships between technical terms used, and changes in emotions.
[0620] The processing flow will be explained below.
[0621] Understood. The specific operation will be explained below, divided into processing steps.
[0622] Step 1:
[0623] A user records audio data during a conference and uploads the audio file to a terminal after the conference ends.
[0624] Step 2:
[0625] The device sends the uploaded audio file to the server.
[0626] Step 3:
[0627] The server sends the received voice data to the emotion engine and generation AI, which converts the voice data into text data and recognizes emotions.
[0628] Step 4:
[0629] The generating AI converts the speech into text data and sends that text data back to the server.
[0630] Step 5:
[0631] The emotion engine recognizes the user's emotion from the voice data and returns the emotion data to the server.
[0632] Step 6:
[0633] The server reprocesses the text data received from the generation AI and the emotion data received from the emotion engine, and sends it back to the generation AI to extract important points and decisions.
[0634] Step 7:
[0635] The generation AI extracts important points and decisions from the text data and automatically generates minutes, which are then sent back to the server.
[0636] Step 8:
[0637] The server stores the generated minutes and emotion data in a database. The minutes are time-stamped and organized in chronological order.
[0638] Step 9:
[0639] A user submits a request specifying a specific date to search for minutes of past meetings.
[0640] Step 10:
[0641] The server receives a request from the user, retrieves the minutes of the specified date and the corresponding emotion data from the database, and transmits them to the terminal.
[0642] Step 11:
[0643] The device receives the minutes and emotion data from the server and displays them to the user, allowing the user to check the emotional trends during the meeting as well as the content of the meeting.
[0644] Step 12:
[0645] The server retrieves the minutes from the database, sends them to the generation AI, and extracts technical terms from the minutes.
[0646] Step 13:
[0647] The generative AI extracts technical terms from the minutes, analyzes their relationships, and generates a keyword map, which is then sent back to the server.
[0648] Step 14:
[0649] The server stores the generated keyword map in a database, and also reflects the emotion data in this map so that it can be displayed to the user.
[0650] Step 15:
[0651] A user submits a request specifying the project name to check the decision history of a specific project.
[0652] Step 16:
[0653] The server receives a request from a user, retrieves the decision-making history and emotion data for the specified project from the database, and sends them to the terminal.
[0654] Step 17:
[0655] The device receives the decision-making history and emotional data from the server and displays it to the user, allowing the user to understand the decision-making process as the project progresses and the emotional trends that occurred during that process.
[0656] The above is the specific processing flow of this system that combines an emotion engine.
[0657] Example 2
[0658] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0659] In today's business environment, efficient meeting progress and recording are crucial. However, accurately recording important discussions and decisions made during meetings, as well as participants' emotions, and storing and retrieving these in a format that can be used later, can be time-consuming. Furthermore, extracting relevant technical terms and emotional changes from minutes and visualizing them can be difficult. For these reasons, there is a need for a way for users to quickly grasp the details of past meeting records and decisions.
[0660] The specific processing by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for inputting voice data or text data and converting the voice data into text data, means for extracting emotion data from the converted text data, means for extracting important points and decisions from the text data and emotion data and automatically generating minutes, means for organizing the generated minutes and emotion data in chronological order and storing them in a database, means for accessing the chronologically organized minutes and emotion data and searching for minutes of a specific date, means for extracting technical terms from the minutes, creating a keyword map, and visualizing changes in emotion, and means for recording past decision-making history and storing it in a traceable form. This allows users to efficiently manage and quickly access important information from meetings.
[0661] "Audio data" refers to digital data of audio recorded during a meeting or conversation.
[0662] "Text data" is digital data of character information converted from audio data.
[0663] "Means of conversion" refers to the technology or device that analyzes voice data and converts it into text data as character information.
[0664] "Emotion data" refers to data that indicates emotional states and changes extracted from text and speech analysis.
[0665] "Extraction means" refers to technology or equipment that extracts key points, decisions, or emotional data from text data or audio data.
[0666] "Means for automatically generating minutes" refers to technology or devices that automatically generate documents summarizing meeting content and decisions based on text data and emotion data.
[0667] A "database" is a digital system or software for systematically storing and managing minutes and emotional data.
[0668] "Searchable means" refers to technology or devices that search for required data based on specific dates or conditions from information stored in a database.
[0669] "Terminology" refers to specialized terms and expressions used in a particular field or industry.
[0670] A "keyword map" is a diagram or chart that visually represents the relationships and interactions of extracted technical terms.
[0671] "Visualization means" refers to technology or devices that visually display extracted data.
[0672] "Decision-making history" is data that records the details and process of decisions made in meetings and projects.
[0673] "Means for storing data in a traceable form" refers to technology or devices that store decision-making history in a database for future reference.
[0674] This invention provides a system for processing audio data and text data from a meeting, automatically generating minutes, and efficiently managing important information. This system includes functions for converting audio data, recognizing emotions, automatically generating minutes, and searching and visualizing data. Specific embodiments of the system are described below.
[0675] Entering data
[0676] Users upload audio data recorded during a meeting to their device, which saves the audio in a standard digital audio format (e.g., MP3 or WAV files), and the device sends the uploaded audio file to the server.
[0677] Speech-to-text conversion
[0678] The server sends the received voice data to a speech recognition service (such as Google Cloud Speech-to-Text API), which is a generating AI, and converts the voice data into text data. The generating AI analyzes the voice and converts the content into text data as character information. The converted text data is then sent back to the server.
[0679] Emotion recognition
[0680] The server then sends the converted text and voice data to an emotion recognition engine (e.g., IBM Watson Tone Analyzer) to recognize the user's emotions. The emotion engine extracts emotion data (e.g., joy, anger, sadness, etc.) from the voice tone and text content. This emotion data is then sent back to the server.
[0681] Automatic generation of meeting minutes
[0682] The server uses a generative AI model (e.g., OpenAI's GPT-3) to extract key points and decisions based on the generated text data and emotion data, and automatically generates meeting minutes. The generated minutes include important discussions and decisions made during the meeting. The minutes and emotion data are organized chronologically by the server and stored in a database.
[0683] Finding and Accessing Data
[0684] A user can submit a search request to access meeting minutes for a specific date, for example, using a specific prompt such as "Show me the meeting minutes for October 10, 2023." The server receives this request and retrieves the corresponding meeting minutes and emotion data from the database. The retrieved data is then sent to the device and displayed to the user.
[0685] Extracting technical terms and generating keyword maps
[0686] The server then sends the generated minutes to a generation AI (e.g., BERT) to extract technical terms from the minutes. The generation AI then analyzes the relationships between the extracted technical terms and generates a keyword map. This keyword map is also stored in a database and displayed visually along with sentiment data.
[0687] Record and track decision history
[0688] The server records the decisions made during the meeting and their process, and stores them in a database along with the minutes. If a user wants to check the decision-making history of a specific project, they can use a specific prompt, such as "Show the decision-making history of project X." The server retrieves the relevant data from the database, sends it to the terminal, and displays it to the user.
[0689] This allows users to quickly and efficiently manage and understand past and present meeting details, decision details, terminology relationships, and emotional changes.
[0690] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0691] Step 1:
[0692] Users upload audio data recorded during a meeting to their device. The audio data is saved in a digital audio file format such as MP3 or WAV. The device then sends the uploaded audio file to the server. The input is the audio data, and the output is the audio file sent to the server.
[0693] Step 2:
[0694] The server sends the received voice data to a generation AI (for example, Google Cloud Speech-to-Text API), which is a voice recognition service. The generation AI analyzes the voice data and generates text data as character information. The input is voice data, and the output from the generation AI is text data. The server receives this text data and proceeds to the next step of processing.
[0695] Step 3:
[0696] The server sends the converted text data and the original voice data to an emotion recognition engine (e.g., IBM Watson Tone Analyzer). The emotion recognition engine extracts emotion data from these data. The input is text data and voice data, and the output is emotion data. The server receives the emotion data and proceeds to the next step of processing.
[0697] Step 4:
[0698] The server uses a generative AI model (e.g., OpenAI's GPT-3) based on the text data and emotion data to extract key points and decisions, and automatically generate meeting minutes. The input is text data and emotion data, and the output is automatically generated meeting minutes. The generated minutes are used for further processing.
[0699] Step 5:
[0700] The server organizes the generated minutes and emotion data in chronological order and stores them in a database (e.g., MySQL or PostgreSQL). The input is the minutes and emotion data, and the output is visualized data stored in the database.
[0701] Step 6:
[0702] To access minutes for a specific date, a user sends a search request through their terminal. For example, they request "Show the meeting records for October 10, 2023." The input is the search request, and the server receives this request and retrieves the corresponding minutes and emotion data from the database. The output is the retrieved minutes and emotion data.
[0703] Step 7:
[0704] The terminal displays the minutes and emotion data obtained from the server to the user through a user interface (e.g., a web browser or a dedicated app). The input is the obtained minutes and emotion data, and the output is the meeting record displayed to the user.
[0705] Step 8:
[0706] The server then sends the generated transcripts to a generation AI (e.g., BERT) to extract technical terms from the transcripts and generate a keyword map. The input is the transcripts, and the output is the generated keyword map. The generated keyword map is also stored in a database and can be visualized along with the sentiment data.
[0707] Step 9:
[0708] The server stores the decision-making process and its details in a database along with the meeting minutes. If a user wants to check the decision-making history of a specific project, for example, they can request, "Show the decision-making history of project X." The server retrieves the relevant data from the database, sends it to the terminal, and displays it to the user. The input is the search request and the decision-making history, and the output is the project's decision-making history displayed to the user.
[0709] (Application example 2)
[0710] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0711] Conventional meeting and online event minutes generation systems require manual minutes creation, which is time-consuming and labor-intensive. Furthermore, it is difficult to accurately capture important information and changes in user emotions, making it difficult to quickly grasp decision-making and analyze emotional trends. This creates the risk that some meeting content may be overlooked, and it is difficult to accurately track the decision-making process.
[0712] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for inputting voice data or text data and converting the voice data into text data, means for extracting important points and decisions from the converted text data and automatically generating minutes, means for organizing the generated minutes in chronological order and saving them in a database, means for accessing the chronologically organized minutes and searching for minutes of a specific date, means for extracting technical terms from the minutes and creating a keyword map, means for recording past decision-making history and storing it in a traceable form, means for visualizing user emotional trends using the generated minutes and emotion data, and means for processing the voice data based on prompts and searching for and displaying specific events or meeting records. This enables the content of meetings and online events to be quickly and efficiently recorded, searched, and analyzed.
[0713] "Audio data" refers to data in which audio is recorded in digital format.
[0714] "Text data" refers to data in which character information is recorded in digital format.
[0715] "Means for converting" refers to an algorithm or machine for converting audio data into text data.
[0716] "Key points" are information or matters that should be given special importance in the content of a meeting or event.
[0717] "Decisions" are matters that were discussed at meetings or events and ultimately decided upon.
[0718] Minutes are documents that record the contents of a meeting or event in chronological order.
[0719] "Chronology" refers to the arrangement of events or data in chronological order.
[0720] A "database" is a storage system that systematically stores data and allows it to be searched and updated.
[0721] "Searchable means" refers to methods and technologies for extracting and viewing data based on specific criteria.
[0722] "Terminology" is a term with a specific meaning used in a particular field or industry.
[0723] A "keyword map" is a graphical tool for visually displaying relevant keywords within minutes or other documents.
[0724] "Decision-making history" is a record of decisions and the processes involved in past meetings and events.
[0725] "Emotion data" is data that represents the user's emotional state and is the result of emotion analysis.
[0726] A "means for visualizing emotion trends" is a method for visually displaying changes in a user's emotions over time.
[0727] A "prompt sentence" is a specific instruction sentence that is input to a generative AI model.
[0728] "Processing means" refers to the technology and devices used to analyze, convert and display audio and text data.
[0729] "Means for searching and displaying events and meeting records" means methods or technologies for searching events and meeting records based on specific criteria and displaying the results.
[0730] This invention is a system that converts audio data from meetings and online events into text data in real time, extracts important points and decisions from the text data, and automatically generates minutes.The system also has the function of visualizing users' emotional trends using the generated minutes and emotion data, and searching and displaying specific event or meeting records based on prompt sentences.
[0731] The system utilizes the following hardware and software:
[0732] Hardware used:
[0733] Microphone: For voice input
[0734] Server: storing and processing audio files
[0735] User device: Device used to display results (smartphone, tablet, PC, etc.)
[0736] Software used:
[0737] speech_recognition library: convert speech to text
[0738] transformers library: sentiment analysis
[0739] Custom database module: storing meeting minutes data
[0740] pipeline (natural language processing): processing emotional data
[0741] Process flow:
[0742] During a meeting, users record audio data using a microphone. This audio data is uploaded to the server via the terminal after the meeting ends. The server converts the uploaded audio file into text data using the speech_recognition library.
[0743] The server then sends the converted text data to the transformers library for sentiment analysis. The resulting sentiment data is then sent back to the server. The server then extracts key points and decisions from the text and sentiment data and automatically generates meeting minutes. The generated minutes are then organized and stored in chronological order using a custom database module.
[0744] Users can access meeting minutes by specifying a specific date. The server can also search for past meeting records and emotion data, and display related events and meeting records on the device based on a specific prompt. For example, by entering the prompt "View the meeting records for the live event on October 10, 2023," the corresponding meeting minutes and emotion data will be displayed.
[0745] Furthermore, the server visualizes the user's emotional trends using the generated minutes and emotional data, allowing users to visually grasp changes in emotions during meetings and events.
[0746] In this way, the content of meetings and online events can be recorded, searched and analyzed quickly and efficiently.
[0747] Examples of prompts:
[0748] "View the recording of the live event on October 10, 2023"
[0749] The prompt allows the user to easily search and view meeting records for a particular date.
[0750] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0751] Step 1:
[0752] During a meeting, a user uses a microphone to record audio. The recorded audio data is saved on the device. The input is audio data, and the output is an audio file.
[0753] Step 2:
[0754] After the conference, the user uploads the recorded audio data to the server using the terminal. In this process, the audio file is the input, and the audio file sent to the server is the output.
[0755] Step 3:
[0756] The server receives the uploaded audio file and converts it into text using the speech_recognition library. The input is the audio file, and the output is the generated text. The server then processes the sound waves to convert them into text.
[0757] Step 4:
[0758] The converted text data is sent by the server to the sentiment analysis function of the transformers library. The input is text data and the output is emotion data. The server recognizes the emotion of the text content and generates emotion data.
[0759] Step 5:
[0760] The server automatically generates minutes by extracting important points and decisions based on the generated text data and emotion data. The input is text data and emotion data, and the output is the automatically generated minutes. The server uses natural language processing to extract important information.
[0761] Step 6:
[0762] The generated minutes are organized in chronological order by the server and saved in a custom database module. The input is the minutes, and the output is the minutes data organized in chronological order. The server stores the data in the database.
[0763] Step 7:
[0764] A user searches for meeting minutes for a specific date on their device by entering the prompt "Show meeting notes for the live event on October 10, 2023." The input is the search prompt, and the output is the meeting minutes for the specified date and emotion data.
[0765] Step 8:
[0766] The server searches the database based on the prompt sentence, obtains the relevant minutes and emotion data, and sends them to the terminal. The input is the prompt sentence and the database query, and the output is the search results: the minutes and emotion data.
[0767] Step 9:
[0768] The user checks the minutes and emotion data displayed on the device. The server generates graphs and dashboards to visually display emotion trends based on the emotion data. The input is emotion data, and the output is visualized emotion trends.
[0769] This series of processing steps enables the content of meetings and online events to be recorded, searched, and analyzed quickly and efficiently.
[0770] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0771] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0772] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0773] [Third embodiment]
[0774] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0775] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[0776] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0777] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0778] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0779] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0780] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0781] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0782] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0783] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0784] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0785] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0786] This invention is a system that can input voice data or text data, automatically generate minutes based on that data, and extract, store, and search important information. An embodiment of this system will be described below in natural language.
[0787] Data Input and Transformation
[0788] Users record audio data during a meeting and upload the audio file to their device after the meeting ends. The device then sends the uploaded audio file to the server. The server processes the received audio data by sending it to a generation AI, which converts it into text data. The generation AI then converts the audio into text and returns the text data to the server.
[0789] Automatic generation of meeting minutes
[0790] The server reprocesses the text data received from the generation AI and extracts important points and decisions. The generation AI automatically generates minutes based on this text data and returns them to the server. The server then stores the generated minutes in a database in chronological order.
[0791] Finding and Accessing Data
[0792] A user requests a search for a specific date to access the minutes of a past meeting. The server receives the request, retrieves the minutes corresponding to that date from the database, and sends them to the terminal. The terminal displays the minutes received from the server to the user.
[0793] Extraction and visualization of technical terms
[0794] After the minutes are generated, the server sends the data to the AI again to extract technical terms from the minutes. The AI analyzes the relationships between these technical terms and generates a keyword map. The generated keyword map is then stored in a database by the server.
[0795] Record and track decision history
[0796] Decisions made during meetings and their progress are recorded along with minutes and stored in a database by the server. When a user wants to check the decision-making history of a particular project, the server retrieves the relevant history from the database and sends it to the device. The device then displays the decision-making history to the user.
[0797] Specific examples
[0798] For example, if a user wants to search for a meeting record for a specific date, they can request, "Show me the meeting record for October 10, 2023." The server processes this request and retrieves the corresponding minutes from the database. The retrieved minutes are then sent to the terminal and displayed to the user.
[0799] Furthermore, if the user wants to check the decision-making history regarding the progress of the project, he or she can instruct "Show the decision-making history of project X." The server retrieves the decision-making history related to project X from the database, sends it to the terminal, and displays it to the user.
[0800] This allows users to quickly and efficiently understand the content of past and present meetings, the history of decision-making, and the relationships between technical terms used.
[0801] The processing flow will be explained below.
[0802] Understood. The program's processing will be explained in the following steps.
[0803] Step 1:
[0804] A user records audio data during a conference and uploads the audio file to a terminal after the conference ends.
[0805] Step 2:
[0806] The device sends the uploaded audio file to the server.
[0807] Step 3:
[0808] The server sends the received voice data to the generation AI, which converts the voice data into text data.
[0809] Step 4:
[0810] The generating AI converts the speech into text data and sends that text data back to the server.
[0811] Step 5:
[0812] The server reprocesses the text data received from the generation AI and sends it back to the generation AI to extract key points and decisions.
[0813] Step 6:
[0814] The generation AI extracts important points and decisions from the text data and automatically generates minutes, which are then sent back to the server.
[0815] Step 7:
[0816] The server stores the generated minutes in a database, assigning a timestamp to the minutes and organizing them in chronological order.
[0817] Step 8:
[0818] A user requests a search for a particular date to access minutes from a past meeting.
[0819] Step 9:
[0820] The server receives a request from the user, retrieves the minutes of the specified date from the database, and transmits them to the terminal.
[0821] Step 10:
[0822] The terminal displays the minutes received from the server to the user.
[0823] Step 11:
[0824] The server retrieves the minutes from the database, sends them to the generation AI, and extracts technical terms from the minutes.
[0825] Step 12:
[0826] The generative AI extracts technical terms from the minutes, analyzes their relationships, and generates a keyword map, which is then sent back to the server.
[0827] Step 13:
[0828] The server stores the generated keyword map in a database.
[0829] Step 14:
[0830] A user requests a search by specifying the project name to check the decision-making history of a particular project.
[0831] Step 15:
[0832] The server receives a request from a user, retrieves the decision-making history of the specified project from the database, and transmits it to the terminal.
[0833] Step 16:
[0834] The terminal displays the decision-making history received from the server to the user.
[0835] The above is the specific processing flow of this system.
[0836] Example 1
[0837] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0838] In today's business environment, it is important to efficiently manage meeting records and decision-making histories. However, manually creating meeting minutes and sorting through technical terms takes time and effort, and there is a high possibility of missing important information or delaying retrieval. There is also a lack of consistent record management to track past decision-making history. A system to solve these issues and improve meeting efficiency is needed.
[0839] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0840] In this invention, the server includes: means for inputting voice data or text data and converting the voice data to text data; means for extracting important points and decisions from the converted text data and automatically generating minutes; means for organizing the generated minutes in chronological order and saving them in a database; means for accessing the chronologically organized minutes and searching for minutes of a specific date; means for extracting technical terms from the minutes and creating a keyword map; means for recording and storing past decision-making history in a traceable form; means for converting voice data to text data using a generative AI model and extracting important information from the text data to automatically generate minutes; and means for searching for minutes of a specified date or project decision-making history and displaying them to the user. This allows for quick and accurate recording of meeting content, facilitating search and history tracking, and improving work efficiency.
[0841] "Audio Data" means human speech or other acoustic signals recorded during a meeting and stored in digital form.
[0842] "Text data" refers to character string information obtained as a result of analyzing voice data, or directly input character information.
[0843] "Conversion means" refers to technical equipment or software for converting audio data into text data, and may for example be a generative AI model.
[0844] Minutes are documents that organize and summarize the contents of a meeting and record the main topics and decisions made.
[0845] "Key points" or "decisions" refer to particularly important information or conclusions from a meeting and are the main contents that should be recorded in the minutes.
[0846] "Generating means" refers to technical devices or software for automatically creating minutes, including, for example, means for creating minutes from text data using a generative AI model.
[0847] "Database" refers to a system for efficiently storing, managing, and searching collected text data and generated minutes.
[0848] "Search means" refers to technical devices or software that search the minutes stored in the database based on specific criteria.
[0849] "Jargon" refers to specialized words and phrases frequently used in a particular field or industry.
[0850] A "keyword map" refers to a diagram or table that visually shows the relationships between technical terms extracted from minutes.
[0851] "Decision-making history" refers to historical information that records the decisions made during a meeting and the process behind them.
[0852] "Storage means" refers to technical devices and software that store decision-making history and generated minutes in a database in a traceable form.
[0853] A "generative AI model" refers to a machine learning model that uses artificial intelligence technology to analyze voice data and convert it into text data, or to generate minutes based on text data.
[0854] "User" refers to an individual or organizational member who operates the system to upload audio data, search meeting minutes, and check decision-making history.
[0855] This invention is a system that can input voice data or text data, automatically generate minutes based on that data, and extract, store, and search important information. This system aims to efficiently manage meeting records and includes the following components:
[0856] Data Input and Transformation
[0857] Users record audio data during a meeting and upload the audio file to their device after the meeting ends. The uploaded audio file is sent from the device to a server. The server then sends the received audio data to a generative AI model, which converts the audio data into text data. The generative AI model can use speech recognition technology such as the Google Cloud Speech-to-Text API. The generated text data is then returned to the server.
[0858] Automatic generation of meeting minutes
[0859] The server reprocesses the text data received from the generative AI model and uses natural language processing (NLP) techniques to extract key points and decisions. Based on the extracted information, the generative AI model automatically generates meeting minutes. The generated minutes are then stored in a database in chronological order by the server. For example, a database management system such as MySQL or MongoDB can be used.
[0860] Finding and Accessing Data
[0861] A user can search for minutes of past meetings by specifying a specific date. For example, a user can instruct "Show the meeting records for October 10, 2023." Upon receiving this request, the server searches the database for the corresponding minutes and sends them to the terminal. The terminal then displays the retrieved minutes to the user.
[0862] Extraction and visualization of technical terms
[0863] After the minutes are generated, the server sends the data to the generative AI model again to extract technical terms from the minutes. The generative AI model analyzes the relationships between these technical terms and generates a keyword map. The generated keyword map is stored in a database by the server and displayed to the user as needed.
[0864] Record and track decision history
[0865] The server records the decisions made during the meeting and their progress along with the minutes, and stores them in a database. If a user wants to check the decision-making history of a specific project, they can instruct the server to "show the decision-making history for project X." The server receives this request, retrieves the relevant history from the database, and sends it to the terminal. The terminal then displays the decision-making history to the user.
[0866] Specific examples
[0867] For example, if a user wants to search for a meeting record for a specific date, they can request, "Show me the meeting record for October 10, 2023." The server processes this request and retrieves the corresponding minutes from the database. The retrieved minutes are then sent to the terminal and displayed to the user.
[0868] Furthermore, if the user wants to check the decision-making history regarding the progress of a project, he or she can instruct, for example, "Show the decision-making history of project X." The server retrieves the decision-making history related to project X from the database, sends it to the terminal, and displays it to the user.
[0869] This allows users to quickly and efficiently grasp the content of past and present meetings, the history of decision-making, and the relationships between technical terms used. Furthermore, the specific operation of this system improves work efficiency and the accuracy of information management.
[0870] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0871] Step 1:
[0872] A user records audio data during a meeting. Specifically, the user collects the audio data using a smartphone or a dedicated recording device. This audio data is stored in a digital format for subsequent processing.
[0873] Step 2:
[0874] After the meeting, the user uploads the recorded audio file to the device. The user accesses the system's web interface using a browser on their computer or smartphone, selects the audio file, and clicks the upload button. This saves the audio data to the device.
[0875] Step 3:
[0876] The device sends the uploaded audio file to the server. Specifically, it uses an HTTP request to send the audio data and user authentication information to the server. The server receives this.
[0877] Step 4:
[0878] The server receives the voice data and sends it to the generative AI model. The server converts the voice data into an appropriate format and sends a speech recognition request to the generative AI model (for example, Google Cloud Speech-to-Text API). The input data is an audio file, and the output data is text data.
[0879] Step 5:
[0880] The generative AI model converts the voice data into text data and returns the text data to the server. The generative AI model analyzes the voice data and outputs it as text data. Specifically, an algorithm is used to extract character string information from the voice. The server receives this text data and proceeds to the next step.
[0881] Step 6:
[0882] The server reprocesses the text data received from the generative AI model. The input data is text data, and the server uses natural language processing (NLP) techniques to extract key points and decisions. Specifically, it analyzes keywords in the text and identifies the necessary information.
[0883] Step 7:
[0884] The generative AI model automatically generates minutes based on the extracted information. The server sends the extracted information to the generative AI model, which then generates the minutes. The input data are key points and decisions, and the output data are the minutes. The generated minutes are well-formed documents suitable for future reference.
[0885] Step 8:
[0886] The server saves the generated minutes in a database in chronological order. The input data is the minutes, and a date tag is added when saving. For example, the minutes can be saved using a database management system such as MySQL or MongoDB. This makes it easy to search and organize the minutes.
[0887] Step 9:
[0888] A user searches for past meeting minutes by specifying a specific date. Specifically, the user enters "Show me the meeting minutes for October 10, 2023" into the system. This request is sent to the server.
[0889] Step 10:
[0890] The server receives a request from a user and retrieves the minutes corresponding to that date from the database. The input data is a search request, and the server searches for the corresponding minutes using an SQL query or similar. The output data is the retrieved minutes.
[0891] Step 11:
[0892] The server sends the acquired minutes to the terminal. Specifically, it returns the minutes data as an HTTP response. The terminal receives this.
[0893] Step 12:
[0894] The terminal displays the minutes received from the server to the user. The user can view the minutes via a browser or other device. The displayed minutes clearly show the contents of the meeting and the decisions made.
[0895] Step 13:
[0896] The server sends the generated minutes to the generative AI model again to extract technical terms. The input data is the minutes, and the generative AI model identifies technical terms and returns a list of technical terms as output data.
[0897] Step 14:
[0898] The generative AI model analyzes the relationships between technical terms and generates a keyword map. The input data is a list of technical terms, and the output data is a keyword map. Specifically, a diagram is generated that visually shows the relationships between technical terms.
[0899] Step 15:
[0900] The server saves the generated keyword map in a database. The input data is a keyword map, which is saved in a database in an appropriate format, allowing for later verification of terminology relationships.
[0901] Step 16:
[0902] A user submits a request to view the decision history of a particular project, for example, by asking the system, "Show me the decision history of project X." This request is sent to the server.
[0903] Step 17:
[0904] The server retrieves the relevant decision-making history from the database and sends it to the terminal. Input data includes project names and dates, and the server searches for the corresponding history and retrieves the decision-making history as output data.
[0905] Step 18:
[0906] The device displays the decision-making history to the user, who can then view it via a browser or other device. The displayed history clearly shows the progress of the project and important decisions.
[0907] (Application example 1)
[0908] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0909] In today's corporate environment, especially in factories and corporate meetings, there is a need to accurately record meeting content and refer to it later. However, manually creating meeting minutes is time-consuming and laborious, and it is difficult to grasp the relationships between important decisions and technical terms. In addition, there is a lack of systems that can process meeting audio data in real time and instantly generate meeting minutes. It is necessary to solve these issues and record and manage meeting content efficiently and accurately.
[0910] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0911] In this invention, the server includes means for recording voice data in real time and instantly converting it into text data, means for extracting important points and decisions from the converted text data in real time, and means for automatically saving the generated minutes and important points in a database. This makes it possible to efficiently record the contents of a meeting in real time and quickly extract and save important information.
[0912] "Audio data" refers to data in which the sounds of a meeting, conversation, etc. are recorded in digital format.
[0913] "Text data" is digital data that expresses voice data as character information.
[0914] The "means for converting voice data into text data" is a function that executes a process of converting voice data into text information using voice recognition technology.
[0915] "Means for extracting key points and decisions" is a process for automatically identifying and extracting the main agenda items and important decisions of a meeting from text data.
[0916] "Means for automatically generating minutes" is a function that creates a systematic record of the meeting contents based on extracted important points and decisions.
[0917] "Means of organizing in chronological order and storing in a database" refers to the process of organizing the generated minutes in chronological order and storing them in a digital database.
[0918] "A means to search for minutes of a specific date" is a function that searches for minutes saved in the past by specifying a date and displays the minutes you need.
[0919] "Means for extracting technical terms and creating and displaying a keyword map" refers to the process of automatically extracting related technical terms from text data, and creating and displaying a keyword map that visually shows the relevance of these terms.
[0920] "A means of recording decisions and storing the history in a traceable form" is a function that records the history of decisions made in meetings and stores them in a form that allows you to check later when and how the decisions were made.
[0921] The "means for recording in real time and instantly converting it into text data" is a process for recording audio data during a meeting on the spot and converting it into text data in real time without stopping.
[0922] "Means for automatically saving to a database" is a function that automatically saves the generated minutes and extracted important information to a database.
[0923] MODE FOR CARRYING OUT THE INVENTION
[0924] This invention is a system that converts voice data from meetings in factories or companies into text data in real time and automatically creates and saves minutes of the meetings. Specific embodiments of the system are described below.
[0925] System Overview
[0926] The system records audio data in real time and instantly converts it into text. It extracts important points and decisions from the generated text, automatically generating and saving meeting minutes. It also extracts technical terms, creates a keyword map, and stores past decision-making history in a traceable format.
[0927] Data Input and Transformation
[0928] During a meeting, users use a hardware device (such as a microphone or smartphone) to record audio data. The recorded audio data is automatically sent to a server. The server processes the received audio data by sending it to a generative AI model that uses speech recognition technology and converting it into text data. The generative AI model uses Hugging Face's pipeline ("automatic-speech-recognition").
[0929] Automatic generation of meeting minutes
[0930] The server reprocesses the text data obtained by speech recognition and extracts key points and decisions in real time using a generative AI summarization model (for example, Hugging Face's pipeline ("summarization")). The automatically generated minutes are organized in chronological order and stored in a database.
[0931] Finding and Accessing Data
[0932] A user can request a search for a specific date to access the minutes of past meetings. The server receives the request from the user, retrieves the minutes corresponding to that date from the database, and sends them to the terminal. The terminal displays the minutes received from the server to the user.
[0933] Extraction and visualization of technical terms
[0934] After the minutes are generated, the server uses the generative AI model again to extract technical terms from the minutes. The generative AI analyzes the relationships between these technical terms and generates a keyword map. The generated keyword map is stored in a database by the server.
[0935] Record and track decision history
[0936] The decision-making process during the meeting is recorded along with the minutes, and the server stores them in a database. When a user wants to check the decision-making history for a specific project, the server retrieves the relevant history from the database, sends it to the user's device, and displays it to the user.
[0937] Specific examples
[0938] For example, if a user wants to search for a meeting record for a specific date, they can request, "Show me the meeting record for October 10, 2023." The server processes this request and retrieves the corresponding minutes from the database. The retrieved minutes are then sent to the terminal and displayed to the user.
[0939] Furthermore, if a user wants to check the decision-making history regarding the progress of a project, they can instruct "Display the decision-making history for project X." The server retrieves the decision-making history related to project X from the database, sends it to the terminal, and displays it to the user. This allows the user to quickly and efficiently understand the content of past and present meetings, the history of decision-making, and the relationships between the technical terms used.
[0940] Prompt Sentence Examples
[0941] Examples of input prompts for generative AI models include:
[0942] "View the meeting minutes for October 10, 2023"
[0943] "Show the decision history for Project X"
[0944] This allows users to use the system efficiently and easily access the information they need.
[0945] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0946] Step 1:
[0947] Users record audio data during a meeting. Audio data is collected using smart devices (smartphones and microphones). The input is in the form of audio data, which is sent to the next step.
[0948] Step 2:
[0949] The device uploads the recorded audio data to the server. The input is the audio data file, and the output is the transfer of the audio data to the server. Specifically, the device selects an audio file and sends it to the server via the Internet.
[0950] Step 3:
[0951] The server sends the received voice data to the generative AI model and converts it into text data. The input is voice data and the output is text data. Specifically, speech recognition is performed using Hugging Face's pipeline ("automatic-speech-recognition").
[0952] Step 4:
[0953] The server reprocesses the generated text data and extracts key points and decisions. The input is the text data, and the output is the extracted key points and decisions. The summary is generated using a generative AI summarization model (Hugging Face's pipeline("summarization")).
[0954] Step 5:
[0955] The server automatically generates minutes based on the extracted content and stores them in a database. The input is important points and decisions, and the output is the automatically generated minutes. Specifically, the minutes are organized in chronological order and stored in a database.
[0956] Step 6:
[0957] A user sends a search request to the server specifying a specific date to search for past meeting minutes. The input is a search query (e.g., "Show meeting minutes from October 10, 2023"), which the server receives.
[0958] Step 7:
[0959] The server processes the request from the user and retrieves the relevant minutes from the database. The input is the search query, and the output is the relevant minutes data. Specifically, the server queries the database based on the date and retrieves the results.
[0960] Step 8:
[0961] The server sends the acquired minutes data to the terminal. The input is the minutes data, and the output is the data transfer to the terminal. Specifically, the server sends the data to the terminal via the Internet.
[0962] Step 9:
[0963] The terminal receives the minutes from the server and displays them to the user. The input is the minutes data, and the output is a visual display to the user. The specific operation is that the minutes are displayed on the terminal's display.
[0964] Step 10:
[0965] The server extracts technical terms from the minutes and generates a keyword map. The input is the minutes data, and the output is a keyword map. A generative AI model is used to analyze and visualize the relationships between technical terms.
[0966] Step 11:
[0967] The server saves the keyword map in a database. The input is the keyword map and the output is the saved data. As a specific operation, the generated keyword map is stored in the database.
[0968] Step 12:
[0969] When a user wants to check the decision-making history of a particular project, the server retrieves the relevant history from the database, sends it to the terminal, and displays it to the user. The input is a database query for the decision-making history, and the output is the decision-making history data and its display. Specifically, the server executes the query, obtains the results, and sends them to the terminal, which then displays the received data to the user.
[0970] This series of processes allows for efficient recording of meeting content in real time, allowing important information to be instantly extracted, stored, and searched.
[0971] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0972] This invention is a system that can input voice data or text data, automatically generate minutes based on that data, and extract, store, and search important information, and it also combines an emotion engine that recognizes the user's emotions. Below, an embodiment of this system is explained in natural language.
[0973] Data Input and Transformation
[0974] Users record audio data during a conference and upload the audio file to their terminal after the conference ends. The terminal then transmits the uploaded audio file to the server.
[0975] Speech-to-text conversion
[0976] The server sends the received voice data to the generation AI, which converts the voice data into text data. The generation AI converts the voice into text and returns the text data to the server.
[0977] Emotion recognition
[0978] The server sends the converted text and voice data to the emotion engine, which then extracts emotion data from the voice and text and returns the emotion data to the server.
[0979] Automatic generation of meeting minutes
[0980] The server reprocesses the text data and emotion data received from the generation AI to extract important points and decisions. The generation AI automatically generates minutes based on this data and returns them to the server. The server then stores the generated minutes in a database in chronological order along with the emotion data.
[0981] Finding and Accessing Data
[0982] To access the minutes of past meetings, a user requests a search by specifying a specific date. The server receives the request from the user, retrieves the minutes and emotion data for the corresponding date from the database, and sends them to the device. The device then displays the minutes and emotion data received from the server to the user.
[0983] Extraction and visualization of technical terms
[0984] After the minutes are generated, the server sends the data to the generation AI again, which extracts technical terms from the minutes. The generation AI analyzes the relationships between these technical terms and generates a keyword map. The generated keyword map is then stored in a database by the server. Emotional data is also reflected in the keyword map, visualizing changes in emotions.
[0985] Record and track decision history
[0986] Decisions made during meetings and their progress are recorded along with minutes and stored in a database by the server. When a user wants to check the decision-making history of a specific project, the server retrieves the relevant history from the database and sends it along with emotion data to the device, which then displays it to the user.
[0987] Specific examples
[0988] For example, if a user wants to search for a meeting record for a specific date, they can request, "Show me the meeting record for October 10, 2023." The server processes this request and retrieves the corresponding meeting minutes and emotion data from the database. The retrieved meeting minutes are sent to the device and displayed to the user. Based on the emotion data, the server also displays the overall emotional changes during the meeting.
[0989] Furthermore, if the user wants to check the decision-making history regarding the progress of the project, they can instruct "Show the decision-making history for project X." The server retrieves the decision-making history and emotion data related to project X from the database, sends them to the terminal, and displays them to the user. By also displaying the emotion data, it is possible to grasp the trend of emotions as the project progresses.
[0990] This allows users to quickly and efficiently grasp the content of past and present meetings, the history of decision-making, the relationships between technical terms used, and changes in emotions.
[0991] The processing flow will be explained below.
[0992] Understood. The specific operation will be explained below, divided into processing steps.
[0993] Step 1:
[0994] A user records audio data during a conference and uploads the audio file to a terminal after the conference ends.
[0995] Step 2:
[0996] The device sends the uploaded audio file to the server.
[0997] Step 3:
[0998] The server sends the received voice data to the emotion engine and generation AI, which converts the voice data into text data and recognizes emotions.
[0999] Step 4:
[1000] The generating AI converts the speech into text data and sends that text data back to the server.
[1001] Step 5:
[1002] The emotion engine recognizes the user's emotion from the voice data and returns the emotion data to the server.
[1003] Step 6:
[1004] The server reprocesses the text data received from the generation AI and the emotion data received from the emotion engine, and sends it back to the generation AI to extract important points and decisions.
[1005] Step 7:
[1006] The generation AI extracts important points and decisions from the text data and automatically generates minutes, which are then sent back to the server.
[1007] Step 8:
[1008] The server stores the generated minutes and emotion data in a database. The minutes are time-stamped and organized in chronological order.
[1009] Step 9:
[1010] A user submits a request specifying a specific date to search for minutes of past meetings.
[1011] Step 10:
[1012] The server receives a request from the user, retrieves the minutes of the specified date and the corresponding emotion data from the database, and transmits them to the terminal.
[1013] Step 11:
[1014] The device receives the minutes and emotion data from the server and displays them to the user, allowing the user to check the emotional trends during the meeting as well as the content of the meeting.
[1015] Step 12:
[1016] The server retrieves the minutes from the database, sends them to the generation AI, and extracts technical terms from the minutes.
[1017] Step 13:
[1018] The generative AI extracts technical terms from the minutes, analyzes their relationships, and generates a keyword map, which is then sent back to the server.
[1019] Step 14:
[1020] The server stores the generated keyword map in a database, and also reflects the emotion data in this map so that it can be displayed to the user.
[1021] Step 15:
[1022] A user submits a request specifying the project name to check the decision-making history of a specific project.
[1023] Step 16:
[1024] The server receives a request from a user, retrieves the decision-making history and emotion data for the specified project from the database, and sends them to the terminal.
[1025] Step 17:
[1026] The device receives the decision-making history and emotional data from the server and displays it to the user, allowing the user to understand the decision-making process as the project progresses and the emotional trends that occurred during that process.
[1027] The above is the specific processing flow of this system that combines an emotion engine.
[1028] Example 2
[1029] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1030] In today's business environment, efficient meeting progress and recording are crucial. However, accurately recording important discussions and decisions made during meetings, as well as participants' emotions, and storing and retrieving these in a format that can be used later, can be time-consuming. Furthermore, extracting relevant technical terms and emotional changes from minutes and visualizing them can be difficult. For these reasons, there is a need for a way for users to quickly grasp the details of past meeting records and decisions.
[1031] The specific processing by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for inputting voice data or text data and converting the voice data into text data, means for extracting emotion data from the converted text data, means for extracting important points and decisions from the text data and emotion data and automatically generating minutes, means for organizing the generated minutes and emotion data in chronological order and storing them in a database, means for accessing the chronologically organized minutes and emotion data and searching for minutes of a specific date, means for extracting technical terms from the minutes, creating a keyword map, and visualizing changes in emotion, and means for recording past decision-making history and storing it in a traceable form. This allows users to efficiently manage and quickly access important information from meetings.
[1032] "Audio data" refers to digital data of audio recorded during a meeting or conversation.
[1033] "Text data" is digital data of character information converted from audio data.
[1034] "Means of conversion" refers to the technology or device that analyzes voice data and converts it into text data as character information.
[1035] "Emotion data" refers to data that indicates emotional states and changes extracted from text and speech analysis.
[1036] "Extraction means" refers to technology or devices that extract key points, decisions, or emotional data from text data or audio data.
[1037] "Means for automatically generating minutes" refers to technology or devices that automatically generate documents summarizing meeting content and decisions based on text data and emotion data.
[1038] A "database" is a digital system or software for systematically storing and managing minutes and emotional data.
[1039] "Searchable means" refers to technology or devices that search for required data based on specific dates or conditions from information stored in a database.
[1040] "Terminology" refers to specialized terms and expressions used in a particular field or industry.
[1041] A "keyword map" is a diagram or chart that visually represents the relationships and interactions of extracted technical terms.
[1042] "Visualization means" refers to technology or devices that visually display extracted data.
[1043] "Decision-making history" is data that records the details and process of decisions made in meetings and projects.
[1044] "Means for storing data in a traceable form" refers to technology or devices that store decision-making history in a database for future reference.
[1045] This invention provides a system for processing audio data and text data from a meeting, automatically generating minutes, and efficiently managing important information. This system includes functions for converting audio data, recognizing emotions, automatically generating minutes, and searching and visualizing data. Specific embodiments of the system are described below.
[1046] Entering data
[1047] Users upload audio data recorded during a meeting to their device, which saves the audio in a standard digital audio format (e.g., MP3 or WAV files), and the device sends the uploaded audio file to the server.
[1048] Speech-to-text conversion
[1049] The server sends the received voice data to a speech recognition service (such as Google Cloud Speech-to-Text API), which is a generating AI, and converts the voice data into text data. The generating AI analyzes the voice and converts the content into text data as character information. The converted text data is then sent back to the server.
[1050] Emotion recognition
[1051] The server then sends the converted text and voice data to an emotion recognition engine (e.g., IBM Watson Tone Analyzer) to recognize the user's emotions. The emotion engine extracts emotion data (e.g., joy, anger, sadness, etc.) from the voice tone and text content. This emotion data is then sent back to the server.
[1052] Automatic generation of meeting minutes
[1053] The server uses a generative AI model (e.g., OpenAI's GPT-3) to extract key points and decisions based on the generated text data and emotion data, and automatically generates meeting minutes. The generated minutes include important discussions and decisions made during the meeting. The minutes and emotion data are organized chronologically by the server and stored in a database.
[1054] Finding and Accessing Data
[1055] A user can submit a search request to access meeting minutes for a specific date, for example, using a specific prompt such as "Show me the meeting minutes for October 10, 2023." The server receives this request and retrieves the corresponding meeting minutes and emotion data from the database. The retrieved data is then sent to the device and displayed to the user.
[1056] Extracting technical terms and generating keyword maps
[1057] The server then sends the generated minutes to a generation AI (e.g., BERT) to extract technical terms from the minutes. The generation AI then analyzes the relationships between the extracted technical terms and generates a keyword map. This keyword map is also stored in a database and displayed visually along with sentiment data.
[1058] Record and track decision history
[1059] The server records the decisions made during the meeting and their process, and stores them in a database along with the minutes. If a user wants to check the decision-making history of a specific project, they can use a specific prompt, such as "Show the decision-making history of project X." The server retrieves the relevant data from the database, sends it to the terminal, and displays it to the user.
[1060] This allows users to quickly and efficiently manage and understand past and present meeting details, decision details, terminology relationships, and emotional changes.
[1061] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1062] Step 1:
[1063] Users upload audio data recorded during a meeting to their device. The audio data is saved in a digital audio file format such as MP3 or WAV. The device then sends the uploaded audio file to the server. The input is the audio data, and the output is the audio file sent to the server.
[1064] Step 2:
[1065] The server sends the received voice data to a generation AI (for example, Google Cloud Speech-to-Text API), which is a voice recognition service. The generation AI analyzes the voice data and generates text data as character information. The input is voice data, and the output from the generation AI is text data. The server receives this text data and proceeds to the next step of processing.
[1066] Step 3:
[1067] The server sends the converted text data and the original voice data to an emotion recognition engine (e.g., IBM Watson Tone Analyzer). The emotion recognition engine extracts emotion data from these data. The input is text data and voice data, and the output is emotion data. The server receives the emotion data and proceeds to the next step of processing.
[1068] Step 4:
[1069] The server uses a generative AI model (e.g., OpenAI's GPT-3) based on the text data and emotion data to extract key points and decisions, and automatically generate meeting minutes. The input is text data and emotion data, and the output is automatically generated meeting minutes. The generated minutes are used for further processing.
[1070] Step 5:
[1071] The server organizes the generated minutes and emotion data in chronological order and stores them in a database (e.g., MySQL or PostgreSQL). The input is the minutes and emotion data, and the output is visualized data stored in the database.
[1072] Step 6:
[1073] To access minutes for a specific date, a user sends a search request through their terminal. For example, they request "Show the meeting records for October 10, 2023." The input is the search request, and the server receives this request and retrieves the corresponding minutes and emotion data from the database. The output is the retrieved minutes and emotion data.
[1074] Step 7:
[1075] The terminal displays the minutes and emotion data obtained from the server to the user through a user interface (e.g., a web browser or a dedicated app). The input is the obtained minutes and emotion data, and the output is the meeting record displayed to the user.
[1076] Step 8:
[1077] The server then sends the generated transcripts to a generation AI (e.g., BERT) to extract technical terms from the transcripts and generate a keyword map. The input is the transcripts, and the output is the generated keyword map. The generated keyword map is also stored in a database and can be visualized along with the sentiment data.
[1078] Step 9:
[1079] The server stores the decision-making process and its details in a database along with the meeting minutes. If a user wants to check the decision-making history of a specific project, for example, they can request, "Show the decision-making history of project X." The server retrieves the relevant data from the database, sends it to the terminal, and displays it to the user. The input is the search request and the decision-making history, and the output is the project's decision-making history displayed to the user.
[1080] (Application example 2)
[1081] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1082] Conventional meeting and online event minutes generation systems require manual minutes creation, which is time-consuming and labor-intensive. Furthermore, it is difficult to accurately capture important information and changes in user emotions, making it difficult to quickly grasp decision-making and analyze emotional trends. This creates the risk that some meeting content may be overlooked, and it is difficult to accurately track the decision-making process.
[1083] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for inputting voice data or text data and converting the voice data into text data, means for extracting important points and decisions from the converted text data and automatically generating minutes, means for organizing the generated minutes in chronological order and saving them in a database, means for accessing the chronologically organized minutes and searching for minutes of a specific date, means for extracting technical terms from the minutes and creating a keyword map, means for recording past decision-making history and storing it in a traceable form, means for visualizing user emotional trends using the generated minutes and emotion data, and means for processing the voice data based on prompts and searching for and displaying specific events or meeting records. This enables the content of meetings and online events to be quickly and efficiently recorded, searched, and analyzed.
[1084] "Audio data" refers to data in which audio is recorded in digital format.
[1085] "Text data" refers to data in which character information is recorded in digital format.
[1086] "Means for converting" refers to an algorithm or machine for converting audio data into text data.
[1087] "Key points" are information or matters that should be given special importance in the content of a meeting or event.
[1088] "Decisions" are matters that were discussed at meetings or events and ultimately decided upon.
[1089] Minutes are documents that record the contents of a meeting or event in chronological order.
[1090] "Chronology" refers to the arrangement of events or data in chronological order.
[1091] A "database" is a storage system that systematically stores data and allows it to be searched and updated.
[1092] "Searchable means" refers to methods and technologies for extracting and viewing data based on specific criteria.
[1093] "Terminology" is a term with a specific meaning used in a particular field or industry.
[1094] A "keyword map" is a graphical tool for visually displaying relevant keywords within minutes or other documents.
[1095] "Decision-making history" is a record of decisions and the processes involved in past meetings and events.
[1096] "Emotion data" is data that represents the user's emotional state and is the result of emotion analysis.
[1097] A "means for visualizing emotion trends" is a method for visually displaying changes in a user's emotions over time.
[1098] A "prompt sentence" is a specific instruction sentence that is input to a generative AI model.
[1099] "Processing means" refers to the technology and devices used to analyze, convert and display audio and text data.
[1100] "Means for searching and displaying events and meeting records" means methods or technologies for searching events and meeting records based on specific criteria and displaying the results.
[1101] This invention is a system that converts audio data from meetings and online events into text data in real time, extracts important points and decisions from the text data, and automatically generates minutes.The system also has the function of visualizing users' emotional trends using the generated minutes and emotion data, and searching and displaying specific event or meeting records based on prompt sentences.
[1102] The system utilizes the following hardware and software:
[1103] Hardware used:
[1104] Microphone: For voice input
[1105] Server: storing and processing audio files
[1106] User device: Device used to display results (smartphone, tablet, PC, etc.)
[1107] Software used:
[1108] speech_recognition library: convert speech to text
[1109] transformers library: sentiment analysis
[1110] Custom database module: storing meeting minutes data
[1111] pipeline (natural language processing): processing emotional data
[1112] Process flow:
[1113] During a meeting, users record audio data using a microphone. This audio data is uploaded to the server via the terminal after the meeting ends. The server converts the uploaded audio file into text data using the speech_recognition library.
[1114] The server then sends the converted text data to the transformers library for sentiment analysis. The resulting sentiment data is then sent back to the server. The server then extracts key points and decisions from the text and sentiment data and automatically generates meeting minutes. The generated minutes are then organized and stored in chronological order using a custom database module.
[1115] Users can access meeting minutes by specifying a specific date. The server can also search for past meeting records and emotion data, and display related events and meeting records on the device based on a specific prompt. For example, by entering the prompt "View the meeting records for the live event on October 10, 2023," the corresponding meeting minutes and emotion data will be displayed.
[1116] Furthermore, the server visualizes the user's emotional trends using the generated minutes and emotional data, allowing users to visually grasp changes in emotions during meetings and events.
[1117] In this way, the content of meetings and online events can be recorded, searched and analyzed quickly and efficiently.
[1118] Examples of prompts:
[1119] "View the recording of the live event on October 10, 2023"
[1120] The prompt allows the user to easily search and view meeting records for a particular date.
[1121] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1122] Step 1:
[1123] During a meeting, a user uses a microphone to record audio. The recorded audio data is saved on the device. The input is audio data, and the output is an audio file.
[1124] Step 2:
[1125] After the conference, the user uploads the recorded audio data to the server using the terminal. In this process, the audio file is the input, and the audio file sent to the server is the output.
[1126] Step 3:
[1127] The server receives the uploaded audio file and converts it into text using the speech_recognition library. The input is the audio file, and the output is the generated text. The server then processes the sound waves to convert them into text.
[1128] Step 4:
[1129] The converted text data is sent by the server to the sentiment analysis function of the transformers library. The input is text data and the output is emotion data. The server recognizes the emotion of the text content and generates emotion data.
[1130] Step 5:
[1131] The server automatically generates minutes by extracting important points and decisions based on the generated text data and emotion data. The input is text data and emotion data, and the output is the automatically generated minutes. The server uses natural language processing to extract important information.
[1132] Step 6:
[1133] The generated minutes are organized in chronological order by the server and saved in a custom database module. The input is the minutes, and the output is the minutes data organized in chronological order. The server stores the data in the database.
[1134] Step 7:
[1135] A user searches for meeting minutes for a specific date on their device by entering the prompt "Show meeting notes for the live event on October 10, 2023." The input is the search prompt, and the output is the meeting minutes for the specified date and emotion data.
[1136] Step 8:
[1137] The server searches the database based on the prompt sentence, obtains the relevant minutes and emotion data, and sends them to the terminal. The input is the prompt sentence and the database query, and the output is the search results: the minutes and emotion data.
[1138] Step 9:
[1139] The user checks the minutes and emotion data displayed on the device. The server generates graphs and dashboards to visually display emotion trends based on the emotion data. The input is emotion data, and the output is visualized emotion trends.
[1140] This series of processing steps enables the content of meetings and online events to be recorded, searched, and analyzed quickly and efficiently.
[1141] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1142] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1143] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1144] [Fourth embodiment]
[1145] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1146] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1147] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1148] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1149] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1150] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1151] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1152] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1153] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1154] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1155] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1156] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1157] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1158] This invention is a system that can input voice data or text data, automatically generate minutes based on that data, and extract, store, and search important information. An embodiment of this system will be described below in natural language.
[1159] Data Input and Transformation
[1160] Users record audio data during a meeting and upload the audio file to their device after the meeting ends. The device then sends the uploaded audio file to the server. The server processes the received audio data by sending it to a generation AI, which converts it into text data. The generation AI then converts the audio into text and returns the text data to the server.
[1161] Automatic generation of meeting minutes
[1162] The server reprocesses the text data received from the generation AI and extracts important points and decisions. The generation AI automatically generates minutes based on this text data and returns them to the server. The server then stores the generated minutes in a database in chronological order.
[1163] Finding and Accessing Data
[1164] A user requests a search for a specific date to access the minutes of a past meeting. The server receives the request, retrieves the minutes corresponding to that date from the database, and sends them to the terminal. The terminal displays the minutes received from the server to the user.
[1165] Extraction and visualization of technical terms
[1166] After the minutes are generated, the server sends the data to the AI again to extract technical terms from the minutes. The AI analyzes the relationships between these technical terms and generates a keyword map. The generated keyword map is then stored in a database by the server.
[1167] Record and track decision history
[1168] Decisions made during meetings and their progress are recorded along with minutes and stored in a database by the server. When a user wants to check the decision-making history of a specific project, the server retrieves the relevant history from the database and sends it to the device. The device then displays the decision-making history to the user.
[1169] Specific examples
[1170] For example, if a user wants to search for a meeting record for a specific date, they can request, "Show me the meeting record for October 10, 2023." The server processes this request and retrieves the corresponding minutes from the database. The retrieved minutes are then sent to the terminal and displayed to the user.
[1171] Furthermore, if the user wants to check the decision-making history regarding the progress of the project, he or she can instruct "Show the decision-making history of project X." The server retrieves the decision-making history related to project X from the database, sends it to the terminal, and displays it to the user.
[1172] This allows users to quickly and efficiently understand the content of past and present meetings, the history of decision-making, and the relationships between technical terms used.
[1173] The processing flow will be explained below.
[1174] Understood. The program's processing will be explained in the following steps.
[1175] Step 1:
[1176] A user records audio data during a conference and uploads the audio file to a terminal after the conference ends.
[1177] Step 2:
[1178] The device sends the uploaded audio file to the server.
[1179] Step 3:
[1180] The server sends the received voice data to the generation AI, which converts the voice data into text data.
[1181] Step 4:
[1182] The generating AI converts the speech into text data and sends that text data back to the server.
[1183] Step 5:
[1184] The server reprocesses the text data received from the generation AI and sends it back to the generation AI to extract key points and decisions.
[1185] Step 6:
[1186] The generation AI extracts important points and decisions from the text data and automatically generates minutes, which are then sent back to the server.
[1187] Step 7:
[1188] The server stores the generated minutes in a database, assigning a timestamp to the minutes and organizing them in chronological order.
[1189] Step 8:
[1190] A user requests a search for a particular date to access minutes from a past meeting.
[1191] Step 9:
[1192] The server receives a request from the user, retrieves the minutes of the specified date from the database, and transmits them to the terminal.
[1193] Step 10:
[1194] The terminal displays the minutes received from the server to the user.
[1195] Step 11:
[1196] The server retrieves the minutes from the database, sends them to the generation AI, and extracts technical terms from the minutes.
[1197] Step 12:
[1198] The generative AI extracts technical terms from the minutes, analyzes their relationships, and generates a keyword map, which is then sent back to the server.
[1199] Step 13:
[1200] The server stores the generated keyword map in a database.
[1201] Step 14:
[1202] A user requests a search by specifying the project name to check the decision-making history of a particular project.
[1203] Step 15:
[1204] The server receives a request from a user, retrieves the decision-making history of the specified project from the database, and transmits it to the terminal.
[1205] Step 16:
[1206] The terminal displays the decision-making history received from the server to the user.
[1207] The above is the specific processing flow of this system.
[1208] Example 1
[1209] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1210] In today's business environment, it is important to efficiently manage meeting records and decision-making histories. However, manually creating meeting minutes and sorting through technical terms takes time and effort, and there is a high possibility of missing important information or delaying retrieval. There is also a lack of consistent record management to track past decision-making history. A system to solve these issues and improve meeting efficiency is needed.
[1211] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1212] In this invention, the server includes: means for inputting voice data or text data and converting the voice data to text data; means for extracting important points and decisions from the converted text data and automatically generating minutes; means for organizing the generated minutes in chronological order and saving them in a database; means for accessing the chronologically organized minutes and searching for minutes of a specific date; means for extracting technical terms from the minutes and creating a keyword map; means for recording and storing past decision-making history in a traceable form; means for converting voice data to text data using a generative AI model and extracting important information from the text data to automatically generate minutes; and means for searching for minutes of a specified date or project decision-making history and displaying them to the user. This allows for quick and accurate recording of meeting content, facilitating search and history tracking, and improving work efficiency.
[1213] "Audio Data" means human speech or other acoustic signals recorded during a meeting and stored in digital form.
[1214] "Text data" refers to character string information obtained as a result of analyzing voice data, or directly input character information.
[1215] "Conversion means" refers to technical equipment or software for converting audio data into text data, and may for example be a generative AI model.
[1216] Minutes are documents that organize and summarize the contents of a meeting and record the main topics and decisions made.
[1217] "Key points" or "decisions" refer to particularly important information or conclusions from a meeting and are the main contents that should be recorded in the minutes.
[1218] "Generating means" refers to technical devices or software for automatically creating minutes, including, for example, means for creating minutes from text data using a generative AI model.
[1219] "Database" refers to a system for efficiently storing, managing, and searching collected text data and generated minutes.
[1220] "Search means" refers to technical devices or software that search the minutes stored in the database based on specific criteria.
[1221] "Jargon" refers to specialized words and phrases frequently used in a particular field or industry.
[1222] A "keyword map" refers to a diagram or table that visually shows the relationships between technical terms extracted from minutes.
[1223] "Decision-making history" refers to historical information that records the decisions made during a meeting and the process behind them.
[1224] "Storage means" refers to technical devices and software that store decision-making history and generated minutes in a database in a traceable form.
[1225] A "generative AI model" refers to a machine learning model that uses artificial intelligence technology to analyze voice data and convert it into text data, or to generate minutes based on text data.
[1226] "User" refers to an individual or organizational member who operates the system to upload audio data, search meeting minutes, and check decision-making history.
[1227] This invention is a system that can input voice data or text data, automatically generate minutes based on that data, and extract, store, and search important information. This system aims to efficiently manage meeting records and includes the following components:
[1228] Data Input and Transformation
[1229] Users record audio data during a meeting and upload the audio file to their device after the meeting ends. The uploaded audio file is sent from the device to a server. The server then sends the received audio data to a generative AI model, which converts the audio data into text data. The generative AI model can use speech recognition technology such as the Google Cloud Speech-to-Text API. The generated text data is then returned to the server.
[1230] Automatic generation of meeting minutes
[1231] The server reprocesses the text data received from the generative AI model and uses natural language processing (NLP) techniques to extract key points and decisions. Based on the extracted information, the generative AI model automatically generates meeting minutes. The generated minutes are then stored in a database in chronological order by the server. For example, a database management system such as MySQL or MongoDB can be used.
[1232] Finding and Accessing Data
[1233] A user can search for minutes of past meetings by specifying a specific date. For example, a user can instruct "Show the meeting records for October 10, 2023." Upon receiving this request, the server searches the database for the corresponding minutes and sends them to the terminal. The terminal then displays the retrieved minutes to the user.
[1234] Extraction and visualization of technical terms
[1235] After the minutes are generated, the server sends the data to the generative AI model again to extract technical terms from the minutes. The generative AI model analyzes the relationships between these technical terms and generates a keyword map. The generated keyword map is stored in a database by the server and displayed to the user as needed.
[1236] Record and track decision history
[1237] The server records the decisions made during the meeting and their progress along with the minutes, and stores them in a database. If a user wants to check the decision-making history of a specific project, they can instruct the server to "show the decision-making history for project X." The server receives this request, retrieves the relevant history from the database, and sends it to the terminal. The terminal then displays the decision-making history to the user.
[1238] Specific examples
[1239] For example, if a user wants to search for a meeting record for a specific date, they can request, "Show me the meeting record for October 10, 2023." The server processes this request and retrieves the corresponding minutes from the database. The retrieved minutes are then sent to the terminal and displayed to the user.
[1240] Furthermore, if the user wants to check the decision-making history regarding the progress of a project, he or she can instruct, for example, "Show the decision-making history of project X." The server retrieves the decision-making history related to project X from the database, sends it to the terminal, and displays it to the user.
[1241] This allows users to quickly and efficiently grasp the content of past and present meetings, the history of decision-making, and the relationships between technical terms used. Furthermore, the specific operation of this system improves work efficiency and the accuracy of information management.
[1242] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1243] Step 1:
[1244] A user records audio data during a meeting. Specifically, the user collects the audio data using a smartphone or a dedicated recording device. This audio data is stored in a digital format for subsequent processing.
[1245] Step 2:
[1246] After the meeting, the user uploads the recorded audio file to the device. The user accesses the system's web interface using a browser on their computer or smartphone, selects the audio file, and clicks the upload button. This saves the audio data to the device.
[1247] Step 3:
[1248] The device sends the uploaded audio file to the server. Specifically, it uses an HTTP request to send the audio data and user authentication information to the server. The server receives this.
[1249] Step 4:
[1250] The server receives the voice data and sends it to the generative AI model. The server converts the voice data into an appropriate format and sends a speech recognition request to the generative AI model (for example, Google Cloud Speech-to-Text API). The input data is an audio file, and the output data is text data.
[1251] Step 5:
[1252] The generative AI model converts the voice data into text data and returns the text data to the server. The generative AI model analyzes the voice data and outputs it as text data. Specifically, an algorithm is used to extract character string information from the voice. The server receives this text data and proceeds to the next step.
[1253] Step 6:
[1254] The server reprocesses the text data received from the generative AI model. The input data is text data, and the server uses natural language processing (NLP) techniques to extract key points and decisions. Specifically, it analyzes keywords in the text and identifies the necessary information.
[1255] Step 7:
[1256] The generative AI model automatically generates minutes based on the extracted information. The server sends the extracted information to the generative AI model, which then generates the minutes. The input data are key points and decisions, and the output data are the minutes. The generated minutes are well-formed documents suitable for future reference.
[1257] Step 8:
[1258] The server saves the generated minutes in a database in chronological order. The input data is the minutes, and a date tag is added when saving. For example, the minutes can be saved using a database management system such as MySQL or MongoDB. This makes it easy to search and organize the minutes.
[1259] Step 9:
[1260] A user searches for past meeting minutes by specifying a specific date. Specifically, the user enters "Show me the meeting minutes for October 10, 2023" into the system. This request is sent to the server.
[1261] Step 10:
[1262] The server receives a request from a user and retrieves the minutes corresponding to that date from the database. The input data is a search request, and the server searches for the corresponding minutes using an SQL query or similar. The output data is the retrieved minutes.
[1263] Step 11:
[1264] The server sends the acquired minutes to the terminal. Specifically, it returns the minutes data as an HTTP response. The terminal receives this.
[1265] Step 12:
[1266] The terminal displays the minutes received from the server to the user. The user can view the minutes via a browser or other device. The displayed minutes clearly show the contents of the meeting and the decisions made.
[1267] Step 13:
[1268] The server sends the generated minutes to the generative AI model again to extract technical terms. The input data is the minutes, and the generative AI model identifies technical terms and returns a list of technical terms as output data.
[1269] Step 14:
[1270] The generative AI model analyzes the relationships between technical terms and generates a keyword map. The input data is a list of technical terms, and the output data is a keyword map. Specifically, a diagram is generated that visually shows the relationships between technical terms.
[1271] Step 15:
[1272] The server saves the generated keyword map in a database. The input data is a keyword map, which is saved in a database in an appropriate format, allowing for later verification of terminology relationships.
[1273] Step 16:
[1274] A user submits a request to view the decision history of a particular project, for example, by asking the system, "Show me the decision history of project X." This request is sent to the server.
[1275] Step 17:
[1276] The server retrieves the relevant decision-making history from the database and sends it to the terminal. Input data includes project names and dates, and the server searches for the corresponding history and retrieves the decision-making history as output data.
[1277] Step 18:
[1278] The device displays the decision-making history to the user, who can then view it via a browser or other device. The displayed history clearly shows the progress of the project and important decisions.
[1279] (Application example 1)
[1280] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1281] In today's corporate environment, especially in factories and corporate meetings, there is a need to accurately record meeting content and refer to it later. However, manually creating meeting minutes is time-consuming and laborious, and it is difficult to grasp the relationships between important decisions and technical terms. In addition, there is a lack of systems that can process meeting audio data in real time and instantly generate meeting minutes. It is necessary to solve these issues and record and manage meeting content efficiently and accurately.
[1282] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1283] In this invention, the server includes means for recording voice data in real time and instantly converting it into text data, means for extracting important points and decisions from the converted text data in real time, and means for automatically saving the generated minutes and important points in a database. This makes it possible to efficiently record the contents of a meeting in real time and quickly extract and save important information.
[1284] "Audio data" refers to data in which the sounds of a meeting, conversation, etc. are recorded in digital format.
[1285] "Text data" is digital data that expresses voice data as character information.
[1286] The "means for converting voice data into text data" is a function that executes a process of converting voice data into text information using voice recognition technology.
[1287] "Means for extracting key points and decisions" is a process for automatically identifying and extracting the main agenda items and important decisions of a meeting from text data.
[1288] "Means for automatically generating minutes" is a function that creates a systematic record of the meeting contents based on extracted important points and decisions.
[1289] "Means of organizing in chronological order and storing in a database" refers to the process of organizing the generated minutes in chronological order and storing them in a digital database.
[1290] "A means to search for minutes of a specific date" is a function that searches for minutes saved in the past by specifying a date and displays the minutes you need.
[1291] "Means for extracting technical terms and creating and displaying a keyword map" refers to the process of automatically extracting related technical terms from text data, and creating and displaying a keyword map that visually shows the relevance of these terms.
[1292] "A means of recording decisions and storing the history in a traceable form" is a function that records the history of decisions made in meetings and stores them in a form that allows you to check later when and how the decisions were made.
[1293] The "means for recording in real time and instantly converting it into text data" is a process for recording audio data during a meeting on the spot and converting it into text data in real time without stopping.
[1294] "Means for automatically saving to a database" is a function that automatically saves the generated minutes and extracted important information to a database.
[1295] MODE FOR CARRYING OUT THE INVENTION
[1296] This invention is a system that converts voice data from meetings in factories or companies into text data in real time and automatically creates and saves minutes of the meetings. Specific embodiments of the system are described below.
[1297] System Overview
[1298] The system records audio data in real time and instantly converts it into text. It extracts important points and decisions from the generated text, automatically generating and saving meeting minutes. It also extracts technical terms, creates a keyword map, and stores past decision-making history in a traceable format.
[1299] Data Input and Transformation
[1300] During a meeting, users use a hardware device (such as a microphone or smartphone) to record audio data. The recorded audio data is automatically sent to a server. The server processes the received audio data by sending it to a generative AI model that uses speech recognition technology and converting it into text data. The generative AI model uses Hugging Face's pipeline ("automatic-speech-recognition").
[1301] Automatic generation of meeting minutes
[1302] The server reprocesses the text data obtained by speech recognition and extracts key points and decisions in real time using a generative AI summarization model (for example, Hugging Face's pipeline ("summarization")). The automatically generated minutes are organized in chronological order and stored in a database.
[1303] Finding and Accessing Data
[1304] A user can request a search for a specific date to access the minutes of past meetings. The server receives the request from the user, retrieves the minutes corresponding to that date from the database, and sends them to the terminal. The terminal displays the minutes received from the server to the user.
[1305] Extraction and visualization of technical terms
[1306] After the minutes are generated, the server uses the generative AI model again to extract technical terms from the minutes. The generative AI analyzes the relationships between these technical terms and generates a keyword map. The generated keyword map is then stored in a database by the server.
[1307] Record and track decision history
[1308] The decision-making process during the meeting is recorded along with the minutes, and the server stores them in a database. When a user wants to check the decision-making history for a specific project, the server retrieves the relevant history from the database, sends it to the user's device, and displays it to the user.
[1309] Specific examples
[1310] For example, if a user wants to search for a meeting record for a specific date, they can request, "Show me the meeting record for October 10, 2023." The server processes this request and retrieves the corresponding minutes from the database. The retrieved minutes are then sent to the terminal and displayed to the user.
[1311] Furthermore, if a user wants to check the decision-making history regarding the progress of a project, they can instruct "Display the decision-making history for project X." The server retrieves the decision-making history related to project X from the database, sends it to the terminal, and displays it to the user. This allows the user to quickly and efficiently understand the content of past and present meetings, the history of decision-making, and the relationships between the technical terms used.
[1312] Prompt Sentence Examples
[1313] Examples of input prompts for generative AI models include:
[1314] "View the meeting minutes for October 10, 2023"
[1315] "Show the decision history for Project X"
[1316] This allows users to use the system efficiently and easily access the information they need.
[1317] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1318] Step 1:
[1319] Users record audio data during a meeting. Audio data is collected using smart devices (smartphones and microphones). The input is in the form of audio data, which is sent to the next step.
[1320] Step 2:
[1321] The device uploads the recorded audio data to the server. The input is the audio data file, and the output is the transfer of the audio data to the server. Specifically, the device selects an audio file and sends it to the server via the Internet.
[1322] Step 3:
[1323] The server sends the received voice data to the generative AI model and converts it into text data. The input is voice data and the output is text data. Specifically, speech recognition is performed using Hugging Face's pipeline ("automatic-speech-recognition").
[1324] Step 4:
[1325] The server reprocesses the generated text data and extracts key points and decisions. The input is the text data, and the output is the extracted key points and decisions. The summary is generated using a generative AI summarization model (Hugging Face's pipeline("summarization")).
[1326] Step 5:
[1327] The server automatically generates minutes based on the extracted content and stores them in a database. The input is important points and decisions, and the output is the automatically generated minutes. Specifically, the minutes are organized in chronological order and stored in a database.
[1328] Step 6:
[1329] A user sends a search request to the server specifying a specific date to search for past meeting minutes. The input is a search query (e.g., "Show meeting minutes from October 10, 2023"), which the server receives.
[1330] Step 7:
[1331] The server processes the request from the user and retrieves the relevant minutes from the database. The input is the search query, and the output is the relevant minutes data. Specifically, the server queries the database based on the date and retrieves the results.
[1332] Step 8:
[1333] The server sends the acquired minutes data to the terminal. The input is the minutes data, and the output is the data transfer to the terminal. Specifically, the server sends the data to the terminal via the Internet.
[1334] Step 9:
[1335] The terminal receives the minutes from the server and displays them to the user. The input is the minutes data, and the output is a visual display to the user. The specific operation is that the minutes are displayed on the terminal's display.
[1336] Step 10:
[1337] The server extracts technical terms from the minutes and generates a keyword map. The input is the minutes data, and the output is a keyword map. A generative AI model is used to analyze and visualize the relationships between technical terms.
[1338] Step 11:
[1339] The server saves the keyword map in a database. The input is the keyword map and the output is the saved data. As a specific operation, the generated keyword map is stored in the database.
[1340] Step 12:
[1341] When a user wants to check the decision-making history of a particular project, the server retrieves the relevant history from the database, sends it to the terminal, and displays it to the user. The input is a database query for the decision-making history, and the output is the decision-making history data and its display. Specifically, the server executes the query, obtains the results, and sends them to the terminal, which then displays the received data to the user.
[1342] This series of processes allows for efficient recording of meeting content in real time, allowing important information to be instantly extracted, stored, and searched.
[1343] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1344] This invention is a system that can input voice data or text data, automatically generate minutes based on that data, and extract, store, and search important information, and it also combines an emotion engine that recognizes the user's emotions. Below, an embodiment of this system will be explained in natural language.
[1345] Data Input and Transformation
[1346] Users record audio data during a conference and upload the audio file to their terminal after the conference ends. The terminal then transmits the uploaded audio file to the server.
[1347] Speech-to-text conversion
[1348] The server sends the received voice data to the generation AI, which converts the voice data into text data. The generation AI converts the voice into text and returns the text data to the server.
[1349] Emotion recognition
[1350] The server sends the converted text and voice data to the emotion engine, which then extracts emotion data from the voice and text and returns the emotion data to the server.
[1351] Automatic generation of meeting minutes
[1352] The server reprocesses the text data and emotion data received from the generation AI to extract important points and decisions. The generation AI automatically generates minutes based on this data and returns them to the server. The server then stores the generated minutes in a database in chronological order along with the emotion data.
[1353] Finding and Accessing Data
[1354] To access the minutes of past meetings, a user requests a search by specifying a specific date. The server receives the request from the user, retrieves the minutes and emotion data for the corresponding date from the database, and sends them to the device. The device then displays the minutes and emotion data received from the server to the user.
[1355] Extraction and visualization of technical terms
[1356] After the minutes are generated, the server sends the data to the generation AI again, which extracts technical terms from the minutes. The generation AI analyzes the relationships between these technical terms and generates a keyword map. The generated keyword map is then stored in a database by the server. Emotional data is also reflected in the keyword map, visualizing changes in emotions.
[1357] Record and track decision history
[1358] Decisions made during meetings and their progress are recorded along with minutes and stored in a database by the server. When a user wants to check the decision-making history of a specific project, the server retrieves the relevant history from the database and sends it along with emotion data to the device, which then displays it to the user.
[1359] Specific examples
[1360] For example, if a user wants to search for a meeting record for a specific date, they can request, "Show me the meeting record for October 10, 2023." The server processes this request and retrieves the corresponding meeting minutes and emotion data from the database. The retrieved meeting minutes are sent to the device and displayed to the user. Based on the emotion data, the server also displays the overall emotional changes during the meeting.
[1361] Furthermore, if the user wants to check the decision-making history regarding the progress of the project, they can instruct "Show the decision-making history for project X." The server retrieves the decision-making history and emotion data related to project X from the database, sends them to the terminal, and displays them to the user. By also displaying the emotion data, it is possible to grasp the trend of emotions as the project progresses.
[1362] This allows users to quickly and efficiently grasp the content of past and present meetings, the history of decision-making, the relationships between technical terms used, and changes in emotions.
[1363] The processing flow will be explained below.
[1364] Understood. The specific operation will be explained below, divided into processing steps.
[1365] Step 1:
[1366] A user records audio data during a conference and uploads the audio file to a terminal after the conference ends.
[1367] Step 2:
[1368] The device sends the uploaded audio file to the server.
[1369] Step 3:
[1370] The server sends the received voice data to the emotion engine and generation AI, which converts the voice data into text data and recognizes emotions.
[1371] Step 4:
[1372] The generating AI converts the speech into text data and sends that text data back to the server.
[1373] Step 5:
[1374] The emotion engine recognizes the user's emotion from the voice data and returns the emotion data to the server.
[1375] Step 6:
[1376] The server reprocesses the text data received from the generation AI and the emotion data received from the emotion engine, and sends it back to the generation AI to extract important points and decisions.
[1377] Step 7:
[1378] The generation AI extracts important points and decisions from the text data and automatically generates minutes, which are then sent back to the server.
[1379] Step 8:
[1380] The server stores the generated minutes and emotion data in a database. The minutes are time-stamped and organized in chronological order.
[1381] Step 9:
[1382] A user submits a request specifying a specific date to search for minutes of past meetings.
[1383] Step 10:
[1384] The server receives a request from the user, retrieves the minutes of the specified date and the corresponding emotion data from the database, and transmits them to the terminal.
[1385] Step 11:
[1386] The device receives the minutes and emotion data from the server and displays them to the user, allowing the user to check the emotional trends during the meeting as well as the content of the meeting.
[1387] Step 12:
[1388] The server retrieves the minutes from the database, sends them to the generation AI, and extracts technical terms from the minutes.
[1389] Step 13:
[1390] The generative AI extracts technical terms from the minutes, analyzes their relationships, and generates a keyword map, which is then sent back to the server.
[1391] Step 14:
[1392] The server stores the generated keyword map in a database, and also reflects the emotion data in this map so that it can be displayed to the user.
[1393] Step 15:
[1394] A user submits a request specifying the project name to check the decision-making history of a specific project.
[1395] Step 16:
[1396] The server receives a request from a user, retrieves the decision-making history and emotion data for the specified project from the database, and sends them to the terminal.
[1397] Step 17:
[1398] The device receives the decision-making history and emotional data from the server and displays it to the user, allowing the user to understand the decision-making process as the project progresses and the emotional trends that occurred during that process.
[1399] The above is the specific processing flow of this system that combines an emotion engine.
[1400] Example 2
[1401] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1402] In today's business environment, efficient meeting progress and recording are crucial. However, accurately recording important discussions and decisions made during meetings, as well as participants' emotions, and storing and retrieving these in a format that can be used later, can be time-consuming. Furthermore, extracting relevant technical terms and emotional changes from minutes and visualizing them can be difficult. For these reasons, there is a need for a way for users to quickly grasp the details of past meeting records and decisions.
[1403] The specific processing by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for inputting voice data or text data and converting the voice data into text data, means for extracting emotion data from the converted text data, means for extracting important points and decisions from the text data and emotion data and automatically generating minutes, means for organizing the generated minutes and emotion data in chronological order and storing them in a database, means for accessing the chronologically organized minutes and emotion data and searching for minutes of a specific date, means for extracting technical terms from the minutes, creating a keyword map, and visualizing changes in emotion, and means for recording past decision-making history and storing it in a traceable form. This allows users to efficiently manage and quickly access important information from meetings.
[1404] "Audio data" refers to digital data of audio recorded during a meeting or conversation.
[1405] "Text data" is digital data of character information converted from audio data.
[1406] "Means of conversion" refers to the technology or device that analyzes voice data and converts it into text data as character information.
[1407] "Emotion data" refers to data that indicates emotional states and changes extracted from text and speech analysis.
[1408] "Extraction means" refers to technology or devices that extract key points, decisions, or emotional data from text data or audio data.
[1409] "Means for automatically generating minutes" refers to technology or devices that automatically generate documents summarizing meeting content and decisions based on text data and emotion data.
[1410] A "database" is a digital system or software for systematically storing and managing minutes and emotional data.
[1411] "Searchable means" refers to technology or devices that search for required data based on specific dates or conditions from information stored in a database.
[1412] "Terminology" refers to specialized terms and expressions used in a particular field or industry.
[1413] A "keyword map" is a diagram or chart that visually represents the relationships and interactions of extracted technical terms.
[1414] "Visualization means" refers to technology or devices that visually display extracted data.
[1415] "Decision-making history" is data that records the details and process of decisions made in meetings and projects.
[1416] "Means for storing data in a traceable form" refers to technology or devices that store decision-making history in a database for future reference.
[1417] This invention provides a system for processing audio data and text data from a meeting, automatically generating minutes, and efficiently managing important information. This system includes functions for converting audio data, recognizing emotions, automatically generating minutes, and searching and visualizing data. Specific embodiments of the system are described below.
[1418] Entering data
[1419] Users upload audio data recorded during a meeting to their device, which saves the audio in a standard digital audio format (e.g., MP3 or WAV files), and the device sends the uploaded audio file to the server.
[1420] Speech-to-text conversion
[1421] The server sends the received voice data to a speech recognition service (such as Google Cloud Speech-to-Text API), which is a generating AI, and converts the voice data into text data. The generating AI analyzes the voice and converts the content into text data as character information. The converted text data is then sent back to the server.
[1422] Emotion recognition
[1423] The server then sends the converted text and voice data to an emotion recognition engine (e.g., IBM Watson Tone Analyzer) to recognize the user's emotions. The emotion engine extracts emotion data (e.g., joy, anger, sadness, etc.) from the voice tone and text content. This emotion data is then sent back to the server.
[1424] Automatic generation of meeting minutes
[1425] The server uses a generative AI model (e.g., OpenAI's GPT-3) to extract key points and decisions based on the generated text data and emotion data, and automatically generates meeting minutes. The generated minutes include important discussions and decisions made during the meeting. The minutes and emotion data are organized chronologically by the server and stored in a database.
[1426] Finding and Accessing Data
[1427] A user can submit a search request to access meeting minutes for a specific date, for example, using a specific prompt such as "Show me the meeting minutes for October 10, 2023." The server receives this request and retrieves the corresponding meeting minutes and emotion data from the database. The retrieved data is then sent to the device and displayed to the user.
[1428] Extracting technical terms and generating keyword maps
[1429] The server then sends the generated minutes to a generation AI (e.g., BERT) to extract technical terms from the minutes. The generation AI then analyzes the relationships between the extracted technical terms and generates a keyword map. This keyword map is also stored in a database and displayed visually along with sentiment data.
[1430] Record and track decision history
[1431] The server records the decisions made during the meeting and their process, and stores them in a database along with the minutes. If a user wants to check the decision-making history of a specific project, they can use a specific prompt, such as "Show the decision-making history of project X." The server retrieves the relevant data from the database, sends it to the terminal, and displays it to the user.
[1432] This allows users to quickly and efficiently manage and understand past and present meeting details, decision details, terminology relationships, and emotional changes.
[1433] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1434] Step 1:
[1435] Users upload audio data recorded during a meeting to their device. The audio data is saved in a digital audio file format such as MP3 or WAV. The device then sends the uploaded audio file to the server. The input is the audio data, and the output is the audio file sent to the server.
[1436] Step 2:
[1437] The server sends the received voice data to a generation AI (for example, Google Cloud Speech-to-Text API), which is a voice recognition service. The generation AI analyzes the voice data and generates text data as character information. The input is voice data, and the output from the generation AI is text data. The server receives this text data and proceeds to the next step of processing.
[1438] Step 3:
[1439] The server sends the converted text data and the original voice data to an emotion recognition engine (e.g., IBM Watson Tone Analyzer). The emotion recognition engine extracts emotion data from these data. The input is text data and voice data, and the output is emotion data. The server receives the emotion data and proceeds to the next step of processing.
[1440] Step 4:
[1441] The server uses a generative AI model (e.g., OpenAI's GPT-3) based on the text data and emotion data to extract key points and decisions, and automatically generate meeting minutes. The input is text data and emotion data, and the output is automatically generated meeting minutes. The generated minutes are used for further processing.
[1442] Step 5:
[1443] The server organizes the generated minutes and emotion data in chronological order and stores them in a database (e.g., MySQL or PostgreSQL). The input is the minutes and emotion data, and the output is visualized data stored in the database.
[1444] Step 6:
[1445] To access minutes for a specific date, a user sends a search request through their terminal. For example, they request "Show the meeting records for October 10, 2023." The input is the search request, and the server receives this request and retrieves the corresponding minutes and emotion data from the database. The output is the retrieved minutes and emotion data.
[1446] Step 7:
[1447] The terminal displays the minutes and emotion data obtained from the server to the user through a user interface (e.g., a web browser or a dedicated app). The input is the obtained minutes and emotion data, and the output is the meeting record displayed to the user.
[1448] Step 8:
[1449] The server then sends the generated transcripts to a generation AI (e.g., BERT) to extract technical terms from the transcripts and generate a keyword map. The input is the transcripts, and the output is the generated keyword map. The generated keyword map is also stored in a database and can be visualized along with the sentiment data.
[1450] Step 9:
[1451] The server stores the decision-making process and its details in a database along with the meeting minutes. If a user wants to check the decision-making history of a specific project, for example, they can request, "Show the decision-making history of project X." The server retrieves the relevant data from the database, sends it to the terminal, and displays it to the user. The input is the search request and the decision-making history, and the output is the project's decision-making history displayed to the user.
[1452] (Application example 2)
[1453] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1454] Conventional meeting and online event minutes generation systems require manual minutes creation, which is time-consuming and labor-intensive. Furthermore, it is difficult to accurately capture important information and changes in user emotions, making it difficult to quickly grasp decision-making and analyze emotional trends. This creates the risk that some meeting content may be overlooked, and it is difficult to accurately track the decision-making process.
[1455] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for inputting voice data or text data and converting the voice data into text data, means for extracting important points and decisions from the converted text data and automatically generating minutes, means for organizing the generated minutes in chronological order and saving them in a database, means for accessing the chronologically organized minutes and searching for minutes of a specific date, means for extracting technical terms from the minutes and creating a keyword map, means for recording past decision-making history and storing it in a traceable form, means for visualizing user emotional trends using the generated minutes and emotion data, and means for processing the voice data based on prompts and searching for and displaying specific events or meeting records. This enables the content of meetings and online events to be quickly and efficiently recorded, searched, and analyzed.
[1456] "Audio data" refers to data in which audio is recorded in digital format.
[1457] "Text data" refers to data in which character information is recorded in digital format.
[1458] "Means for converting" refers to an algorithm or machine for converting audio data into text data.
[1459] "Key points" are information or matters that should be given special importance in the content of a meeting or event.
[1460] "Decisions" are matters that were discussed at meetings or events and ultimately decided upon.
[1461] Minutes are documents that record the contents of a meeting or event in chronological order.
[1462] "Chronology" refers to the arrangement of events or data in chronological order.
[1463] A "database" is a storage system that systematically stores data and allows it to be searched and updated.
[1464] "Searchable means" refers to methods and technologies for extracting and viewing data based on specific criteria.
[1465] "Terminology" is a term with a specific meaning used in a particular field or industry.
[1466] A "keyword map" is a graphical tool for visually displaying relevant keywords within minutes or other documents.
[1467] "Decision-making history" is a record of decisions and the processes involved in past meetings and events.
[1468] "Emotion data" is data that represents the user's emotional state and is the result of emotion analysis.
[1469] A "means for visualizing emotion trends" is a method for visually displaying changes in a user's emotions over time.
[1470] A "prompt sentence" is a specific instruction sentence that is input to a generative AI model.
[1471] "Processing means" refers to the technology and devices used to analyze, convert and display audio and text data.
[1472] "Means for searching and displaying events and meeting records" means methods or technologies for searching events and meeting records based on specific criteria and displaying the results.
[1473] This invention is a system that converts audio data from meetings and online events into text data in real time, extracts important points and decisions from the text data, and automatically generates minutes.The system also has the function of visualizing users' emotional trends using the generated minutes and emotion data, and searching and displaying specific event or meeting records based on prompt sentences.
[1474] The system utilizes the following hardware and software:
[1475] Hardware used:
[1476] Microphone: For voice input
[1477] Server: storing and processing audio files
[1478] User device: Device used to display results (smartphone, tablet, PC, etc.)
[1479] Software used:
[1480] speech_recognition library: convert speech to text
[1481] transformers library: sentiment analysis
[1482] Custom database module: storing meeting minutes data
[1483] pipeline (natural language processing): processing emotional data
[1484] Process flow:
[1485] During a meeting, users record audio data using a microphone. This audio data is uploaded to the server via the terminal after the meeting ends. The server converts the uploaded audio file into text data using the speech_recognition library.
[1486] The server then sends the converted text data to the transformers library for sentiment analysis. The resulting sentiment data is then sent back to the server. The server then extracts key points and decisions from the text and sentiment data and automatically generates meeting minutes. The generated minutes are then organized and stored in chronological order using a custom database module.
[1487] Users can access meeting minutes by specifying a specific date. The server can also search for past meeting records and emotion data, and display related events and meeting records on the device based on a specific prompt. For example, by entering the prompt "View the meeting records for the live event on October 10, 2023," the corresponding meeting minutes and emotion data will be displayed.
[1488] Furthermore, the server visualizes the user's emotional trends using the generated minutes and emotional data, allowing users to visually grasp changes in emotions during meetings and events.
[1489] In this way, the content of meetings and online events can be recorded, searched and analyzed quickly and efficiently.
[1490] Examples of prompts:
[1491] "View the recording of the live event on October 10, 2023"
[1492] The prompt allows the user to easily search and view meeting records for a particular date.
[1493] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1494] Step 1:
[1495] During a meeting, a user uses a microphone to record audio. The recorded audio data is saved on the device. The input is audio data, and the output is an audio file.
[1496] Step 2:
[1497] After the conference, the user uploads the recorded audio data to the server using the terminal. In this process, the audio file is the input, and the audio file sent to the server is the output.
[1498] Step 3:
[1499] The server receives the uploaded audio file and converts it into text using the speech_recognition library. The input is the audio file, and the output is the generated text. The server then processes the sound waves to convert them into text.
[1500] Step 4:
[1501] The converted text data is sent by the server to the sentiment analysis function of the transformers library. The input is text data and the output is emotion data. The server recognizes the emotion of the text content and generates emotion data.
[1502] Step 5:
[1503] The server automatically generates minutes by extracting important points and decisions based on the generated text data and emotion data. The input is text data and emotion data, and the output is the automatically generated minutes. The server uses natural language processing to extract important information.
[1504] Step 6:
[1505] The generated minutes are organized in chronological order by the server and saved in a custom database module. The input is the minutes, and the output is the minutes data organized in chronological order. The server stores the data in the database.
[1506] Step 7:
[1507] A user searches for meeting minutes for a specific date on their device by entering the prompt "Show meeting notes for the live event on October 10, 2023." The input is the search prompt, and the output is the meeting minutes for the specified date and emotion data.
[1508] Step 8:
[1509] The server searches the database based on the prompt sentence, obtains the relevant minutes and emotion data, and sends them to the terminal. The input is the prompt sentence and the database query, and the output is the search results: the minutes and emotion data.
[1510] Step 9:
[1511] The user checks the minutes and emotion data displayed on the device. The server generates graphs and dashboards to visually display emotion trends based on the emotion data. The input is emotion data, and the output is visualized emotion trends.
[1512] This series of processing steps enables the content of meetings and online events to be recorded, searched, and analyzed quickly and efficiently.
[1513] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1514] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1515] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1516] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1517] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1518] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1519] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1520] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1521] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1522] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1523] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1524] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1525] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1526] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1527] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1528] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1529] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1530] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1531] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1532] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1533] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1534] The following is further disclosed regarding the above embodiment.
[1535] Understood. Below are the draft claims.
[1536] (Claim 1)
[1537] a means for inputting voice data or text data and converting the voice data into text data;
[1538] A means to extract important points and decisions from the converted text data and automatically generate minutes of the meeting;
[1539] A means to organize the generated minutes in chronological order and store them in a database;
[1540] A means to access minutes organized chronologically and search for minutes for a specific date;
[1541] A method for extracting technical terms from minutes and creating a keyword map;
[1542] A system that includes a means to record and store past decision-making history in a traceable form.
[1543] (Claim 2)
[1544] 2. The system according to claim 1, wherein audio data during a meeting is recorded and the audio data is used to generate minutes of the meeting.
[1545] (Claim 3)
[1546] The system of claim 1 visualizes the relationships between technical terms based on the generated minutes and displays them as a keyword map.
[1547] "Example 1"
[1548] (Claim 1)
[1549] means for inputting voice data or text data and converting the voice data into text data;
[1550] A means to extract important points and decisions from the converted text data and automatically generate minutes of the meeting;
[1551] A means to organize the generated minutes in chronological order and store them in a database;
[1552] A means to access minutes organized chronologically and search for minutes for a specific date;
[1553] A method for extracting technical terms from minutes and creating a keyword map;
[1554] A means of recording and storing past decision-making history in a traceable form;
[1555] A means for automatically generating meeting minutes by converting voice data into text data using a generative AI model and extracting important information from the text data;
[1556] A means for searching and displaying to a user meeting minutes or project decision history for a specified date;
[1557] A system including:
[1558] (Claim 2)
[1559] 2. The system according to claim 1, wherein audio data during a meeting is recorded and the audio data is used to generate minutes of the meeting.
[1560] (Claim 3)
[1561] The system of claim 1 visualizes the relationships between technical terms based on the generated minutes and displays them as a keyword map.
[1562] "Application Example 1"
[1563] (Claim 1)
[1564] means for inputting voice data or text data and converting the voice data into text data;
[1565] A means to extract important points and decisions from the converted text data and automatically generate minutes of the meeting;
[1566] A means to organize the generated minutes in chronological order and store them in a database;
[1567] A means to search for minutes for a specific date;
[1568] A means to extract technical terms from minutes and create and display the relationships between terms as a keyword map;
[1569] In addition to the means to record decisions and keep a traceable history of them,
[1570] A means to record voice data in real time and instantly convert it into text data;
[1571] A means to extract key points and decisions in real time,
[1572] The system includes a means for automatically storing the generated minutes and key points in a database.
[1573] (Claim 2)
[1574] 2. The system according to claim 1, wherein audio data during a meeting is recorded in real time and the audio data is used to generate minutes of the meeting.
[1575] (Claim 3)
[1576] The system of claim 1 visualizes the relationships between technical terms based on the generated minutes and displays them as a keyword map.
[1577] "Example 2: Combining Emotion Engines"
[1578] (Claim 1)
[1579] means for inputting voice data or text data and converting the voice data into text data;
[1580] means for extracting emotion data from the converted text data;
[1581] A means for extracting important points and decisions from text data and emotion data and automatically generating minutes;
[1582] A means to organize the generated minutes and emotion data in chronological order and store them in a database;
[1583] A means to access chronologically organized minutes and sentiment data, and to search for minutes for a specific date;
[1584] A method to extract technical terms from minutes, create a keyword map, and visualize changes in emotions.
[1585] A system that includes a means to record and store past decision-making history in a traceable form.
[1586] (Claim 2)
[1587] 2. The system according to claim 1, wherein audio data during a meeting is recorded and the audio data is used to generate minutes and extract emotion data.
[1588] (Claim 3)
[1589] The system of claim 1 visualizes the relationships between technical terms and changes in sentiment based on the generated minutes and displays them as a keyword map.
[1590] "Application example 2 when combining emotion engines"
[1591] (Claim 1)
[1592] a means for inputting voice data or text data and converting the voice data into text data;
[1593] A means to extract important points and decisions from the converted text data and automatically generate minutes of the meeting;
[1594] A means to organize the generated minutes in chronological order and store them in a database;
[1595] A means to access minutes organized chronologically and search for minutes for a specific date;
[1596] A method for extracting technical terms from minutes and creating a keyword map;
[1597] A means of recording and storing past decision-making history in a traceable form;
[1598] A means for visualizing user emotional trends using the generated minutes and emotional data;
[1599] The system includes a means for processing audio data based on prompts to search for and display specific events or meeting records.
[1600] (Claim 2)
[1601] 2. The system according to claim 1, wherein audio data during a meeting is recorded and the audio data is used to generate minutes of the meeting.
[1602] (Claim 3)
[1603] The system of claim 1 visualizes the relationships between technical terms based on the generated minutes and displays them as a keyword map. [Explanation of symbols]
[1604] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. a means for inputting voice data or text data and converting the voice data into text data; A means to extract important points and decisions from the converted text data and automatically generate minutes of the meeting; A means to organize the generated minutes in chronological order and store them in a database; A means to access minutes organized chronologically and search for minutes for a specific date; A method for extracting technical terms from minutes and creating a keyword map; A system that includes a means to record and store past decision-making history in a traceable form.
2. 2. The system according to claim 1, wherein audio data during a meeting is recorded and the audio data is used to generate minutes of the meeting.
3. The system according to claim 1, wherein the system visualizes the relationships between technical terms based on the generated minutes and displays them as a keyword map.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A