system
The system automates meeting minute generation, capturing key points and emotions, reducing manual effort and improving meeting efficiency and communication.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-11-13
- Publication Date
- 2026-05-25
AI Technical Summary
The burden of manually creating meeting minutes is significant, and there is a lack of systems that can accurately capture and summarize important content while considering the emotions and tone of participants.
A system that includes audio acquisition, text conversion, summary generation, formatting, and transmission capabilities to automatically generate meeting minutes, incorporating sentiment analysis to reflect the emotions and tone of participants.
Reduces the time and effort required for manual minute-taking, ensures accurate capture of key points, and provides insights into meeting dynamics through emotion analysis, enhancing organizational communication and decision-making.
Smart Images

Figure 2026085750000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, the method including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance as a response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] An object of the present invention is to reduce the burden when creating minutes during a meeting and improve accuracy and efficiency. Specifically, an object is to solve the problem of being unable to concentrate on the speech and missing important content in the minutes, and to provide a technology for supporting the effective progress of the meeting.
Means for Solving the Problems
[0005] This invention provides a system that enables the automatic summarization of meeting content and the generation of meeting minutes by comprising an audio acquisition means for acquiring audio, a text conversion means for converting this audio into text data, a summary generation means for summarizing this text data, and a formatting means for formatting the generated summary. As a result, users are freed from the burden of creating meeting minutes and can record important meeting content without omission.
[0006] "Voice acquisition means" refers to a function or device used to collect audio from meetings and other similar events.
[0007] "Character conversion means" refers to a function or device that analyzes acquired audio and converts it into text data.
[0008] "Summary generation means" refers to a function or device that extracts important information from converted text data and summarizes it concisely.
[0009] "Formatting means" refers to a function or device that arranges the generated summary into a standard format and presents it in an easily understandable manner.
[0010] "Transmission means" refers to a function or device that sends formatted data to a specific destination or device.
[0011] "Correction means" refers to a function or device that allows users to review generated summaries and data and edit or modify them as needed. [Brief explanation of the drawing]
[0012] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3]It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which a plurality of emotions are mapped. [Figure 10] It shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when an emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when an emotion engine is combined.
Mode for Carrying Out the Invention
[0013] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0014] First, the language used in the following description will be explained.
[0015] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0016] In the following embodiments, the numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0017] In the following embodiments, the numbered storage is one or more non-volatile storage devices that store various programs, various parameters, and the like. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.
[0018] In the following embodiments, the numbered communication I / F (Interface) is an interface including a communication processor, an antenna, and the like. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0019] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0020] [First Embodiment]
[0021] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0022] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0023] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0024] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0025] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0026] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0027] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0028] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0029] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0030] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0031] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0032] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0033] This invention provides a system that automatically generates summarized meeting minutes using audio data acquired during a meeting, and allows users to easily review them. The following methods are envisioned for carrying out this invention.
[0034] The server receives audio data transmitted from the terminals during the meeting. The terminals are equipped with audio acquisition capabilities and begin recording at the start of the meeting. The recorded audio is sent to the server at regular time intervals or in fixed data volumes.
[0035] The server converts the transmitted audio data into text format using a text conversion means. This converted text is then processed by a summary generation means within the server to generate a summary that extracts the important points of the content.
[0036] Subsequently, the server formats the summary obtained by the summary generation means into a predetermined format using the formatting means. For example, it visually organizes the content by dividing it into bullet points or paragraphs.
[0037] The formatted summary is sent to the user's device via a transmission method. The user reviews the transmitted minutes on their device, corrects any errors using the correction method if necessary, and saves or shares it as the final minutes.
[0038] As a concrete example, a user can start recording a meeting and immediately receive automatically generated meeting minutes on their device after the meeting ends. These minutes cover the key points of the meeting, allowing the user to quickly grasp the content and share it with their supervisor or team members as needed.
[0039] Thus, the system of the present invention eliminates the significant time and effort required for the conventional manual creation of meeting minutes, and allows for more effective capture of the key points of a meeting.
[0040] The following describes the processing flow.
[0041] Step 1:
[0042] The device prepares to capture participants' speech in real time using the microphone, after the user launches the application and starts recording the meeting.
[0043] Step 2:
[0044] The device sends recorded audio data to the server at regular time intervals. This data is treated as an audio stream and is designed to enable rapid processing.
[0045] Step 3:
[0046] The server converts the received audio data into text data using a text conversion device. This process utilizes speech recognition technology to accurately transcribe speech into text.
[0047] Step 4:
[0048] The server inputs the transcribed text into a summary generation system, which extracts key information from the meeting and generates a concise summary. This process utilizes generative AI, which analyzes the context and content of the statements made.
[0049] Step 5:
[0050] The server formats the generated summary into a standard meeting minutes format using formatting tools. This provides the necessary structure to make it easy for users to read.
[0051] Step 6:
[0052] The server sends formatted meeting minutes to the user's terminal using a transmission method. Transmission can occur in real time or after the meeting has ended.
[0053] Step 7:
[0054] Users review the meeting minutes sent to their own devices. If there are any errors in the minutes, they can edit them directly using the correction tools to finalize them.
[0055] Step 8:
[0056] Users can share revised meeting minutes with colleagues and supervisors as needed, or save them as an archive. This ensures that the meeting content is accurately preserved and can be referenced later.
[0057] (Example 1)
[0058] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0059] In modern meetings, much information is exchanged verbally, and it is essential to efficiently and accurately grasp and record the key points. However, manually summarizing meeting content is time-consuming and labor-intensive, so a system is needed to automate this process and improve efficiency.
[0060] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0061] In this invention, the server includes means for acquiring audio information, speech recognition means for converting the audio information into text information, and summary generation means for processing the acquired text information based on a summary generation algorithm and extracting key points. This makes it possible to automatically transcribe meeting audio data into text, extract key points from it, and easily create a highly readable summary.
[0062] "Means for acquiring audio information" refers to devices or systems for electronically recording speech, such as in meetings, and includes microphones and communication devices.
[0063] "Speech recognition means" refers to the technology or software that processes acquired speech information and converts it into corresponding text information, and generally uses a speech recognition engine.
[0064] "Summary generation means" refers to algorithms or systems that extract important points from textual information, condense the amount of information, and summarize the important content.
[0065] "Formatting methods" refer to the processes and techniques used to arrange the generated summary into a more visually appealing and easily understandable format, and include conversion to charts, graphs, and bullet points.
[0066] "Communication methods" refer to technologies and protocols for transmitting formatted information to information devices in remote locations, and include network communication.
[0067] An "information device" is an electronic device used by users to receive, display, and edit information, and typically includes computers and smartphones.
[0068] "Display means" refers to a device or technology that visually presents information received on an information device to a user, and includes displays and monitors.
[0069] "Editing means" refers to technologies or software that enable users to modify or add to information they have received.
[0070] This system is designed to automatically transcribe audio information into text and summarize it. Specifically, it can process audio information acquired in meetings and other similar settings in real time or after recording.
[0071] The terminal is equipped with a device for acquiring audio, thereby collecting audio information during the meeting. The hardware uses a microphone and is operated in combination with software that can save the audio in digital format.
[0072] The server receives audio data transmitted from the terminal and converts this audio data into text data using speech recognition technology. In this step, it is possible to use common technologies as the speech recognition engine; for example, a common speech recognition API could be used.
[0073] The converted text data is processed on the server by a summarization generation mechanism, which extracts the key points and creates a summary. An example of a prompt when generating a summary using a generative AI model is, "Summarize the key points of the following text." A generative AI model can be used as the summarization algorithm.
[0074] The server then uses formatting tools to arrange the generated summary into a visually easy-to-understand format. For example, it improves readability by dividing it into bullet points or paragraphs. This formatted summary is then transmitted to the user's terminal via communication tools.
[0075] Users can review the summary received on their device and modify its content as needed using the editing tools. The final summary is saved by the user and can be shared with other members as needed.
[0076] Thus, this system provides a convenient and effective means of accurately summarizing, efficiently recording, and sharing meeting content.
[0077] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0078] Step 1:
[0079] The terminal uses an audio acquisition device to record the meeting audio in real time in order to obtain audio information. The input is the audio data spoken during the meeting, which is output as a digital audio file via the terminal's microphone. This audio data is saved in an audio file format (e.g., WAV, MP3).
[0080] Step 2:
[0081] The terminal transmits voice data to the server at regular time intervals or according to the amount of data. The input is the voice data stored in step 1, and the output is the voice packets transferred to the server over the network. This ensures that the voice data is delivered to the server in real time.
[0082] Step 3:
[0083] The server converts the received audio data into text data using speech recognition. The input here is digital audio data transmitted from the terminal, and the output converted by the speech recognition engine is text data. For example, a speech recognition API is used for this process.
[0084] Step 4:
[0085] The server summarizes the converted character data using a summarization generation mechanism. The input is the text data obtained in step 3, and the server uses a generation AI model to extract the key points and output them as a summarized text. The specific prompt used is, "Please summarize the key points of the following text."
[0086] Step 5:
[0087] The server formats the generated summary into a visually understandable format using formatting tools. The input is the summary text generated in step 4, and the output is the formatted visual summary. For example, it can be made easier to understand by using bullet points or dividing it into paragraphs.
[0088] Step 6:
[0089] The server sends the formatted summary to the user's terminal. The input is the formatted summary data from step 5, and the output is the summary data delivered to the user's terminal via the network. This allows the user to immediately review the summary.
[0090] Step 7:
[0091] The user reviews the summary received on their device and makes corrections as needed using the editing tools. The input is the summary data received in step 6, and the final meeting minutes text edited by the user is output. The revised meeting minutes may be saved and shared with other members.
[0092] (Application Example 1)
[0093] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0094] Conventional autonomous vehicles lacked sufficient systems to effectively grasp ambient sounds and emergency information in real time and support appropriate decision-making. Therefore, there is a need to quickly provide useful information obtained from ambient sounds to assist the driver.
[0095] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0096] In this invention, the server includes a voice collection means, an information conversion means, an information summarization means, an environmental analysis means, and an information presentation means. This makes it possible to quickly extract important information from ambient sounds and notify the driver in real time.
[0097] A "sound collection device" is a device that converts ambient sounds into electrical signals and records them as digital data.
[0098] "Information conversion means" refers to a process or device that converts collected audio data into textual information or other forms of symbolic information.
[0099] An "information summarization tool" is a technology or device that extracts important elements from converted symbolic information and reconstructs it into concise and to-the-point information.
[0100] A "structuring method" is a technique or system for organizing summarized information visually or in documents and making it presentable in a standardized format.
[0101] "Environmental analysis means" refers to a device or technology that has the function of detecting ambient sounds and selecting and analyzing important information from them.
[0102] "Information presentation means" refers to interfaces or systems for providing users with analyzed and structured information visually or audibly.
[0103] "Transmission means" refers to a communication channel or technology for transmitting formed data to a specific device.
[0104] A "correction tool" is a technology that provides an interface or functions for users to edit or correct the information presented to them.
[0105] The following system configurations are possible as embodiments for carrying out this invention.
[0106] The server acquires ambient sounds using voice acquisition means installed in the in-vehicle device. A high-sensitivity microphone is used in this process. The server then uses speech recognition software (e.g., Google® Speech-to-Text API) as an information conversion means to convert the collected voice data into text-based symbolic information.
[0107] As a means of information summarization, a generative AI model is used to extract important content from the converted text information. This summarized information is then visually organized using structuring methods and formatted in a way that is easily understandable to the user.
[0108] Furthermore, as an environmental analysis tool, the server incorporates a function to analyze audio data around the vehicle and filter out emergency and traffic information. As a result of this analysis, important information is extracted and provided to the driver in real time via an information display device. For example, if the sound of an emergency vehicle siren is detected, a warning can be immediately issued to the driver. Information is presented using an on-screen display or a voice assistant.
[0109] As a concrete example, the system notifies the driver in real time that "an emergency vehicle is approaching" and prompts them to take appropriate action. An example of a prompt message to the generating AI model in this case could be an instruction such as, "Acquire ambient audio data while driving, summarize the emergency information, and inform me."
[0110] This system allows the server to provide users with a safe and comfortable operating environment.
[0111] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0112] Step 1:
[0113] The server acquires ambient sound using the audio acquisition means of the in-vehicle device. At this time, the audio data is input from the microphone as an analog signal. Because a high-sensitivity microphone is used, a wide range of audio data, including ambient sounds and conversations, can be accurately captured.
[0114] Step 2:
[0115] The server uses speech recognition software to convert the acquired analog audio signal into digital text format. This conversion process transforms the input audio signal into string data. As a result, the audio content is output as visualized symbolic information.
[0116] Step 3:
[0117] The server uses a generative AI model to summarize the converted text information. This method involves supplying converted text as input and extracting important information. In this summarization process, essential information is extracted from longer texts, and a short summary containing only the key points is output.
[0118] Step 4:
[0119] The server organizes the extracted summaries using structuring methods and converts them into a visually clear format. The goal is to format the input summary data into paragraphs, bullet points, etc., and output it in a format that is easy for the user to understand.
[0120] Step 5:
[0121] The server uses environmental analysis tools to select important information from surrounding audio data. Audio data to be analyzed is provided as input, and important sounds such as emergency signals are identified from it. The selected data is classified as requiring immediate attention.
[0122] Step 6:
[0123] The device notifies the driver of important information in real time through information display means. The information displayed is visualized or spoken digitally via a voice assistant or display. For example, information such as "Emergency vehicle approaching" is displayed immediately to support the driver's actions.
[0124] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0125] This invention provides a system that not only generates summarized meeting minutes using audio data acquired during a meeting, but also analyzes the emotions of the speakers and reflects them in the meeting minutes and emotion reports. The following methods can be considered to implement this invention.
[0126] The terminal records speech during the meeting using an audio acquisition device. This audio data is transmitted to the server in real time. The server converts the received audio into text data using a text conversion device. The converted text is processed by a summary generation device, which extracts important information from the meeting and generates a summary.
[0127] Here, the present invention performs a further advanced analysis using an emotion engine. The server recognizes the speaker's emotions from the text data and analyzes what emotions were expressed. Based on this emotion data, the summary generation means adjusts the tone of the meeting minutes to reflect the emotions of the participants.
[0128] The formatting means formats the summary, including the sentiment analysis results, and this is sent to the user's terminal via the transmission means. The user can review the received meeting minutes on their terminal and make corrections as needed. Furthermore, the sentiment report generation means analyzes the sentiment trends of the entire meeting and compiles this into an easy-to-understand report.
[0129] As a concrete example, a user can start recording a meeting and immediately receive a summarized transcript on their device after the meeting ends. This transcript reflects not only the main points made by the meeting participants but also an analysis of the emotions conveyed in their statements. This allows the user to gain insights into the atmosphere of the meeting and the reactions of the participants, which can be used to improve future meetings and project management.
[0130] This system allows for the capture of meeting dynamics that could not be fully understood through traditional meeting minute-taking methods, contributing to improved organizational communication and decision-making.
[0131] The following describes the processing flow.
[0132] Step 1:
[0133] The device launches an application that allows the user to start recording the meeting. It uses the microphone to capture the speech of meeting participants in real time.
[0134] Step 2:
[0135] The device sends the acquired audio data to the server at regular intervals. The audio data is processed in a streaming format.
[0136] Step 3:
[0137] The server converts the received audio data into text data using speech recognition technology and a text conversion method. By considering pronunciation and linguistic characteristics during the conversion process, it achieves highly accurate text conversion.
[0138] Step 4:
[0139] The server passes the converted text data to the sentiment engine, which analyzes keywords and context within the text to recognize the speaker's emotions. This process utilizes advanced natural language processing techniques to ensure accurate sentiment evaluation.
[0140] Step 5:
[0141] The server uses a summarization generation mechanism to extract key information from the text data of the meeting and create a concise summary. The tone and emphasis of the summary are adjusted by providing sentiment data received from the sentiment engine.
[0142] Step 6:
[0143] The server formats the meeting minutes, including the summary and sentiment analysis results, using a formatting method. These minutes are structured using headings and bullet points.
[0144] Step 7:
[0145] The server delivers the formatted meeting minutes to the user's terminal using a transmission method. The user can then immediately view the meeting minutes on their terminal.
[0146] Step 8:
[0147] Users can read the received meeting minutes through the terminal interface and make any necessary corrections or additional edits. Once the corrections are complete, they can save or share the meeting minutes.
[0148] Step 9:
[0149] The server uses an emotion report generation mechanism to generate a report summarizing the overall emotional trends of the meeting. This includes temporal fluctuations in emotions and an overall emotional assessment of all participants.
[0150] Step 10:
[0151] Users can review sentiment reports provided on their devices and receive feedback that reflects the atmosphere of the meeting and the reactions of the participants. This information can then be used to plan future meetings and adjust communication strategies.
[0152] (Example 2)
[0153] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0154] In modern meetings, creating meeting minutes is crucial for accurately reflecting the content of the meeting, but manually creating detailed minutes is time-consuming and laborious. Furthermore, because the emotions and tone of voice of participants during the meeting are not recorded, it is difficult to grasp the overall atmosphere of the meeting.
[0155] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0156] In this invention, the server includes means for acquiring audio information, means for converting the acquired audio information into text information, means for analyzing the converted text information and extracting important content, and means for associating the generated summary with emotional information using an emotional analysis engine. This makes it possible to automatically generate meeting minutes that summarize the meeting content while reflecting the emotions of the participants.
[0157] "Means for acquiring audio information" refers to devices and methods for recording speech in a meeting, which involve acquiring audio as digital data using microphones or recording devices.
[0158] "Means for converting audio information into text information" refers to technologies that analyze acquired audio data and convert it into corresponding strings, and are implemented using speech recognition engines or similar technologies.
[0159] "Methods for analyzing text information and extracting important content" refer to methods for selecting key points and important information from converted text, and are implemented using natural language processing technology.
[0160] "Methods for associating emotional information with an emotional analysis engine" refers to the process of analyzing the speaker's emotions from the text content and reflecting this in the summary, and this is done through an emotional analysis algorithm.
[0161] "Formatting methods" refer to techniques and methods for organizing generated summaries and sentiment information in an easily understandable way and arranging them into a defined format.
[0162] "Means of transmitting to an information terminal" refers to devices or protocols used to transmit processed data to an external terminal via a network, and utilizes communication technologies for data transfer.
[0163] "User-modifiable means" refers to methods and interfaces that provide an environment in which users can view and edit the received summary data and make modifications as needed.
[0164] This invention is a system for automatically generating meeting minutes, which includes everything from acquiring audio information and converting it to text, generating summaries, performing sentiment analysis, formatting the summaries, and providing them to the user. This system uses multiple means to extract value-added information from meeting records.
[0165] The terminal uses an input device such as a microphone to acquire audio and collects speech during the meeting in real time. This audio data is compressed and sent to the server via a communication method. The server uses speech recognition technology to convert the audio data into text data. A common speech recognition API can be used for this technology.
[0166] Next, the server uses a generative AI model to extract key information from the text data and generate a summary. Because natural language processing techniques are used in this process, the generated summary accurately captures the essence of the meeting. Furthermore, an emotion analysis engine is used to analyze the speaker's emotions from the text data and reflect this in the summary. In this step, an emotion analysis algorithm is applied, quantifying emotions based on factors such as the tone of speech.
[0167] The formatted summary is presented in a visually easy-to-understand format. This format utilizes data representation formats such as Markdown or HTML to establish a consistent summary format. Finally, the server sends the formatted summary and sentiment analysis information to the user's device. The user can review the received summary and make corrections as needed.
[0168] As a concrete example, a user starts recording a meeting using an application. After the meeting ends, the system automatically processes the audio data and sends a summarized meeting transcript along with the results of a sentiment analysis of the speakers to the user's device. The user can then use this information to consider ways to improve future meetings or projects.
[0169] An example of a prompt might be: "Summarize the key information from the meeting recording sequentially, including the speakers' emotions, and create a report. Provide it in a highly visual format."
[0170] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0171] Step 1:
[0172] The terminal uses an audio acquisition device to record all speech during the meeting in real time. The input is an analog audio signal acquired through a microphone, which is then converted and stored as digital audio data. The digital audio data is efficiently compressed and transmitted to the server according to the communication protocol.
[0173] Step 2:
[0174] The server processes the received digital audio data using a speech recognition engine and converts it into text data. The input is a compressed audio file, and speech recognition technology is used to analyze each phrase and word and convert it into corresponding text. The output is in text format and is temporarily stored for further processing.
[0175] Step 3:
[0176] The server uses a generative AI model to extract important content from the transformed text data and generate a summary. The input is the text data obtained in the previous step, and natural language processing algorithms are applied to select the key points and essential information of the meeting. The output is the summarized text, which is used in the next processing step.
[0177] Step 4:
[0178] The server analyzes text data using an emotion analysis engine to identify the speaker's emotional information. The input is text data after summary generation, and emotion analysis technology is used to extract and analyze emotional keywords and context. The output is quantified emotion information, which is added to the summary and reflected in the meeting minutes.
[0179] Step 5:
[0180] The server formats summaries and sentiment information in a visually organized manner. Input consists of processed summary text and sentiment data, which are then converted into a readable format. Output is delivered to the user as a Markdown or HTML document via a transmission method.
[0181] Step 6:
[0182] The user receives formatted meeting minutes on their terminal and reviews the information provided. Input consists of various formatted data sent from the server, and the user can view the content using the application and modify it as needed. Output is the modified meeting minutes, which can be used as a reference for future meetings or projects.
[0183] (Application Example 2)
[0184] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0185] In recent years, improving communication among workers and enhancing work efficiency have become crucial issues in factories and workplaces. Furthermore, there is a need to understand workers' emotions and health conditions in real time to improve the safety and efficiency of the work environment. However, current systems struggle to automatically perform real-time emotion analysis and propose safety-enhancing measures based on that analysis.
[0186] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0187] In this invention, the server includes voice acquisition means, text conversion means, emotion analysis means, and suggestion generation means. This makes it possible to instantly analyze the emotional state of workers from voice data acquired within the factory and provide appropriate suggestions to improve safety and efficiency.
[0188] A "speech acquisition means" is a device that collects speech data from the environment and provides it in an analyzable format.
[0189] A "text conversion means" is a device that processes audio data and converts it into text data that represents that audio.
[0190] A "summary generation tool" is a device that extracts important information from text data and summarizes it in a concise form.
[0191] A "formatting tool" is a device that has the function of preparing summary data into a format that is easy to read and organize for output.
[0192] An "emotion analysis tool" is a device that identifies a worker's emotional state from voice or text data and analyzes the type and intensity of that emotion.
[0193] A "proposal generation method" is a system that has the function of proposing specific actions to improve work efficiency and safety based on analyzed emotional information.
[0194] "Transmission means" refers to a device or system that provides the function of transmitting formatted summaries or proposals to the user's terminal.
[0195] "Correction mechanisms" refer to features that allow users to modify or update submitted summaries and suggestions.
[0196] The system implementing this invention mainly consists of three elements: a server, a terminal, and a user. The server collects voice data in real time from factories and work sites through voice acquisition means. This voice data is converted into text data by text conversion means. On the server, the essence of important information is extracted from this text data by summary generation means and compiled into a concise summary.
[0197] Next, the server uses sentiment analysis means to analyze the worker's emotional state from the text data. Based on this analysis, the suggestion generation means generates specific suggestions to improve the safety and efficiency of the work. The generated summary and suggestions are organized by the formatting means and transmitted to the user's terminal via the transmission means.
[0198] On the terminal, users can not only review these summaries and suggestions, but also make corrections as needed using the correction tools. This facilitates smoother communication at the work site and optimizes the work environment.
[0199] For example, if text data such as "I'm tired today" or "It seems I'm slowing down my work pace" is analyzed on a terminal and negative emotions are detected, the server can immediately generate a suggestion to take a break and communicate it to the user. This enables a quick response and proper management of human resources.
[0200] An example of a prompt for a generative AI model might be an instruction such as, "Analyze the text data extracted from the conversation and generate suggestions to improve safety and efficiency in the workplace."
[0201] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0202] Step 1:
[0203] The server uses voice acquisition methods to collect audio data in real time from factories and work sites. The acquired audio data is stored in digital format. The input is an audio signal collected by a microphone, and the output is digital audio data. Specifically, it actively performs noise cancellation to improve the quality of the collected data.
[0204] Step 2:
[0205] The server uses a text conversion mechanism to convert the acquired audio data into text data. Here, ASR (Automatic Speech Recognition) technology is used to analyze the audio signal and output the results as text data. The input is previously processed audio data, and the output is text data that reflects the audio content. Specifically, it analyzes each audio sequence and generates its text representation.
[0206] Step 3:
[0207] The server uses a summarization generation mechanism to extract important information from the generated text data and create a summary. Using natural language processing techniques, it identifies key points in the text and outputs a concise summary. The input is text data, and the output is a summarized, concise text. Specifically, it uses morphological analysis to extract important nouns and verbs and generates content based on them.
[0208] Step 4:
[0209] The server utilizes sentiment analysis tools to identify the worker's emotional state from text data. Through a sentiment analysis engine, it classifies linguistic features within the text, resulting in sentiment data. The input is text data, and the output is sentiment data indicating the type and intensity of the emotion. Specifically, it classifies text into three categories: positive, negative, and neutral.
[0210] Step 5:
[0211] The server, using a proposal generation mechanism, formulates proposals to improve the safety and efficiency of the work environment based on sentiment data. In this process, sentiment data is analyzed, optimal countermeasures are devised, and output as text. The input is sentiment data, and the output is text proposing specific action plans. Specifically, it refers to past cases in the database and proposes solutions for similar cases.
[0212] Step 6:
[0213] The server uses formatting tools to convert summaries and proposals into a neat and organized format. Here, the output text is formatted and maintained for readability. Since the input is summaries and proposals, the output is text in a neat format. Specifically, the text is edited according to the predetermined formatting settings.
[0214] Step 7:
[0215] The server transmits formatted summaries and suggestions to the user's terminal via a transmission method. Digital communication technology is used to deliver these texts to the terminal. The input is formatted text, and the output is display data on the user's terminal. Specifically, it is converted into data packets and transmitted over the network.
[0216] Step 8:
[0217] The terminal displays the received summary and suggestions, allowing the user to review the content. Furthermore, the user can modify this information through editing tools and make optimized decisions as needed. The input is text data received from the server, and the output is the modified text. Specifically, the text is displayed on the screen, and editing is enabled on the interface.
[0218] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0219] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0220] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0221] [Second Embodiment]
[0222] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0223] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0224] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0225] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0226] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0227] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0228] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0229] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0230] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0231] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0232] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0233] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0234] This invention provides a system that automatically generates summarized meeting minutes using audio data acquired during a meeting, and allows users to easily review them. The following methods are envisioned for carrying out this invention.
[0235] The server receives audio data transmitted from the terminals during the meeting. The terminals are equipped with audio acquisition capabilities and begin recording at the start of the meeting. The recorded audio is sent to the server at regular time intervals or in fixed data volumes.
[0236] The server converts the transmitted audio data into text format using a text conversion means. This converted text is then processed by a summary generation means within the server to generate a summary that extracts the important points of the content.
[0237] Subsequently, the server formats the summary obtained by the summary generation means into a predetermined format using the formatting means. For example, it visually organizes the content by dividing it into bullet points or paragraphs.
[0238] The formatted summary is sent to the user's device via a transmission method. The user reviews the transmitted minutes on their device, corrects any errors using the correction method if necessary, and saves or shares it as the final minutes.
[0239] As a concrete example, a user can start recording a meeting and immediately receive automatically generated meeting minutes on their device after the meeting ends. These minutes cover the key points of the meeting, allowing the user to quickly grasp the content and share it with their supervisor or team members as needed.
[0240] Thus, the system of the present invention eliminates the significant time and effort required for the conventional manual creation of meeting minutes, and allows for more effective capture of the key points of a meeting.
[0241] The following describes the processing flow.
[0242] Step 1:
[0243] The device prepares to capture participants' speech in real time using the microphone, after the user launches the application and starts recording the meeting.
[0244] Step 2:
[0245] The device sends recorded audio data to the server at regular time intervals. This data is treated as an audio stream and is designed to enable rapid processing.
[0246] Step 3:
[0247] The server converts the received audio data into text data using a text conversion device. This process utilizes speech recognition technology to accurately transcribe speech into text.
[0248] Step 4:
[0249] The server inputs the transcribed text into a summary generation system, which extracts key information from the meeting and generates a concise summary. This process utilizes generative AI, which analyzes the context and content of the statements made.
[0250] Step 5:
[0251] The server formats the generated summary into a standard meeting minutes format using formatting tools. This provides the necessary structure to make it easy for users to read.
[0252] Step 6:
[0253] The server sends formatted meeting minutes to the user's terminal using a transmission method. Transmission can occur in real time or after the meeting has ended.
[0254] Step 7:
[0255] Users review the meeting minutes sent to their own devices. If there are any errors in the minutes, they can edit them directly using the correction tools to finalize them.
[0256] Step 8:
[0257] Users can share revised meeting minutes with colleagues and supervisors as needed, or save them as an archive. This ensures that the meeting content is accurately preserved and can be referenced later.
[0258] (Example 1)
[0259] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0260] In modern meetings, much information is exchanged verbally, and it is essential to efficiently and accurately grasp and record the key points. However, manually summarizing meeting content is time-consuming and labor-intensive, so a system is needed to automate this process and improve efficiency.
[0261] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0262] In this invention, the server includes means for acquiring audio information, speech recognition means for converting the audio information into text information, and summary generation means for processing the acquired text information based on a summary generation algorithm and extracting key points. This makes it possible to automatically transcribe meeting audio data into text, extract key points from it, and easily create a highly readable summary.
[0263] "Means for acquiring audio information" refers to devices or systems for electronically recording speech, such as in meetings, and includes microphones and communication devices.
[0264] "Speech recognition means" refers to the technology or software that processes acquired speech information and converts it into corresponding text information, and generally uses a speech recognition engine.
[0265] "Summary generation means" refers to algorithms or systems that extract important points from textual information, condense the amount of information, and summarize the important content.
[0266] "Formatting methods" refer to the processes and techniques used to arrange the generated summary into a more visually appealing and easily understandable format, and include conversion to charts, graphs, and bullet points.
[0267] "Communication methods" refer to technologies and protocols for transmitting formatted information to information devices in remote locations, and include network communication.
[0268] An "information device" is an electronic device used by users to receive, display, and edit information, and typically includes computers and smartphones.
[0269] "Display means" refers to a device or technology that visually presents information received on an information device to a user, and includes displays and monitors.
[0270] "Editing means" refers to technologies or software that enable users to modify or add to information they have received.
[0271] This system is designed to automatically transcribe audio information into text and summarize it. Specifically, it can process audio information acquired in meetings and other similar settings in real time or after recording.
[0272] The terminal is equipped with a device for acquiring audio, thereby collecting audio information during the meeting. The hardware uses a microphone and is operated in combination with software that can save the audio in digital format.
[0273] The server receives audio data transmitted from the terminal and converts this audio data into text data using speech recognition technology. In this step, it is possible to use common technologies as the speech recognition engine; for example, a common speech recognition API could be used.
[0274] The converted text data is processed on the server by a summarization generation mechanism, which extracts the key points and creates a summary. An example of a prompt when generating a summary using a generative AI model is, "Summarize the key points of the following text." A generative AI model can be used as the summarization algorithm.
[0275] The server then uses formatting tools to arrange the generated summary into a visually easy-to-understand format. For example, it improves readability by dividing it into bullet points or paragraphs. This formatted summary is then transmitted to the user's terminal via communication tools.
[0276] Users can review the summary received on their device and modify its content as needed using the editing tools. The final summary is saved by the user and can be shared with other members as needed.
[0277] Thus, this system provides a convenient and effective means of accurately summarizing, efficiently recording, and sharing meeting content.
[0278] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0279] Step 1:
[0280] The terminal uses an audio acquisition device to record the meeting audio in real time in order to obtain audio information. The input is the audio data spoken during the meeting, which is output as a digital audio file via the terminal's microphone. This audio data is saved in an audio file format (e.g., WAV, MP3).
[0281] Step 2:
[0282] The terminal transmits voice data to the server at regular time intervals or according to the amount of data. The input is the voice data stored in step 1, and the output is the voice packets transferred to the server over the network. This ensures that the voice data is delivered to the server in real time.
[0283] Step 3:
[0284] The server converts the received voice data into character data using voice recognition means. The input here is the digital voice data transmitted from the terminal, and the output converted by the voice recognition engine is text data. For this process, for example, a voice recognition API is used.
[0285] Step 4:
[0286] The server summarizes the converted character data using summarization means. The input is the text data obtained in Step 3, and by utilizing a generation AI model, key points are extracted and output as a summary text. As a specific prompt sentence, "Please summarize the important points of the following text." is used.
[0287] Step 5:
[0288] The server formats the generated summary into a visually easy-to-understand format by formatting means. The input is the summary text generated in Step 4, and the output is the formatted visual summary format. For example, it is made easier to understand by making it into a list or separating it into paragraphs.
[0289] Step 6:
[0290] The server transmits the formatted summary to the user's terminal. The input is the summary data formatted in Step 5, and the output is the summary data distributed to the user's terminal through the network. As a result, the user can immediately check the summary. <
[0294] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0295] Conventional autonomous vehicles lacked sufficient systems to effectively grasp ambient sounds and emergency information in real time and support appropriate decision-making. Therefore, there is a need to quickly provide useful information obtained from ambient sounds to assist the driver.
[0296] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0297] In this invention, the server includes a voice collection means, an information conversion means, an information summarization means, an environmental analysis means, and an information presentation means. This makes it possible to quickly extract important information from ambient sounds and notify the driver in real time.
[0298] A "sound collection device" is a device that converts ambient sounds into electrical signals and records them as digital data.
[0299] "Information conversion means" refers to a process or device that converts collected audio data into textual information or other forms of symbolic information.
[0300] An "information summarization tool" is a technology or device that extracts important elements from converted symbolic information and reconstructs it into concise and to-the-point information.
[0301] A "structuring method" is a technique or system for organizing summarized information visually or in documents and making it presentable in a standardized format.
[0302] "Environmental analysis means" refers to a device or technology that has the function of detecting ambient sounds and selecting and analyzing important information from them.
[0303] "Information presentation means" refers to an interface or system for visually or auditorily providing analyzed and structured information to users.
[0304] "Transmission means" refers to a communication channel or technology for transmitting formed data to a specific device.
[0305] "Modification means" refers to a technology equipped with an interface or function for editing and correcting the content of information presented to the user.
[0306] As a form for implementing this invention, the following system configuration can be considered.
[0307] The server acquires ambient sound using the voice collection means installed in the in-vehicle device. In this process, a high-sensitivity microphone is utilized. Next, the server uses speech recognition software (e.g., Google Speech-to-Text API) as information conversion means to convert the collected voice data into symbolic information in text format.
[0308] ?As information summarization means, a generative AI model is utilized to extract important content from the converted text information. This summarized information is visually organized by the structuring means and formatted into a form that can be easily understood by the user.
[0309] Furthermore, as environmental analysis means, a function for analyzing the voice data around the vehicle and selecting emergency information and traffic information is incorporated into the server. As a result of this analysis, important information is extracted and provided to the driver in real time via the information presentation means. For example, when the siren sound of an emergency vehicle is detected, a warning can be immediately issued to the driver. Information presentation is performed using an on-screen display or a voice assistant.
[0310] As a concrete example, the system notifies the driver in real time that "an emergency vehicle is approaching" and prompts them to take appropriate action. An example of a prompt message to the generating AI model in this case could be an instruction such as, "Acquire ambient audio data while driving, summarize the emergency information, and inform me."
[0311] This system allows the server to provide users with a safe and comfortable operating environment.
[0312] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0313] Step 1:
[0314] The server acquires ambient sounds using the audio acquisition method of the in-vehicle device. At this time, the audio data is input from the microphone as an analog signal. Because a high-sensitivity microphone is used, a wide range of audio data, including ambient sounds and conversations, can be accurately captured.
[0315] Step 2:
[0316] The server uses speech recognition software to convert the acquired analog audio signal into digital text format. This conversion process transforms the input audio signal into string data. As a result, the audio content is output as visualized symbolic information.
[0317] Step 3:
[0318] The server uses a generative AI model to summarize the converted text information. This method involves supplying converted text as input and extracting important information. In this summarization process, essential information is extracted from longer texts, and a short summary containing only the key points is output.
[0319] Step 4:
[0320] The server organizes the extracted summaries using structuring methods and converts them into a visually clear format. The goal is to format the input summary data into paragraphs, bullet points, etc., and output it in a format that is easy for the user to understand.
[0321] Step 5:
[0322] The server uses environmental analysis tools to select important information from surrounding audio data. Audio data to be analyzed is provided as input, and important sounds such as emergency signals are identified from it. The selected data is classified as requiring immediate attention.
[0323] Step 6:
[0324] The device notifies the driver of important information in real time through information display means. The information displayed is visualized or spoken digitally via a voice assistant or display. For example, information such as "Emergency vehicle approaching" is displayed immediately to support the driver's actions.
[0325] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0326] This invention provides a system that not only generates summarized meeting minutes using audio data acquired during a meeting, but also analyzes the emotions of the speakers and reflects them in the meeting minutes and emotion reports. The following methods can be considered to implement this invention.
[0327] The terminal records speech during the meeting using an audio acquisition device. This audio data is transmitted to the server in real time. The server converts the received audio into text data using a text conversion device. The converted text is processed by a summary generation device, which extracts important information from the meeting and generates a summary.
[0328] Here, the present invention performs a further advanced analysis using an emotion engine. The server recognizes the speaker's emotions from the text data and analyzes what emotions were expressed. Based on this emotion data, the summary generation means adjusts the tone of the meeting minutes to reflect the emotions of the participants.
[0329] The formatting means formats the summary, including the sentiment analysis results, and this is sent to the user's terminal via the transmission means. The user can review the received meeting minutes on their terminal and make corrections as needed. Furthermore, the sentiment report generation means analyzes the sentiment trends of the entire meeting and compiles this into an easy-to-understand report.
[0330] As a concrete example, a user can start recording a meeting and immediately receive a summarized transcript on their device after the meeting ends. This transcript reflects not only the main points made by the meeting participants but also an analysis of the emotions conveyed in their statements. This allows the user to gain insights into the atmosphere of the meeting and the reactions of the participants, which can be used to improve future meetings and project management.
[0331] This system allows for the capture of meeting dynamics that could not be fully understood through traditional meeting minute-taking methods, contributing to improved organizational communication and decision-making.
[0332] The following describes the processing flow.
[0333] Step 1:
[0334] The device launches an application that allows the user to start recording the meeting. It uses the microphone to capture the speech of meeting participants in real time.
[0335] Step 2:
[0336] The device sends the acquired audio data to the server at regular intervals. The audio data is processed in a streaming format.
[0337] Step 3:
[0338] The server converts the received audio data into text data using speech recognition technology and a text conversion method. By considering pronunciation and linguistic characteristics during the conversion process, it achieves highly accurate text conversion.
[0339] Step 4:
[0340] The server passes the converted text data to the sentiment engine, which analyzes keywords and context within the text to recognize the speaker's emotions. This process utilizes advanced natural language processing techniques to ensure accurate sentiment evaluation.
[0341] Step 5:
[0342] The server uses a summarization generation mechanism to extract key information from the text data of the meeting and create a concise summary. The tone and emphasis of the summary are adjusted by providing sentiment data received from the sentiment engine.
[0343] Step 6:
[0344] The server formats the meeting minutes, including the summary and sentiment analysis results, using a formatting method. These minutes are structured using headings and bullet points.
[0345] Step 7:
[0346] The server delivers the formatted meeting minutes to the user's terminal using a transmission method. The user can immediately view the meeting minutes on their terminal.
[0347] Step 8:
[0348] Users can read the received meeting minutes through the terminal interface and make any necessary corrections or additional edits. Once the corrections are complete, they can save or share the meeting minutes.
[0349] Step 9:
[0350] The server uses an emotion report generation mechanism to generate a report summarizing the overall emotional trends of the meeting. This includes temporal fluctuations in emotions and an overall emotional assessment of all participants.
[0351] Step 10:
[0352] Users can review sentiment reports provided on their devices and receive feedback that reflects the atmosphere of the meeting and the reactions of the participants. This information can then be used to plan future meetings and adjust communication strategies.
[0353] (Example 2)
[0354] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0355] In modern meetings, creating meeting minutes is crucial for accurately reflecting the content of the meeting, but manually creating detailed minutes is time-consuming and laborious. Furthermore, because the emotions and tone of voice of participants during the meeting are not recorded, it is difficult to grasp the overall atmosphere of the meeting.
[0356] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0357] In this invention, the server includes means for acquiring audio information, means for converting the acquired audio information into text information, means for analyzing the converted text information and extracting important content, and means for associating the generated summary with emotional information using an emotional analysis engine. This makes it possible to automatically generate meeting minutes that summarize the meeting content while reflecting the emotions of the participants.
[0358] "Means for acquiring audio information" refers to devices and methods for recording speech in a meeting, which involve acquiring audio as digital data using microphones or recording devices.
[0359] "Means for converting audio information into text information" refers to technologies that analyze acquired audio data and convert it into corresponding strings, and are implemented using speech recognition engines or similar technologies.
[0360] "Methods for analyzing text information and extracting important content" refer to methods for selecting key points and important information from converted text, and are implemented using natural language processing technology.
[0361] "Methods for associating emotional information with an emotional analysis engine" refers to the process of analyzing the speaker's emotions from the text content and reflecting this in the summary, and this is done through an emotional analysis algorithm.
[0362] "Formatting methods" refer to techniques and methods for organizing generated summaries and sentiment information in an easily understandable way and arranging them into a defined format.
[0363] "Means of transmitting to an information terminal" refers to devices or protocols used to transmit processed data to an external terminal via a network, and utilizes communication technologies for data transfer.
[0364] "User-modifiable means" refers to methods and interfaces that provide an environment in which users can view and edit the received summary data and make modifications as needed.
[0365] This invention is a system for automatically generating meeting minutes, which includes everything from acquiring audio information and converting it to text, generating summaries, performing sentiment analysis, formatting the summaries, and providing them to the user. This system uses multiple means to extract value-added information from meeting records.
[0366] The terminal uses an input device such as a microphone to acquire audio and collects speech during the meeting in real time. This audio data is compressed and sent to the server via a communication method. The server uses speech recognition technology to convert the audio data into text data. A common speech recognition API can be used for this technology.
[0367] Next, the server uses a generative AI model to extract key information from the text data and generate a summary. Because natural language processing techniques are used in this process, the generated summary accurately captures the essence of the meeting. Furthermore, an emotion analysis engine is used to analyze the speaker's emotions from the text data and reflect this in the summary. In this step, an emotion analysis algorithm is applied, quantifying emotions based on factors such as the tone of speech.
[0368] The formatted summary is presented in a visually easy-to-understand format. This format utilizes data representation formats such as Markdown or HTML to establish a consistent summary format. Finally, the server sends the formatted summary and sentiment analysis information to the user's device. The user can review the received summary and make corrections as needed.
[0369] As a concrete example, a user starts recording a meeting using an application. After the meeting ends, the system automatically processes the audio data and sends a summarized meeting transcript along with the results of a sentiment analysis of the speakers to the user's device. The user can then use this information to consider ways to improve future meetings or projects.
[0370] An example of a prompt might be: "Summarize the key information from the meeting recording sequentially, including the speakers' emotions, and create a report. Provide it in a highly visual format."
[0371] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0372] Step 1:
[0373] The terminal uses an audio acquisition device to record all speech during the meeting in real time. The input is an analog audio signal acquired through a microphone, which is then converted and stored as digital audio data. The digital audio data is efficiently compressed and transmitted to the server according to the communication protocol.
[0374] Step 2:
[0375] The server processes the received digital audio data using a speech recognition engine and converts it into text data. The input is a compressed audio file, and speech recognition technology is used to analyze each phrase and word and convert it into corresponding text. The output is in text format and is temporarily stored for further processing.
[0376] Step 3:
[0377] The server uses a generative AI model to extract important content from the transformed text data and generate a summary. The input is the text data obtained in the previous step, and natural language processing algorithms are applied to select the key points and essential information of the meeting. The output is the summarized text, which is used in the next processing step.
[0378] Step 4:
[0379] The server analyzes text data using an emotion analysis engine to identify the speaker's emotional information. The input is text data after summary generation, and emotion analysis technology is used to extract and analyze emotional keywords and context. The output is quantified emotion information, which is added to the summary and reflected in the meeting minutes.
[0380] Step 5:
[0381] The server formats summaries and sentiment information in a visually organized manner. Input consists of processed summary text and sentiment data, which are then converted into a readable format. Output is delivered to the user as a Markdown or HTML document via a transmission method.
[0382] Step 6:
[0383] The user receives formatted meeting minutes on their terminal and reviews the information provided. Input consists of various formatted data sent from the server, and the user can view the content using the application and modify it as needed. Output is the modified meeting minutes, which can be used as a reference for future meetings or projects.
[0384] (Application Example 2)
[0385] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0386] In recent years, improving communication among workers and enhancing work efficiency have become crucial issues in factories and workplaces. Furthermore, there is a need to understand workers' emotions and health conditions in real time to improve the safety and efficiency of the work environment. However, current systems struggle to automatically perform real-time emotion analysis and propose safety-enhancing measures based on that analysis.
[0387] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0388] In this invention, the server includes voice acquisition means, text conversion means, emotion analysis means, and suggestion generation means. This makes it possible to instantly analyze the emotional state of workers from voice data acquired within the factory and provide appropriate suggestions to improve safety and efficiency.
[0389] A "voice acquisition means" is a device that collects voice data from the environment and provides it in an analyzable format.
[0390] A "text conversion means" is a device that processes audio data and converts it into text data that represents that audio.
[0391] A "summary generation means" is a device that extracts important information from text data and summarizes it in a concise form.
[0392] A "formatting tool" is a device that has the function of preparing summary data into a format that is easy to read and organize for output.
[0393] An "emotion analysis tool" is a device that identifies a worker's emotional state from voice or text data and analyzes the type and intensity of that emotion.
[0394] A "proposal generation method" is a system that has the function of proposing specific actions to improve work efficiency and safety based on analyzed emotional information.
[0395] "Transmission means" refers to a device or system that provides the function of transmitting formatted summaries or proposals to the user's terminal.
[0396] "Correction mechanisms" refer to features that allow users to modify or update submitted summaries and suggestions.
[0397] The system implementing this invention mainly consists of three elements: a server, a terminal, and a user. The server collects voice data in real time from factories and work sites through voice acquisition means. This voice data is converted into text data by text conversion means. On the server, the essence of important information is extracted from this text data by summary generation means and compiled into a concise summary.
[0398] Next, the server uses sentiment analysis means to analyze the worker's emotional state from the text data. Based on this analysis, the suggestion generation means generates specific suggestions to improve the safety and efficiency of the work. The generated summary and suggestions are organized by the formatting means and transmitted to the user's terminal via the transmission means.
[0399] On the terminal, users can not only review these summaries and suggestions, but also make corrections as needed using the correction tools. This facilitates smoother communication at the work site and optimizes the work environment.
[0400] For example, if text data such as "I'm tired today" or "It seems I'm slowing down my work pace" is analyzed on a terminal and negative emotions are detected, the server can immediately generate a suggestion to take a break and communicate it to the user. This enables a quick response and proper management of human resources.
[0401] An example of a prompt for a generative AI model might be an instruction such as, "Analyze the text data extracted from the conversation and generate suggestions to improve safety and efficiency in the workplace."
[0402] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0403] Step 1:
[0404] The server uses voice acquisition methods to collect audio data in real time from factories and work sites. The acquired audio data is stored in digital format. The input is an audio signal collected by a microphone, and the output is digital audio data. Specifically, it actively performs noise cancellation to improve the quality of the collected data.
[0405] Step 2:
[0406] The server uses a text conversion mechanism to convert the acquired audio data into text data. Here, ASR (Automatic Speech Recognition) technology is used to analyze the audio signal and output the results as text data. The input is previously processed audio data, and the output is text data that reflects the audio content. Specifically, it analyzes each audio sequence and generates its text representation.
[0407] Step 3:
[0408] The server uses a summarization generation mechanism to extract important information from the generated text data and create a summary. Using natural language processing techniques, it identifies key points in the text and outputs a concise summary. The input is text data, and the output is a summarized, concise text. Specifically, it uses morphological analysis to extract important nouns and verbs and generates content based on them.
[0409] Step 4:
[0410] The server utilizes sentiment analysis tools to identify the worker's emotional state from text data. Through a sentiment analysis engine, it classifies linguistic features within the text, resulting in sentiment data. The input is text data, and the output is sentiment data indicating the type and intensity of the emotion. Specifically, it classifies text into three categories: positive, negative, and neutral.
[0411] Step 5:
[0412] The server, using a proposal generation mechanism, formulates proposals to improve the safety and efficiency of the work environment based on sentiment data. In this process, sentiment data is analyzed, optimal countermeasures are devised, and output as text. The input is sentiment data, and the output is text proposing specific action plans. Specifically, it refers to past cases in the database and proposes solutions for similar cases.
[0413] Step 6:
[0414] The server uses formatting tools to convert summaries and proposals into a neat and organized format. Here, the output text is formatted and maintained for readability. Since the input is summaries and proposals, the output is text in a neat format. Specifically, the text is edited according to the predetermined formatting settings.
[0415] Step 7:
[0416] The server transmits formatted summaries and suggestions to the user's terminal via a transmission method. Digital communication technology is used to deliver these texts to the terminal. The input is formatted text, and the output is display data on the user's terminal. Specifically, it is converted into data packets and transmitted over the network.
[0417] Step 8:
[0418] The terminal displays the received summary and suggestions, allowing the user to review the content. Furthermore, the user can modify this information through editing tools and make optimized decisions as needed. The input is text data received from the server, and the output is the modified text. Specifically, the text is displayed on the screen, and editing is enabled on the interface.
[0419] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0420] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0421] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0422] [Third Embodiment]
[0423] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0424] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0425] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0426] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0427] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0428] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0429] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0430] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0431] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0432] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0433] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0434] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0435] This invention provides a system that automatically generates summarized meeting minutes using audio data acquired during a meeting, and allows users to easily review them. The following methods are envisioned for carrying out this invention.
[0436] The server receives audio data transmitted from the terminals during the meeting. The terminals are equipped with audio acquisition capabilities and begin recording at the start of the meeting. The recorded audio is sent to the server at regular time intervals or in fixed data volumes.
[0437] The server converts the transmitted audio data into text format using a text conversion means. This converted text is then processed by a summary generation means within the server to generate a summary that extracts the important points of the content.
[0438] Subsequently, the server formats the summary obtained by the summary generation means into a predetermined format using the formatting means. For example, it visually organizes the content by dividing it into bullet points or paragraphs.
[0439] The formatted summary is sent to the user's device via a transmission method. The user reviews the transmitted minutes on their device, corrects any errors using the correction method if necessary, and saves or shares it as the final minutes.
[0440] As a concrete example, a user can start recording a meeting and immediately receive automatically generated meeting minutes on their device after the meeting ends. These minutes cover the key points of the meeting, allowing the user to quickly grasp the content and share it with their supervisor or team members as needed.
[0441] Thus, the system of the present invention eliminates the significant time and effort required for the conventional manual creation of meeting minutes, and allows for more effective capture of the key points of a meeting.
[0442] The following describes the processing flow.
[0443] Step 1:
[0444] The device prepares to capture participants' speech in real time using the microphone, after the user launches the application and starts recording the meeting.
[0445] Step 2:
[0446] The device sends recorded audio data to the server at regular time intervals. This data is treated as an audio stream and is designed to enable rapid processing.
[0447] Step 3:
[0448] The server converts the received audio data into text data using a text conversion device. This process utilizes speech recognition technology to accurately transcribe speech into text.
[0449] Step 4:
[0450] The server inputs the transcribed text into a summary generation system, which extracts key information from the meeting and generates a concise summary. This process utilizes generative AI, which analyzes the context and content of the statements made.
[0451] Step 5:
[0452] The server formats the generated summary into a standard meeting minutes format using formatting tools. This provides the necessary structure to make it easy for users to read.
[0453] Step 6:
[0454] The server sends formatted meeting minutes to the user's terminal using a transmission method. Transmission can occur in real time or after the meeting has ended.
[0455] Step 7:
[0456] Users review the meeting minutes sent to their own devices. If there are any errors in the minutes, they can edit them directly using the correction tools to finalize them.
[0457] Step 8:
[0458] Users can share revised meeting minutes with colleagues and supervisors as needed, or save them as an archive. This ensures that the meeting content is accurately preserved and can be referenced later.
[0459] (Example 1)
[0460] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0461] In modern meetings, much information is exchanged verbally, and it is essential to efficiently and accurately grasp and record the key points. However, manually summarizing meeting content is time-consuming and labor-intensive, so a system is needed to automate this process and improve efficiency.
[0462] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0463] In this invention, the server includes means for acquiring audio information, speech recognition means for converting the audio information into text information, and summary generation means for processing the acquired text information based on a summary generation algorithm and extracting key points. This makes it possible to automatically transcribe meeting audio data into text, extract key points from it, and easily create a highly readable summary.
[0464] "Means for acquiring audio information" refers to devices or systems for electronically recording speech, such as in meetings, and includes microphones and communication devices.
[0465] "Speech recognition means" refers to the technology or software that processes acquired speech information and converts it into corresponding text information, and generally uses a speech recognition engine.
[0466] "Summary generation means" refers to algorithms or systems that extract important points from textual information, condense the amount of information, and summarize the important content.
[0467] "Formatting methods" refer to the processes and techniques used to arrange the generated summary into a more visually appealing and easily understandable format, and include conversion to charts, graphs, and bullet points.
[0468] "Communication methods" refer to technologies and protocols for transmitting formatted information to information devices in remote locations, and include network communication.
[0469] An "information device" is an electronic device used by users to receive, display, and edit information, and typically includes computers and smartphones.
[0470] "Display means" refers to a device or technology that visually presents information received on an information device to a user, and includes displays and monitors.
[0471] "Editing means" refers to technologies or software that enable users to modify or add to information they have received.
[0472] This system is designed to automatically transcribe audio information into text and summarize it. Specifically, it can process audio information acquired in meetings and other similar settings in real time or after recording.
[0473] The terminal is equipped with a device for acquiring audio, thereby collecting audio information during the meeting. The hardware uses a microphone and is operated in combination with software that can save the audio in digital format.
[0474] The server receives audio data transmitted from the terminal and converts this audio data into text data using speech recognition technology. In this step, it is possible to use common technologies as the speech recognition engine; for example, a common speech recognition API could be used.
[0475] The converted text data is processed on the server by a summarization generation mechanism, which extracts the key points and creates a summary. An example of a prompt when generating a summary using a generative AI model is, "Summarize the key points of the following text." A generative AI model can be used as the summarization algorithm.
[0476] The server then uses formatting tools to arrange the generated summary into a visually easy-to-understand format. For example, it improves readability by dividing it into bullet points or paragraphs. This formatted summary is then transmitted to the user's terminal via communication tools.
[0477] Users can review the summary received on their device and modify its content as needed using the editing tools. The final summary is saved by the user and can be shared with other members as needed.
[0478] Thus, this system provides a convenient and effective means of accurately summarizing, efficiently recording, and sharing meeting content.
[0479] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0480] Step 1:
[0481] The terminal uses an audio acquisition device to record the meeting audio in real time in order to obtain audio information. The input is the audio data spoken during the meeting, which is output as a digital audio file via the terminal's microphone. This audio data is saved in an audio file format (e.g., WAV, MP3).
[0482] Step 2:
[0483] The terminal transmits voice data to the server at regular time intervals or according to the amount of data. The input is the voice data stored in step 1, and the output is the voice packets transferred to the server over the network. This ensures that the voice data is delivered to the server in real time.
[0484] Step 3:
[0485] The server converts the received audio data into text data using speech recognition. The input here is digital audio data transmitted from the terminal, and the output converted by the speech recognition engine is text data. For example, a speech recognition API is used for this process.
[0486] Step 4:
[0487] The server summarizes the converted character data using a summarization generation mechanism. The input is the text data obtained in step 3, and the server uses a generation AI model to extract the key points and output them as a summarized text. The specific prompt used is, "Please summarize the key points of the following text."
[0488] Step 5:
[0489] The server formats the generated summary into a visually understandable format using formatting tools. The input is the summary text generated in step 4, and the output is the formatted visual summary. For example, it can be made easier to understand by using bullet points or dividing it into paragraphs.
[0490] Step 6:
[0491] The server sends the formatted summary to the user's terminal. The input is the formatted summary data from step 5, and the output is the summary data delivered to the user's terminal via the network. This allows the user to immediately review the summary.
[0492] Step 7:
[0493] The user reviews the summary received on their device and makes corrections as needed using the editing tools. The input is the summary data received in step 6, and the final meeting minutes text edited by the user is output. The revised meeting minutes may be saved and shared with other members.
[0494] (Application Example 1)
[0495] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0496] Conventional autonomous vehicles lacked sufficient systems to effectively grasp ambient sounds and emergency information in real time and support appropriate decision-making. Therefore, there is a need to quickly provide useful information obtained from ambient sounds to assist the driver.
[0497] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0498] In this invention, the server includes a voice collection means, an information conversion means, an information summarization means, an environmental analysis means, and an information presentation means. This makes it possible to quickly extract important information from ambient sounds and notify the driver in real time.
[0499] A "sound collection device" is a device that converts ambient sounds into electrical signals and records them as digital data.
[0500] "Information conversion means" refers to a process or device that converts collected audio data into textual information or other forms of symbolic information.
[0501] An "information summarization tool" is a technology or device that extracts important elements from converted symbolic information and reconstructs it into concise and to-the-point information.
[0502] A "structuring method" is a technique or system for organizing summarized information visually or in documents and making it presentable in a standardized format.
[0503] "Environmental analysis means" refers to a device or technology that has the function of detecting ambient sounds and selecting and analyzing important information from them.
[0504] "Information presentation means" refers to interfaces or systems for providing users with analyzed and structured information visually or audibly.
[0505] "Transmission means" refers to a communication channel or technology for transmitting formed data to a specific device.
[0506] A "correction tool" is a technology that provides an interface or functions for users to edit or correct the information presented to them.
[0507] The following system configurations are possible as embodiments for carrying out this invention.
[0508] The server acquires ambient sounds using voice acquisition means installed in the in-vehicle device. A high-sensitivity microphone is used in this process. The server then uses speech recognition software (e.g., Google Speech-to-Text API) as an information conversion means to convert the collected audio data into text-based symbolic information.
[0509] As a means of information summarization, a generative AI model is used to extract important content from the converted text information. This summarized information is then visually organized using structuring methods and formatted in a way that is easily understandable to the user.
[0510] Furthermore, as an environmental analysis tool, the server incorporates a function to analyze audio data around the vehicle and filter out emergency and traffic information. As a result of this analysis, important information is extracted and provided to the driver in real time via an information display device. For example, if the sound of an emergency vehicle siren is detected, a warning can be immediately issued to the driver. Information is presented using an on-screen display or a voice assistant.
[0511] As a concrete example, the system notifies the driver in real time that "an emergency vehicle is approaching" and prompts them to take appropriate action. An example of a prompt message to the generating AI model in this case could be an instruction such as, "Acquire ambient audio data while driving, summarize the emergency information, and inform me."
[0512] This system allows the server to provide users with a safe and comfortable operating environment.
[0513] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0514] Step 1:
[0515] The server acquires ambient sound using the audio acquisition means of the in-vehicle device. At this time, the audio data is input from the microphone as an analog signal. Because a high-sensitivity microphone is used, a wide range of audio data, including ambient sounds and conversations, can be accurately captured.
[0516] Step 2:
[0517] The server uses speech recognition software to convert the acquired analog audio signal into digital text format. This conversion process transforms the input audio signal into string data. As a result, the audio content is output as visualized symbolic information.
[0518] Step 3:
[0519] The server uses a generative AI model to summarize the converted text information. This method involves supplying converted text as input and extracting important information. In this summarization process, essential information is extracted from longer texts, and a short summary containing only the key points is output.
[0520] Step 4:
[0521] The server organizes the extracted summaries using structuring methods and converts them into a visually clear format. The goal is to format the input summary data into paragraphs, bullet points, etc., and output it in a format that is easy for the user to understand.
[0522] Step 5:
[0523] The server uses environmental analysis tools to select important information from surrounding audio data. Audio data to be analyzed is provided as input, and important sounds such as emergency signals are identified from it. The selected data is classified as requiring immediate attention.
[0524] Step 6:
[0525] The device notifies the driver of important information in real time through information display means. The information displayed is visualized or spoken digitally via a voice assistant or display. For example, information such as "Emergency vehicle approaching" is displayed immediately to support the driver's actions.
[0526] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0527] This invention provides a system that not only generates summarized meeting minutes using audio data acquired during a meeting, but also analyzes the emotions of the speakers and reflects them in the meeting minutes and emotion reports. The following methods can be considered to implement this invention.
[0528] The terminal records speech during the meeting using an audio acquisition device. This audio data is transmitted to the server in real time. The server converts the received audio into text data using a text conversion device. The converted text is processed by a summary generation device, which extracts important information from the meeting and generates a summary.
[0529] Here, the present invention performs a further advanced analysis using an emotion engine. The server recognizes the speaker's emotions from the text data and analyzes what emotions were expressed. Based on this emotion data, the summary generation means adjusts the tone of the meeting minutes to reflect the emotions of the participants.
[0530] The formatting means formats the summary, including the sentiment analysis results, and this is sent to the user's terminal via the transmission means. The user can review the received meeting minutes on their terminal and make corrections as needed. Furthermore, the sentiment report generation means analyzes the sentiment trends of the entire meeting and compiles this into an easy-to-understand report.
[0531] As a concrete example, a user can start recording a meeting and immediately receive a summarized transcript on their device after the meeting ends. This transcript reflects not only the main points made by the meeting participants but also an analysis of the emotions conveyed in their statements. This allows the user to gain insights into the atmosphere of the meeting and the reactions of the participants, which can be used to improve future meetings and project management.
[0532] This system allows for the capture of meeting dynamics that could not be fully understood through traditional meeting minute-taking methods, contributing to improved organizational communication and decision-making.
[0533] The following describes the processing flow.
[0534] Step 1:
[0535] The device launches an application that allows the user to start recording the meeting. It uses the microphone to capture the speech of meeting participants in real time.
[0536] Step 2:
[0537] The device sends the acquired audio data to the server at regular intervals. The audio data is processed in a streaming format.
[0538] Step 3:
[0539] The server converts the received audio data into text data using speech recognition technology and a text conversion method. By considering pronunciation and linguistic characteristics during the conversion process, it achieves highly accurate text conversion.
[0540] Step 4:
[0541] The server passes the converted text data to the sentiment engine, which analyzes keywords and context within the text to recognize the speaker's emotions. This process utilizes advanced natural language processing techniques to ensure accurate sentiment evaluation.
[0542] Step 5:
[0543] The server uses a summarization generation mechanism to extract key information from the text data of the meeting and create a concise summary. The tone and emphasis of the summary are adjusted by providing sentiment data received from the sentiment engine.
[0544] Step 6:
[0545] The server formats the meeting minutes, including the summary and sentiment analysis results, using a formatting method. These minutes are structured using headings and bullet points.
[0546] Step 7:
[0547] The server delivers the formatted meeting minutes to the user's terminal using a transmission method. The user can then immediately view the meeting minutes on their terminal.
[0548] Step 8:
[0549] Users can read the received meeting minutes through the terminal interface and make any necessary corrections or additional edits. Once the corrections are complete, they can save or share the meeting minutes.
[0550] Step 9:
[0551] The server uses an emotion report generation mechanism to generate a report summarizing the overall emotional trends of the meeting. This includes temporal fluctuations in emotions and an overall emotional assessment of all participants.
[0552] Step 10:
[0553] Users can review sentiment reports provided on their devices and receive feedback that reflects the atmosphere of the meeting and the reactions of the participants. This information can then be used to plan future meetings and adjust communication strategies.
[0554] (Example 2)
[0555] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0556] In modern meetings, creating meeting minutes is crucial for accurately reflecting the content of the meeting, but manually creating detailed minutes is time-consuming and laborious. Furthermore, because the emotions and tone of voice of participants during the meeting are not recorded, it is difficult to grasp the overall atmosphere of the meeting.
[0557] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0558] In this invention, the server includes means for acquiring audio information, means for converting the acquired audio information into text information, means for analyzing the converted text information and extracting important content, and means for associating the generated summary with emotional information using an emotional analysis engine. This makes it possible to automatically generate meeting minutes that summarize the meeting content while reflecting the emotions of the participants.
[0559] "Means for acquiring audio information" refers to devices and methods for recording speech in a meeting, which involve acquiring audio as digital data using microphones or recording devices.
[0560] "Means for converting audio information into text information" refers to technologies that analyze acquired audio data and convert it into corresponding strings, and are implemented using speech recognition engines or similar technologies.
[0561] "Methods for analyzing text information and extracting important content" refer to methods for selecting key points and important information from converted text, and are implemented using natural language processing technology.
[0562] "Methods for associating emotional information with an emotional analysis engine" refers to the process of analyzing the speaker's emotions from the text content and reflecting this in the summary, and this is done through an emotional analysis algorithm.
[0563] "Formatting methods" refer to techniques and methods for organizing generated summaries and sentiment information in an easily understandable way and arranging them into a defined format.
[0564] "Means of transmitting to an information terminal" refers to devices or protocols used to transmit processed data to an external terminal via a network, and utilizes communication technologies for data transfer.
[0565] "User-modifiable means" refers to methods and interfaces that provide an environment in which users can view and edit the received summary data and make modifications as needed.
[0566] This invention is a system for automatically generating meeting minutes, which includes everything from acquiring audio information and converting it to text, generating summaries, performing sentiment analysis, formatting the summaries, and providing them to the user. This system uses multiple means to extract value-added information from meeting records.
[0567] The terminal uses an input device such as a microphone to acquire audio and collects speech during the meeting in real time. This audio data is compressed and sent to the server via a communication method. The server uses speech recognition technology to convert the audio data into text data. A common speech recognition API can be used for this technology.
[0568] Next, the server uses a generative AI model to extract key information from the text data and generate a summary. Because natural language processing techniques are used in this process, the generated summary accurately captures the essence of the meeting. Furthermore, an emotion analysis engine is used to analyze the speaker's emotions from the text data and reflect this in the summary. In this step, an emotion analysis algorithm is applied, quantifying emotions based on factors such as the tone of speech.
[0569] The formatted summary is presented in a visually easy-to-understand format. This format utilizes data representation formats such as Markdown or HTML to establish a consistent summary format. Finally, the server sends the formatted summary and sentiment analysis information to the user's device. The user can review the received summary and make corrections as needed.
[0570] As a concrete example, a user starts recording a meeting using an application. After the meeting ends, the system automatically processes the audio data and sends a summarized meeting transcript along with the results of a sentiment analysis of the speakers to the user's device. The user can then use this information to consider ways to improve future meetings or projects.
[0571] An example of a prompt might be: "Summarize the key information from the meeting recording sequentially, including the speakers' emotions, and create a report. Provide it in a highly visual format."
[0572] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0573] Step 1:
[0574] The terminal uses an audio acquisition device to record all speech during the meeting in real time. The input is an analog audio signal acquired through a microphone, which is then converted and stored as digital audio data. The digital audio data is efficiently compressed and transmitted to the server according to the communication protocol.
[0575] Step 2:
[0576] The server processes the received digital audio data using a speech recognition engine and converts it into text data. The input is a compressed audio file, and speech recognition technology is used to analyze each phrase and word and convert it into corresponding text. The output is in text format and is temporarily stored for further processing.
[0577] Step 3:
[0578] The server uses a generative AI model to extract important content from the transformed text data and generate a summary. The input is the text data obtained in the previous step, and natural language processing algorithms are applied to select the key points and essential information of the meeting. The output is the summarized text, which is used in the next processing step.
[0579] Step 4:
[0580] The server analyzes text data using an emotion analysis engine to identify the speaker's emotional information. The input is text data after summary generation, and emotion analysis technology is used to extract and analyze emotional keywords and context. The output is quantified emotion information, which is added to the summary and reflected in the meeting minutes.
[0581] Step 5:
[0582] The server formats summaries and sentiment information in a visually organized manner. Input consists of processed summary text and sentiment data, which are then converted into a readable format. Output is delivered to the user as a Markdown or HTML document via a transmission method.
[0583] Step 6:
[0584] The user receives formatted meeting minutes on their terminal and reviews the information provided. Input consists of various formatted data sent from the server, and the user can view the content using the application and modify it as needed. Output is the modified meeting minutes, which can be used as a reference for future meetings or projects.
[0585] (Application Example 2)
[0586] Next, we will explain Application Example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0587] In recent years, improving communication among workers and enhancing work efficiency have become crucial issues in factories and workplaces. Furthermore, there is a need to understand workers' emotions and health conditions in real time to improve the safety and efficiency of the work environment. However, current systems struggle to automatically perform real-time emotion analysis and propose safety-enhancing measures based on that analysis.
[0588] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0589] In this invention, the server includes voice acquisition means, text conversion means, emotion analysis means, and suggestion generation means. This makes it possible to instantly analyze the emotional state of workers from voice data acquired within the factory and provide appropriate suggestions to improve safety and efficiency.
[0590] A "speech acquisition means" is a device that collects speech data from the environment and provides it in an analyzable format.
[0591] A "text conversion means" is a device that processes audio data and converts it into text data that represents that audio.
[0592] A "summary generation tool" is a device that extracts important information from text data and summarizes it in a concise form.
[0593] A "formatting tool" is a device that has the function of preparing summary data into a format that is easy to read and organize for output.
[0594] An "emotion analysis tool" is a device that identifies a worker's emotional state from voice or text data and analyzes the type and intensity of that emotion.
[0595] A "proposal generation method" is a system that has the function of proposing specific actions to improve work efficiency and safety based on analyzed emotional information.
[0596] "Transmission means" refers to a device or system that provides the function of transmitting formatted summaries or proposals to the user's terminal.
[0597] "Correction mechanisms" refer to features that allow users to modify or update submitted summaries and suggestions.
[0598] The system implementing this invention mainly consists of three elements: a server, a terminal, and a user. The server collects voice data in real time from factories and work sites through voice acquisition means. This voice data is converted into text data by text conversion means. On the server, the essence of important information is extracted from this text data by summary generation means and compiled into a concise summary.
[0599] Next, the server uses sentiment analysis means to analyze the worker's emotional state from the text data. Based on this analysis, the suggestion generation means generates specific suggestions to improve the safety and efficiency of the work. The generated summary and suggestions are organized by the formatting means and transmitted to the user's terminal via the transmission means.
[0600] On the terminal, users can not only review these summaries and suggestions, but also make corrections as needed using the correction tools. This facilitates smoother communication at the work site and optimizes the work environment.
[0601] For example, if text data such as "I'm tired today" or "It seems I'm slowing down my work pace" is analyzed on a terminal and negative emotions are detected, the server can immediately generate a suggestion to take a break and communicate it to the user. This enables a quick response and proper management of human resources.
[0602] An example of a prompt for a generative AI model might be an instruction such as, "Analyze the text data extracted from the conversation and generate suggestions to improve safety and efficiency in the workplace."
[0603] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0604] Step 1:
[0605] The server uses voice acquisition methods to collect audio data in real time from factories and work sites. The acquired audio data is stored in digital format. The input is an audio signal collected by a microphone, and the output is digital audio data. Specifically, it actively performs noise cancellation to improve the quality of the collected data.
[0606] Step 2:
[0607] The server uses a text conversion mechanism to convert the acquired audio data into text data. Here, ASR (Automatic Speech Recognition) technology is used to analyze the audio signal and output the results as text data. The input is previously processed audio data, and the output is text data that reflects the audio content. Specifically, it analyzes each audio sequence and generates its text representation.
[0608] Step 3:
[0609] The server uses a summarization generation mechanism to extract important information from the generated text data and create a summary. Using natural language processing techniques, it identifies key points in the text and outputs a concise summary. The input is text data, and the output is a summarized, concise text. Specifically, it uses morphological analysis to extract important nouns and verbs and generates content based on them.
[0610] Step 4:
[0611] The server utilizes sentiment analysis tools to identify the worker's emotional state from text data. Through a sentiment analysis engine, it classifies linguistic features within the text, resulting in sentiment data. The input is text data, and the output is sentiment data indicating the type and intensity of the emotion. Specifically, it classifies text into three categories: positive, negative, and neutral.
[0612] Step 5:
[0613] The server, using a proposal generation mechanism, formulates proposals to improve the safety and efficiency of the work environment based on sentiment data. In this process, sentiment data is analyzed, optimal countermeasures are devised, and output as text. The input is sentiment data, and the output is text proposing specific action plans. Specifically, it refers to past cases in the database and proposes solutions for similar cases.
[0614] Step 6:
[0615] The server uses formatting tools to convert summaries and proposals into a neat and organized format. Here, the output text is formatted and maintained for readability. Since the input is summaries and proposals, the output is text in a neat format. Specifically, the text is edited according to the predetermined formatting settings.
[0616] Step 7:
[0617] The server transmits formatted summaries and suggestions to the user's terminal via a transmission method. Digital communication technology is used to deliver these texts to the terminal. The input is formatted text, and the output is display data on the user's terminal. Specifically, it is converted into data packets and transmitted over the network.
[0618] Step 8:
[0619] The terminal displays the received summary and suggestions, allowing the user to review the content. Furthermore, the user can modify this information through editing tools and make optimized decisions as needed. The input is text data received from the server, and the output is the modified text. Specifically, the text is displayed on the screen, and editing is enabled on the interface.
[0620] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0621] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0622] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[0623] [Fourth Embodiment]
[0624] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0625] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0626] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0627] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0628] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0629] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0630] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0631] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0632] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0633] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0634] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0635] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0636] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0637] This invention provides a system that automatically generates summarized meeting minutes using audio data acquired during a meeting, and allows users to easily review them. The following methods are envisioned for carrying out this invention.
[0638] The server receives audio data transmitted from the terminals during the meeting. The terminals are equipped with audio acquisition capabilities and begin recording at the start of the meeting. The recorded audio is sent to the server at regular time intervals or in fixed data volumes.
[0639] The server converts the transmitted audio data into text format using a text conversion means. This converted text is then processed by a summary generation means within the server to generate a summary that extracts the important points of the content.
[0640] Subsequently, the server formats the summary obtained by the summary generation means into a predetermined format using the formatting means. For example, it visually organizes the content by dividing it into bullet points or paragraphs.
[0641] The formatted summary is sent to the user's device via a transmission method. The user reviews the transmitted minutes on their device, corrects any errors using the correction method if necessary, and saves or shares it as the final minutes.
[0642] As a concrete example, a user can start recording a meeting and immediately receive automatically generated meeting minutes on their device after the meeting ends. These minutes cover the key points of the meeting, allowing the user to quickly grasp the content and share it with their supervisor or team members as needed.
[0643] Thus, the system of the present invention eliminates the significant time and effort required for the conventional manual creation of meeting minutes, and allows for more effective capture of the key points of a meeting.
[0644] The following describes the processing flow.
[0645] Step 1:
[0646] The device prepares to capture participants' speech in real time using the microphone, after the user launches the application and starts recording the meeting.
[0647] Step 2:
[0648] The device sends recorded audio data to the server at regular time intervals. This data is treated as an audio stream and is designed to enable rapid processing.
[0649] Step 3:
[0650] The server converts the received audio data into text data using a text conversion device. This process utilizes speech recognition technology to accurately transcribe speech into text.
[0651] Step 4:
[0652] The server inputs the transcribed text into a summary generation system, which extracts key information from the meeting and generates a concise summary. This process utilizes generative AI, which analyzes the context and content of the statements made.
[0653] Step 5:
[0654] The server formats the generated summary into a standard meeting minutes format using formatting tools. This provides the necessary structure to make it easy for users to read.
[0655] Step 6:
[0656] The server sends formatted meeting minutes to the user's terminal using a transmission method. Transmission can occur in real time or after the meeting has ended.
[0657] Step 7:
[0658] Users review the meeting minutes sent to their own devices. If there are any errors in the minutes, they can edit them directly using the correction tools to finalize them.
[0659] Step 8:
[0660] Users can share revised meeting minutes with colleagues and supervisors as needed, or save them as an archive. This ensures that the meeting content is accurately preserved and can be referenced later.
[0661] (Example 1)
[0662] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0663] In modern meetings, much information is exchanged verbally, and it is essential to efficiently and accurately grasp and record the key points. However, manually summarizing meeting content is time-consuming and labor-intensive, so a system is needed to automate this process and improve efficiency.
[0664] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0665] In this invention, the server includes means for acquiring audio information, speech recognition means for converting the audio information into text information, and summary generation means for processing the acquired text information based on a summary generation algorithm and extracting key points. This makes it possible to automatically transcribe meeting audio data into text, extract key points from it, and easily create a highly readable summary.
[0666] "Means for acquiring audio information" refers to devices or systems for electronically recording speech, such as in meetings, and includes microphones and communication devices.
[0667] "Speech recognition means" refers to the technology or software that processes acquired speech information and converts it into corresponding text information, and generally uses a speech recognition engine.
[0668] "Summary generation means" refers to algorithms or systems that extract important points from textual information, condense the amount of information, and summarize the important content.
[0669] "Formatting methods" refer to the processes and techniques used to arrange the generated summary into a more visually appealing and easily understandable format, and include conversion to charts, graphs, and bullet points.
[0670] "Communication methods" refer to technologies and protocols for transmitting formatted information to information devices in remote locations, and include network communication.
[0671] An "information device" is an electronic device used by users to receive, display, and edit information, and typically includes computers and smartphones.
[0672] "Display means" refers to a device or technology that visually presents information received on an information device to a user, and includes displays and monitors.
[0673] "Editing means" refers to technologies or software that enable users to modify or add to information they have received.
[0674] This system is designed to automatically transcribe audio information into text and summarize it. Specifically, it can process audio information acquired in meetings and other similar settings in real time or after recording.
[0675] The terminal is equipped with a device for acquiring audio, thereby collecting audio information during the meeting. The hardware uses a microphone and is operated in combination with software that can save the audio in digital format.
[0676] The server receives audio data transmitted from the terminal and converts this audio data into text data using speech recognition technology. In this step, it is possible to use common technologies as the speech recognition engine; for example, a common speech recognition API could be used.
[0677] The converted text data is processed on the server by a summarization generation mechanism, which extracts the key points and creates a summary. An example of a prompt when generating a summary using a generative AI model is, "Summarize the key points of the following text." A generative AI model can be used as the summarization algorithm.
[0678] The server then uses formatting tools to arrange the generated summary into a visually easy-to-understand format. For example, it improves readability by dividing it into bullet points or paragraphs. This formatted summary is then transmitted to the user's terminal via communication tools.
[0679] Users can review the summary received on their device and modify its content as needed using the editing tools. The final summary is saved by the user and can be shared with other members as needed.
[0680] Thus, this system provides a convenient and effective means of accurately summarizing, efficiently recording, and sharing meeting content.
[0681] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0682] Step 1:
[0683] The terminal uses an audio acquisition device to record the meeting audio in real time in order to obtain audio information. The input is the audio data spoken during the meeting, which is output as a digital audio file via the terminal's microphone. This audio data is saved in an audio file format (e.g., WAV, MP3).
[0684] Step 2:
[0685] The terminal transmits voice data to the server at regular time intervals or according to the amount of data. The input is the voice data stored in step 1, and the output is the voice packets transferred to the server over the network. This ensures that the voice data is delivered to the server in real time.
[0686] Step 3:
[0687] The server converts the received audio data into text data using speech recognition. The input here is digital audio data transmitted from the terminal, and the output converted by the speech recognition engine is text data. For example, a speech recognition API is used for this process.
[0688] Step 4:
[0689] The server summarizes the converted character data using a summarization generation mechanism. The input is the text data obtained in step 3, and the server uses a generation AI model to extract the key points and output them as a summarized text. The specific prompt used is, "Please summarize the key points of the following text."
[0690] Step 5:
[0691] The server formats the generated summary into a visually understandable format using formatting tools. The input is the summary text generated in step 4, and the output is the formatted visual summary. For example, it can be made easier to understand by using bullet points or dividing it into paragraphs.
[0692] Step 6:
[0693] The server sends the formatted summary to the user's terminal. The input is the formatted summary data from step 5, and the output is the summary data delivered to the user's terminal via the network. This allows the user to immediately review the summary.
[0694] Step 7:
[0695] The user reviews the summary received on their device and makes corrections as needed using the editing tools. The input is the summary data received in step 6, and the final meeting minutes text edited by the user is output. The revised meeting minutes may be saved and shared with other members.
[0696] (Application Example 1)
[0697] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0698] Conventional autonomous vehicles lacked sufficient systems to effectively grasp ambient sounds and emergency information in real time and support appropriate decision-making. Therefore, there is a need to quickly provide useful information obtained from ambient sounds to assist the driver.
[0699] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0700] In this invention, the server includes a voice collection means, an information conversion means, an information summarization means, an environmental analysis means, and an information presentation means. This makes it possible to quickly extract important information from ambient sounds and notify the driver in real time.
[0701] A "sound collection device" is a device that converts ambient sounds into electrical signals and records them as digital data.
[0702] "Information conversion means" refers to a process or device that converts collected audio data into textual information or other forms of symbolic information.
[0703] An "information summarization tool" is a technology or device that extracts important elements from converted symbolic information and reconstructs it into concise and to-the-point information.
[0704] A "structuring method" is a technique or system for organizing summarized information visually or in documents and making it presentable in a standardized format.
[0705] "Environmental analysis means" refers to a device or technology that has the function of detecting ambient sounds and selecting and analyzing important information from them.
[0706] "Information presentation means" refers to interfaces or systems for providing users with analyzed and structured information visually or audibly.
[0707] "Transmission means" refers to a communication channel or technology for transmitting formed data to a specific device.
[0708] A "correction tool" is a technology that provides an interface or functions for users to edit or correct the information presented to them.
[0709] The following system configurations are possible as embodiments for carrying out this invention.
[0710] The server acquires ambient sounds using voice acquisition means installed in the in-vehicle device. A high-sensitivity microphone is used in this process. The server then uses speech recognition software (e.g., Google Speech-to-Text API) as an information conversion means to convert the collected audio data into text-based symbolic information.
[0711] As a means of information summarization, a generative AI model is used to extract important content from the converted text information. This summarized information is then visually organized using structuring methods and formatted in a way that is easily understandable to the user.
[0712] Furthermore, as an environmental analysis tool, the server incorporates a function to analyze audio data around the vehicle and filter out emergency and traffic information. As a result of this analysis, important information is extracted and provided to the driver in real time via an information display device. For example, if the sound of an emergency vehicle siren is detected, a warning can be immediately issued to the driver. Information is presented using an on-screen display or a voice assistant.
[0713] As a concrete example, the system notifies the driver in real time that "an emergency vehicle is approaching" and prompts them to take appropriate action. An example of a prompt message to the generating AI model in this case could be an instruction such as, "Acquire ambient audio data while driving, summarize the emergency information, and inform me."
[0714] This system allows the server to provide users with a safe and comfortable operating environment.
[0715] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0716] Step 1:
[0717] The server acquires ambient sound using the audio acquisition means of the in-vehicle device. At this time, the audio data is input from the microphone as an analog signal. Because a high-sensitivity microphone is used, a wide range of audio data, including ambient sounds and conversations, can be accurately captured.
[0718] Step 2:
[0719] The server uses speech recognition software to convert the acquired analog audio signal into digital text format. This conversion process transforms the input audio signal into string data. As a result, the audio content is output as visualized symbolic information.
[0720] Step 3:
[0721] The server uses a generative AI model to summarize the converted text information. This method involves supplying converted text as input and extracting important information. In this summarization process, essential information is extracted from longer texts, and a short summary containing only the key points is output.
[0722] Step 4:
[0723] The server organizes the extracted summaries using structuring methods and converts them into a visually clear format. The goal is to format the input summary data into paragraphs, bullet points, etc., and output it in a format that is easy for the user to understand.
[0724] Step 5:
[0725] The server uses environmental analysis tools to select important information from surrounding audio data. Audio data to be analyzed is provided as input, and important sounds such as emergency signals are identified from it. The selected data is classified as requiring immediate attention.
[0726] Step 6:
[0727] The device notifies the driver of important information in real time through information display means. The information displayed is visualized or spoken digitally via a voice assistant or display. For example, information such as "Emergency vehicle approaching" is displayed immediately to support the driver's actions.
[0728] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0729] This invention provides a system that not only generates summarized meeting minutes using audio data acquired during a meeting, but also analyzes the emotions of the speakers and reflects them in the meeting minutes and emotion reports. The following methods can be considered to implement this invention.
[0730] The terminal records speech during the meeting using an audio acquisition device. This audio data is transmitted to the server in real time. The server converts the received audio into text data using a text conversion device. The converted text is processed by a summary generation device, which extracts important information from the meeting and generates a summary.
[0731] Here, the present invention performs a further advanced analysis using an emotion engine. The server recognizes the speaker's emotions from the text data and analyzes what emotions were expressed. Based on this emotion data, the summary generation means adjusts the tone of the meeting minutes to reflect the emotions of the participants.
[0732] The formatting means formats the summary, including the sentiment analysis results, and this is sent to the user's terminal via the transmission means. The user can review the received meeting minutes on their terminal and make corrections as needed. Furthermore, the sentiment report generation means analyzes the sentiment trends of the entire meeting and compiles this into an easy-to-understand report.
[0733] As a concrete example, a user can start recording a meeting and immediately receive a summarized transcript on their device after the meeting ends. This transcript reflects not only the main points made by the meeting participants but also an analysis of the emotions conveyed in their statements. This allows the user to gain insights into the atmosphere of the meeting and the reactions of the participants, which can be used to improve future meetings and project management.
[0734] This system allows for the capture of meeting dynamics that could not be fully understood through traditional meeting minute-taking methods, contributing to improved organizational communication and decision-making.
[0735] The following describes the processing flow.
[0736] Step 1:
[0737] The device launches an application that allows the user to start recording the meeting. It uses the microphone to capture the speech of meeting participants in real time.
[0738] Step 2:
[0739] The device sends the acquired audio data to the server at regular intervals. The audio data is processed in a streaming format.
[0740] Step 3:
[0741] The server converts the received audio data into text data using speech recognition technology and a text conversion method. By considering pronunciation and linguistic characteristics during the conversion process, it achieves highly accurate text conversion.
[0742] Step 4:
[0743] The server passes the converted text data to the sentiment engine, which analyzes keywords and context within the text to recognize the speaker's emotions. This process utilizes advanced natural language processing techniques to ensure accurate sentiment evaluation.
[0744] Step 5:
[0745] The server uses a summarization generation mechanism to extract key information from the text data of the meeting and create a concise summary. The tone and emphasis of the summary are adjusted by providing sentiment data received from the sentiment engine.
[0746] Step 6:
[0747] The server formats the meeting minutes, including the summary and sentiment analysis results, using a formatting method. These minutes are structured using headings and bullet points.
[0748] Step 7:
[0749] The server delivers the formatted meeting minutes to the user's terminal using a transmission method. The user can then immediately view the meeting minutes on their terminal.
[0750] Step 8:
[0751] Users can read the received meeting minutes through the terminal interface and make any necessary corrections or additional edits. Once the corrections are complete, they can save or share the meeting minutes.
[0752] Step 9:
[0753] The server uses an emotion report generation mechanism to generate a report summarizing the overall emotional trends of the meeting. This includes temporal fluctuations in emotions and an overall emotional assessment of all participants.
[0754] Step 10:
[0755] Users can review sentiment reports provided on their devices and receive feedback that reflects the atmosphere of the meeting and the reactions of the participants. This information can then be used to plan future meetings and adjust communication strategies.
[0756] (Example 2)
[0757] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0758] In modern meetings, creating meeting minutes is crucial for accurately reflecting the content of the meeting, but manually creating detailed minutes is time-consuming and laborious. Furthermore, because the emotions and tone of voice of participants during the meeting are not recorded, it is difficult to grasp the overall atmosphere of the meeting.
[0759] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0760] In this invention, the server includes means for acquiring audio information, means for converting the acquired audio information into text information, means for analyzing the converted text information and extracting important content, and means for associating the generated summary with emotional information using an emotional analysis engine. This makes it possible to automatically generate meeting minutes that summarize the meeting content while reflecting the emotions of the participants.
[0761] "Means for acquiring audio information" refers to devices and methods for recording speech in a meeting, which involve acquiring audio as digital data using microphones or recording devices.
[0762] "Means for converting audio information into text information" refers to technologies that analyze acquired audio data and convert it into corresponding strings, and are implemented using speech recognition engines or similar technologies.
[0763] "Methods for analyzing text information and extracting important content" refer to methods for selecting key points and important information from converted text, and are implemented using natural language processing technology.
[0764] "Methods for associating emotional information with an emotional analysis engine" refers to the process of analyzing the speaker's emotions from the text content and reflecting this in the summary, and this is done through an emotional analysis algorithm.
[0765] "Formatting methods" refer to techniques and methods for organizing generated summaries and sentiment information in an easily understandable way and arranging them into a defined format.
[0766] "Means of transmitting to an information terminal" refers to devices or protocols used to transmit processed data to an external terminal via a network, and utilizes communication technologies for data transfer.
[0767] "User-modifiable means" refers to methods and interfaces that provide an environment in which users can view and edit the received summary data and make modifications as needed.
[0768] This invention is a system for automatically generating meeting minutes, which includes everything from acquiring audio information and converting it to text, generating summaries, performing sentiment analysis, formatting the summaries, and providing them to the user. This system uses multiple means to extract value-added information from meeting records.
[0769] The terminal uses an input device such as a microphone to acquire audio and collects speech during the meeting in real time. This audio data is compressed and sent to the server via a communication method. The server uses speech recognition technology to convert the audio data into text data. A common speech recognition API can be used for this technology.
[0770] Next, the server uses a generative AI model to extract key information from the text data and generate a summary. Because natural language processing techniques are used in this process, the generated summary accurately captures the essence of the meeting. Furthermore, an emotion analysis engine is used to analyze the speaker's emotions from the text data and reflect this in the summary. In this step, an emotion analysis algorithm is applied, quantifying emotions based on factors such as the tone of speech.
[0771] The formatted summary is presented in a visually easy-to-understand format. This format utilizes data representation formats such as Markdown or HTML to establish a consistent summary format. Finally, the server sends the formatted summary and sentiment analysis information to the user's device. The user can review the received summary and make corrections as needed.
[0772] As a concrete example, a user starts recording a meeting using an application. After the meeting ends, the system automatically processes the audio data and sends a summarized meeting transcript along with the results of a sentiment analysis of the speakers to the user's device. The user can then use this information to consider ways to improve future meetings or projects.
[0773] An example of a prompt might be: "Summarize the key information from the meeting recording sequentially, including the speakers' emotions, and create a report. Provide it in a highly visual format."
[0774] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0775] Step 1:
[0776] The terminal uses an audio acquisition device to record all speech during the meeting in real time. The input is an analog audio signal acquired through a microphone, which is then converted and stored as digital audio data. The digital audio data is efficiently compressed and transmitted to the server according to the communication protocol.
[0777] Step 2:
[0778] The server processes the received digital audio data using a speech recognition engine and converts it into text data. The input is a compressed audio file, and speech recognition technology is used to analyze each phrase and word and convert it into corresponding text. The output is in text format and is temporarily stored for further processing.
[0779] Step 3:
[0780] The server uses a generative AI model to extract important content from the transformed text data and generate a summary. The input is the text data obtained in the previous step, and natural language processing algorithms are applied to select the key points and essential information of the meeting. The output is the summarized text, which is used in the next processing step.
[0781] Step 4:
[0782] The server analyzes text data using an emotion analysis engine to identify the speaker's emotional information. The input is text data after summary generation, and emotion analysis technology is used to extract and analyze emotional keywords and context. The output is quantified emotion information, which is added to the summary and reflected in the meeting minutes.
[0783] Step 5:
[0784] The server formats summaries and sentiment information in a visually organized manner. Input consists of processed summary text and sentiment data, which are then converted into a readable format. Output is delivered to the user as a Markdown or HTML document via a transmission method.
[0785] Step 6:
[0786] The user receives formatted meeting minutes on their terminal and reviews the information provided. Input consists of various formatted data sent from the server, and the user can view the content using the application and modify it as needed. Output is the modified meeting minutes, which can be used as a reference for future meetings or projects.
[0787] (Application Example 2)
[0788] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0789] In recent years, improving communication among workers and enhancing work efficiency have become crucial issues in factories and workplaces. Furthermore, there is a need to understand workers' emotions and health conditions in real time to improve the safety and efficiency of the work environment. However, current systems struggle to automatically perform real-time emotion analysis and propose safety-enhancing measures based on that analysis.
[0790] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0791] In this invention, the server includes voice acquisition means, text conversion means, emotion analysis means, and suggestion generation means. This makes it possible to instantly analyze the emotional state of workers from voice data acquired within the factory and provide appropriate suggestions to improve safety and efficiency.
[0792] A "speech acquisition means" is a device that collects speech data from the environment and provides it in an analyzable format.
[0793] A "text conversion means" is a device that processes audio data and converts it into text data that represents that audio.
[0794] A "summary generation tool" is a device that extracts important information from text data and summarizes it in a concise form.
[0795] A "formatting tool" is a device that has the function of preparing summary data into a format that is easy to read and organize for output.
[0796] An "emotion analysis tool" is a device that identifies a worker's emotional state from voice or text data and analyzes the type and intensity of that emotion.
[0797] A "proposal generation method" is a system that has the function of proposing specific actions to improve work efficiency and safety based on analyzed emotional information.
[0798] "Transmission means" refers to a device or system that provides the function of transmitting formatted summaries or proposals to the user's terminal.
[0799] "Correction mechanisms" refer to features that allow users to modify or update submitted summaries and suggestions.
[0800] The system implementing this invention mainly consists of three elements: a server, a terminal, and a user. The server collects voice data in real time from factories and work sites through voice acquisition means. This voice data is converted into text data by text conversion means. On the server, the essence of important information is extracted from this text data by summary generation means and compiled into a concise summary.
[0801] Next, the server uses sentiment analysis means to analyze the worker's emotional state from the text data. Based on this analysis, the suggestion generation means generates specific suggestions to improve the safety and efficiency of the work. The generated summary and suggestions are organized by the formatting means and transmitted to the user's terminal via the transmission means.
[0802] On the terminal, users can not only review these summaries and suggestions, but also make corrections as needed using the correction tools. This facilitates smoother communication at the work site and optimizes the work environment.
[0803] For example, if text data such as "I'm tired today" or "It seems I'm slowing down my work pace" is analyzed on a terminal and negative emotions are detected, the server can immediately generate a suggestion to take a break and communicate it to the user. This enables a quick response and proper management of human resources.
[0804] An example of a prompt for a generative AI model might be an instruction such as, "Analyze the text data extracted from the conversation and generate suggestions to improve safety and efficiency in the workplace."
[0805] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0806] Step 1:
[0807] The server uses voice acquisition methods to collect audio data in real time from factories and work sites. The acquired audio data is stored in digital format. The input is an audio signal collected by a microphone, and the output is digital audio data. Specifically, it actively performs noise cancellation to improve the quality of the collected data.
[0808] Step 2:
[0809] The server uses a text conversion mechanism to convert the acquired audio data into text data. Here, ASR (Automatic Speech Recognition) technology is used to analyze the audio signal and output the results as text data. The input is previously processed audio data, and the output is text data that reflects the audio content. Specifically, it analyzes each audio sequence and generates its text representation.
[0810] Step 3:
[0811] The server uses a summarization generation mechanism to extract important information from the generated text data and create a summary. Using natural language processing techniques, it identifies key points in the text and outputs a concise summary. The input is text data, and the output is a summarized, concise text. Specifically, it uses morphological analysis to extract important nouns and verbs and generates content based on them.
[0812] Step 4:
[0813] The server utilizes sentiment analysis tools to identify the worker's emotional state from text data. Through a sentiment analysis engine, it classifies linguistic features within the text, resulting in sentiment data. The input is text data, and the output is sentiment data indicating the type and intensity of the emotion. Specifically, it classifies text into three categories: positive, negative, and neutral.
[0814] Step 5:
[0815] The server, using a proposal generation mechanism, formulates proposals to improve the safety and efficiency of the work environment based on sentiment data. In this process, sentiment data is analyzed, optimal countermeasures are devised, and output as text. The input is sentiment data, and the output is text proposing specific action plans. Specifically, it refers to past cases in the database and proposes solutions for similar cases.
[0816] Step 6:
[0817] The server uses formatting tools to convert summaries and proposals into a neat and organized format. Here, the output text is formatted and maintained for readability. Since the input is summaries and proposals, the output is text in a neat format. Specifically, the text is edited according to the predetermined formatting settings.
[0818] Step 7:
[0819] The server transmits formatted summaries and suggestions to the user's terminal via a transmission method. Digital communication technology is used to deliver these texts to the terminal. The input is formatted text, and the output is display data on the user's terminal. Specifically, it is converted into data packets and transmitted over the network.
[0820] Step 8:
[0821] The terminal displays the received summary and suggestions, allowing the user to review the content. Furthermore, the user can modify this information through editing tools and make optimized decisions as needed. The input is text data received from the server, and the output is the modified text. Specifically, the text is displayed on the screen, and editing is enabled on the interface.
[0822] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0823] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0824] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[0825] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0826] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0827] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0828] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0829] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0830] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0831] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0832] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[0833] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[0834] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0835] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0836] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0837] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0838] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0839] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0840] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0841] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0842] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[0843] The following is further disclosed regarding the embodiments described above.
[0844] (Claim 1)
[0845] A means of acquiring sound,
[0846] A character conversion means that converts the audio data acquired by the audio acquisition means into character data,
[0847] A summary generation means for summarizing the character data obtained by the character conversion means,
[0848] A formatting means for formatting the summary generated by the summary generation means,
[0849] A system that includes this.
[0850] (Claim 2)
[0851] The system according to claim 1, further comprising a transmission means for transmitting a summary formatted by the formatting means to a terminal.
[0852] (Claim 3)
[0853] The system according to claim 1, further comprising a terminal for displaying a summary transmitted by the transmission means and a modification means for allowing a user to modify the summary.
[0854] "Example 1"
[0855] (Claim 1)
[0856] Means for acquiring audio information,
[0857] A speech recognition means for converting the audio information into text information,
[0858] A summary generation means that processes the text information obtained by the speech recognition means based on a summary generation algorithm and extracts important points,
[0859] A formatting means for visually organizing the summary generated by the summary generation means,
[0860] An information processing system that includes this.
[0861] (Claim 2)
[0862] The information processing system according to claim 1, further comprising a communication means for transmitting the summary formatted by the formatting means to an information device.
[0863] (Claim 3)
[0864] The information processing system according to claim 1, further comprising a display means for displaying a summary transmitted by the communication means, and an editing means for allowing a user to modify the summary.
[0865] "Application Example 1"
[0866] (Claim 1)
[0867] Audio collection means and
[0868] An information conversion means that converts the audio information collected by the audio collection means into symbolic information,
[0869] Information summarization means for summarizing symbolic information obtained by the information conversion means,
[0870] A structuring means for structuring the summary generated by the information summarizing means,
[0871] An environmental analysis tool that collects ambient sound data and selects important information,
[0872] Information presentation means for notifying information obtained by the environmental analysis means,
[0873] A system that includes this.
[0874] (Claim 2)
[0875] The system according to claim 1, further comprising a transmission means for transmitting a summary structured by the structuring means to a device.
[0876] (Claim 3)
[0877] The system according to claim 1, further comprising a device for displaying a summary transmitted by the transmission means and a modification means for allowing a user to modify the summary.
[0878] "Example 2 of combining an emotion engine"
[0879] (Claim 1)
[0880] Means for acquiring audio information,
[0881] A means for converting acquired audio information into text information,
[0882] A means of analyzing the converted text information and extracting important content,
[0883] A means of generating a summary based on the extracted information,
[0884] A means of associating the generated summary with emotional information using an emotional analysis engine,
[0885] A means of formatting associated summaries,
[0886] A system that includes this.
[0887] (Claim 2)
[0888] The system according to claim 1, further comprising means for transmitting formatted summaries and sentiment information to an information terminal.
[0889] (Claim 3)
[0890] The system according to claim 1, further comprising means for displaying a summary transmitted on an information terminal and for a user to modify the summary.
[0891] "Application example 2 when combining with an emotional engine"
[0892] (Claim 1)
[0893] A means of acquiring sound,
[0894] A character conversion means that converts the audio data acquired by the audio acquisition means into character data,
[0895] A summary generation means for summarizing the character data obtained by the character conversion means,
[0896] A formatting means for formatting the summary generated by the summary generation means,
[0897] An emotion analysis means for analyzing emotional states from the character data,
[0898] A proposal generation means that generates proposals to improve safety and efficiency based on the analyzed emotional state,
[0899] A system that includes this.
[0900] (Claim 2)
[0901] The system according to claim 1, further comprising a transmission means for transmitting a formatted summary and a proposal generated by the proposal generation means to a terminal.
[0902] (Claim 3)
[0903] The system according to claim 1, further comprising a terminal for displaying summaries and suggestions transmitted by the transmission means, and a modification means for which a user can modify them. [Explanation of symbols]
[0904] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. A means of acquiring sound, A character conversion means that converts the audio data acquired by the audio acquisition means into character data, A summary generation means for summarizing the character data obtained by the character conversion means, A formatting means for formatting the summary generated by the summary generation means, A system that includes this.
2. The system according to claim 1, further comprising a transmission means for transmitting a summary formatted by the formatting means to a terminal.
3. The system according to claim 1, further comprising a terminal for displaying a summary transmitted by the transmission means and a modification means for allowing a user to modify the summary.