system
The system addresses the challenge of accurately capturing and sharing meeting content by converting audio to text, extracting key points, and generating visual summaries, thereby improving communication efficiency.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-09-30
- Publication Date
- 2026-04-09
AI Technical Summary
In meetings, it is difficult to accurately grasp and share the content and key points, leading to cognitive discrepancies and inefficient communication, with current technologies failing to effectively convert audio data into text, extract key points, and generate visual summaries.
A system that captures audio data, converts it into text, extracts key points using natural language processing, and generates visual summaries and diagrams to facilitate understanding.
The system reduces misunderstandings and enhances efficient information sharing by visually presenting important information and key points, allowing users to easily grasp meeting content and workflows.
Smart Images

Figure 2026062279000001_ABST
Abstract
Description
Technical Field
[0005] ,
[0001] The technology of this disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to the description of the chatbot's character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance that responds to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In meetings (MTGs) in enterprises and organizations, there is a problem that it is difficult to accurately grasp and share the content and key points of discussions. Specifically, cognitive discrepancies are likely to occur among participants, and what the speaker wants to convey is often not accurately transmitted and is often not visually imaginable. As a result, wasted time occurs and efficient communication is hindered. Also, it is difficult to accurately review the content of the discussion after the meeting, which may affect subsequent activities. To solve such problems, a system that organizes the content during the meeting into key points and makes it easy to visually imagine is required.
Means for Solving the Problems
[0005] <00000This invention provides a system that includes means for capturing audio data and converting it into text data, means for extracting key points from the text data, and means for displaying the extracted key points as a summary. Specifically, the system captures audio data spoken by the user on a terminal, sends this audio data to a server, and converts it into text data. Then, using natural language processing technology, it extracts key points from the text data and displays them to the user as a summary. Furthermore, when explaining specific use cases or business flows in a meeting, the system incorporates a function to analyze the text recognized from the audio data and automatically generate an image diagram to support visual understanding. In addition, at the end of the meeting, it generates a summary as a single visual by combining the important points of the discussion with related illustrations and provides it to the user. In this way, the system reduces misunderstandings in meetings and realizes efficient information sharing.
[0006] "Voice data" refers to data that records user speech and other voice inputs in digital format.
[0007] "Capturing" refers to the act of acquiring and recording audio data using an input device such as a microphone.
[0008] "Text data" refers to digital data obtained by converting audio data into written text.
[0009] "Converting" refers to the process of changing audio data into text data, often using a speech recognition engine.
[0010] "Extracting key points" is the act of analyzing text data and picking out important information and features.
[0011] "Displaying as a summary" means presenting the extracted key points to the user in a concise and visually organized manner.
[0012] A "use case" refers to an example of how a system or process is used or the application scenario it is used in.
[0013] A "business process flow" is a sequence of steps or processes that show how a particular business operation proceeds.
[0014] "Automatically generating an image diagram" refers to the act of automatically creating visual diagrams or flowcharts using a program based on analyzed text data.
[0015] "Graphic recording" is a method of visually recording the content of discussions and explanations, and combining key points with related illustrations to create a single visual representation. [Brief explanation of the drawing]
[0016] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] It is a sequence diagram showing the flow of processing of the data processing system in Embodiment 1. [Figure 12] It is a sequence diagram showing the flow of processing of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the flow of processing of the data processing system in Embodiment 2 when the emotion engine is combined. [Figure 14] It is a sequence diagram showing the flow of processing of the data processing system in Application Example 2 when the emotion engine is combined.
Mode for Carrying Out the Invention
[0017] Hereinafter, an example of an embodiment of the system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0018] First, the terms used in the following description will be explained. <00oooo102>>
[0019] In the following embodiments, a numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0020] In the following embodiments, a numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0021] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.
[0022] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0023] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0024] [First Embodiment]
[0025] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0026] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0027] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0028] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0029] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0030] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0031] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0032] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0033] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0034] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0035] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0036] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0037] This invention is a system designed to reduce misunderstandings in meetings and enable efficient information sharing. The system captures audio data, converts it into text data, and extracts key points, visually and concisely presenting important information and key points spoken by the user. The specific processing and functions are described below.
[0038] Specific processes and functions of the system
[0039] 1. Capture audio data
[0040] When a user speaks during a meeting, the device captures their voice using the microphone. The captured audio data is temporarily stored locally. This audio data is then converted into text data in the next step.
[0041] 2. Converting speech to text
[0042] Audio data stored on the device is sent to a server via the network. The server converts the audio into text data using a speech recognition engine (e.g., a standard speech recognition API). This process converts the audio data into text format, making subsequent processing easier.
[0043] 3. Key points extraction
[0044] The server inputs the converted text data into a natural language processing (NLP) model (e.g., a BERT model) for analysis. The NLP model extracts the main points from the text and identifies key information. The extracted points are concisely summarized, clearly indicating the main agenda and conclusions of the meeting.
[0045] 4. Displaying the main points
[0046] The server sends summarized key points back to the terminal. The terminal visually displays the summarized points to the user. This display is done in the form of a text box or pop-up window, allowing the user to quickly see the important points.
[0047] 5. Automatic generation of image diagrams
[0048] When a user verbally describes a use case or business workflow, the audio data is also captured and converted into text data. Based on the converted text data, the server analyzes the use case or business workflow recognized by the system and automatically generates a visual diagram. The generated diagram is displayed on the user's terminal in a format that is easy for the user to understand visually.
[0049] 6. Summary of meeting content at the end
[0050] At the end of a meeting, users verbally summarize its content. This audio data is also captured and converted into text. The server uses this text data to organize the key points and combines them with relevant illustrations to generate a single visual summary. This visual summary visually summarizes the meeting's discussion and can be displayed on the device as a PDF or image for easy reference.
[0051] Specific example
[0052] Examples of automatic summary generation
[0053] When a user says, "The features of this product are its low cost and high performance. It also has excellent energy-saving capabilities," the audio data is captured and converted into text data. The server extracts the key points, "low cost, high performance, and energy saving," and sends them back to the terminal for display.
[0054] Examples of automated image generation
[0055] When a user describes a user visiting a website, logging in, searching for products, and adding them to their cart, the audio data is converted to text data, and the server generates a use case diagram like the following:
[0056] [User] --> [Access Site] --> [Login] --> [Search Products] --> [Add to Cart]
[0057] This diagram will be displayed on the device and will be in a format that is easy for the user to understand visually.
[0058] Examples of content summaries
[0059] When a user summarizes the meeting at the end by saying, "Today we talked about the features of the new product and the marketing strategy for it," this audio data is converted into text data, and the server generates a visual summary like the following:
[0060] Features of the new product
[0061] Low cost
[0062] high performance
[0063] Energy saving
[0064] Marketing Strategy
[0065] Online advertising
[0066] social media
[0067] This visual summary will be displayed on the device as a graphic recording in PDF or image format.
[0068] As described above, this system automatically captures speech during meetings, converts it into text data, extracts key points, and displays them visually, allowing users to easily grasp important information. Furthermore, it generates visual diagrams of use cases and business workflows, providing a visualized version of the content at the end of the meeting, thereby reducing misunderstandings and enabling efficient information sharing.
[0069] The following describes the processing flow.
[0070] 1. Processing steps for automatic key point generation
[0071] Step 1:
[0072] User: "I will verbally state the important points."
[0073] Step 2:
[0074] Terminal: Captures the user's voice using the microphone and temporarily saves it as an audio file locally.
[0075] Step 3:
[0076] Terminal: Sends saved audio data to the server via the network.
[0077] Step 4:
[0078] Server: Inputs the received audio data into the speech recognition engine and converts the audio into text data.
[0079] Step 5:
[0080] Server: Inputs the converted text data into a natural language processing (NLP) model and performs analysis.
[0081] Step 6:
[0082] Server: Uses NLP models to extract key points and essential information from text data.
[0083] Step 7:
[0084] Server: Organizes the extracted key points into a summary and condenses them into concise text.
[0085] Step 8:
[0086] Server: Sends the generated summary back to the terminal.
[0087] Step 9:
[0088] Terminal: Displays summarized key information to the user. Specifically, it uses formats such as text boxes and pop-up windows.
[0089] 2. Processing steps for automatic image diagram generation
[0090] Step 1:
[0091] User: "I will verbally explain the use cases and business processes."
[0092] Step 2:
[0093] Terminal: Captures the user's voice using the microphone and temporarily saves it as an audio file locally.
[0094] Step 3:
[0095] Terminal: Sends saved audio data to the server via the network.
[0096] Step 4:
[0097] Server: Inputs the received audio data into the speech recognition engine and converts the audio into text data.
[0098] Step 5:
[0099] Server: Inputs the converted text data into a natural language processing (NLP) model and performs analysis.
[0100] Step 6:
[0101] Server: Uses an NLP model to extract information about use cases and business processes from text data.
[0102] Step 7:
[0103] Server: Based on the extracted information, it automatically generates an image diagram using the business flow design library.
[0104] Step 8:
[0105] Server: Sends the generated image back to the terminal as image data.
[0106] Step 9:
[0107] Terminal: Displays an image diagram to the user. Specifically, it uses a graphical viewer or image display window.
[0108] 3. Steps for summarizing the meeting content at the end of the meeting
[0109] Step 1:
[0110] User: "Summarize the meeting content verbally."
[0111] Step 2:
[0112] Terminal: Captures the user's voice using the microphone and temporarily saves it as an audio file locally.
[0113] Step 3:
[0114] Terminal: Sends saved audio data to the server via the network.
[0115] Step 4:
[0116] Server: Inputs the received audio data into the speech recognition engine and converts the audio into text data.
[0117] Step 5:
[0118] Server: Inputs the converted text data into a natural language processing (NLP) model and performs analysis.
[0119] Step 6:
[0120] Server: Uses NLP models to extract key points and essential information from text data.
[0121] Step 7:
[0122] Server: Organizes extracted points and generates a visual summary combining them with relevant illustrations.
[0123] Step 8:
[0124] Server: Sends the generated visual summary back to the terminal as a PDF or image file.
[0125] Step 9:
[0126] Terminal: Displays a visual summary to the user. Specifically, it uses a PDF viewer or an image display window.
[0127] The above outlines the specific processing steps of this system, which is a mechanism for efficiently and visually organizing and presenting comments and content from meetings to the user.
[0128] (Example 1)
[0129] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0130] In today's business environment, many meetings are conducted remotely, but misunderstandings often occur during voice communication. Therefore, there is a need for efficient ways to share meeting content and refer to it later. There is also a demand for easily visualizing and reviewing verbal explanations of use cases and business flows. However, current technologies often separate speech recognition and key point extraction, and automatic generation of diagrams is limited to certain special cases. This makes it difficult to accurately grasp the key points of a meeting.
[0131] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0132] In this invention, the server includes means for capturing audio data, means for converting the captured audio data into text data using conversion means, and means for extracting key points from the text data. This makes it possible to convert audio data into text, extract key points, and display them visually.
[0133] "Means for capturing audio data" refers to methods for acquiring the voice spoken by a user as digital data through an input device such as a microphone.
[0134] "Means for converting captured audio data into text data using conversion means" refers to means for converting captured audio data into text information using speech recognition technology.
[0135] "Methods for extracting key points from text data" refers to methods that utilize natural language processing techniques to select and extract important information from text data.
[0136] "Means of displaying as a summary" refers to a method of visualizing the extracted key points and presenting them to the user.
[0137] "Means for recognizing use cases and business flows" refers to methods for analyzing text data to understand and recognize specific usage scenarios and business procedures.
[0138] "Methods for generating image diagrams" refer to methods for automatically creating visually easy-to-understand diagrams based on recognized use cases and business flows.
[0139] A "means for generating visual summaries" is a method for creating a visually easy-to-understand summary by combining points extracted from text data with relevant illustrations.
[0140] This invention is a system for reducing misunderstandings in meetings and achieving efficient information sharing. This system visually and concisely presents important information and key points spoken by users through a series of processes: capturing audio data, converting it into text data, and extracting key points. Specific embodiments for carrying out this invention are described below.
[0141] Audio data capture
[0142] When a user speaks during a meeting, the device captures their voice using its microphone. The captured audio data is temporarily stored locally. Suitable devices for this purpose include computers and smartphones with built-in microphones.
[0143] Sending audio data
[0144] The terminal compresses the captured audio data and sends it to the server using a secure communication protocol (e.g., HTTPS). Compressing the audio data reduces the communication load.
[0145] Speech-to-text conversion
[0146] The server converts the received audio data into text data using a standard speech recognition engine (e.g., Google® Cloud Speech-to-Text API). This conversion makes the audio data into text format, which facilitates the next processing step.
[0147] Extracting key points
[0148] The server uses a generative AI model (e.g., the BERT model) to extract key points from the transformed text data. Natural language processing techniques are used to select important information from the text data.
[0149] Return and display of key points
[0150] The server converts the extracted key points into JSON format and sends them back to the terminal. The terminal visually displays the received key points in a text box or pop-up window on the screen. This allows the user to quickly confirm the important points.
[0151] Automatic generation of image diagrams
[0152] When a user verbally explains a use case or business workflow, the audio data is captured and sent to the server. The server converts the audio into text data and uses an AI model for illustration generation (e.g., Diagrams.net API) to automatically generate diagrams of the use case or business workflow. The generated diagrams are displayed on the terminal in a format that is easy for the user to understand visually.
[0153] Summary of the meeting's conclusion
[0154] At the end of a meeting, users verbally summarize its content. This audio data is also captured and converted into text. The server uses this text data to organize the key points and combines them with relevant illustrations to generate a visual summary. The generated visual summary is displayed on the device in PDF or image format for easy reference.
[0155] Specific example
[0156] Examples of automatic summary generation
[0157] When a user says, "The features of this product are its low cost and high performance. It also has excellent energy-saving capabilities," the audio data is captured and converted into text data. The server extracts the key points, "low cost, high performance, and energy saving," and sends them back to the terminal for display.
[0158] Examples of automated image generation
[0159] When a user describes a user visiting a website, logging in, searching for products, and adding them to their cart, the audio data is converted to text data, and the server generates a use case diagram like the following:
[0160] Users access the site, log in, search for products, and add them to their cart.
[0161] Example of a summary of meeting content at the end
[0162] When a user summarizes the meeting at the end by saying, "Today we talked about the features of the new product and the marketing strategy for it," this audio data is converted into text data, and the server generates a visual summary like the following:
[0163] Features of the new product
[0164] Low cost
[0165] high performance
[0166] Energy saving
[0167] Marketing Strategy
[0168] Online advertising
[0169] social media
[0170] The generated visual summary will be displayed on the device in PDF format.
[0171] As described above, this system includes a series of processes to reduce misunderstandings during meetings and enable efficient information sharing. High accuracy and efficiency are achieved by using specific hardware and software in each step. Furthermore, the specific operation of the system is explained through concrete examples.
[0172] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0173] Step 1: Capture audio data
[0174] When a user speaks during a meeting, the device captures their voice using its microphone. The input is the user's voice, and the output is digitized audio data. Specifically, the device's microphone detects the voice in real time and temporarily saves it as an audio file (e.g., in WAV format).
[0175] Step 2: Sending the audio data
[0176] The device compresses the captured audio data and sends it to the server using a secure communication protocol (e.g., HTTPS). The input is the captured audio data, and the output is the audio data sent to the server. Specifically, the device compresses the audio data (e.g., converts it to AAC format) and sends a POST request to the API endpoint.
[0177] Step 3: Convert speech to text
[0178] The server converts the received audio data into text data using a standard speech recognition engine (e.g., Google Cloud Speech-to-Text API). The input is the received audio data, and the output is text data. Specifically, the server sends an API request to the speech recognition service and retrieves the result in text format (e.g., JSON format).
[0179] Step 4: Extracting the key points
[0180] The server uses a generative AI model (e.g., a BERT model) to extract key points from the transformed text data. The input is text data, and the output is a list of extracted key points. Specifically, the server inputs text data into the model and organizes the output key points in list format.
[0181] Step 5: Return and display the key points.
[0182] The server converts the extracted key points into JSON data and sends it back to the terminal. The terminal displays the received key points on the screen in the form of text boxes or pop-up windows. The input is the extracted key points in JSON data, and the output is the visual information presented to the user. Specifically, the terminal parses the JSON data received from the server and updates the UI using HTML / CSS to display it on the screen.
[0183] Step 6: Automatic generation of image diagrams
[0184] When a user verbally describes a use case or business process, the audio data is captured and sent to the server. The server converts the audio into text data and uses an illustration generation AI model (e.g., Diagrams.net API) to automatically generate image diagrams of the use case or business process. The input is text data, and the output is the generated image diagram. Specifically, the server analyzes the text data and inputs it as prompt text into the generation AI model. For example, a prompt text might be: "The user accesses the site, logs in, searches for a product, and adds it to the cart."
[0185] Step 7: Summary of meeting content at the end
[0186] At the end of a meeting, users verbally summarize its content. This audio data is also captured and converted into text data. The server uses this text data to organize the key points and combines them with relevant illustrations to generate a visual summary. The input is the audio summary data from the end of the meeting, and the output is the visual summary. Specifically, the server combines the text data and illustrations, generates a visual summary using a PDF creation library (e.g., Apache® PDFBox), and sends it to the terminal. The PDF is then displayed on the terminal.
[0187] (Application Example 1)
[0188] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0189] In modern factories, frequent misunderstandings in information sharing during meetings and production flow setup are a major challenge. This is especially true in environments where verbal explanations are frequently used, where information is often misinterpreted or vaguely remembered. Furthermore, the difficulty in visually grasping meeting points and workflows hinders efficient information sharing. This can lead to delays and errors, potentially reducing productivity.
[0190] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0191] In this invention, the server includes means for capturing audio data, means for converting the captured audio data into text data, means for extracting key points from the text data, means for displaying the extracted key points as a summary, means for recognizing use cases and business flows, means for generating image diagrams based on the recognized use cases and business flows, means for displaying the generated summary and image diagrams on factory monitors or autonomous device displays, and means for extracting key points at the end of a meeting and generating a related visual summary. This makes it possible to improve the efficiency of information sharing by accurately transcribing audio information during a meeting into text and extracting key points. In addition, by visually displaying business flows and use cases, all participants can proceed with their work based on the same understanding, reducing information discrepancies. At the end of a meeting, the key points can be visualized and summarized to prevent information from being overlooked.
[0192] "Audio data" refers to information recorded in digital format.
[0193] "Means of capturing" refers to methods and devices for acquiring audio as digital data.
[0194] "Means of converting to text data" refers to methods and technologies for converting audio data into text-based data.
[0195] "Methods for extracting key points" refer to techniques for identifying and extracting particularly important parts from texts or conversations.
[0196] "Means of displaying as a summary" refers to a method of visually presenting extracted important information in a concise format.
[0197] A "use case" refers to a specific example of use or procedure under particular circumstances.
[0198] A "business process flow" is a diagram that shows the series of steps and procedures involved in performing a business task.
[0199] "Means of recognition" refer to the techniques and methods for identifying and understanding specific patterns or information.
[0200] "Means for generating image diagrams" refers to methods and techniques for creating visual shapes and charts based on text data and analysis results.
[0201] A "factory monitor" refers to a display device installed within a factory for displaying information.
[0202] An "autonomous device display" is a display device installed in a device that operates automatically.
[0203] "Meeting conclusions" refers to the main points or conclusions summarized at the end of a meeting.
[0204] A "visual summary" is a concise summary that visually presents key points and important information using diagrams and charts.
[0205] This invention is a system designed to streamline meetings and production flow setup in factories and reduce information sharing discrepancies. Through a series of processes—capturing audio data, converting it to text data, and extracting key points—this system visually and concisely presents important information and workflows spoken by users. The specific processing and functions of the system are described below.
[0206] 1. Capture audio data
[0207] When a user speaks during a meeting, the device captures their voice using its microphone. The captured audio data is temporarily stored locally. This audio data is then converted into text data in the next step. This process is achieved by using a device equipped with a microphone and recording capabilities.
[0208] 2. Converting speech to text
[0209] Audio data stored on the device is sent to a server via the network. The server converts the audio into text data using a speech recognition engine. Specific software used in this process includes the Google Speech-to-Text API. This converts the audio data into text format, making subsequent processing easier.
[0210] 3. Key points extraction
[0211] The server analyzes the converted text data using a natural language processing (NLP) model. The BERT model is a possible NLP model used here. The server extracts the main points from the text and identifies important information. This point extraction clearly identifies the main topics and conclusions discussed in the meeting.
[0212] 4. Displaying the main points
[0213] The server sends summarized key information back to the terminal. The terminal visually displays the summarized key points to the user. This display is shown in the form of text boxes or pop-up windows on monitors or autonomous device displays within the factory. This allows the user to quickly identify the important points.
[0214] 5. Automatic generation of image diagrams
[0215] When a user verbally describes a use case or business workflow, the audio data is also captured and converted into text data. Based on this converted text data, the server analyzes the recognized use case or business workflow and automatically generates a visual diagram. A generative AI model may be used in this process. The generated diagram is displayed on the user's device in a format that is easy for them to understand visually.
[0216] 6. Summary of meeting content at the end
[0217] At the end of a meeting, users verbally summarize its content. This audio data is also captured and converted into text. The server uses this text data to organize the key points and generates a single visual summary along with relevant images. This visual summary is displayed on the device as a graphic recording in PDF or image format, allowing users to easily refer to it.
[0218] Specific example
[0219] For example, if a factory worker says, "The procedure for setting up the new production line is to first transport the parts, then assemble them, and then conduct quality checks," the key points generated will be "parts transport, assembly, quality check." Based on this, the server will generate a flowchart and display an image showing "parts transport -> assembly -> quality check." Also, if the meeting concludes with a summary such as, "Today we talked about the procedure for setting up the new line and the points to note," a visual summary of those key points will be generated.
[0220] Example of a prompt
[0221] "The following audio data concerns setting up a new production line. Please extract the key points and generate the production line setup procedure. 'The procedure for setting up the new production line is to first transport the parts, then assemble them, and then conduct quality control.'"
[0222] By following the above steps, it is possible to utilize this method in meetings and production flow settings on the factory floor, significantly improving the efficiency of information sharing.
[0223] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0224] Step 1:
[0225] Audio data capture
[0226] The device uses a microphone to capture audio during the meeting. The input is the audio signal from the meeting, and the output is digital audio data. This audio data is temporarily stored in local storage. Specifically, the microphone built into the device picks up the audio signal, converts it to a digital format, and stores it.
[0227] Step 2:
[0228] Speech-to-text conversion
[0229] The device sends the captured audio data to the server over the network. The server converts the audio data into text data using the Google Speech-to-Text API. The input is digital audio data, and the output is text data. This is achieved by the server passing the audio data to the API and receiving the resulting text data.
[0230] Step 3:
[0231] Extracting key points
[0232] The server inputs the converted text data into a natural language processing (NLP) model (such as the BERT model) to analyze and extract key points. The input is text data, and the output is text data of the extracted key points. This is achieved by the server using the NLP model to analyze the data and extract important information.
[0233] Step 4:
[0234] Displaying the main points
[0235] The server returns summarized key information to the terminal. The terminal displays this information on factory monitors or autonomous device displays. The input is the extracted key points as text data, and the output is the key points displayed visually. The data is displayed visually using the terminal's display control software.
[0236] Step 5:
[0237] Recognition of use cases and business processes
[0238] When a user verbally describes a use case or business process, the audio data is captured and converted to text. The server analyzes the converted text data to recognize the use case or business process. The input consists of the audio data and the converted text data, and the output is the text data of the recognized use case or business process.
[0239] Step 6:
[0240] Automatic generation of image diagrams
[0241] The server generates visual diagrams based on recognized use cases and business workflows. The input is recognized text data, and the output is a visual diagram. A generation AI model is used to convert the data into flowcharts and charts for display.
[0242] Step 7:
[0243] Display of an image diagram
[0244] The server sends the generated image back to the terminal. The terminal displays this image on a factory monitor or an autonomous device's display. The input is the image, and the output is the visually displayed image.
[0245] Step 8:
[0246] Summary of the meeting's conclusion
[0247] At the end of a meeting, the user verbally summarizes its contents. This audio data is also captured and converted into text data. The server uses this text data to organize the key points and combines them with relevant images to generate a visual summary. The input is the user's summary audio and text data, and the output is the visual summary.
[0248] Step 9:
[0249] Display a visual summary
[0250] The server sends the generated visual summary back to the terminal. The terminal displays this in PDF or image format on factory monitors or autonomous machine displays. The input is the visual summary data, and the output is the visually displayed visual summary.
[0251] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0252] This invention is a system designed to reduce misunderstandings in meetings and enable efficient information sharing. Furthermore, it supports more accessible and understandable communication by recognizing user emotions and displaying or generating information based on those emotions. This system combines a series of processes—capturing audio data, converting it to text data, and extracting key points—with an emotion engine to visually and concisely present important information and key points expressed by the user. The specific processing and functions are described below.
[0253] Specific processes and functions of the system
[0254] 1. Capture audio data
[0255] When a user speaks during a meeting, the device captures their voice using the microphone. The captured audio data is temporarily stored locally. This audio data is then converted into text data in the next step.
[0256] 2. Converting speech to text
[0257] The audio data stored on the device is sent to the server via the network. The server converts the audio into text data using a speech recognition engine. This process simplifies subsequent processing because the audio data is converted into text format.
[0258] 3. Key points extraction
[0259] The server inputs the converted text data into a natural language processing (NLP) model for analysis. The NLP model extracts the key points from the text and identifies important information. The extracted points are concisely summarized, clearly indicating the main agenda and conclusions of the meeting.
[0260] 4. Emotion recognition
[0261] The server uses an emotion engine to recognize the user's emotions based on text and audio data analyzed by an NLP model. The emotion engine identifies the user's emotions based on factors such as voice tone and facial expression analysis, and evaluates that emotional information.
[0262] 5. Displaying the main points
[0263] The server sends summarized key points and recognized sentiment information back to the terminal. The terminal visually displays the summarized key points to the user in an appropriately adjusted format. This display is presented in the form of a text box or pop-up window, and is done in a way that is sensitive to the user's emotions.
[0264] 6. Automatic generation of image diagrams
[0265] When a user verbally describes a use case or business workflow, the audio data is also captured and converted into text data. Based on the converted text data, the server analyzes the use case or business workflow recognized by the system and automatically generates a visual diagram. The generated diagram is appropriately adjusted based on the user's feelings and displayed on the terminal in a visually easy-to-understand format.
[0266] 7. Summary of meeting content at the end
[0267] At the end of a meeting, users verbally summarize its content. This audio data is also captured and converted into text. The server uses this text data to organize the key points and combines them with relevant illustrations to generate a single visual summary. This visual summary is tailored to the user's emotions and provides a visual representation of the meeting's discussion. It can be displayed on the device as a PDF or image for easy reference.
[0268] Specific example
[0269] Examples of automatic summary generation
[0270] When a user says, "The features of this product are low cost and high performance. It also has excellent energy efficiency," the audio data is captured and converted into text data. The server extracts the key points "low cost, high performance, and energy efficiency," and an emotion engine recognizes the user's emotions. If the emotion engine determines that "the features of this product are very important," a summary is sent back to the terminal and displayed to the user.
[0271] Examples of automated image generation
[0272] When a user describes a user visiting a website, logging in, searching for products, and adding them to their cart, the audio data is converted to text data, and the server generates a use case diagram like the following:
[0273] [User] --> [Access Site] --> [Login] --> [Search Products] --> [Add to Cart]
[0274] If the emotion engine determines that the user is excited, the colors and fonts of the diagram are adjusted and displayed in a visually easy-to-understand format.
[0275] Examples of content summaries
[0276] When a user summarizes the meeting at the end by saying, "Today we talked about the features of the new product and the marketing strategy for it," this audio data is converted into text data, and the server generates a visual summary like the following:
[0277] Features of the new product
[0278] Low cost
[0279] high performance
[0280] Energy saving
[0281] Marketing Strategy
[0282] Online Advertising
[0283] Social Media
[0284] When it is recognized by the emotion engine that "the user is satisfied", the summary is displayed in colors and designs that reflect the user's sense of satisfaction.
[0285] As described above, this system can automatically capture the speech during a meeting, convert it into text data, extract key points, and display them to the user in combination with an emotion recognition function, enabling deeper insights and flexible communication. In addition, by generating visual images to illustrate use cases and business processes and providing a visualization of the content at the end of the meeting, it reduces recognition discrepancies and realizes efficient information sharing.
[0286] The processing flow will be described below.
[0287] 1. Processing Steps for Automatic Key Point Generation and Emotion Recognition
[0288] Step 1:
[0289] User: "Verbally state important points."
[0290] Step 2:
[0291] Terminal: Capture the user's voice with a microphone and temporarily save it locally as an audio file.
[0292] Step 3:
[0293] Terminal: Transmit the saved audio data to the server via the network.
[0294] Step 4:
[0295] Server: Input the received voice data into the speech recognition engine and convert the voice into text data
[0296] Step 5:
[0297] Server: Input the converted text data into the natural language processing (NLP) model for analysis
[0298] Step 6:
[0299] Server: Use the NLP model to extract important points and key points from the text data
[0300] Step 7:
[0301] Server: At the same time, input the voice data and text data into the emotion engine to recognize the user's emotion
[0302] Step 8:
[0303] Server: Combine the extracted key points and recognized emotion information to generate a summary
[0304] Step 9:
[0305] Server: Return the generated summary and emotion information to the terminal
[0306] Step 10:
[0307] Terminal: Based on the summarized key point information and emotion information, display it to the user in a visually adjusted format. For example, if the emotion is "gratitude", adjust the color and font of the text
[0308] 2. Processing steps for automatic image diagram generation and emotion recognition
[0309] Step 1:
[0310] User: "Verbally describe the use case and business process"
[0311] Step 2:
[0312] Terminal: Captures the user's voice using the microphone and temporarily saves it as an audio file locally.
[0313] Step 3:
[0314] Terminal: Sends saved audio data to the server via the network.
[0315] Step 4:
[0316] Server: Inputs the received audio data into the speech recognition engine and converts the audio into text data.
[0317] Step 5:
[0318] Server: Inputs the converted text data into a natural language processing (NLP) model and performs analysis.
[0319] Step 6:
[0320] Server: Uses an NLP model to extract information about use cases and business processes from text data.
[0321] Step 7:
[0322] Server: Simultaneously inputs voice and text data into the emotion engine to recognize the user's emotions.
[0323] Step 8:
[0324] Server: Based on extracted use cases and business flows, it automatically generates image diagrams using the business flow design library.
[0325] Step 9:
[0326] Server: Adjusts the generated image diagram appropriately based on recognized emotional information (e.g., color scheme and layout).
[0327] Step 10:
[0328] Server: Sends the adjusted image back to the terminal as image data.
[0329] Step 11:
[0330] Terminal: Displays an image to the user. For example, if the emotion is "excitement," the layout and graphical elements of the image will become more vivid.
[0331] 3. Summary of meeting content and processing steps for emotional recognition at the end of the meeting
[0332] Step 1:
[0333] User: "Summarize the meeting content verbally."
[0334] Step 2:
[0335] Terminal: Captures the user's voice using the microphone and temporarily saves it as an audio file locally.
[0336] Step 3:
[0337] Terminal: Sends saved audio data to the server via the network.
[0338] Step 4:
[0339] Server: Inputs the received audio data into the speech recognition engine and converts the audio into text data.
[0340] Step 5:
[0341] Server: Inputs the converted text data into a natural language processing (NLP) model and performs analysis.
[0342] Step 6:
[0343] Server: Uses NLP models to extract key points and essential information from text data.
[0344] Step 7:
[0345] Server: Simultaneously inputs voice and text data into the emotion engine to recognize the user's emotions.
[0346] Step 8:
[0347] Server: Organizes extracted points and combines them with relevant illustrations to generate a single visual summary.
[0348] Step 9:
[0349] Server: Appropriately adjusts the generated visual summary based on recognized sentiment information (e.g., color scheme and layout).
[0350] Step 10:
[0351] Server: Sends the adjusted visual summary back to the terminal as a PDF or image file.
[0352] Step 11:
[0353] Terminal: Displays a visual summary to the user. For example, if the emotion is "satisfied," the visual's color scheme will be warm tones.
[0354] The above outlines the specific processing steps of this system, which is a mechanism for efficiently and visually organizing and presenting meeting comments and content to the user. By using emotion recognition to display and generate information in accordance with the user's emotions, more effective communication becomes possible.
[0355] (Example 2)
[0356] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0357] Traditional meeting systems struggle to accurately grasp the key points of what is said and display them visually in an easy-to-understand manner. Furthermore, they often fail to consider user emotions when displaying information, leading to one-way communication. Additionally, they are inefficient at providing concise visual summaries of meeting discussions, resulting in low information sharing efficiency.
[0358] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0359] In this invention, the server includes means for capturing audio data, means for converting the captured audio data into text data, means for utilizing a natural language processing model to extract key points from the text data, means for using an emotion recognition engine to recognize the extracted key points and the user's emotions, means for displaying the key points based on the recognized emotion information, means for automatically generating an image diagram by converting the audio data into text data and performing analysis, means for adjusting and displaying the generated image diagram based on the user's emotions, and means for visualizing and displaying the summarized content at the end of the meeting. This enables accurate understanding of what the user said, information display that takes emotions into account, and reduces misunderstandings and efficient information sharing by visually summarizing the meeting discussion.
[0360] "Audio data" refers to the waveform information of sounds that record what a user says.
[0361] "Capturing" refers to the act of recording and acquiring information such as audio data.
[0362] "Text data" refers to audio data converted into written information.
[0363] A "natural language processing model" is a model that uses machine learning techniques to analyze, understand, and generate human language.
[0364] "Key points" refer to important information or content extracted from text data.
[0365] An "emotion recognition engine" is a device or software that analyzes voice data and text data to identify the user's emotions.
[0366] An "image diagram" refers to a visual explanatory diagram or flowchart generated based on text data.
[0367] A "summary" is a way of putting long texts or complex information into a concise format.
[0368] A "visual summary" is a visual representation of summarized information, and may include graphics, diagrams, and illustrations.
[0369] A "meeting" refers to a gathering of multiple users to discuss and share information.
[0370] This invention is a system designed to reduce misunderstandings in meetings and enable efficient information sharing. Furthermore, it supports more accessible and understandable communication by recognizing user emotions and displaying or generating information based on those emotions. This system combines a series of processes—capturing audio data, converting it into text data, and extracting key points—with an emotion recognition engine to visually and concisely present important information and key points expressed by the user.
[0371] Hardware and software to be used
[0372] Device: A device equipped with a microphone for the user to speak. This includes PCs, smartphones, tablets, etc.
[0373] Server: A backend system for converting and analyzing audio data. High processing power is desirable.
[0374] Speech recognition engine: Uses speech recognition services such as Google Cloud Speech-to-Text.
[0375] Natural language processing models: Advanced NLP models such as BERT.
[0376] Emotion recognition engine: Emotion analysis systems such as IBM Watson®.
[0377] System Description
[0378] 1. Capture audio data
[0379] When a user speaks during a meeting, the device captures their voice using the microphone. The captured audio data is temporarily stored in local storage and converted into text data in the next step.
[0380] 2. Converting speech to text
[0381] The device sends locally stored audio data to the server over the network. The server uses Google Cloud Speech-to-Text's speech recognition engine to convert the audio data into text data.
[0382] 3. Key points extraction
[0383] The server inputs text data into the BERT natural language processing (NLP) model for analysis. The NLP model extracts the main points of the text and identifies important information.
[0384] 4. Emotion recognition
[0385] The server inputs text and audio data, analyzed by an NLP model, into IBM Watson's emotion recognition engine to identify the user's emotions. The emotion recognition engine analyzes the tone and metadata of the voice to obtain information about the user's emotions.
[0386] 5. Displaying the main points
[0387] The server returns summarized key information and recognized sentiment information to the terminal. The terminal visually displays this information in an appropriately formatted way and presents it to the user. The display format is provided as a text box or a pop-up window.
[0388] 6. Automatic generation of image diagrams
[0389] When a user verbally describes a use case or workflow, the audio data is captured and converted into text. The server analyzes the converted text data and automatically generates a visual diagram. The generated diagram is then displayed on the device, with its colors and fonts adjusted based on the user's mood.
[0390] 7. Summary of meeting content at the end
[0391] At the end of the meeting, users verbally summarize the content. This audio data is also converted into text data, and the server extracts the key points and generates a visual summary along with relevant illustrations. The visual summary is displayed on the device in PDF or image format.
[0392] Specific example
[0393] 1. Specific examples of automatic summary generation
[0394] When a user says, "The features of this product are low cost and high performance. It also has excellent energy efficiency," the audio data is captured and converted into text data. The server extracts the key points "low cost, high performance, and energy efficiency," and an emotion recognition engine recognizes the user's emotions. If the emotion engine determines that "the features of this product are very important," a summary is sent back to the terminal and displayed to the user.
[0395] 2. Specific Examples of Automatic Image Generation
[0396] When a user describes a user visiting a website, logging in, searching for products, and adding them to their cart, the audio data is converted to text data, and the server generates a use case diagram like the following:
[0397] [User] --> [Access Site] --> [Login] --> [Search Products] --> [Add to Cart]
[0398] If the emotion engine determines that the user is excited, the colors and fonts of the diagram are adjusted and displayed in a visually easy-to-understand format.
[0399] 3. Specific examples of summaries at the end of a meeting
[0400] When a user summarizes the meeting at the end, saying, "Today we talked about the features of the new product and the marketing strategy for it," the audio data is converted to text data, and the server generates a visual summary like the following:
[0401] Features of the new product
[0402] Low cost
[0403] high performance
[0404] Energy saving
[0405] Marketing Strategy
[0406] Online advertising
[0407] social media
[0408] If the emotion engine recognizes that "the user is satisfied," the summary will be displayed with colors and a design that reflects the user's level of satisfaction.
[0409] The above describes the detailed embodiments for carrying out the invention. This invention automatically captures speech during meetings, converts it into text data, extracts key points, and combines this with sentiment recognition to provide deeper insights and more flexible communication. Furthermore, it generates visual diagrams of use cases and business flow explanations, and provides a visualized version of the content at the end of the meeting, thereby reducing misunderstandings and enabling efficient information sharing.
[0410] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0411] Step 1:
[0412] Audio data capture
[0413] When a user speaks during a meeting, the device captures their voice using the microphone. The input is the user's speech, and the output is the captured audio data. The device temporarily saves this audio data to local storage. As a concrete example, if a user says, "I'm going to talk about the design of the new product," that voice is captured and saved locally.
[0414] Step 2:
[0415] Speech-to-text conversion
[0416] The device sends the stored audio data to the server. The input is locally stored audio data, and the output is text data. The server uses the Google Cloud Speech-to-Text speech recognition engine to convert the audio data into text data. As a concrete example, the audio data "I'll talk about the design of the new product" is converted into the text data "I'll talk about the design of the new product".
[0417] Step 3:
[0418] Key points extraction
[0419] The server inputs the converted text data into the BERT natural language processing (NLP) model for analysis. The input is text data, and the output is key information. The BERT model extracts the important parts of the text and summarizes them. As a concrete example, from the text data "I will talk about the design of the new product," the key point "new product design" is extracted.
[0420] Step 4:
[0421] emotion recognition
[0422] The server inputs text and audio data, analyzed by the BERT model, into the emotion recognition engine. The input is the analyzed text and audio data, and the output is emotion information. The emotion recognition engine analyzes and identifies the user's emotions. As a concrete example of its operation, it identifies that the user is "excited" based on the text "I'm going to talk about the design of the new product" and the tone of voice.
[0423] Step 5:
[0424] Displaying the main points
[0425] The server sends back key information and recognized sentiment information to the terminal. The input is key information and sentiment information, and the output is the key information displayed visually. The terminal displays the key information to the user in an appropriate format. As a concrete example of operation, the key information "New product design" and the sentiment information "User is excited" are displayed in a pop-up window.
[0426] Step 6:
[0427] Automatic generation of image diagrams
[0428] When a user verbally describes a use case or workflow, the audio data is also captured and converted into text data. The input is the user's utterance, and the output is a visual diagram. The server generates a visual diagram based on the text data. As a concrete example, the following use case diagram is generated from the statement, "A user accesses the site, logs in, searches for a product, and adds it to their cart":
[0429] [User] --> [Access Site] --> [Login] --> [Search Products] --> [Add to Cart]
[0430] The colors and fonts of the images are adjusted based on emotions and displayed on the device.
[0431] Step 7:
[0432] Summary of the meeting's conclusion
[0433] At the end of the meeting, the user verbally summarizes the content. This audio data is also converted into text data, and the server extracts the key points and generates a visual summary along with relevant illustrations. The input is the user's summary, and the output is the visual summary. As a concrete example, the following visual summary is generated from the statement, "Today we talked about the features of the new product and the marketing strategy for it":
[0434] Features of the new product
[0435] Low cost
[0436] high performance
[0437] Energy saving
[0438] Marketing Strategy
[0439] Online advertising
[0440] social media
[0441] The visual summary will be displayed on your device in PDF format.
[0442] (Application Example 2)
[0443] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0444] Current autonomous vehicles face challenges in facilitating smooth communication between the driver and the vehicle system, as well as in understanding the driver's emotions in real time and providing appropriate information accordingly. Furthermore, there is a lack of mechanisms to support safe driving by appropriately managing the driver's emotional state. This can lead to delays in driver understanding and reaction, potentially increasing the risk of traffic accidents.
[0445] In Application Example 2, the specific processing performed by the specific processing unit 290 of the data processing device 12 is realized by the following means. In this invention, the server includes means for capturing audio data, means for converting the captured audio data into text data, means for extracting key points from the text data, means for recognizing emotions from the analyzed text data, and means for visually displaying information based on the driver's emotions. This enables smooth communication between the driver and the vehicle system, and by grasping the driver's emotional state in real time, it becomes possible to speed up the driver's understanding and reaction, and support safe driving.
[0446] "Means for capturing audio data" refers to a function that collects the driver's speech and surrounding sounds using microphones or other means, and saves them as digital data.
[0447] "Means for converting captured audio data into text data" refers to a function that uses speech recognition technology to convert collected audio data into text information.
[0448] "Methods for extracting key points from text data" refers to functions that use natural language processing technology to identify and summarize important information and topics from converted text.
[0449] "Means for displaying extracted key points as a summary" refers to a function for displaying extracted key points in a format that is easy for the driver to visually understand, on a display or other display device.
[0450] "Means for recognizing emotions from analyzed text data" refers to a function that automatically analyzes and recognizes the driver's emotional state (e.g., stress, satisfaction, excitement, etc.) based on text data and audio data.
[0451] "Means of visually displaying information based on the driver's emotions" refers to a function that adjusts the format and content of the displayed information according to the recognized emotional information, and provides appropriate visual feedback to the driver.
[0452] "Means for analyzing text data and recognizing use cases and business flows" refers to a function that uses natural language processing technology to extract specific procedures and processes from text data and identify them as use cases and business flows.
[0453] "Means for generating image diagrams based on recognized use cases and business flows" refers to a function that converts extracted process information into visual formats such as diagrams and flowcharts.
[0454] "Means for displaying the generated image diagram" refers to a function for displaying the generated visual diagram on a display or other display device.
[0455] "A means of generating a summary as a single visual by combining extracted points and related illustrations" refers to a function that extracts important points and combines them with related illustrations and diagrams to visually represent them as a single summary.
[0456] "Means for adjusting and displaying summaries based on perceived emotions" refers to a function for appropriately adjusting extracted key points and visual summaries based on the driver's emotional state and displaying them visually.
[0457] This invention relates to a system for autonomous vehicles that grasps the driver's emotional state in real time and enables smooth communication between the driver and the vehicle system. This system assists the driver by combining voice data capture, voice-to-text conversion, key point extraction, emotion recognition, and visual display of information.
[0458] Hardware and software configuration
[0459] hardware
[0460] Microphone: A device used to capture the driver's voice and surrounding sounds.
[0461] Display: A device that visually displays summary information and emotion-based feedback to the driver.
[0462] Vehicle ECU (Electronic Control Unit): A computing unit for real-time data processing.
[0463] software
[0464] Speech recognition engine (e.g., Google Speech Recognition API): Converts speech data into text data.
[0465] Natural language processing engines (e.g., the summarization model in the transformers library): Extract key points from text data.
[0466] Emotion recognition engine (e.g., the sentiment-analysis model in the transformers library): Recognizes emotions from the driver's statements.
[0467] Display control software: Software used to properly display information on a display.
[0468] Overall system processing
[0469] The server temporarily stores the audio data captured by the microphone locally. Then, it converts the audio data into text data using a speech recognition engine. This text data is analyzed by a natural language processing engine to extract key points. These extracted points, along with the driver's emotions, are analyzed by an emotion recognition engine, and feedback is provided to the driver based on this information. The feedback is displayed visually through the display.
[0470] Specific example
[0471] 1. Audio capture and text conversion:
[0472] When the driver says, "Turn right and then be careful at the intersection. A left turn follows," the voice is captured by the microphone. This voice data is then converted into text by a speech recognition engine, which reads, "Turn right and then be careful at the intersection. A left turn follows."
[0473] 2. Extracting key points and recognizing emotions:
[0474] Text data is input into a natural language processing engine (e.g., a summarization model), and key points such as "turn right, pay attention to the intersection, turn left" are extracted. Furthermore, an emotion recognition engine determines that "stress has been detected."
[0475] 3. Display of information:
[0476] Based on the extracted key points and emotion recognition results, information such as "Turn right, pay attention to intersection, turn left" is highlighted on the display. Furthermore, if stress is detected, messages and information encouraging relaxation are provided.
[0477] Example of a prompt:
[0478] "User comment: After turning right, be careful at the intersection. A left turn follows.
[0479] This allows drivers to receive necessary information and appropriate feedback in real time, enabling them to continue driving safely and effectively. This system contributes to improved traffic safety and reduced driver stress.
[0480] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0481] Step 1:
[0482] The server captures the driver's speech and surrounding sounds using a microphone. The audio data is input and temporarily stored locally in digital format.
[0483] Step 2:
[0484] The server sends the captured audio data to a speech recognition engine, where it is converted into text data. The input is audio data, and the output is text data. The speech recognition engine (e.g., Google Speech Recognition API) analyzes the audio waveform and generates the corresponding text.
[0485] Step 3:
[0486] The server inputs text data into a natural language processing engine and extracts key points. The input data is text data, and the output data is key point information. The natural language processing engine (e.g., the summarization model in the transformers library) analyzes the content of the text and extracts the important information by summarizing it.
[0487] Step 4:
[0488] The server inputs the extracted text data into an emotion recognition engine to recognize the driver's emotions. The input data is the extracted text data, and the output data is emotion information. The emotion recognition engine (e.g., the sentiment-analysis model from the transformers library) analyzes the emotions contained in the text and identifies emotional states such as stress, excitement, and satisfaction.
[0489] Step 5:
[0490] The server integrates extracted key information and recognized sentiment information and transmits it to the display device in a visual format. Input data consists of key information and sentiment information, while output data is visual feedback. Display control software determines the display format based on this data and provides appropriate information to the driver.
[0491] As a concrete example, if the driver says, "Turn right, then be careful at the intersection. Then turn left," in step 1, the voice data is captured, and in step 2, the speech recognition engine converts it into text data. In step 3, the key points "turn right, be careful at the intersection, turn left" are extracted, and in step 4, "stress" is recognized. Finally, in step 5, the important points are highlighted on the display, and if stress is detected, a message encouraging relaxation is added.
[0492] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0493] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0494] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0495] [Second Embodiment]
[0496] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0497] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0498] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0499] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0500] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0501] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0502] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0503] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0504] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0505] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0506] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0507] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0508] This invention is a system designed to reduce misunderstandings in meetings and enable efficient information sharing. The system captures audio data, converts it into text data, and extracts key points, visually and concisely presenting important information and key points spoken by the user. The specific processing and functions are described below.
[0509] Specific processes and functions of the system
[0510] 1. Capture audio data
[0511] When a user speaks during a meeting, the device captures their voice using the microphone. The captured audio data is temporarily stored locally. This audio data is then converted into text data in the next step.
[0512] 2. Converting speech to text
[0513] Audio data stored on the device is sent to a server via the network. The server converts the audio into text data using a speech recognition engine (e.g., a standard speech recognition API). This process converts the audio data into text format, making subsequent processing easier.
[0514] 3. Key points extraction
[0515] The server inputs the converted text data into a natural language processing (NLP) model (e.g., a BERT model) for analysis. The NLP model extracts the main points from the text and identifies key information. The extracted points are concisely summarized, clearly indicating the main agenda and conclusions of the meeting.
[0516] 4. Displaying the main points
[0517] The server sends summarized key points back to the terminal. The terminal visually displays the summarized points to the user. This display is done in the form of a text box or pop-up window, allowing the user to quickly see the important points.
[0518] 5. Automatic generation of image diagrams
[0519] When a user verbally describes a use case or business workflow, the audio data is also captured and converted into text data. Based on the converted text data, the server analyzes the use case or business workflow recognized by the system and automatically generates a visual diagram. The generated diagram is displayed on the user's terminal in a format that is easy for the user to understand visually.
[0520] 6. Summary of meeting content at the end
[0521] At the end of a meeting, users verbally summarize its content. This audio data is also captured and converted into text. The server uses this text data to organize the key points and combines them with relevant illustrations to generate a single visual summary. This visual summary visually summarizes the meeting's discussion and can be displayed on the device as a PDF or image for easy reference.
[0522] Specific example
[0523] Examples of automatic summary generation
[0524] When a user says, "The features of this product are its low cost and high performance. It also has excellent energy-saving capabilities," the audio data is captured and converted into text data. The server extracts the key points, "low cost, high performance, and energy saving," and sends them back to the terminal for display.
[0525] Examples of automated image generation
[0526] When a user describes a user visiting a website, logging in, searching for products, and adding them to their cart, the audio data is converted to text data, and the server generates a use case diagram like the following:
[0527] [User] --> [Access Site] --> [Login] --> [Search Products] --> [Add to Cart]
[0528] This diagram will be displayed on the device and will be in a format that is easy for the user to understand visually.
[0529] Examples of content summaries
[0530] When a user summarizes the meeting at the end by saying, "Today we talked about the features of the new product and the marketing strategy for it," this audio data is converted into text data, and the server generates a visual summary like the following:
[0531] Features of the new product
[0532] Low cost
[0533] high performance
[0534] Energy saving
[0535] Marketing Strategy
[0536] Online advertising
[0537] social media
[0538] This visual summary will be displayed on the device as a graphic recording in PDF or image format.
[0539] As described above, this system automatically captures speech during meetings, converts it into text data, extracts key points, and displays them visually, allowing users to easily grasp important information. Furthermore, it generates visual diagrams of use cases and business workflows, providing a visualized version of the content at the end of the meeting, thereby reducing misunderstandings and enabling efficient information sharing.
[0540] The following describes the processing flow.
[0541] 1. Processing steps for automatic key point generation
[0542] Step 1:
[0543] User: "I will verbally state the important points."
[0544] Step 2:
[0545] Terminal: Captures the user's voice using the microphone and temporarily saves it as an audio file locally.
[0546] Step 3:
[0547] Terminal: Sends saved audio data to the server via the network.
[0548] Step 4:
[0549] Server: Inputs the received audio data into the speech recognition engine and converts the audio into text data.
[0550] Step 5:
[0551] Server: Inputs the converted text data into a natural language processing (NLP) model and performs analysis.
[0552] Step 6:
[0553] Server: Uses NLP models to extract key points and essential information from text data.
[0554] Step 7:
[0555] Server: Organizes the extracted key points into a summary and condenses them into concise text.
[0556] Step 8:
[0557] Server: Sends the generated summary back to the terminal.
[0558] Step 9:
[0559] Terminal: Displays summarized key information to the user. Specifically, it uses formats such as text boxes and pop-up windows.
[0560] 2. Processing steps for automatic image diagram generation
[0561] Step 1:
[0562] User: "I will verbally explain the use cases and business processes."
[0563] Step 2:
[0564] Terminal: Captures the user's voice using the microphone and temporarily saves it as an audio file locally.
[0565] Step 3:
[0566] Terminal: Sends saved audio data to the server via the network.
[0567] Step 4:
[0568] Server: Inputs the received audio data into the speech recognition engine and converts the audio into text data.
[0569] Step 5:
[0570] Server: Inputs the converted text data into a natural language processing (NLP) model and performs analysis.
[0571] Step 6:
[0572] Server: Uses an NLP model to extract information about use cases and business processes from text data.
[0573] Step 7:
[0574] Server: Based on the extracted information, it automatically generates an image diagram using the business flow design library.
[0575] Step 8:
[0576] Server: Sends the generated image back to the terminal as image data.
[0577] Step 9:
[0578] Terminal: Displays an image diagram to the user. Specifically, it uses a graphical viewer or image display window.
[0579] 3. Steps for summarizing the meeting content at the end of the meeting
[0580] Step 1:
[0581] User: "Summarize the meeting content verbally."
[0582] Step 2:
[0583] Terminal: Captures the user's voice using the microphone and temporarily saves it as an audio file locally.
[0584] Step 3:
[0585] Terminal: Sends saved audio data to the server via the network.
[0586] Step 4:
[0587] Server: Inputs the received audio data into the speech recognition engine and converts the audio into text data.
[0588] Step 5:
[0589] Server: Inputs the converted text data into a natural language processing (NLP) model and performs analysis.
[0590] Step 6:
[0591] Server: Uses NLP models to extract key points and essential information from text data.
[0592] Step 7:
[0593] Server: Organizes extracted points and generates a visual summary combining them with relevant illustrations.
[0594] Step 8:
[0595] Server: Sends the generated visual summary back to the terminal as a PDF or image file.
[0596] Step 9:
[0597] Terminal: Displays a visual summary to the user. Specifically, it uses a PDF viewer or an image display window.
[0598] The above outlines the specific processing steps of this system, which is a mechanism for efficiently and visually organizing and presenting comments and content from meetings to the user.
[0599] (Example 1)
[0600] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0601] In today's business environment, many meetings are conducted remotely, but misunderstandings often occur during voice communication. Therefore, there is a need for efficient ways to share meeting content and refer to it later. There is also a demand for easily visualizing and reviewing verbal explanations of use cases and business flows. However, current technologies often separate speech recognition and key point extraction, and automatic generation of diagrams is limited to certain special cases. This makes it difficult to accurately grasp the key points of a meeting.
[0602] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0603] In this invention, the server includes means for capturing audio data, means for converting the captured audio data into text data using conversion means, and means for extracting key points from the text data. This makes it possible to convert audio data into text, extract key points, and display them visually.
[0604] "Means for capturing audio data" refers to methods for acquiring the voice spoken by a user as digital data through an input device such as a microphone.
[0605] "Means for converting captured audio data into text data using conversion means" refers to means for converting captured audio data into text information using speech recognition technology.
[0606] "Methods for extracting key points from text data" refers to methods that utilize natural language processing techniques to select and extract important information from text data.
[0607] "Means of displaying as a summary" refers to a method of visualizing the extracted key points and presenting them to the user.
[0608] "Means for recognizing use cases and business flows" refers to methods for analyzing text data to understand and recognize specific usage scenarios and business procedures.
[0609] "Methods for generating image diagrams" refer to methods for automatically creating visually easy-to-understand diagrams based on recognized use cases and business flows.
[0610] A "means for generating visual summaries" is a method for creating a visually easy-to-understand summary by combining points extracted from text data with relevant illustrations.
[0611] This invention is a system for reducing misunderstandings in meetings and achieving efficient information sharing. This system visually and concisely presents important information and key points spoken by users through a series of processes: capturing audio data, converting it into text data, and extracting key points. Specific embodiments for carrying out this invention are described below.
[0612] Audio data capture
[0613] When a user speaks during a meeting, the device captures their voice using its microphone. The captured audio data is temporarily stored locally. Suitable devices for this purpose include computers and smartphones with built-in microphones.
[0614] Sending audio data
[0615] The terminal compresses the captured audio data and sends it to the server using a secure communication protocol (e.g., HTTPS). Compressing the audio data reduces the communication load.
[0616] Speech-to-text conversion
[0617] The server converts the received audio data into text data using a standard speech recognition engine (e.g., Google Cloud Speech-to-Text API). This conversion makes the audio data into text format, which simplifies the next processing step.
[0618] Extracting key points
[0619] The server uses a generative AI model (e.g., the BERT model) to extract key points from the transformed text data. Natural language processing techniques are used to select important information from the text data.
[0620] Return and display of key points
[0621] The server converts the extracted key points into JSON format and sends them back to the terminal. The terminal visually displays the received key points in a text box or pop-up window on the screen. This allows the user to quickly confirm the important points.
[0622] Automatic generation of image diagrams
[0623] When a user verbally explains a use case or business workflow, the audio data is captured and sent to the server. The server converts the audio into text data and uses an AI model for illustration generation (e.g., Diagrams.net API) to automatically generate diagrams of the use case or business workflow. The generated diagrams are displayed on the terminal in a format that is easy for the user to understand visually.
[0624] Summary of the meeting's conclusion
[0625] At the end of a meeting, users verbally summarize its content. This audio data is also captured and converted into text. The server uses this text data to organize the key points and combines them with relevant illustrations to generate a visual summary. The generated visual summary is displayed on the device in PDF or image format for easy reference.
[0626] Specific example
[0627] Examples of automatic summary generation
[0628] When a user says, "The features of this product are its low cost and high performance. It also has excellent energy-saving capabilities," the audio data is captured and converted into text data. The server extracts the key points, "low cost, high performance, and energy saving," and sends them back to the terminal for display.
[0629] Examples of automated image generation
[0630] When a user describes a user visiting a website, logging in, searching for products, and adding them to their cart, the audio data is converted to text data, and the server generates a use case diagram like the following:
[0631] Users access the site, log in, search for products, and add them to their cart.
[0632] Example of a summary of meeting content at the end
[0633] When a user summarizes the meeting at the end by saying, "Today we talked about the features of the new product and the marketing strategy for it," this audio data is converted into text data, and the server generates a visual summary like the following:
[0634] Features of the new product
[0635] Low cost
[0636] high performance
[0637] Energy saving
[0638] Marketing Strategy
[0639] Online advertising
[0640] social media
[0641] The generated visual summary will be displayed on the device in PDF format.
[0642] As described above, this system includes a series of processes to reduce misunderstandings during meetings and enable efficient information sharing. High accuracy and efficiency are achieved by using specific hardware and software in each step. Furthermore, the specific operation of the system is explained through concrete examples.
[0643] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0644] Step 1: Capture audio data
[0645] When a user speaks during a meeting, the device captures their voice using its microphone. The input is the user's voice, and the output is digitized audio data. Specifically, the device's microphone detects the voice in real time and temporarily saves it as an audio file (e.g., in WAV format).
[0646] Step 2: Sending the audio data
[0647] The device compresses the captured audio data and sends it to the server using a secure communication protocol (e.g., HTTPS). The input is the captured audio data, and the output is the audio data sent to the server. Specifically, the device compresses the audio data (e.g., converts it to AAC format) and sends a POST request to the API endpoint.
[0648] Step 3: Convert speech to text
[0649] The server converts the received audio data into text data using a standard speech recognition engine (e.g., Google Cloud Speech-to-Text API). The input is the received audio data, and the output is text data. Specifically, the server sends an API request to the speech recognition service and retrieves the result in text format (e.g., JSON format).
[0650] Step 4: Extracting the key points
[0651] The server uses a generative AI model (e.g., a BERT model) to extract key points from the transformed text data. The input is text data, and the output is a list of extracted key points. Specifically, the server inputs text data into the model and organizes the output key points in list format.
[0652] Step 5: Return and display the key points.
[0653] The server converts the extracted key points into JSON data and sends it back to the terminal. The terminal displays the received key points on the screen in the form of text boxes or pop-up windows. The input is the extracted key points in JSON data, and the output is the visual information presented to the user. Specifically, the terminal parses the JSON data received from the server and updates the UI using HTML / CSS to display it on the screen.
[0654] Step 6: Automatic generation of image diagrams
[0655] When a user verbally describes a use case or business process, the audio data is captured and sent to the server. The server converts the audio into text data and uses an illustration generation AI model (e.g., Diagrams.net API) to automatically generate image diagrams of the use case or business process. The input is text data, and the output is the generated image diagram. Specifically, the server analyzes the text data and inputs it as prompt text into the generation AI model. For example, a prompt text might be: "The user accesses the site, logs in, searches for a product, and adds it to the cart."
[0656] Step 7: Summary of meeting content at the end
[0657] At the end of a meeting, users verbally summarize its content. This audio data is also captured and converted into text data. The server uses this text data to organize the key points and combines them with relevant illustrations to generate a visual summary. The input is the audio summary data from the end of the meeting, and the output is the visual summary. Specifically, the server combines the text data and illustrations, generates a visual summary using a PDF creation library (e.g., Apache PDFBox), and sends it to the terminal. The PDF is then displayed on the terminal.
[0658] (Application Example 1)
[0659] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0660] In modern factories, frequent misunderstandings in information sharing during meetings and production flow setup are a major challenge. This is especially true in environments where verbal explanations are frequently used, where information is often misinterpreted or vaguely remembered. Furthermore, the difficulty in visually grasping meeting points and workflows hinders efficient information sharing. This can lead to delays and errors, potentially reducing productivity.
[0661] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0662] In this invention, the server includes means for capturing audio data, means for converting the captured audio data into text data, means for extracting key points from the text data, means for displaying the extracted key points as a summary, means for recognizing use cases and business flows, means for generating image diagrams based on the recognized use cases and business flows, means for displaying the generated summary and image diagrams on factory monitors or autonomous device displays, and means for extracting key points at the end of a meeting and generating a related visual summary. This makes it possible to improve the efficiency of information sharing by accurately transcribing audio information during a meeting into text and extracting key points. In addition, by visually displaying business flows and use cases, all participants can proceed with their work based on the same understanding, reducing information discrepancies. At the end of a meeting, the key points can be visualized and summarized to prevent information from being overlooked.
[0663] "Audio data" refers to information recorded in digital format.
[0664] "Means of capturing" refers to methods and devices for acquiring audio as digital data.
[0665] "Means of converting to text data" refers to methods and technologies for converting audio data into text-based data.
[0666] "Methods for extracting key points" refer to techniques for identifying and extracting particularly important parts from texts or conversations.
[0667] "Means of displaying as a summary" refers to a method of visually presenting extracted important information in a concise format.
[0668] A "use case" refers to a specific example of use or procedure under particular circumstances.
[0669] A "business process flow" is a diagram that shows the series of steps and procedures involved in performing a business task.
[0670] "Means of recognition" refer to the techniques and methods for identifying and understanding specific patterns or information.
[0671] "Means for generating image diagrams" refers to methods and techniques for creating visual shapes and charts based on text data and analysis results.
[0672] A "factory monitor" refers to a display device installed within a factory for displaying information.
[0673] An "autonomous device display" is a display device installed in a device that operates automatically.
[0674] "Meeting conclusions" refers to the main points or conclusions summarized at the end of a meeting.
[0675] A "visual summary" is a concise summary that visually presents key points and important information using diagrams and charts.
[0676] This invention is a system designed to streamline meetings and production flow setup in factories and reduce information sharing discrepancies. Through a series of processes—capturing audio data, converting it to text data, and extracting key points—this system visually and concisely presents important information and workflows spoken by users. The specific processing and functions of the system are described below.
[0677] 1. Capture audio data
[0678] When a user speaks during a meeting, the device captures their voice using its microphone. The captured audio data is temporarily stored locally. This audio data is then converted into text data in the next step. This process is achieved by using a device equipped with a microphone and recording capabilities.
[0679] 2. Converting speech to text
[0680] Audio data stored on the device is sent to a server via the network. The server converts the audio into text data using a speech recognition engine. Specific software used in this process includes the Google Speech-to-Text API. This converts the audio data into text format, making subsequent processing easier.
[0681] 3. Key points extraction
[0682] The server analyzes the converted text data using a natural language processing (NLP) model. The BERT model is a possible NLP model used here. The server extracts the main points from the text and identifies important information. This point extraction clearly identifies the main topics and conclusions discussed in the meeting.
[0683] 4. Displaying the main points
[0684] The server sends summarized key information back to the terminal. The terminal visually displays the summarized key points to the user. This display is shown in the form of text boxes or pop-up windows on monitors or autonomous device displays within the factory. This allows the user to quickly identify the important points.
[0685] 5. Automatic generation of image diagrams
[0686] When a user verbally describes a use case or business workflow, the audio data is also captured and converted into text data. Based on this converted text data, the server analyzes the recognized use case or business workflow and automatically generates a visual diagram. A generative AI model may be used in this process. The generated diagram is displayed on the user's device in a format that is easy for them to understand visually.
[0687] 6. Summary of meeting content at the end
[0688] At the end of a meeting, users verbally summarize its content. This audio data is also captured and converted into text. The server uses this text data to organize the key points and generates a single visual summary along with relevant images. This visual summary is displayed on the device as a graphic recording in PDF or image format, allowing users to easily refer to it.
[0689] Specific example
[0690] For example, if a factory worker says, "The procedure for setting up the new production line is to first transport the parts, then assemble them, and then conduct quality checks," the key points generated will be "parts transport, assembly, quality check." Based on this, the server will generate a flowchart and display an image showing "parts transport -> assembly -> quality check." Also, if the meeting concludes with a summary such as, "Today we talked about the procedure for setting up the new line and the points to note," a visual summary of those key points will be generated.
[0691] Example of a prompt
[0692] "The following audio data concerns setting up a new production line. Please extract the key points and generate the production line setup procedure. 'The procedure for setting up the new production line is to first transport the parts, then assemble them, and then conduct quality control.'"
[0693] By following the above steps, it is possible to utilize this method in meetings and production flow settings on the factory floor, significantly improving the efficiency of information sharing.
[0694] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0695] Step 1:
[0696] Audio data capture
[0697] The device uses a microphone to capture audio during the meeting. The input is the audio signal from the meeting, and the output is digital audio data. This audio data is temporarily stored in local storage. Specifically, the microphone built into the device picks up the audio signal, converts it to a digital format, and stores it.
[0698] Step 2:
[0699] Speech-to-text conversion
[0700] The device sends the captured audio data to the server over the network. The server converts the audio data into text data using the Google Speech-to-Text API. The input is digital audio data, and the output is text data. This is achieved by the server passing the audio data to the API and receiving the resulting text data.
[0701] Step 3:
[0702] Extracting key points
[0703] The server inputs the converted text data into a natural language processing (NLP) model (such as the BERT model) to analyze and extract key points. The input is text data, and the output is text data of the extracted key points. This is achieved by the server using the NLP model to analyze the data and extract important information.
[0704] Step 4:
[0705] Displaying the main points
[0706] The server returns summarized key information to the terminal. The terminal displays this information on factory monitors or autonomous device displays. The input is the extracted key points as text data, and the output is the key points displayed visually. The data is displayed visually using the terminal's display control software.
[0707] Step 5:
[0708] Recognition of use cases and business processes
[0709] When a user verbally describes a use case or business process, the audio data is captured and converted to text. The server analyzes the converted text data to recognize the use case or business process. The input consists of the audio data and the converted text data, and the output is the text data of the recognized use case or business process.
[0710] Step 6:
[0711] Automatic generation of image diagrams
[0712] The server generates visual diagrams based on recognized use cases and business workflows. The input is recognized text data, and the output is a visual diagram. A generation AI model is used to convert the data into flowcharts and charts for display.
[0713] Step 7:
[0714] Display of an image diagram
[0715] The server sends the generated image back to the terminal. The terminal displays this image on a factory monitor or an autonomous device's display. The input is the image, and the output is the visually displayed image.
[0716] Step 8:
[0717] Summary of the meeting's conclusion
[0718] At the end of a meeting, the user verbally summarizes its contents. This audio data is also captured and converted into text data. The server uses this text data to organize the key points and combines them with relevant images to generate a visual summary. The input is the user's summary audio and text data, and the output is the visual summary.
[0719] Step 9:
[0720] Display a visual summary
[0721] The server sends the generated visual summary back to the terminal. The terminal displays this in PDF or image format on factory monitors or autonomous machine displays. The input is the visual summary data, and the output is the visually displayed visual summary.
[0722] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0723] This invention is a system designed to reduce misunderstandings in meetings and enable efficient information sharing. Furthermore, it supports more accessible and understandable communication by recognizing user emotions and displaying or generating information based on those emotions. This system combines a series of processes—capturing audio data, converting it to text data, and extracting key points—with an emotion engine to visually and concisely present important information and key points expressed by the user. The specific processing and functions are described below.
[0724] Specific processes and functions of the system
[0725] 1. Capture audio data
[0726] When a user speaks during a meeting, the device captures their voice using the microphone. The captured audio data is temporarily stored locally. This audio data is then converted into text data in the next step.
[0727] 2. Converting speech to text
[0728] The audio data stored on the device is sent to the server via the network. The server converts the audio into text data using a speech recognition engine. This process simplifies subsequent processing because the audio data is converted into text format.
[0729] 3. Key points extraction
[0730] The server inputs the converted text data into a natural language processing (NLP) model for analysis. The NLP model extracts the key points from the text and identifies important information. The extracted points are concisely summarized, clearly indicating the main agenda and conclusions of the meeting.
[0731] 4. Emotion recognition
[0732] The server uses an emotion engine to recognize the user's emotions based on text and audio data analyzed by an NLP model. The emotion engine identifies the user's emotions based on factors such as voice tone and facial expression analysis, and evaluates that emotional information.
[0733] 5. Displaying the main points
[0734] The server sends summarized key points and recognized sentiment information back to the terminal. The terminal visually displays the summarized key points to the user in an appropriately adjusted format. This display is presented in the form of a text box or pop-up window, and is done in a way that is sensitive to the user's emotions.
[0735] 6. Automatic generation of image diagrams
[0736] When a user verbally describes a use case or business workflow, the audio data is also captured and converted into text data. Based on the converted text data, the server analyzes the use case or business workflow recognized by the system and automatically generates a visual diagram. The generated diagram is appropriately adjusted based on the user's feelings and displayed on the terminal in a visually easy-to-understand format.
[0737] 7. Summary of meeting content at the end
[0738] At the end of a meeting, users verbally summarize its content. This audio data is also captured and converted into text. The server uses this text data to organize the key points and combines them with relevant illustrations to generate a single visual summary. This visual summary is tailored to the user's emotions and provides a visual representation of the meeting's discussion. It can be displayed on the device as a PDF or image for easy reference.
[0739] Specific example
[0740] Examples of automatic summary generation
[0741] When a user says, "The features of this product are low cost and high performance. It also has excellent energy efficiency," the audio data is captured and converted into text data. The server extracts the key points "low cost, high performance, and energy efficiency," and an emotion engine recognizes the user's emotions. If the emotion engine determines that "the features of this product are very important," a summary is sent back to the terminal and displayed to the user.
[0742] Examples of automated image generation
[0743] When a user describes a user visiting a website, logging in, searching for products, and adding them to their cart, the audio data is converted to text data, and the server generates a use case diagram like the following:
[0744] [User] --> [Access Site] --> [Login] --> [Search Products] --> [Add to Cart]
[0745] If the emotion engine determines that the user is excited, the colors and fonts of the diagram are adjusted and displayed in a visually easy-to-understand format.
[0746] Examples of content summaries
[0747] When a user summarizes the meeting at the end by saying, "Today we talked about the features of the new product and the marketing strategy for it," this audio data is converted into text data, and the server generates a visual summary like the following:
[0748] Features of the new product
[0749] Low cost
[0750] high performance
[0751] Energy saving
[0752] Marketing Strategy
[0753] Online advertising
[0754] social media
[0755] If the emotion engine recognizes that "the user is satisfied," the summary will be displayed with colors and a design that reflects the user's level of satisfaction.
[0756] As described above, this system automatically captures speech during meetings, converts it into text data, extracts key points, and displays them to the user in combination with sentiment recognition, enabling deeper insights and more flexible communication. Furthermore, it generates visual diagrams of use cases and business flow explanations, and provides a visualized version of the content at the end of the meeting, thereby reducing misunderstandings and facilitating efficient information sharing.
[0757] The following describes the processing flow.
[0758] 1. Processing steps for automatic summary generation and sentiment recognition
[0759] Step 1:
[0760] User: "I will verbally state the important points."
[0761] Step 2:
[0762] Terminal: Captures the user's voice using the microphone and temporarily saves it as an audio file locally.
[0763] Step 3:
[0764] Terminal: Sends saved audio data to the server via the network.
[0765] Step 4:
[0766] Server: Inputs the received audio data into the speech recognition engine and converts the audio into text data.
[0767] Step 5:
[0768] Server: Inputs the converted text data into a natural language processing (NLP) model and performs analysis.
[0769] Step 6:
[0770] Server: Uses NLP models to extract key points and essential information from text data.
[0771] Step 7:
[0772] Server: Simultaneously inputs voice and text data into the emotion engine to recognize the user's emotions.
[0773] Step 8:
[0774] Server: Combines extracted key points and recognized sentiment information to generate a summary.
[0775] Step 9:
[0776] Server: Sends the generated summary and sentiment information back to the terminal.
[0777] Step 10:
[0778] Terminal: Displays summarized key information and emotional information to the user in a visually adjusted format. For example, if the emotion is "emotional," the text color and font are adjusted.
[0779] 2. Image diagram automatic generation and emotion recognition processing steps
[0780] Step 1:
[0781] User: "I will verbally explain the use cases and business processes."
[0782] Step 2:
[0783] Terminal: Captures the user's voice using the microphone and temporarily saves it as an audio file locally.
[0784] Step 3:
[0785] Terminal: Sends saved audio data to the server via the network.
[0786] Step 4:
[0787] Server: Inputs the received audio data into the speech recognition engine and converts the audio into text data.
[0788] Step 5:
[0789] Server: Inputs the converted text data into a natural language processing (NLP) model and performs analysis.
[0790] Step 6:
[0791] Server: Uses an NLP model to extract information about use cases and business processes from text data.
[0792] Step 7:
[0793] Server: Simultaneously inputs voice and text data into the emotion engine to recognize the user's emotions.
[0794] Step 8:
[0795] Server: Based on extracted use cases and business flows, it automatically generates image diagrams using the business flow design library.
[0796] Step 9:
[0797] Server: Adjusts the generated image diagram appropriately based on recognized emotional information (e.g., color scheme and layout).
[0798] Step 10:
[0799] Server: Sends the adjusted image back to the terminal as image data.
[0800] Step 11:
[0801] Terminal: Displays an image to the user. For example, if the emotion is "excitement," the layout and graphical elements of the image will become more vivid.
[0802] 3. Summary of meeting content and processing steps for emotional recognition at the end of the meeting
[0803] Step 1:
[0804] User: "Summarize the meeting content verbally."
[0805] Step 2:
[0806] Terminal: Captures the user's voice using the microphone and temporarily saves it as an audio file locally.
[0807] Step 3:
[0808] Terminal: Sends saved audio data to the server via the network.
[0809] Step 4:
[0810] Server: Inputs the received audio data into the speech recognition engine and converts the audio into text data.
[0811] Step 5:
[0812] Server: Inputs the converted text data into a natural language processing (NLP) model and performs analysis.
[0813] Step 6:
[0814] Server: Uses NLP models to extract key points and essential information from text data.
[0815] Step 7:
[0816] Server: Simultaneously inputs voice and text data into the emotion engine to recognize the user's emotions.
[0817] Step 8:
[0818] Server: Organizes extracted points and combines them with relevant illustrations to generate a single visual summary.
[0819] Step 9:
[0820] Server: Appropriately adjusts the generated visual summary based on recognized sentiment information (e.g., color scheme and layout).
[0821] Step 10:
[0822] Server: Sends the adjusted visual summary back to the terminal as a PDF or image file.
[0823] Step 11:
[0824] Terminal: Displays a visual summary to the user. For example, if the emotion is "satisfied," the visual's color scheme will be warm tones.
[0825] The above outlines the specific processing steps of this system, which is a mechanism for efficiently and visually organizing and presenting meeting comments and content to the user. By using emotion recognition to display and generate information in accordance with the user's emotions, more effective communication becomes possible.
[0826] (Example 2)
[0827] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0828] Traditional meeting systems struggle to accurately grasp the key points of what is said and display them visually in an easy-to-understand manner. Furthermore, they often fail to consider user emotions when displaying information, leading to one-way communication. Additionally, they are inefficient at providing concise visual summaries of meeting discussions, resulting in low information sharing efficiency.
[0829] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0830] In this invention, the server includes means for capturing audio data, means for converting the captured audio data into text data, means for utilizing a natural language processing model to extract key points from the text data, means for using an emotion recognition engine to recognize the extracted key points and the user's emotions, means for displaying the key points based on the recognized emotion information, means for automatically generating an image diagram by converting the audio data into text data and performing analysis, means for adjusting and displaying the generated image diagram based on the user's emotions, and means for visualizing and displaying the summarized content at the end of the meeting. This enables accurate understanding of what the user said, information display that takes emotions into account, and reduces misunderstandings and efficient information sharing by visually summarizing the meeting discussion.
[0831] "Audio data" refers to the waveform information of sounds that record what a user says.
[0832] "Capturing" refers to the act of recording and acquiring information such as audio data.
[0833] "Text data" refers to audio data converted into written information.
[0834] A "natural language processing model" is a model that uses machine learning techniques to analyze, understand, and generate human language.
[0835] "Key points" refer to important information or content extracted from text data.
[0836] An "emotion recognition engine" is a device or software that analyzes voice data and text data to identify the user's emotions.
[0837] An "image diagram" refers to a visual explanatory diagram or flowchart generated based on text data.
[0838] A "summary" is a way of putting long texts or complex information into a concise format.
[0839] A "visual summary" is a visual representation of summarized information, and may include graphics, diagrams, and illustrations.
[0840] A "meeting" refers to a gathering of multiple users to discuss and share information.
[0841] This invention is a system designed to reduce misunderstandings in meetings and enable efficient information sharing. Furthermore, it supports more accessible and understandable communication by recognizing user emotions and displaying or generating information based on those emotions. This system combines a series of processes—capturing audio data, converting it into text data, and extracting key points—with an emotion recognition engine to visually and concisely present important information and key points expressed by the user.
[0842] Hardware and software to be used
[0843] Device: A device equipped with a microphone for the user to speak. This includes PCs, smartphones, tablets, etc.
[0844] Server: A backend system for converting and analyzing audio data. High processing power is desirable.
[0845] Speech recognition engine: Uses speech recognition services such as Google Cloud Speech-to-Text.
[0846] Natural language processing models: Advanced NLP models such as BERT.
[0847] Emotion recognition engine: An emotion analysis system such as IBM Watson.
[0848] System Description
[0849] 1. Capture audio data
[0850] When a user speaks during a meeting, the device captures their voice using the microphone. The captured audio data is temporarily stored in local storage and converted into text data in the next step.
[0851] 2. Converting speech to text
[0852] The device sends locally stored audio data to the server over the network. The server uses Google Cloud Speech-to-Text's speech recognition engine to convert the audio data into text data.
[0853] 3. Key points extraction
[0854] The server inputs text data into the BERT natural language processing (NLP) model for analysis. The NLP model extracts the main points of the text and identifies important information.
[0855] 4. Emotion recognition
[0856] The server inputs text and audio data, analyzed by an NLP model, into IBM Watson's emotion recognition engine to identify the user's emotions. The emotion recognition engine analyzes the tone and metadata of the voice to obtain information about the user's emotions.
[0857] 5. Displaying the main points
[0858] The server returns summarized key information and recognized sentiment information to the terminal. The terminal visually displays this information in an appropriately formatted way and presents it to the user. The display format is provided as a text box or a pop-up window.
[0859] 6. Automatic generation of image diagrams
[0860] When a user verbally describes a use case or workflow, the audio data is captured and converted into text. The server analyzes the converted text data and automatically generates a visual diagram. The generated diagram is then displayed on the device, with its colors and fonts adjusted based on the user's mood.
[0861] 7. Summary of meeting content at the end
[0862] At the end of the meeting, users verbally summarize the content. This audio data is also converted into text data, and the server extracts the key points and generates a visual summary along with relevant illustrations. The visual summary is displayed on the device in PDF or image format.
[0863] Specific example
[0864] 1. Specific examples of automatic summary generation
[0865] When a user says, "The features of this product are low cost and high performance. It also has excellent energy efficiency," the audio data is captured and converted into text data. The server extracts the key points "low cost, high performance, and energy efficiency," and an emotion recognition engine recognizes the user's emotions. If the emotion engine determines that "the features of this product are very important," a summary is sent back to the terminal and displayed to the user.
[0866] 2. Specific Examples of Automatic Image Generation
[0867] When a user describes a user visiting a website, logging in, searching for products, and adding them to their cart, the audio data is converted to text data, and the server generates a use case diagram like the following:
[0868] [User] --> [Access Site] --> [Login] --> [Search Products] --> [Add to Cart]
[0869] If the emotion engine determines that the user is excited, the colors and fonts of the diagram are adjusted and displayed in a visually easy-to-understand format.
[0870] 3. Specific examples of summaries at the end of a meeting
[0871] When a user summarizes the meeting at the end, saying, "Today we talked about the features of the new product and the marketing strategy for it," the audio data is converted to text data, and the server generates a visual summary like the following:
[0872] Features of the new product
[0873] Low cost
[0874] high performance
[0875] Energy saving
[0876] Marketing Strategy
[0877] Online advertising
[0878] social media
[0879] If the emotion engine recognizes that "the user is satisfied," the summary will be displayed with colors and a design that reflects the user's level of satisfaction.
[0880] The above describes the detailed embodiments for carrying out the invention. This invention automatically captures speech during meetings, converts it into text data, extracts key points, and combines this with sentiment recognition to provide deeper insights and more flexible communication. Furthermore, it generates visual diagrams of use cases and business flow explanations, and provides a visualized version of the content at the end of the meeting, thereby reducing misunderstandings and enabling efficient information sharing.
[0881] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0882] Step 1:
[0883] Audio data capture
[0884] When a user speaks during a meeting, the device captures their voice using the microphone. The input is the user's speech, and the output is the captured audio data. The device temporarily saves this audio data to local storage. As a concrete example, if a user says, "I'm going to talk about the design of the new product," that voice is captured and saved locally.
[0885] Step 2:
[0886] Speech-to-text conversion
[0887] The device sends the stored audio data to the server. The input is locally stored audio data, and the output is text data. The server uses the Google Cloud Speech-to-Text speech recognition engine to convert the audio data into text data. As a concrete example, the audio data "I'll talk about the design of the new product" is converted into the text data "I'll talk about the design of the new product".
[0888] Step 3:
[0889] Key points extraction
[0890] The server inputs the converted text data into the BERT natural language processing (NLP) model for analysis. The input is text data, and the output is key information. The BERT model extracts the important parts of the text and summarizes them. As a concrete example, from the text data "I will talk about the design of the new product," the key point "new product design" is extracted.
[0891] Step 4:
[0892] emotion recognition
[0893] The server inputs text and audio data, analyzed by the BERT model, into the emotion recognition engine. The input is the analyzed text and audio data, and the output is emotion information. The emotion recognition engine analyzes and identifies the user's emotions. As a concrete example of its operation, it identifies that the user is "excited" based on the text "I'm going to talk about the design of the new product" and the tone of voice.
[0894] Step 5:
[0895] Displaying the main points
[0896] The server sends back key information and recognized sentiment information to the terminal. The input is key information and sentiment information, and the output is the key information displayed visually. The terminal displays the key information to the user in an appropriate format. As a concrete example of operation, the key information "New product design" and the sentiment information "User is excited" are displayed in a pop-up window.
[0897] Step 6:
[0898] Automatic generation of image diagrams
[0899] When a user verbally describes a use case or workflow, the audio data is also captured and converted into text data. The input is the user's utterance, and the output is a visual diagram. The server generates a visual diagram based on the text data. As a concrete example, the following use case diagram is generated from the statement, "A user accesses the site, logs in, searches for a product, and adds it to their cart":
[0900] [User] --> [Access Site] --> [Login] --> [Search Products] --> [Add to Cart]
[0901] The colors and fonts of the images are adjusted based on emotions and displayed on the device.
[0902] Step 7:
[0903] Summary of the meeting's conclusion
[0904] At the end of the meeting, the user verbally summarizes the content. This audio data is also converted into text data, and the server extracts the key points and generates a visual summary along with relevant illustrations. The input is the user's summary, and the output is the visual summary. As a concrete example, the following visual summary is generated from the statement, "Today we talked about the features of the new product and the marketing strategy for it":
[0905] Features of the new product
[0906] Low cost
[0907] high performance
[0908] Energy saving
[0909] Marketing Strategy
[0910] Online advertising
[0911] social media
[0912] The visual summary will be displayed on your device in PDF format.
[0913] (Application Example 2)
[0914] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0915] Current autonomous vehicles face challenges in facilitating smooth communication between the driver and the vehicle system, as well as in understanding the driver's emotions in real time and providing appropriate information accordingly. Furthermore, there is a lack of mechanisms to support safe driving by appropriately managing the driver's emotional state. This can lead to delays in driver understanding and reaction, potentially increasing the risk of traffic accidents.
[0916] In Application Example 2, the specific processing performed by the specific processing unit 290 of the data processing device 12 is realized by the following means. In this invention, the server includes means for capturing audio data, means for converting the captured audio data into text data, means for extracting key points from the text data, means for recognizing emotions from the analyzed text data, and means for visually displaying information based on the driver's emotions. This enables smooth communication between the driver and the vehicle system, and by grasping the driver's emotional state in real time, it becomes possible to speed up the driver's understanding and reaction, and support safe driving.
[0917] "Means for capturing audio data" refers to a function that collects the driver's speech and surrounding sounds using microphones or other means, and saves them as digital data.
[0918] "Means for converting captured audio data into text data" refers to a function that uses speech recognition technology to convert collected audio data into text information.
[0919] "Methods for extracting key points from text data" refers to functions that use natural language processing technology to identify and summarize important information and topics from converted text.
[0920] "Means for displaying extracted key points as a summary" refers to a function for displaying extracted key points in a format that is easy for the driver to visually understand, on a display or other display device.
[0921] "Means for recognizing emotions from analyzed text data" refers to a function that automatically analyzes and recognizes the driver's emotional state (e.g., stress, satisfaction, excitement, etc.) based on text data and audio data.
[0922] "Means of visually displaying information based on the driver's emotions" refers to a function that adjusts the format and content of the displayed information according to the recognized emotional information, and provides appropriate visual feedback to the driver.
[0923] "Means for analyzing text data and recognizing use cases and business flows" refers to a function that uses natural language processing technology to extract specific procedures and processes from text data and identify them as use cases and business flows.
[0924] "Means for generating image diagrams based on recognized use cases and business flows" refers to a function that converts extracted process information into visual formats such as diagrams and flowcharts.
[0925] "Means for displaying the generated image diagram" refers to a function for displaying the generated visual diagram on a display or other display device.
[0926] "A means of generating a summary as a single visual by combining extracted points and related illustrations" refers to a function that extracts important points and combines them with related illustrations and diagrams to visually represent them as a single summary.
[0927] "Means for adjusting and displaying summaries based on perceived emotions" refers to a function for appropriately adjusting extracted key points and visual summaries based on the driver's emotional state and displaying them visually.
[0928] This invention relates to a system for autonomous vehicles that grasps the driver's emotional state in real time and enables smooth communication between the driver and the vehicle system. This system assists the driver by combining voice data capture, voice-to-text conversion, key point extraction, emotion recognition, and visual display of information.
[0929] Hardware and software configuration
[0930] hardware
[0931] Microphone: A device used to capture the driver's voice and surrounding sounds.
[0932] Display: A device that visually displays summary information and emotion-based feedback to the driver.
[0933] Vehicle ECU (Electronic Control Unit): A computing unit for real-time data processing.
[0934] software
[0935] Speech recognition engine (e.g., Google Speech Recognition API): Converts speech data into text data.
[0936] Natural language processing engines (e.g., the summarization model in the transformers library): Extract key points from text data.
[0937] Emotion recognition engine (e.g., the sentiment-analysis model in the transformers library): Recognizes emotions from the driver's statements.
[0938] Display control software: Software used to properly display information on a display.
[0939] Overall system processing
[0940] The server temporarily stores the audio data captured by the microphone locally. Then, it converts the audio data into text data using a speech recognition engine. This text data is analyzed by a natural language processing engine to extract key points. These extracted points, along with the driver's emotions, are analyzed by an emotion recognition engine, and feedback is provided to the driver based on this information. The feedback is displayed visually through the display.
[0941] Specific example
[0942] 1. Audio capture and text conversion:
[0943] When the driver says, "Turn right and then be careful at the intersection. A left turn follows," the voice is captured by the microphone. This voice data is then converted into text by a speech recognition engine, which reads, "Turn right and then be careful at the intersection. A left turn follows."
[0944] 2. Extracting key points and recognizing emotions:
[0945] Text data is input into a natural language processing engine (e.g., a summarization model), and key points such as "turn right, pay attention to the intersection, turn left" are extracted. Furthermore, an emotion recognition engine determines that "stress has been detected."
[0946] 3. Display of information:
[0947] Based on the extracted key points and emotion recognition results, information such as "Turn right, pay attention to intersection, turn left" is highlighted on the display. Furthermore, if stress is detected, messages and information encouraging relaxation are provided.
[0948] Example of a prompt:
[0949] "User comment: After turning right, be careful at the intersection. A left turn follows.
[0950] This allows drivers to receive necessary information and appropriate feedback in real time, enabling them to continue driving safely and effectively. This system contributes to improved traffic safety and reduced driver stress.
[0951] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0952] Step 1:
[0953] The server captures the driver's speech and surrounding sounds using a microphone. The audio data is input and temporarily stored locally in digital format.
[0954] Step 2:
[0955] The server sends the captured audio data to a speech recognition engine, where it is converted into text data. The input is audio data, and the output is text data. The speech recognition engine (e.g., Google Speech Recognition API) analyzes the audio waveform and generates the corresponding text.
[0956] Step 3:
[0957] The server inputs text data into a natural language processing engine and extracts key points. The input data is text data, and the output data is key point information. The natural language processing engine (e.g., the summarization model in the transformers library) analyzes the content of the text and extracts the important information by summarizing it.
[0958] Step 4:
[0959] The server inputs the extracted text data into an emotion recognition engine to recognize the driver's emotions. The input data is the extracted text data, and the output data is emotion information. The emotion recognition engine (e.g., the sentiment-analysis model from the transformers library) analyzes the emotions contained in the text and identifies emotional states such as stress, excitement, and satisfaction.
[0960] Step 5:
[0961] The server integrates extracted key information and recognized sentiment information and transmits it to the display device in a visual format. Input data consists of key information and sentiment information, while output data is visual feedback. Display control software determines the display format based on this data and provides appropriate information to the driver.
[0962] As a concrete example, if the driver says, "Turn right, then be careful at the intersection. Then turn left," in step 1, the voice data is captured, and in step 2, the speech recognition engine converts it into text data. In step 3, the key points "turn right, be careful at the intersection, turn left" are extracted, and in step 4, "stress" is recognized. Finally, in step 5, the important points are highlighted on the display, and if stress is detected, a message encouraging relaxation is added.
[0963] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0964] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0965] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0966] [Third Embodiment]
[0967] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0968] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0969] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0970] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0971] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0972] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0973] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0974] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0975] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0976] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0977] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0978] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0979] This invention is a system designed to reduce misunderstandings in meetings and enable efficient information sharing. The system captures audio data, converts it into text data, and extracts key points, visually and concisely presenting important information and key points spoken by the user. The specific processing and functions are described below.
[0980] Specific processes and functions of the system
[0981] 1. Capture audio data
[0982] When a user speaks during a meeting, the device captures their voice using the microphone. The captured audio data is temporarily stored locally. This audio data is then converted into text data in the next step.
[0983] 2. Converting speech to text
[0984] Audio data stored on the device is sent to a server via the network. The server converts the audio into text data using a speech recognition engine (e.g., a standard speech recognition API). This process converts the audio data into text format, making subsequent processing easier.
[0985] 3. Key points extraction
[0986] The server inputs the converted text data into a natural language processing (NLP) model (e.g., a BERT model) for analysis. The NLP model extracts the main points from the text and identifies key information. The extracted points are concisely summarized, clearly indicating the main agenda and conclusions of the meeting.
[0987] 4. Displaying the main points
[0988] The server sends summarized key points back to the terminal. The terminal visually displays the summarized points to the user. This display is done in the form of a text box or pop-up window, allowing the user to quickly see the important points.
[0989] 5. Automatic generation of image diagrams
[0990] When a user verbally describes a use case or business workflow, the audio data is also captured and converted into text data. Based on the converted text data, the server analyzes the use case or business workflow recognized by the system and automatically generates a visual diagram. The generated diagram is displayed on the user's terminal in a format that is easy for the user to understand visually.
[0991] 6. Summary of meeting content at the end
[0992] At the end of a meeting, users verbally summarize its content. This audio data is also captured and converted into text. The server uses this text data to organize the key points and combines them with relevant illustrations to generate a single visual summary. This visual summary visually summarizes the meeting's discussion and can be displayed on the device as a PDF or image for easy reference.
[0993] Specific example
[0994] Examples of automatic summary generation
[0995] When a user says, "The features of this product are its low cost and high performance. It also has excellent energy-saving capabilities," the audio data is captured and converted into text data. The server extracts the key points, "low cost, high performance, and energy saving," and sends them back to the terminal for display.
[0996] Examples of automated image generation
[0997] When a user describes a user visiting a website, logging in, searching for products, and adding them to their cart, the audio data is converted to text data, and the server generates a use case diagram like the following:
[0998] [User] --> [Access Site] --> [Login] --> [Search Products] --> [Add to Cart]
[0999] This diagram will be displayed on the device and will be in a format that is easy for the user to understand visually.
[1000] Examples of content summaries
[1001] When a user summarizes the meeting at the end by saying, "Today we talked about the features of the new product and the marketing strategy for it," this audio data is converted into text data, and the server generates a visual summary like the following:
[1002] Features of the new product
[1003] Low cost
[1004] high performance
[1005] Energy saving
[1006] Marketing Strategy
[1007] Online advertising
[1008] social media
[1009] This visual summary will be displayed on the device as a graphic recording in PDF or image format.
[1010] As described above, this system automatically captures speech during meetings, converts it into text data, extracts key points, and displays them visually, allowing users to easily grasp important information. Furthermore, it generates visual diagrams of use cases and business workflows, providing a visualized version of the content at the end of the meeting, thereby reducing misunderstandings and enabling efficient information sharing.
[1011] The following describes the processing flow.
[1012] 1. Processing steps for automatic key point generation
[1013] Step 1:
[1014] User: "I will verbally state the important points."
[1015] Step 2:
[1016] Terminal: Captures the user's voice using the microphone and temporarily saves it as an audio file locally.
[1017] Step 3:
[1018] Terminal: Sends saved audio data to the server via the network.
[1019] Step 4:
[1020] Server: Inputs the received audio data into the speech recognition engine and converts the audio into text data.
[1021] Step 5:
[1022] Server: Inputs the converted text data into a natural language processing (NLP) model and performs analysis.
[1023] Step 6:
[1024] Server: Uses NLP models to extract key points and essential information from text data.
[1025] Step 7:
[1026] Server: Organizes the extracted key points into a summary and condenses them into concise text.
[1027] Step 8:
[1028] Server: Sends the generated summary back to the terminal.
[1029] Step 9:
[1030] Terminal: Displays summarized key information to the user. Specifically, it uses formats such as text boxes and pop-up windows.
[1031] 2. Processing steps for automatic image diagram generation
[1032] Step 1:
[1033] User: "I will verbally explain the use cases and business processes."
[1034] Step 2:
[1035] Terminal: Captures the user's voice using the microphone and temporarily saves it as an audio file locally.
[1036] Step 3:
[1037] Terminal: Sends saved audio data to the server via the network.
[1038] Step 4:
[1039] Server: Inputs the received audio data into the speech recognition engine and converts the audio into text data.
[1040] Step 5:
[1041] Server: Inputs the converted text data into a natural language processing (NLP) model and performs analysis.
[1042] Step 6:
[1043] Server: Uses an NLP model to extract information about use cases and business processes from text data.
[1044] Step 7:
[1045] Server: Based on the extracted information, it automatically generates an image diagram using the business flow design library.
[1046] Step 8:
[1047] Server: Sends the generated image back to the terminal as image data.
[1048] Step 9:
[1049] Terminal: Displays an image diagram to the user. Specifically, it uses a graphical viewer or image display window.
[1050] 3. Steps for summarizing the meeting content at the end of the meeting
[1051] Step 1:
[1052] User: "Summarize the meeting content verbally."
[1053] Step 2:
[1054] Terminal: Captures the user's voice using the microphone and temporarily saves it as an audio file locally.
[1055] Step 3:
[1056] Terminal: Sends saved audio data to the server via the network.
[1057] Step 4:
[1058] Server: Inputs the received audio data into the speech recognition engine and converts the audio into text data.
[1059] Step 5:
[1060] Server: Inputs the converted text data into a natural language processing (NLP) model and performs analysis.
[1061] Step 6:
[1062] Server: Uses NLP models to extract key points and essential information from text data.
[1063] Step 7:
[1064] Server: Organizes extracted points and generates a visual summary combining them with relevant illustrations.
[1065] Step 8:
[1066] Server: Sends the generated visual summary back to the terminal as a PDF or image file.
[1067] Step 9:
[1068] Terminal: Displays a visual summary to the user. Specifically, it uses a PDF viewer or an image display window.
[1069] The above outlines the specific processing steps of this system, which is a mechanism for efficiently and visually organizing and presenting comments and content from meetings to the user.
[1070] (Example 1)
[1071] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[1072] In today's business environment, many meetings are conducted remotely, but misunderstandings often occur during voice communication. Therefore, there is a need for efficient ways to share meeting content and refer to it later. There is also a demand for easily visualizing and reviewing verbal explanations of use cases and business flows. However, current technologies often separate speech recognition and key point extraction, and automatic generation of diagrams is limited to certain special cases. This makes it difficult to accurately grasp the key points of a meeting.
[1073] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[1074] In this invention, the server includes means for capturing audio data, means for converting the captured audio data into text data using conversion means, and means for extracting key points from the text data. This makes it possible to convert audio data into text, extract key points, and display them visually.
[1075] "Means for capturing audio data" refers to methods for acquiring the voice spoken by a user as digital data through an input device such as a microphone.
[1076] "Means for converting captured audio data into text data using conversion means" refers to means for converting captured audio data into text information using speech recognition technology.
[1077] "Methods for extracting key points from text data" refers to methods that utilize natural language processing techniques to select and extract important information from text data.
[1078] "Means of displaying as a summary" refers to a method of visualizing the extracted key points and presenting them to the user.
[1079] "Means for recognizing use cases and business flows" refers to methods for analyzing text data to understand and recognize specific usage scenarios and business procedures.
[1080] "Methods for generating image diagrams" refer to methods for automatically creating visually easy-to-understand diagrams based on recognized use cases and business flows.
[1081] A "means for generating visual summaries" is a method for creating a visually easy-to-understand summary by combining points extracted from text data with relevant illustrations.
[1082] This invention is a system for reducing misunderstandings in meetings and achieving efficient information sharing. This system visually and concisely presents important information and key points spoken by users through a series of processes: capturing audio data, converting it into text data, and extracting key points. Specific embodiments for carrying out this invention are described below.
[1083] Audio data capture
[1084] When a user speaks during a meeting, the device captures their voice using its microphone. The captured audio data is temporarily stored locally. Suitable devices for this purpose include computers and smartphones with built-in microphones.
[1085] Sending audio data
[1086] The terminal compresses the captured audio data and sends it to the server using a secure communication protocol (e.g., HTTPS). Compressing the audio data reduces the communication load.
[1087] Speech-to-text conversion
[1088] The server converts the received audio data into text data using a standard speech recognition engine (e.g., Google Cloud Speech-to-Text API). This conversion makes the audio data into text format, which simplifies the next processing step.
[1089] Extracting key points
[1090] The server uses a generative AI model (e.g., the BERT model) to extract key points from the transformed text data. Natural language processing techniques are used to select important information from the text data.
[1091] Return and display of key points
[1092] The server converts the extracted key points into JSON format and sends them back to the terminal. The terminal visually displays the received key points in a text box or pop-up window on the screen. This allows the user to quickly confirm the important points.
[1093] Automatic generation of image diagrams
[1094] When a user verbally explains a use case or business workflow, the audio data is captured and sent to the server. The server converts the audio into text data and uses an AI model for illustration generation (e.g., Diagrams.net API) to automatically generate diagrams of the use case or business workflow. The generated diagrams are displayed on the terminal in a format that is easy for the user to understand visually.
[1095] Summary of the meeting's conclusion
[1096] At the end of a meeting, users verbally summarize its content. This audio data is also captured and converted into text. The server uses this text data to organize the key points and combines them with relevant illustrations to generate a visual summary. The generated visual summary is displayed on the device in PDF or image format for easy reference.
[1097] Specific example
[1098] Examples of automatic summary generation
[1099] When a user says, "The features of this product are its low cost and high performance. It also has excellent energy-saving capabilities," the audio data is captured and converted into text data. The server extracts the key points, "low cost, high performance, and energy saving," and sends them back to the terminal for display.
[1100] Examples of automated image generation
[1101] When a user describes a user visiting a website, logging in, searching for products, and adding them to their cart, the audio data is converted to text data, and the server generates a use case diagram like the following:
[1102] Users access the site, log in, search for products, and add them to their cart.
[1103] Example of a summary of meeting content at the end
[1104] When a user summarizes the meeting at the end by saying, "Today we talked about the features of the new product and the marketing strategy for it," this audio data is converted into text data, and the server generates a visual summary like the following:
[1105] Features of the new product
[1106] Low cost
[1107] high performance
[1108] Energy saving
[1109] Marketing Strategy
[1110] Online advertising
[1111] social media
[1112] The generated visual summary will be displayed on the device in PDF format.
[1113] As described above, this system includes a series of processes to reduce misunderstandings during meetings and enable efficient information sharing. High accuracy and efficiency are achieved by using specific hardware and software in each step. Furthermore, the specific operation of the system is explained through concrete examples.
[1114] The flow of the specific processing in Example 1 will be explained using Figure 11.
[1115] Step 1: Capture audio data
[1116] When a user speaks during a meeting, the device captures their voice using its microphone. The input is the user's voice, and the output is digitized audio data. Specifically, the device's microphone detects the voice in real time and temporarily saves it as an audio file (e.g., in WAV format).
[1117] Step 2: Sending the audio data
[1118] The device compresses the captured audio data and sends it to the server using a secure communication protocol (e.g., HTTPS). The input is the captured audio data, and the output is the audio data sent to the server. Specifically, the device compresses the audio data (e.g., converts it to AAC format) and sends a POST request to the API endpoint.
[1119] Step 3: Convert speech to text
[1120] The server converts the received audio data into text data using a standard speech recognition engine (e.g., Google Cloud Speech-to-Text API). The input is the received audio data, and the output is text data. Specifically, the server sends an API request to the speech recognition service and retrieves the result in text format (e.g., JSON format).
[1121] Step 4: Extracting the key points
[1122] The server uses a generative AI model (e.g., a BERT model) to extract key points from the transformed text data. The input is text data, and the output is a list of extracted key points. Specifically, the server inputs text data into the model and organizes the output key points in list format.
[1123] Step 5: Return and display the key points.
[1124] The server converts the extracted key points into JSON data and sends it back to the terminal. The terminal displays the received key points on the screen in the form of text boxes or pop-up windows. The input is the extracted key points in JSON data, and the output is the visual information presented to the user. Specifically, the terminal parses the JSON data received from the server and updates the UI using HTML / CSS to display it on the screen.
[1125] Step 6: Automatic generation of image diagrams
[1126] When a user verbally describes a use case or business process, the audio data is captured and sent to the server. The server converts the audio into text data and uses an illustration generation AI model (e.g., Diagrams.net API) to automatically generate image diagrams of the use case or business process. The input is text data, and the output is the generated image diagram. Specifically, the server analyzes the text data and inputs it as prompt text into the generation AI model. For example, a prompt text might be: "The user accesses the site, logs in, searches for a product, and adds it to the cart."
[1127] Step 7: Summary of meeting content at the end
[1128] At the end of a meeting, users verbally summarize its content. This audio data is also captured and converted into text data. The server uses this text data to organize the key points and combines them with relevant illustrations to generate a visual summary. The input is the audio summary data from the end of the meeting, and the output is the visual summary. Specifically, the server combines the text data and illustrations, generates a visual summary using a PDF creation library (e.g., Apache PDFBox), and sends it to the terminal. The PDF is then displayed on the terminal.
[1129] (Application Example 1)
[1130] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[1131] In modern factories, frequent misunderstandings in information sharing during meetings and production flow setup are a major challenge. This is especially true in environments where verbal explanations are frequently used, where information is often misinterpreted or vaguely remembered. Furthermore, the difficulty in visually grasping meeting points and workflows hinders efficient information sharing. This can lead to delays and errors, potentially reducing productivity.
[1132] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[1133] In this invention, the server includes means for capturing audio data, means for converting the captured audio data into text data, means for extracting key points from the text data, means for displaying the extracted key points as a summary, means for recognizing use cases and business flows, means for generating image diagrams based on the recognized use cases and business flows, means for displaying the generated summary and image diagrams on factory monitors or autonomous device displays, and means for extracting key points at the end of a meeting and generating a related visual summary. This makes it possible to improve the efficiency of information sharing by accurately transcribing audio information during a meeting into text and extracting key points. In addition, by visually displaying business flows and use cases, all participants can proceed with their work based on the same understanding, reducing information discrepancies. At the end of a meeting, the key points can be visualized and summarized to prevent information from being overlooked.
[1134] "Audio data" refers to information recorded in digital format.
[1135] "Means of capturing" refers to methods and devices for acquiring audio as digital data.
[1136] "Means of converting to text data" refers to methods and technologies for converting audio data into text-based data.
[1137] "Methods for extracting key points" refer to techniques for identifying and extracting particularly important parts from texts or conversations.
[1138] "Means of displaying as a summary" refers to a method of visually presenting extracted important information in a concise format.
[1139] A "use case" refers to a specific example of use or procedure under particular circumstances.
[1140] A "business process flow" is a diagram that shows the series of steps and procedures involved in performing a business task.
[1141] "Means of recognition" refer to the techniques and methods for identifying and understanding specific patterns or information.
[1142] "Means for generating image diagrams" refers to methods and techniques for creating visual shapes and charts based on text data and analysis results.
[1143] A "factory monitor" refers to a display device installed within a factory for displaying information.
[1144] An "autonomous device display" is a display device installed in a device that operates automatically.
[1145] "Meeting conclusions" refers to the main points or conclusions summarized at the end of a meeting.
[1146] A "visual summary" is a concise summary that visually presents key points and important information using diagrams and charts.
[1147] This invention is a system designed to streamline meetings and production flow setup in factories and reduce information sharing discrepancies. Through a series of processes—capturing audio data, converting it to text data, and extracting key points—this system visually and concisely presents important information and workflows spoken by users. The specific processing and functions of the system are described below.
[1148] 1. Capture audio data
[1149] When a user speaks during a meeting, the device captures their voice using its microphone. The captured audio data is temporarily stored locally. This audio data is then converted into text data in the next step. This process is achieved by using a device equipped with a microphone and recording capabilities.
[1150] 2. Converting speech to text
[1151] Audio data stored on the device is sent to a server via the network. The server converts the audio into text data using a speech recognition engine. Specific software used in this process includes the Google Speech-to-Text API. This converts the audio data into text format, making subsequent processing easier.
[1152] 3. Key points extraction
[1153] The server analyzes the converted text data using a natural language processing (NLP) model. The BERT model is a possible NLP model used here. The server extracts the main points from the text and identifies important information. This point extraction clearly identifies the main topics and conclusions discussed in the meeting.
[1154] 4. Displaying the main points
[1155] The server sends summarized key information back to the terminal. The terminal visually displays the summarized key points to the user. This display is shown in the form of text boxes or pop-up windows on monitors or autonomous device displays within the factory. This allows the user to quickly identify the important points.
[1156] 5. Automatic generation of image diagrams
[1157] When a user verbally describes a use case or business workflow, the audio data is also captured and converted into text data. Based on this converted text data, the server analyzes the recognized use case or business workflow and automatically generates a visual diagram. A generative AI model may be used in this process. The generated diagram is displayed on the user's device in a format that is easy for them to understand visually.
[1158] 6. Summary of meeting content at the end
[1159] At the end of a meeting, users verbally summarize its content. This audio data is also captured and converted into text. The server uses this text data to organize the key points and generates a single visual summary along with relevant images. This visual summary is displayed on the device as a graphic recording in PDF or image format, allowing users to easily refer to it.
[1160] Specific example
[1161] For example, if a factory worker says, "The procedure for setting up the new production line is to first transport the parts, then assemble them, and then conduct quality checks," the key points generated will be "parts transport, assembly, quality check." Based on this, the server will generate a flowchart and display an image showing "parts transport -> assembly -> quality check." Also, if the meeting concludes with a summary such as, "Today we talked about the procedure for setting up the new line and the points to note," a visual summary of those key points will be generated.
[1162] Example of a prompt
[1163] "The following audio data concerns setting up a new production line. Please extract the key points and generate the production line setup procedure. 'The procedure for setting up the new production line is to first transport the parts, then assemble them, and then conduct quality control.'"
[1164] By following the above steps, it is possible to utilize this method in meetings and production flow settings on the factory floor, significantly improving the efficiency of information sharing.
[1165] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[1166] Step 1:
[1167] Audio data capture
[1168] The device uses a microphone to capture audio during the meeting. The input is the audio signal from the meeting, and the output is digital audio data. This audio data is temporarily stored in local storage. Specifically, the microphone built into the device picks up the audio signal, converts it to a digital format, and stores it.
[1169] Step 2:
[1170] Speech-to-text conversion
[1171] The device sends the captured audio data to the server over the network. The server converts the audio data into text data using the Google Speech-to-Text API. The input is digital audio data, and the output is text data. This is achieved by the server passing the audio data to the API and receiving the resulting text data.
[1172] Step 3:
[1173] Extracting key points
[1174] The server inputs the converted text data into a natural language processing (NLP) model (such as the BERT model) to analyze and extract key points. The input is text data, and the output is text data of the extracted key points. This is achieved by the server using the NLP model to analyze the data and extract important information.
[1175] Step 4:
[1176] Displaying the main points
[1177] The server returns summarized key information to the terminal. The terminal displays this information on factory monitors or autonomous device displays. The input is the extracted key points as text data, and the output is the key points displayed visually. The data is displayed visually using the terminal's display control software.
[1178] Step 5:
[1179] Recognition of use cases and business processes
[1180] When a user verbally describes a use case or business process, the audio data is captured and converted to text. The server analyzes the converted text data to recognize the use case or business process. The input consists of the audio data and the converted text data, and the output is the text data of the recognized use case or business process.
[1181] Step 6:
[1182] Automatic generation of image diagrams
[1183] The server generates visual diagrams based on recognized use cases and business workflows. The input is recognized text data, and the output is a visual diagram. A generation AI model is used to convert the data into flowcharts and charts for display.
[1184] Step 7:
[1185] Display of an image diagram
[1186] The server sends the generated image back to the terminal. The terminal displays this image on a factory monitor or an autonomous device's display. The input is the image, and the output is the visually displayed image.
[1187] Step 8:
[1188] Summary of the meeting's conclusion
[1189] At the end of a meeting, the user verbally summarizes its contents. This audio data is also captured and converted into text data. The server uses this text data to organize the key points and combines them with relevant images to generate a visual summary. The input is the user's summary audio and text data, and the output is the visual summary.
[1190] Step 9:
[1191] Display a visual summary
[1192] The server sends the generated visual summary back to the terminal. The terminal displays this in PDF or image format on factory monitors or autonomous machine displays. The input is the visual summary data, and the output is the visually displayed visual summary.
[1193] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[1194] This invention is a system designed to reduce misunderstandings in meetings and enable efficient information sharing. Furthermore, it supports more accessible and understandable communication by recognizing user emotions and displaying or generating information based on those emotions. This system combines a series of processes—capturing audio data, converting it to text data, and extracting key points—with an emotion engine to visually and concisely present important information and key points expressed by the user. The specific processing and functions are described below.
[1195] Specific processes and functions of the system
[1196] 1. Capture audio data
[1197] When a user speaks during a meeting, the device captures their voice using the microphone. The captured audio data is temporarily stored locally. This audio data is then converted into text data in the next step.
[1198] 2. Converting speech to text
[1199] The audio data stored on the device is sent to the server via the network. The server converts the audio into text data using a speech recognition engine. This process simplifies subsequent processing because the audio data is converted into text format.
[1200] 3. Key points extraction
[1201] The server inputs the converted text data into a natural language processing (NLP) model for analysis. The NLP model extracts the key points from the text and identifies important information. The extracted points are concisely summarized, clearly indicating the main agenda and conclusions of the meeting.
[1202] 4. Emotion recognition
[1203] The server uses an emotion engine to recognize the user's emotions based on text and audio data analyzed by an NLP model. The emotion engine identifies the user's emotions based on factors such as voice tone and facial expression analysis, and evaluates that emotional information.
[1204] 5. Displaying the main points
[1205] The server sends summarized key points and recognized sentiment information back to the terminal. The terminal visually displays the summarized key points to the user in an appropriately adjusted format. This display is presented in the form of a text box or pop-up window, and is done in a way that is sensitive to the user's emotions.
[1206] 6. Automatic generation of image diagrams
[1207] When a user verbally describes a use case or business workflow, the audio data is also captured and converted into text data. Based on the converted text data, the server analyzes the use case or business workflow recognized by the system and automatically generates a visual diagram. The generated diagram is appropriately adjusted based on the user's feelings and displayed on the terminal in a visually easy-to-understand format.
[1208] 7. Summary of meeting content at the end
[1209] At the end of a meeting, users verbally summarize its content. This audio data is also captured and converted into text. The server uses this text data to organize the key points and combines them with relevant illustrations to generate a single visual summary. This visual summary is tailored to the user's emotions and provides a visual representation of the meeting's discussion. It can be displayed on the device as a PDF or image for easy reference.
[1210] Specific example
[1211] Examples of automatic summary generation
[1212] When a user says, "The features of this product are low cost and high performance. It also has excellent energy efficiency," the audio data is captured and converted into text data. The server extracts the key points "low cost, high performance, and energy efficiency," and an emotion engine recognizes the user's emotions. If the emotion engine determines that "the features of this product are very important," a summary is sent back to the terminal and displayed to the user.
[1213] Examples of automated image generation
[1214] When a user describes a user visiting a website, logging in, searching for products, and adding them to their cart, the audio data is converted to text data, and the server generates a use case diagram like the following:
[1215] [User] --> [Access Site] --> [Login] --> [Search Products] --> [Add to Cart]
[1216] If the emotion engine determines that the user is excited, the colors and fonts of the diagram are adjusted and displayed in a visually easy-to-understand format.
[1217] Examples of content summaries
[1218] When a user summarizes the meeting at the end by saying, "Today we talked about the features of the new product and the marketing strategy for it," this audio data is converted into text data, and the server generates a visual summary like the following:
[1219] Features of the new product
[1220] Low cost
[1221] high performance
[1222] Energy saving
[1223] Marketing Strategy
[1224] Online advertising
[1225] social media
[1226] If the emotion engine recognizes that "the user is satisfied," the summary will be displayed with colors and a design that reflects the user's level of satisfaction.
[1227] As described above, this system automatically captures speech during meetings, converts it into text data, extracts key points, and displays them to the user in combination with sentiment recognition, enabling deeper insights and more flexible communication. Furthermore, it generates visual diagrams of use cases and business flow explanations, and provides a visualized version of the content at the end of the meeting, thereby reducing misunderstandings and facilitating efficient information sharing.
[1228] The following describes the processing flow.
[1229] 1. Processing steps for automatic summary generation and sentiment recognition
[1230] Step 1:
[1231] User: "I will verbally state the important points."
[1232] Step 2:
[1233] Terminal: Captures the user's voice using the microphone and temporarily saves it as an audio file locally.
[1234] Step 3:
[1235] Terminal: Sends saved audio data to the server via the network.
[1236] Step 4:
[1237] Server: Inputs the received audio data into the speech recognition engine and converts the audio into text data.
[1238] Step 5:
[1239] Server: Inputs the converted text data into a natural language processing (NLP) model and performs analysis.
[1240] Step 6:
[1241] Server: Uses NLP models to extract key points and essential information from text data.
[1242] Step 7:
[1243] Server: Simultaneously inputs voice and text data into the emotion engine to recognize the user's emotions.
[1244] Step 8:
[1245] Server: Combines extracted key points and recognized sentiment information to generate a summary.
[1246] Step 9:
[1247] Server: Sends the generated summary and sentiment information back to the terminal.
[1248] Step 10:
[1249] Terminal: Displays summarized key information and emotional information to the user in a visually adjusted format. For example, if the emotion is "emotional," the text color and font are adjusted.
[1250] 2. Image diagram automatic generation and emotion recognition processing steps
[1251] Step 1:
[1252] User: "I will verbally explain the use cases and business processes."
[1253] Step 2:
[1254] Terminal: Captures the user's voice using the microphone and temporarily saves it as an audio file locally.
[1255] Step 3:
[1256] Terminal: Sends saved audio data to the server via the network.
[1257] Step 4:
[1258] Server: Inputs the received audio data into the speech recognition engine and converts the audio into text data.
[1259] Step 5:
[1260] Server: Inputs the converted text data into a natural language processing (NLP) model and performs analysis.
[1261] Step 6:
[1262] Server: Uses an NLP model to extract information about use cases and business processes from text data.
[1263] Step 7:
[1264] Server: Simultaneously inputs voice and text data into the emotion engine to recognize the user's emotions.
[1265] Step 8:
[1266] Server: Based on extracted use cases and business flows, it automatically generates image diagrams using the business flow design library.
[1267] Step 9:
[1268] Server: Adjusts the generated image diagram appropriately based on recognized emotional information (e.g., color scheme and layout).
[1269] Step 10:
[1270] Server: Sends the adjusted image back to the terminal as image data.
[1271] Step 11:
[1272] Terminal: Displays an image to the user. For example, if the emotion is "excitement," the layout and graphical elements of the image will become more vivid.
[1273] 3. Summary of meeting content and processing steps for emotional recognition at the end of the meeting
[1274] Step 1:
[1275] User: "Summarize the meeting content verbally."
[1276] Step 2:
[1277] Terminal: Captures the user's voice using the microphone and temporarily saves it as an audio file locally.
[1278] Step 3:
[1279] Terminal: Sends saved audio data to the server via the network.
[1280] Step 4:
[1281] Server: Inputs the received audio data into the speech recognition engine and converts the audio into text data.
[1282] Step 5:
[1283] Server: Inputs the converted text data into a natural language processing (NLP) model and performs analysis.
[1284] Step 6:
[1285] Server: Uses NLP models to extract key points and essential information from text data.
[1286] Step 7:
[1287] Server: Simultaneously inputs voice and text data into the emotion engine to recognize the user's emotions.
[1288] Step 8:
[1289] Server: Organizes extracted points and combines them with relevant illustrations to generate a single visual summary.
[1290] Step 9:
[1291] Server: Appropriately adjusts the generated visual summary based on recognized sentiment information (e.g., color scheme and layout).
[1292] Step 10:
[1293] Server: Sends the adjusted visual summary back to the terminal as a PDF or image file.
[1294] Step 11:
[1295] Terminal: Displays a visual summary to the user. For example, if the emotion is "satisfied," the visual's color scheme will be warm tones.
[1296] The above outlines the specific processing steps of this system, which is a mechanism for efficiently and visually organizing and presenting meeting comments and content to the user. By using emotion recognition to display and generate information in accordance with the user's emotions, more effective communication becomes possible.
[1297] (Example 2)
[1298] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[1299] Traditional meeting systems struggle to accurately grasp the key points of what is said and display them visually in an easy-to-understand manner. Furthermore, they often fail to consider user emotions when displaying information, leading to one-way communication. Additionally, they are inefficient at providing concise visual summaries of meeting discussions, resulting in low information sharing efficiency.
[1300] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[1301] In this invention, the server includes means for capturing audio data, means for converting the captured audio data into text data, means for utilizing a natural language processing model to extract key points from the text data, means for using an emotion recognition engine to recognize the extracted key points and the user's emotions, means for displaying the key points based on the recognized emotion information, means for automatically generating an image diagram by converting the audio data into text data and performing analysis, means for adjusting and displaying the generated image diagram based on the user's emotions, and means for visualizing and displaying the summarized content at the end of the meeting. This enables accurate understanding of what the user said, information display that takes emotions into account, and reduces misunderstandings and efficient information sharing by visually summarizing the meeting discussion.
[1302] "Audio data" refers to the waveform information of sounds that record what a user says.
[1303] "Capturing" refers to the act of recording and acquiring information such as audio data.
[1304] "Text data" refers to audio data converted into written information.
[1305] A "natural language processing model" is a model that uses machine learning techniques to analyze, understand, and generate human language.
[1306] "Key points" refer to important information or content extracted from text data.
[1307] An "emotion recognition engine" is a device or software that analyzes voice data and text data to identify the user's emotions.
[1308] An "image diagram" refers to a visual explanatory diagram or flowchart generated based on text data.
[1309] A "summary" is a way of putting long texts or complex information into a concise format.
[1310] A "visual summary" is a visual representation of summarized information, and may include graphics, diagrams, and illustrations.
[1311] A "meeting" refers to a gathering of multiple users to discuss and share information.
[1312] This invention is a system designed to reduce misunderstandings in meetings and enable efficient information sharing. Furthermore, it supports more accessible and understandable communication by recognizing user emotions and displaying or generating information based on those emotions. This system combines a series of processes—capturing audio data, converting it into text data, and extracting key points—with an emotion recognition engine to visually and concisely present important information and key points expressed by the user.
[1313] Hardware and software to be used
[1314] Device: A device equipped with a microphone for the user to speak. This includes PCs, smartphones, tablets, etc.
[1315] Server: A backend system for converting and analyzing audio data. High processing power is desirable.
[1316] Speech recognition engine: Uses speech recognition services such as Google Cloud Speech-to-Text.
[1317] Natural language processing models: Advanced NLP models such as BERT.
[1318] Emotion recognition engine: An emotion analysis system such as IBM Watson.
[1319] System Description
[1320] 1. Capture audio data
[1321] When a user speaks during a meeting, the device captures their voice using the microphone. The captured audio data is temporarily stored in local storage and converted into text data in the next step.
[1322] 2. Converting speech to text
[1323] The device sends locally stored audio data to the server over the network. The server uses Google Cloud Speech-to-Text's speech recognition engine to convert the audio data into text data.
[1324] 3. Key points extraction
[1325] The server inputs text data into the BERT natural language processing (NLP) model for analysis. The NLP model extracts the main points of the text and identifies important information.
[1326] 4. Emotion recognition
[1327] The server inputs text and audio data, analyzed by an NLP model, into IBM Watson's emotion recognition engine to identify the user's emotions. The emotion recognition engine analyzes the tone and metadata of the voice to obtain information about the user's emotions.
[1328] 5. Displaying the main points
[1329] The server returns summarized key information and recognized sentiment information to the terminal. The terminal visually displays this information in an appropriately formatted way and presents it to the user. The display format is provided as a text box or a pop-up window.
[1330] 6. Automatic generation of image diagrams
[1331] When a user verbally describes a use case or workflow, the audio data is captured and converted into text. The server analyzes the converted text data and automatically generates a visual diagram. The generated diagram is then displayed on the device, with its colors and fonts adjusted based on the user's mood.
[1332] 7. Summary of meeting content at the end
[1333] At the end of the meeting, users verbally summarize the content. This audio data is also converted into text data, and the server extracts the key points and generates a visual summary along with relevant illustrations. The visual summary is displayed on the device in PDF or image format.
[1334] Specific example
[1335] 1. Specific examples of automatic summary generation
[1336] When a user says, "The features of this product are low cost and high performance. It also has excellent energy efficiency," the audio data is captured and converted into text data. The server extracts the key points "low cost, high performance, and energy efficiency," and an emotion recognition engine recognizes the user's emotions. If the emotion engine determines that "the features of this product are very important," a summary is sent back to the terminal and displayed to the user.
[1337] 2. Specific Examples of Automatic Image Generation
[1338] When a user describes a user visiting a website, logging in, searching for products, and adding them to their cart, the audio data is converted to text data, and the server generates a use case diagram like the following:
[1339] [User] --> [Access Site] --> [Login] --> [Search Products] --> [Add to Cart]
[1340] If the emotion engine determines that the user is excited, the colors and fonts of the diagram are adjusted and displayed in a visually easy-to-understand format.
[1341] 3. Specific examples of summaries at the end of a meeting
[1342] When a user summarizes the meeting at the end, saying, "Today we talked about the features of the new product and the marketing strategy for it," the audio data is converted to text data, and the server generates a visual summary like the following:
[1343] Features of the new product
[1344] Low cost
[1345] high performance
[1346] Energy saving
[1347] Marketing Strategy
[1348] Online advertising
[1349] social media
[1350] If the emotion engine recognizes that "the user is satisfied," the summary will be displayed with colors and a design that reflects the user's level of satisfaction.
[1351] The above describes the detailed embodiments for carrying out the invention. This invention automatically captures speech during meetings, converts it into text data, extracts key points, and combines this with sentiment recognition to provide deeper insights and more flexible communication. Furthermore, it generates visual diagrams of use cases and business flow explanations, and provides a visualized version of the content at the end of the meeting, thereby reducing misunderstandings and enabling efficient information sharing.
[1352] The flow of the specific processing in Example 2 will be explained using Figure 13.
[1353] Step 1:
[1354] Audio data capture
[1355] When a user speaks during a meeting, the device captures their voice using the microphone. The input is the user's speech, and the output is the captured audio data. The device temporarily saves this audio data to local storage. As a concrete example, if a user says, "I'm going to talk about the design of the new product," that voice is captured and saved locally.
[1356] Step 2:
[1357] Speech-to-text conversion
[1358] The device sends the stored audio data to the server. The input is locally stored audio data, and the output is text data. The server uses the Google Cloud Speech-to-Text speech recognition engine to convert the audio data into text data. As a concrete example, the audio data "I'll talk about the design of the new product" is converted into the text data "I'll talk about the design of the new product".
[1359] Step 3:
[1360] Key points extraction
[1361] The server inputs the converted text data into the BERT natural language processing (NLP) model for analysis. The input is text data, and the output is key information. The BERT model extracts the important parts of the text and summarizes them. As a concrete example, from the text data "I will talk about the design of the new product," the key point "new product design" is extracted.
[1362] Step 4:
[1363] emotion recognition
[1364] The server inputs text and audio data, analyzed by the BERT model, into the emotion recognition engine. The input is the analyzed text and audio data, and the output is emotion information. The emotion recognition engine analyzes and identifies the user's emotions. As a concrete example of its operation, it identifies that the user is "excited" based on the text "I'm going to talk about the design of the new product" and the tone of voice.
[1365] Step 5:
[1366] Displaying the main points
[1367] The server sends back key information and recognized sentiment information to the terminal. The input is key information and sentiment information, and the output is the key information displayed visually. The terminal displays the key information to the user in an appropriate format. As a concrete example of operation, the key information "New product design" and the sentiment information "User is excited" are displayed in a pop-up window.
[1368] Step 6:
[1369] Automatic generation of image diagrams
[1370] When a user verbally describes a use case or workflow, the audio data is also captured and converted into text data. The input is the user's utterance, and the output is a visual diagram. The server generates a visual diagram based on the text data. As a concrete example, the following use case diagram is generated from the statement, "A user accesses the site, logs in, searches for a product, and adds it to their cart":
[1371] [User] --> [Access Site] --> [Login] --> [Search Products] --> [Add to Cart]
[1372] The colors and fonts of the images are adjusted based on emotions and displayed on the device.
[1373] Step 7:
[1374] Summary of the meeting's conclusion
[1375] At the end of the meeting, the user verbally summarizes the content. This audio data is also converted into text data, and the server extracts the key points and generates a visual summary along with relevant illustrations. The input is the user's summary, and the output is the visual summary. As a concrete example, the following visual summary is generated from the statement, "Today we talked about the features of the new product and the marketing strategy for it":
[1376] Features of the new product
[1377] Low cost
[1378] high performance
[1379] Energy saving
[1380] Marketing Strategy
[1381] Online advertising
[1382] social media
[1383] The visual summary will be displayed on your device in PDF format.
[1384] (Application Example 2)
[1385] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[1386] Current autonomous vehicles face challenges in facilitating smooth communication between the driver and the vehicle system, as well as in understanding the driver's emotions in real time and providing appropriate information accordingly. Furthermore, there is a lack of mechanisms to support safe driving by appropriately managing the driver's emotional state. This can lead to delays in driver understanding and reaction, potentially increasing the risk of traffic accidents.
[1387] In Application Example 2, the specific processing performed by the specific processing unit 290 of the data processing device 12 is realized by the following means. In this invention, the server includes means for capturing audio data, means for converting the captured audio data into text data, means for extracting key points from the text data, means for recognizing emotions from the analyzed text data, and means for visually displaying information based on the driver's emotions. This enables smooth communication between the driver and the vehicle system, and by grasping the driver's emotional state in real time, it becomes possible to speed up the driver's understanding and reaction, and support safe driving.
[1388] "Means for capturing audio data" refers to a function that collects the driver's speech and surrounding sounds using microphones or other means, and saves them as digital data.
[1389] "Means for converting captured audio data into text data" refers to a function that uses speech recognition technology to convert collected audio data into text information.
[1390] "Methods for extracting key points from text data" refers to functions that use natural language processing technology to identify and summarize important information and topics from converted text.
[1391] "Means for displaying extracted key points as a summary" refers to a function for displaying extracted key points in a format that is easy for the driver to visually understand, on a display or other display device.
[1392] "Means for recognizing emotions from analyzed text data" refers to a function that automatically analyzes and recognizes the driver's emotional state (e.g., stress, satisfaction, excitement, etc.) based on text data and audio data.
[1393] "Means of visually displaying information based on the driver's emotions" refers to a function that adjusts the format and content of the displayed information according to the recognized emotional information, and provides appropriate visual feedback to the driver.
[1394] "Means for analyzing text data and recognizing use cases and business flows" refers to a function that uses natural language processing technology to extract specific procedures and processes from text data and identify them as use cases and business flows.
[1395] "Means for generating image diagrams based on recognized use cases and business flows" refers to a function that converts extracted process information into visual formats such as diagrams and flowcharts.
[1396] "Means for displaying the generated image diagram" refers to a function for displaying the generated visual diagram on a display or other display device.
[1397] "A means of generating a summary as a single visual by combining extracted points and related illustrations" refers to a function that extracts important points and combines them with related illustrations and diagrams to visually represent them as a single summary.
[1398] "Means for adjusting and displaying summaries based on perceived emotions" refers to a function for appropriately adjusting extracted key points and visual summaries based on the driver's emotional state and displaying them visually.
[1399] This invention relates to a system for autonomous vehicles that grasps the driver's emotional state in real time and enables smooth communication between the driver and the vehicle system. This system assists the driver by combining voice data capture, voice-to-text conversion, key point extraction, emotion recognition, and visual display of information.
[1400] Hardware and software configuration
[1401] hardware
[1402] Microphone: A device used to capture the driver's voice and surrounding sounds.
[1403] Display: A device that visually displays summary information and emotion-based feedback to the driver.
[1404] Vehicle ECU (Electronic Control Unit): A computing unit for real-time data processing.
[1405] software
[1406] Speech recognition engine (e.g., Google Speech Recognition API): Converts speech data into text data.
[1407] Natural language processing engines (e.g., the summarization model in the transformers library): Extract key points from text data.
[1408] Emotion recognition engine (e.g., the sentiment-analysis model in the transformers library): Recognizes emotions from the driver's statements.
[1409] Display control software: Software used to properly display information on a display.
[1410] Overall system processing
[1411] The server temporarily stores the audio data captured by the microphone locally. Then, it converts the audio data into text data using a speech recognition engine. This text data is analyzed by a natural language processing engine to extract key points. These extracted points, along with the driver's emotions, are analyzed by an emotion recognition engine, and feedback is provided to the driver based on this information. The feedback is displayed visually through the display.
[1412] Specific example
[1413] 1. Audio capture and text conversion:
[1414] When the driver says, "Turn right and then be careful at the intersection. A left turn follows," the voice is captured by the microphone. This voice data is then converted into text by a speech recognition engine, which reads, "Turn right and then be careful at the intersection. A left turn follows."
[1415] 2. Extracting key points and recognizing emotions:
[1416] Text data is input into a natural language processing engine (e.g., a summarization model), and key points such as "turn right, pay attention to the intersection, turn left" are extracted. Furthermore, an emotion recognition engine determines that "stress has been detected."
[1417] 3. Display of information:
[1418] Based on the extracted key points and emotion recognition results, information such as "Turn right, pay attention to intersection, turn left" is highlighted on the display. Furthermore, if stress is detected, messages and information encouraging relaxation are provided.
[1419] Example of a prompt:
[1420] "User comment: After turning right, be careful at the intersection. A left turn follows.
[1421] This allows drivers to receive necessary information and appropriate feedback in real time, enabling them to continue driving safely and effectively. This system contributes to improved traffic safety and reduced driver stress.
[1422] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[1423] Step 1:
[1424] The server captures the driver's speech and surrounding sounds using a microphone. The audio data is input and temporarily stored locally in digital format.
[1425] Step 2:
[1426] The server sends the captured audio data to a speech recognition engine, where it is converted into text data. The input is audio data, and the output is text data. The speech recognition engine (e.g., Google Speech Recognition API) analyzes the audio waveform and generates the corresponding text.
[1427] Step 3:
[1428] The server inputs text data into a natural language processing engine and extracts the key points. The input data is text data, and the output data is key point information. The natural language processing engine (e.g., the summarization model in the transformers library) analyzes the content of the text and extracts the important information by summarizing it.
[1429] Step 4:
[1430] The server inputs the extracted text data into an emotion recognition engine to recognize the driver's emotions. The input data is the extracted text data, and the output data is emotion information. The emotion recognition engine (e.g., the sentiment-analysis model from the transformers library) analyzes the emotions contained in the text and identifies emotional states such as stress, excitement, and satisfaction.
[1431] Step 5:
[1432] The server integrates extracted key information and recognized sentiment information and transmits it to the display device in a visual format. Input data consists of key information and sentiment information, while output data is visual feedback. Display control software determines the display format based on this data and provides appropriate information to the driver.
[1433] As a concrete example, if the driver says, "Turn right, then be careful at the intersection. Then turn left," in step 1, the voice data is captured, and in step 2, the speech recognition engine converts it into text data. In step 3, the key points "turn right, be careful at the intersection, turn left" are extracted, and in step 4, "stress" is recognized. Finally, in step 5, the important points are highlighted on the display, and if stress is detected, a message encouraging relaxation is added.
[1434] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[1435] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1436] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[1437] [Fourth Embodiment]
[1438] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[1439] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1440] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1441] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[1442] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[1443] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[1444] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[1445] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[1446] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[1447] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1448] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1449] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[1450] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1451] This invention is a system designed to reduce misunderstandings in meetings and enable efficient information sharing. The system captures audio data, converts it into text data, and extracts key points, visually and concisely presenting important information and key points spoken by the user. The specific processing and functions are described below.
[1452] Specific processes and functions of the system
[1453] 1. Capture audio data
[1454] When a user speaks during a meeting, the device captures their voice using the microphone. The captured audio data is temporarily stored locally. This audio data is then converted into text data in the next step.
[1455] 2. Converting speech to text
[1456] Audio data stored on the device is sent to a server via the network. The server converts the audio into text data using a speech recognition engine (e.g., a standard speech recognition API). This process converts the audio data into text format, making subsequent processing easier.
[1457] 3. Key points extraction
[1458] The server inputs the converted text data into a natural language processing (NLP) model (e.g., a BERT model) for analysis. The NLP model extracts the main points from the text and identifies key information. The extracted points are concisely summarized, clearly indicating the main agenda and conclusions of the meeting.
[1459] 4. Displaying the main points
[1460] The server sends summarized key points back to the terminal. The terminal visually displays the summarized points to the user. This display is done in the form of a text box or pop-up window, allowing the user to quickly see the important points.
[1461] 5. Automatic generation of image diagrams
[1462] When a user verbally describes a use case or business workflow, the audio data is also captured and converted into text data. Based on the converted text data, the server analyzes the use case or business workflow recognized by the system and automatically generates a visual diagram. The generated diagram is displayed on the user's terminal in a format that is easy for the user to understand visually.
[1463] 6. Summary of meeting content at the end
[1464] At the end of a meeting, users verbally summarize its content. This audio data is also captured and converted into text. The server uses this text data to organize the key points and combines them with relevant illustrations to generate a single visual summary. This visual summary visually summarizes the meeting's discussion and can be displayed on the device as a PDF or image for easy reference.
[1465] Specific example
[1466] Examples of automatic summary generation
[1467] When a user says, "The features of this product are its low cost and high performance. It also has excellent energy-saving capabilities," the audio data is captured and converted into text data. The server extracts the key points, "low cost, high performance, and energy saving," and sends them back to the terminal for display.
[1468] Examples of automated image generation
[1469] When a user describes a user visiting a website, logging in, searching for products, and adding them to their cart, the audio data is converted to text data, and the server generates a use case diagram like the following:
[1470] [User] --> [Access Site] --> [Login] --> [Search Products] --> [Add to Cart]
[1471] This diagram will be displayed on the device and will be in a format that is easy for the user to understand visually.
[1472] Examples of content summaries
[1473] When a user summarizes the meeting at the end by saying, "Today we talked about the features of the new product and the marketing strategy for it," this audio data is converted into text data, and the server generates a visual summary like the following:
[1474] Features of the new product
[1475] Low cost
[1476] high performance
[1477] Energy saving
[1478] Marketing Strategy
[1479] Online advertising
[1480] social media
[1481] This visual summary will be displayed on the device as a graphic recording in PDF or image format.
[1482] As described above, this system automatically captures speech during meetings, converts it into text data, extracts key points, and displays them visually, allowing users to easily grasp important information. Furthermore, it generates visual diagrams of use cases and business workflows, providing a visualized version of the content at the end of the meeting, thereby reducing misunderstandings and enabling efficient information sharing.
[1483] The following describes the processing flow.
[1484] 1. Processing steps for automatic key point generation
[1485] Step 1:
[1486] User: "I will verbally state the important points."
[1487] Step 2:
[1488] Terminal: Captures the user's voice using the microphone and temporarily saves it as an audio file locally.
[1489] Step 3:
[1490] Terminal: Sends saved audio data to the server via the network.
[1491] Step 4:
[1492] Server: Inputs the received audio data into the speech recognition engine and converts the audio into text data.
[1493] Step 5:
[1494] Server: Inputs the converted text data into a natural language processing (NLP) model and performs analysis.
[1495] Step 6:
[1496] Server: Uses NLP models to extract key points and essential information from text data.
[1497] Step 7:
[1498] Server: Organizes the extracted key points into a summary and condenses them into concise text.
[1499] Step 8:
[1500] Server: Sends the generated summary back to the terminal.
[1501] Step 9:
[1502] Terminal: Displays summarized key information to the user. Specifically, it uses formats such as text boxes and pop-up windows.
[1503] 2. Processing steps for automatic image diagram generation
[1504] Step 1:
[1505] User: "I will verbally explain the use cases and business processes."
[1506] Step 2:
[1507] Terminal: Captures the user's voice using the microphone and temporarily saves it as an audio file locally.
[1508] Step 3:
[1509] Terminal: Sends saved audio data to the server via the network.
[1510] Step 4:
[1511] Server: Inputs the received audio data into the speech recognition engine and converts the audio into text data.
[1512] Step 5:
[1513] Server: Inputs the converted text data into a natural language processing (NLP) model and performs analysis.
[1514] Step 6:
[1515] Server: Uses an NLP model to extract information about use cases and business processes from text data.
[1516] Step 7:
[1517] Server: Based on the extracted information, it automatically generates an image diagram using the business flow design library.
[1518] Step 8:
[1519] Server: Sends the generated image back to the terminal as image data.
[1520] Step 9:
[1521] Terminal: Displays an image diagram to the user. Specifically, it uses a graphical viewer or image display window.
[1522] 3. Steps for summarizing the meeting content at the end of the meeting
[1523] Step 1:
[1524] User: "Summarize the meeting content verbally."
[1525] Step 2:
[1526] Terminal: Captures the user's voice using the microphone and temporarily saves it as an audio file locally.
[1527] Step 3:
[1528] Terminal: Sends saved audio data to the server via the network.
[1529] Step 4:
[1530] Server: Inputs the received audio data into the speech recognition engine and converts the audio into text data.
[1531] Step 5:
[1532] Server: Inputs the converted text data into a natural language processing (NLP) model and performs analysis.
[1533] Step 6:
[1534] Server: Uses NLP models to extract key points and essential information from text data.
[1535] Step 7:
[1536] Server: Organizes extracted points and generates a visual summary combining them with relevant illustrations.
[1537] Step 8:
[1538] Server: Sends the generated visual summary back to the terminal as a PDF or image file.
[1539] Step 9:
[1540] Terminal: Displays a visual summary to the user. Specifically, it uses a PDF viewer or an image display window.
[1541] The above outlines the specific processing steps of this system, which is a mechanism for efficiently and visually organizing and presenting comments and content from meetings to the user.
[1542] (Example 1)
[1543] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1544] In today's business environment, many meetings are conducted remotely, but misunderstandings often occur during voice communication. Therefore, there is a need for efficient ways to share meeting content and refer to it later. There is also a demand for easily visualizing and reviewing verbal explanations of use cases and business flows. However, current technologies often separate speech recognition and key point extraction, and automatic generation of diagrams is limited to certain special cases. This makes it difficult to accurately grasp the key points of a meeting.
[1545] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[1546] In this invention, the server includes means for capturing audio data, means for converting the captured audio data into text data using conversion means, and means for extracting key points from the text data. This makes it possible to convert audio data into text, extract key points, and display them visually.
[1547] "Means for capturing audio data" refers to methods for acquiring the voice spoken by a user as digital data through an input device such as a microphone.
[1548] "Means for converting captured audio data into text data using conversion means" refers to means for converting captured audio data into text information using speech recognition technology.
[1549] "Methods for extracting key points from text data" refers to methods that utilize natural language processing techniques to select and extract important information from text data.
[1550] "Means of displaying as a summary" refers to a method of visualizing the extracted key points and presenting them to the user.
[1551] "Means for recognizing use cases and business flows" refers to methods for analyzing text data to understand and recognize specific usage scenarios and business procedures.
[1552] "Methods for generating image diagrams" refer to methods for automatically creating visually easy-to-understand diagrams based on recognized use cases and business flows.
[1553] A "means for generating visual summaries" is a method for creating a visually easy-to-understand summary by combining points extracted from text data with relevant illustrations.
[1554] This invention is a system for reducing misunderstandings in meetings and achieving efficient information sharing. This system visually and concisely presents important information and key points spoken by users through a series of processes: capturing audio data, converting it into text data, and extracting key points. Specific embodiments for carrying out this invention are described below.
[1555] Audio data capture
[1556] When a user speaks during a meeting, the device captures their voice using its microphone. The captured audio data is temporarily stored locally. Suitable devices for this purpose include computers and smartphones with built-in microphones.
[1557] Sending audio data
[1558] The terminal compresses the captured audio data and sends it to the server using a secure communication protocol (e.g., HTTPS). Compressing the audio data reduces the communication load.
[1559] Speech-to-text conversion
[1560] The server converts the received audio data into text data using a standard speech recognition engine (e.g., Google Cloud Speech-to-Text API). This conversion makes the audio data into text format, which simplifies the next processing step.
[1561] Extracting key points
[1562] The server uses a generative AI model (e.g., the BERT model) to extract key points from the transformed text data. Natural language processing techniques are used to select important information from the text data.
[1563] Return and display of key points
[1564] The server converts the extracted key points into JSON format and sends them back to the terminal. The terminal visually displays the received key points in a text box or pop-up window on the screen. This allows the user to quickly confirm the important points.
[1565] Automatic generation of image diagrams
[1566] When a user verbally explains a use case or business workflow, the audio data is captured and sent to the server. The server converts the audio into text data and uses an AI model for illustration generation (e.g., Diagrams.net API) to automatically generate diagrams of the use case or business workflow. The generated diagrams are displayed on the terminal in a format that is easy for the user to understand visually.
[1567] Summary of the meeting's conclusion
[1568] At the end of a meeting, users verbally summarize its content. This audio data is also captured and converted into text. The server uses this text data to organize the key points and combines them with relevant illustrations to generate a visual summary. The generated visual summary is displayed on the device in PDF or image format for easy reference.
[1569] Specific example
[1570] Examples of automatic summary generation
[1571] When a user says, "The features of this product are its low cost and high performance. It also has excellent energy-saving capabilities," the audio data is captured and converted into text data. The server extracts the key points, "low cost, high performance, and energy saving," and sends them back to the terminal for display.
[1572] Examples of automated image generation
[1573] When a user describes a user visiting a website, logging in, searching for products, and adding them to their cart, the audio data is converted to text data, and the server generates a use case diagram like the following:
[1574] Users access the site, log in, search for products, and add them to their cart.
[1575] Example of a summary of meeting content at the end
[1576] When a user summarizes the meeting at the end by saying, "Today we talked about the features of the new product and the marketing strategy for it," this audio data is converted into text data, and the server generates a visual summary like the following:
[1577] Features of the new product
[1578] Low cost
[1579] high performance
[1580] Energy saving
[1581] Marketing Strategy
[1582] Online advertising
[1583] social media
[1584] The generated visual summary will be displayed on the device in PDF format.
[1585] As described above, this system includes a series of processes to reduce misunderstandings during meetings and enable efficient information sharing. High accuracy and efficiency are achieved by using specific hardware and software in each step. Furthermore, the specific operation of the system is explained through concrete examples.
[1586] The flow of the specific processing in Example 1 will be explained using Figure 11.
[1587] Step 1: Capture audio data
[1588] When a user speaks during a meeting, the device captures their voice using its microphone. The input is the user's voice, and the output is digitized audio data. Specifically, the device's microphone detects the voice in real time and temporarily saves it as an audio file (e.g., in WAV format).
[1589] Step 2: Sending the audio data
[1590] The device compresses the captured audio data and sends it to the server using a secure communication protocol (e.g., HTTPS). The input is the captured audio data, and the output is the audio data sent to the server. Specifically, the device compresses the audio data (e.g., converts it to AAC format) and sends a POST request to the API endpoint.
[1591] Step 3: Convert speech to text
[1592] The server converts the received audio data into text data using a standard speech recognition engine (e.g., Google Cloud Speech-to-Text API). The input is the received audio data, and the output is text data. Specifically, the server sends an API request to the speech recognition service and retrieves the result in text format (e.g., JSON format).
[1593] Step 4: Extracting the key points
[1594] The server uses a generative AI model (e.g., a BERT model) to extract key points from the transformed text data. The input is text data, and the output is a list of extracted key points. Specifically, the server inputs text data into the model and organizes the output key points in list format.
[1595] Step 5: Return and display the key points.
[1596] The server converts the extracted key points into JSON data and sends it back to the terminal. The terminal displays the received key points on the screen in the form of text boxes or pop-up windows. The input is the extracted key points in JSON data, and the output is the visual information presented to the user. Specifically, the terminal parses the JSON data received from the server and updates the UI using HTML / CSS to display it on the screen.
[1597] Step 6: Automatic generation of image diagrams
[1598] When a user verbally describes a use case or business process, the audio data is captured and sent to the server. The server converts the audio into text data and uses an illustration generation AI model (e.g., Diagrams.net API) to automatically generate image diagrams of the use case or business process. The input is text data, and the output is the generated image diagram. Specifically, the server analyzes the text data and inputs it as prompt text into the generation AI model. For example, a prompt text might be: "The user accesses the site, logs in, searches for a product, and adds it to the cart."
[1599] Step 7: Summary of meeting content at the end
[1600] At the end of a meeting, users verbally summarize its content. This audio data is also captured and converted into text data. The server uses this text data to organize the key points and combines them with relevant illustrations to generate a visual summary. The input is the audio summary data from the end of the meeting, and the output is the visual summary. Specifically, the server combines the text data and illustrations, generates a visual summary using a PDF creation library (e.g., Apache PDFBox), and sends it to the terminal. The PDF is then displayed on the terminal.
[1601] (Application Example 1)
[1602] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1603] In modern factories, frequent misunderstandings in information sharing during meetings and production flow setup are a major challenge. This is especially true in environments where verbal explanations are frequently used, where information is often misinterpreted or vaguely remembered. Furthermore, the difficulty in visually grasping meeting points and workflows hinders efficient information sharing. This can lead to delays and errors, potentially reducing productivity.
[1604] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[1605] In this invention, the server includes means for capturing audio data, means for converting the captured audio data into text data, means for extracting key points from the text data, means for displaying the extracted key points as a summary, means for recognizing use cases and business flows, means for generating image diagrams based on the recognized use cases and business flows, means for displaying the generated summary and image diagrams on factory monitors or autonomous device displays, and means for extracting key points at the end of a meeting and generating a related visual summary. This makes it possible to improve the efficiency of information sharing by accurately transcribing audio information during a meeting into text and extracting key points. In addition, by visually displaying business flows and use cases, all participants can proceed with their work based on the same understanding, reducing information discrepancies. At the end of a meeting, the key points can be visualized and summarized to prevent information from being overlooked.
[1606] "Audio data" refers to information recorded in digital format.
[1607] "Means of capturing" refers to methods and devices for acquiring audio as digital data.
[1608] "Means of converting to text data" refers to methods and technologies for converting audio data into text-based data.
[1609] "Methods for extracting key points" refer to techniques for identifying and extracting particularly important parts from texts or conversations.
[1610] "Means of displaying as a summary" refers to a method of visually presenting extracted important information in a concise format.
[1611] A "use case" refers to a specific example of use or procedure under particular circumstances.
[1612] A "business process flow" is a diagram that shows the series of steps and procedures involved in performing a business task.
[1613] "Means of recognition" refer to the techniques and methods for identifying and understanding specific patterns or information.
[1614] "Means for generating image diagrams" refers to methods and techniques for creating visual shapes and charts based on text data and analysis results.
[1615] A "factory monitor" refers to a display device installed within a factory for displaying information.
[1616] An "autonomous device display" is a display device installed in a device that operates automatically.
[1617] "Meeting conclusions" refers to the main points or conclusions summarized at the end of a meeting.
[1618] A "visual summary" is a concise summary that visually presents key points and important information using diagrams and charts.
[1619] This invention is a system designed to streamline meetings and production flow setup in factories and reduce information sharing discrepancies. Through a series of processes—capturing audio data, converting it to text data, and extracting key points—this system visually and concisely presents important information and workflows spoken by users. The specific processing and functions of the system are described below.
[1620] 1. Capture audio data
[1621] When a user speaks during a meeting, the device captures their voice using its microphone. The captured audio data is temporarily stored locally. This audio data is then converted into text data in the next step. This process is achieved by using a device equipped with a microphone and recording capabilities.
[1622] 2. Converting speech to text
[1623] Audio data stored on the device is sent to a server via the network. The server converts the audio into text data using a speech recognition engine. Specific software used in this process includes the Google Speech-to-Text API. This converts the audio data into text format, making subsequent processing easier.
[1624] 3. Key points extraction
[1625] The server analyzes the converted text data using a natural language processing (NLP) model. The BERT model is a possible NLP model used here. The server extracts the main points from the text and identifies important information. This point extraction clearly identifies the main topics and conclusions discussed in the meeting.
[1626] 4. Displaying the main points
[1627] The server sends summarized key information back to the terminal. The terminal visually displays the summarized key points to the user. This display is shown in the form of text boxes or pop-up windows on monitors or autonomous device displays within the factory. This allows the user to quickly identify the important points.
[1628] 5. Automatic generation of image diagrams
[1629] When a user verbally describes a use case or business workflow, the audio data is also captured and converted into text data. Based on this converted text data, the server analyzes the recognized use case or business workflow and automatically generates a visual diagram. A generative AI model may be used in this process. The generated diagram is displayed on the user's device in a format that is easy for them to understand visually.
[1630] 6. Summary of meeting content at the end
[1631] At the end of a meeting, users verbally summarize its content. This audio data is also captured and converted into text. The server uses this text data to organize the key points and generates a single visual summary along with relevant images. This visual summary is displayed on the device as a graphic recording in PDF or image format, allowing users to easily refer to it.
[1632] Specific example
[1633] For example, if a factory worker says, "The procedure for setting up the new production line is to first transport the parts, then assemble them, and then conduct quality checks," the key points generated will be "parts transport, assembly, quality check." Based on this, the server will generate a flowchart and display an image showing "parts transport -> assembly -> quality check." Also, if the meeting concludes with a summary such as, "Today we talked about the procedure for setting up the new line and the points to note," a visual summary of those key points will be generated.
[1634] Example of a prompt
[1635] "The following audio data concerns setting up a new production line. Please extract the key points and generate the production line setup procedure. 'The procedure for setting up the new production line is to first transport the parts, then assemble them, and then conduct quality control.'"
[1636] By following the above steps, it is possible to utilize this method in meetings and production flow settings on the factory floor, significantly improving the efficiency of information sharing.
[1637] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[1638] Step 1:
[1639] Audio data capture
[1640] The device uses a microphone to capture audio during the meeting. The input is the audio signal from the meeting, and the output is digital audio data. This audio data is temporarily stored in local storage. Specifically, the microphone built into the device picks up the audio signal, converts it to a digital format, and stores it.
[1641] Step 2:
[1642] Speech-to-text conversion
[1643] The device sends the captured audio data to the server over the network. The server converts the audio data into text data using the Google Speech-to-Text API. The input is digital audio data, and the output is text data. This is achieved by the server passing the audio data to the API and receiving the resulting text data.
[1644] Step 3:
[1645] Extracting key points
[1646] The server inputs the converted text data into a natural language processing (NLP) model (such as the BERT model) to analyze and extract key points. The input is text data, and the output is text data of the extracted key points. This is achieved by the server using the NLP model to analyze the data and extract important information.
[1647] Step 4:
[1648] Displaying the main points
[1649] The server returns summarized key information to the terminal. The terminal displays this information on factory monitors or autonomous device displays. The input is the extracted key points as text data, and the output is the key points displayed visually. The data is displayed visually using the terminal's display control software.
[1650] Step 5:
[1651] Recognition of use cases and business processes
[1652] When a user verbally describes a use case or business process, the audio data is captured and converted to text. The server analyzes the converted text data to recognize the use case or business process. The input consists of the audio data and the converted text data, and the output is the text data of the recognized use case or business process.
[1653] Step 6:
[1654] Automatic generation of image diagrams
[1655] The server generates visual diagrams based on recognized use cases and business workflows. The input is recognized text data, and the output is a visual diagram. A generation AI model is used to convert the data into flowcharts and charts for display.
[1656] Step 7:
[1657] Display of an image diagram
[1658] The server sends the generated image back to the terminal. The terminal displays this image on a factory monitor or an autonomous device's display. The input is the image, and the output is the visually displayed image.
[1659] Step 8:
[1660] Summary of the meeting's conclusion
[1661] At the end of a meeting, the user verbally summarizes its contents. This audio data is also captured and converted into text data. The server uses this text data to organize the key points and combines them with relevant images to generate a visual summary. The input is the user's summary audio and text data, and the output is the visual summary.
[1662] Step 9:
[1663] Display a visual summary
[1664] The server sends the generated visual summary back to the terminal. The terminal displays this in PDF or image format on factory monitors or autonomous machine displays. The input is the visual summary data, and the output is the visually displayed visual summary.
[1665] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[1666] This invention is a system designed to reduce misunderstandings in meetings and enable efficient information sharing. Furthermore, it supports more accessible and understandable communication by recognizing user emotions and displaying or generating information based on those emotions. This system combines a series of processes—capturing audio data, converting it to text data, and extracting key points—with an emotion engine to visually and concisely present important information and key points expressed by the user. The specific processing and functions are described below.
[1667] Specific processes and functions of the system
[1668] 1. Capture audio data
[1669] When a user speaks during a meeting, the device captures their voice using the microphone. The captured audio data is temporarily stored locally. This audio data is then converted into text data in the next step.
[1670] 2. Converting speech to text
[1671] The audio data stored on the device is sent to the server via the network. The server converts the audio into text data using a speech recognition engine. This process simplifies subsequent processing because the audio data is converted into text format.
[1672] 3. Key points extraction
[1673] The server inputs the converted text data into a natural language processing (NLP) model for analysis. The NLP model extracts the key points from the text and identifies important information. The extracted points are concisely summarized, clearly indicating the main agenda and conclusions of the meeting.
[1674] 4. Emotion recognition
[1675] The server uses an emotion engine to recognize the user's emotions based on text and audio data analyzed by an NLP model. The emotion engine identifies the user's emotions based on factors such as voice tone and facial expression analysis, and evaluates that emotional information.
[1676] 5. Displaying the main points
[1677] The server sends summarized key points and recognized sentiment information back to the terminal. The terminal visually displays the summarized key points to the user in an appropriately adjusted format. This display is presented in the form of a text box or pop-up window, and is done in a way that is sensitive to the user's emotions.
[1678] 6. Automatic generation of image diagrams
[1679] When a user verbally describes a use case or business workflow, the audio data is also captured and converted into text data. Based on the converted text data, the server analyzes the use case or business workflow recognized by the system and automatically generates a visual diagram. The generated diagram is appropriately adjusted based on the user's feelings and displayed on the terminal in a visually easy-to-understand format.
[1680] 7. Summary of meeting content at the end
[1681] At the end of a meeting, users verbally summarize its content. This audio data is also captured and converted into text. The server uses this text data to organize the key points and combines them with relevant illustrations to generate a single visual summary. This visual summary is tailored to the user's emotions and provides a visual representation of the meeting's discussion. It can be displayed on the device as a PDF or image for easy reference.
[1682] Specific example
[1683] Examples of automatic summary generation
[1684] When a user says, "The features of this product are low cost and high performance. It also has excellent energy efficiency," the audio data is captured and converted into text data. The server extracts the key points "low cost, high performance, and energy efficiency," and an emotion engine recognizes the user's emotions. If the emotion engine determines that "the features of this product are very important," a summary is sent back to the terminal and displayed to the user.
[1685] Examples of automated image generation
[1686] When a user describes a user visiting a website, logging in, searching for products, and adding them to their cart, the audio data is converted to text data, and the server generates a use case diagram like the following:
[1687] [User] --> [Access Site] --> [Login] --> [Search Products] --> [Add to Cart]
[1688] If the emotion engine determines that the user is excited, the colors and fonts of the diagram are adjusted and displayed in a visually easy-to-understand format.
[1689] Examples of content summaries
[1690] When a user summarizes the meeting at the end by saying, "Today we talked about the features of the new product and the marketing strategy for it," this audio data is converted into text data, and the server generates a visual summary like the following:
[1691] Features of the new product
[1692] Low cost
[1693] high performance
[1694] Energy saving
[1695] Marketing Strategy
[1696] Online advertising
[1697] social media
[1698] If the emotion engine recognizes that "the user is satisfied," the summary will be displayed with colors and a design that reflects the user's level of satisfaction.
[1699] As described above, this system automatically captures speech during meetings, converts it into text data, extracts key points, and displays them to the user in combination with sentiment recognition, enabling deeper insights and more flexible communication. Furthermore, it generates visual diagrams of use cases and business flow explanations, and provides a visualized version of the content at the end of the meeting, thereby reducing misunderstandings and facilitating efficient information sharing.
[1700] The following describes the processing flow.
[1701] 1. Processing steps for automatic summary generation and sentiment recognition
[1702] Step 1:
[1703] User: "I will verbally state the important points."
[1704] Step 2:
[1705] Terminal: Captures the user's voice using the microphone and temporarily saves it as an audio file locally.
[1706] Step 3:
[1707] Terminal: Sends saved audio data to the server via the network.
[1708] Step 4:
[1709] Server: Inputs the received audio data into the speech recognition engine and converts the audio into text data.
[1710] Step 5:
[1711] Server: Inputs the converted text data into a natural language processing (NLP) model and performs analysis.
[1712] Step 6:
[1713] Server: Uses NLP models to extract key points and essential information from text data.
[1714] Step 7:
[1715] Server: Simultaneously inputs voice and text data into the emotion engine to recognize the user's emotions.
[1716] Step 8:
[1717] Server: Combines extracted key points and recognized sentiment information to generate a summary.
[1718] Step 9:
[1719] Server: Sends the generated summary and sentiment information back to the terminal.
[1720] Step 10:
[1721] Terminal: Displays summarized key information and emotional information to the user in a visually adjusted format. For example, if the emotion is "emotional," the text color and font are adjusted.
[1722] 2. Image diagram automatic generation and emotion recognition processing steps
[1723] Step 1:
[1724] User: "I will verbally explain the use cases and business processes."
[1725] Step 2:
[1726] Terminal: Captures the user's voice using the microphone and temporarily saves it as an audio file locally.
[1727] Step 3:
[1728] Terminal: Sends saved audio data to the server via the network.
[1729] Step 4:
[1730] Server: Inputs the received audio data into the speech recognition engine and converts the audio into text data.
[1731] Step 5:
[1732] Server: Inputs the converted text data into a natural language processing (NLP) model and performs analysis.
[1733] Step 6:
[1734] Server: Uses an NLP model to extract information about use cases and business processes from text data.
[1735] Step 7:
[1736] Server: Simultaneously inputs voice and text data into the emotion engine to recognize the user's emotions.
[1737] Step 8:
[1738] Server: Based on extracted use cases and business flows, it automatically generates image diagrams using the business flow design library.
[1739] Step 9:
[1740] Server: Adjusts the generated image diagram appropriately based on recognized emotional information (e.g., color scheme and layout).
[1741] Step 10:
[1742] Server: Sends the adjusted image back to the terminal as image data.
[1743] Step 11:
[1744] Terminal: Displays an image to the user. For example, if the emotion is "excitement," the layout and graphical elements of the image will become more vivid.
[1745] 3. Summary of meeting content and processing steps for emotional recognition at the end of the meeting
[1746] Step 1:
[1747] User: "Summarize the meeting content verbally."
[1748] Step 2:
[1749] Terminal: Captures the user's voice using the microphone and temporarily saves it as an audio file locally.
[1750] Step 3:
[1751] Terminal: Sends saved audio data to the server via the network.
[1752] Step 4:
[1753] Server: Inputs the received audio data into the speech recognition engine and converts the audio into text data.
[1754] Step 5:
[1755] Server: Inputs the converted text data into a natural language processing (NLP) model and performs analysis.
[1756] Step 6:
[1757] Server: Uses NLP models to extract key points and essential information from text data.
[1758] Step 7:
[1759] Server: Simultaneously inputs voice and text data into the emotion engine to recognize the user's emotions.
[1760] Step 8:
[1761] Server: Organizes extracted points and combines them with relevant illustrations to generate a single visual summary.
[1762] Step 9:
[1763] Server: Appropriately adjusts the generated visual summary based on recognized sentiment information (e.g., color scheme and layout).
[1764] Step 10:
[1765] Server: Sends the adjusted visual summary back to the terminal as a PDF or image file.
[1766] Step 11:
[1767] Terminal: Displays a visual summary to the user. For example, if the emotion is "satisfied," the visual's color scheme will be warm tones.
[1768] The above outlines the specific processing steps of this system, which is a mechanism for efficiently and visually organizing and presenting meeting comments and content to the user. By using emotion recognition to display and generate information in accordance with the user's emotions, more effective communication becomes possible.
[1769] (Example 2)
[1770] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1771] Traditional meeting systems struggle to accurately grasp the key points of what is said and display them visually in an easy-to-understand manner. Furthermore, they often fail to consider user emotions when displaying information, leading to one-way communication. Additionally, they are inefficient at providing concise visual summaries of meeting discussions, resulting in low information sharing efficiency.
[1772] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[1773] In this invention, the server includes means for capturing audio data, means for converting the captured audio data into text data, means for utilizing a natural language processing model to extract key points from the text data, means for using an emotion recognition engine to recognize the extracted key points and the user's emotions, means for displaying the key points based on the recognized emotion information, means for automatically generating an image diagram by converting the audio data into text data and performing analysis, means for adjusting and displaying the generated image diagram based on the user's emotions, and means for visualizing and displaying the summarized content at the end of the meeting. This enables accurate understanding of what the user said, information display that takes emotions into account, and reduces misunderstandings and efficient information sharing by visually summarizing the meeting discussion.
[1774] "Audio data" refers to the waveform information of sounds that record what a user says.
[1775] "Capturing" refers to the act of recording and acquiring information such as audio data.
[1776] "Text data" refers to audio data converted into written information.
[1777] A "natural language processing model" is a model that uses machine learning techniques to analyze, understand, and generate human language.
[1778] "Key points" refer to important information or content extracted from text data.
[1779] An "emotion recognition engine" is a device or software that analyzes voice data and text data to identify the user's emotions.
[1780] An "image diagram" refers to a visual explanatory diagram or flowchart generated based on text data.
[1781] A "summary" is a way of putting long texts or complex information into a concise format.
[1782] A "visual summary" is a visual representation of summarized information, and may include graphics, diagrams, and illustrations.
[1783] A "meeting" refers to a gathering of multiple users to discuss and share information.
[1784] This invention is a system designed to reduce misunderstandings in meetings and enable efficient information sharing. Furthermore, it supports more accessible and understandable communication by recognizing user emotions and displaying or generating information based on those emotions. This system combines a series of processes—capturing audio data, converting it into text data, and extracting key points—with an emotion recognition engine to visually and concisely present important information and key points expressed by the user.
[1785] Hardware and software to be used
[1786] Device: A device equipped with a microphone for the user to speak. This includes PCs, smartphones, tablets, etc.
[1787] Server: A backend system for converting and analyzing audio data. High processing power is desirable.
[1788] Speech recognition engine: Uses speech recognition services such as Google Cloud Speech-to-Text.
[1789] Natural language processing models: Advanced NLP models such as BERT.
[1790] Emotion recognition engine: An emotion analysis system such as IBM Watson.
[1791] System Description
[1792] 1. Capture audio data
[1793] When a user speaks during a meeting, the device captures their voice using the microphone. The captured audio data is temporarily stored in local storage and converted into text data in the next step.
[1794] 2. Converting speech to text
[1795] The device sends locally stored audio data to the server over the network. The server uses Google Cloud Speech-to-Text's speech recognition engine to convert the audio data into text data.
[1796] 3. Key points extraction
[1797] The server inputs text data into the BERT natural language processing (NLP) model for analysis. The NLP model extracts the main points of the text and identifies important information.
[1798] 4. Emotion recognition
[1799] The server inputs text and audio data, analyzed by an NLP model, into IBM Watson's emotion recognition engine to identify the user's emotions. The emotion recognition engine analyzes the tone and metadata of the voice to obtain information about the user's emotions.
[1800] 5. Displaying the main points
[1801] The server returns summarized key information and recognized sentiment information to the terminal. The terminal visually displays this information in an appropriately formatted way and presents it to the user. The display format is provided as a text box or a pop-up window.
[1802] 6. Automatic generation of image diagrams
[1803] When a user verbally describes a use case or workflow, the audio data is captured and converted into text. The server analyzes the converted text data and automatically generates a visual diagram. The generated diagram is then displayed on the device, with its colors and fonts adjusted based on the user's mood.
[1804] 7. Summary of meeting content at the end
[1805] At the end of the meeting, users verbally summarize the content. This audio data is also converted into text data, and the server extracts the key points and generates a visual summary along with relevant illustrations. The visual summary is displayed on the device in PDF or image format.
[1806] Specific example
[1807] 1. Specific examples of automatic summary generation
[1808] When a user says, "The features of this product are low cost and high performance. It also has excellent energy efficiency," the audio data is captured and converted into text data. The server extracts the key points "low cost, high performance, and energy efficiency," and an emotion recognition engine recognizes the user's emotions. If the emotion engine determines that "the features of this product are very important," a summary is sent back to the terminal and displayed to the user.
[1809] 2. Specific Examples of Automatic Image Generation
[1810] When a user describes a user visiting a website, logging in, searching for products, and adding them to their cart, the audio data is converted to text data, and the server generates a use case diagram like the following:
[1811] [User] --> [Access Site] --> [Login] --> [Search Products] --> [Add to Cart]
[1812] If the emotion engine determines that the user is excited, the colors and fonts of the diagram are adjusted and displayed in a visually easy-to-understand format.
[1813] 3. Specific examples of summaries at the end of a meeting
[1814] When a user summarizes the meeting at the end, saying, "Today we talked about the features of the new product and the marketing strategy for it," the audio data is converted to text data, and the server generates a visual summary like the following:
[1815] Features of the new product
[1816] Low cost
[1817] high performance
[1818] Energy saving
[1819] Marketing Strategy
[1820] Online advertising
[1821] social media
[1822] If the emotion engine recognizes that "the user is satisfied," the summary will be displayed with colors and a design that reflects the user's level of satisfaction.
[1823] The above describes the detailed embodiments for carrying out the invention. This invention automatically captures speech during meetings, converts it into text data, extracts key points, and combines this with sentiment recognition to provide deeper insights and more flexible communication. Furthermore, it generates visual diagrams of use cases and business flow explanations, and provides a visualized version of the content at the end of the meeting, thereby reducing misunderstandings and enabling efficient information sharing.
[1824] The flow of the specific processing in Example 2 will be explained using Figure 13.
[1825] Step 1:
[1826] Audio data capture
[1827] When a user speaks during a meeting, the device captures their voice using the microphone. The input is the user's speech, and the output is the captured audio data. The device temporarily saves this audio data to local storage. As a concrete example, if a user says, "I'm going to talk about the design of the new product," that voice is captured and saved locally.
[1828] Step 2:
[1829] Speech-to-text conversion
[1830] The device sends the stored audio data to the server. The input is locally stored audio data, and the output is text data. The server uses the Google Cloud Speech-to-Text speech recognition engine to convert the audio data into text data. As a concrete example, the audio data "I'll talk about the design of the new product" is converted into the text data "I'll talk about the design of the new product".
[1831] Step 3:
[1832] Key points extraction
[1833] The server inputs the converted text data into the BERT natural language processing (NLP) model for analysis. The input is text data, and the output is key information. The BERT model extracts the important parts of the text and summarizes them. As a concrete example, from the text data "I will talk about the design of the new product," the key point "new product design" is extracted.
[1834] Step 4:
[1835] emotion recognition
[1836] The server inputs text and audio data, analyzed by the BERT model, into the emotion recognition engine. The input is the analyzed text and audio data, and the output is emotion information. The emotion recognition engine analyzes and identifies the user's emotions. As a concrete example of its operation, it identifies that the user is "excited" based on the text "I'm going to talk about the design of the new product" and the tone of voice.
[1837] Step 5:
[1838] Displaying the main points
[1839] The server sends back key information and recognized sentiment information to the terminal. The input is key information and sentiment information, and the output is the key information displayed visually. The terminal displays the key information to the user in an appropriate format. As a concrete example of operation, the key information "New product design" and the sentiment information "User is excited" are displayed in a pop-up window.
[1840] Step 6:
[1841] Automatic generation of image diagrams
[1842] When a user verbally describes a use case or workflow, the audio data is also captured and converted into text data. The input is the user's utterance, and the output is a visual diagram. The server generates a visual diagram based on the text data. As a concrete example, the following use case diagram is generated from the statement, "A user accesses the site, logs in, searches for a product, and adds it to their cart":
[1843] [User] --> [Access Site] --> [Login] --> [Search Products] --> [Add to Cart]
[1844] The colors and fonts of the images are adjusted based on emotions and displayed on the device.
[1845] Step 7:
[1846] Summary of the meeting's conclusion
[1847] At the end of the meeting, the user verbally summarizes the content. This audio data is also converted into text data, and the server extracts the key points and generates a visual summary along with relevant illustrations. The input is the user's summary, and the output is the visual summary. As a concrete example, the following visual summary is generated from the statement, "Today we talked about the features of the new product and the marketing strategy for it":
[1848] Features of the new product
[1849] Low cost
[1850] high performance
[1851] Energy saving
[1852] Marketing Strategy
[1853] Online advertising
[1854] social media
[1855] The visual summary will be displayed on your device in PDF format.
[1856] (Application Example 2)
[1857] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1858] Current autonomous vehicles face challenges in facilitating smooth communication between the driver and the vehicle system, as well as in understanding the driver's emotions in real time and providing appropriate information accordingly. Furthermore, there is a lack of mechanisms to support safe driving by appropriately managing the driver's emotional state. This can lead to delays in driver understanding and reaction, potentially increasing the risk of traffic accidents.
[1859] In Application Example 2, the specific processing performed by the specific processing unit 290 of the data processing device 12 is realized by the following means. In this invention, the server includes means for capturing audio data, means for converting the captured audio data into text data, means for extracting key points from the text data, means for recognizing emotions from the analyzed text data, and means for visually displaying information based on the driver's emotions. This enables smooth communication between the driver and the vehicle system, and by grasping the driver's emotional state in real time, it becomes possible to speed up the driver's understanding and reaction, and support safe driving.
[1860] "Means for capturing audio data" refers to a function that collects the driver's speech and surrounding sounds using microphones and other means, and saves them as digital data.
[1861] "Means for converting captured audio data into text data" refers to a function that uses speech recognition technology to convert collected audio data into text information.
[1862] "Methods for extracting key points from text data" refers to functions that use natural language processing technology to identify and summarize important information and topics from converted text.
[1863] "Means for displaying extracted key points as a summary" refers to a function for displaying extracted key points in a format that is easy for the driver to visually understand, on a display or other display device.
[1864] "Means for recognizing emotions from analyzed text data" refers to a function that automatically analyzes and recognizes the driver's emotional state (e.g., stress, satisfaction, excitement, etc.) based on text data and audio data.
[1865] "Means of visually displaying information based on the driver's emotions" refers to a function that adjusts the format and content of the displayed information according to the recognized emotional information, and provides appropriate visual feedback to the driver.
[1866] "Means for analyzing text data and recognizing use cases and business flows" refers to a function that uses natural language processing technology to extract specific procedures and processes from text data and identify them as use cases and business flows.
[1867] "Means for generating image diagrams based on recognized use cases and business flows" refers to a function that converts extracted process information into visual formats such as diagrams and flowcharts.
[1868] "Means for displaying the generated image diagram" refers to a function for displaying the generated visual diagram on a display or other display device.
[1869] "A means of generating a summary as a single visual by combining extracted points and related illustrations" refers to a function that extracts important points and combines them with related illustrations and diagrams to visually represent them as a single summary.
[1870] "Means for adjusting and displaying summaries based on perceived emotions" refers to a function for appropriately adjusting extracted key points and visual summaries based on the driver's emotional state and displaying them visually.
[1871] This invention relates to a system for autonomous vehicles that grasps the driver's emotional state in real time and enables smooth communication between the driver and the vehicle system. This system assists the driver by combining voice data capture, voice-to-text conversion, key point extraction, emotion recognition, and visual display of information.
[1872] Hardware and software configuration
[1873] hardware
[1874] Microphone: A device used to capture the driver's voice and surrounding sounds.
[1875] Display: A device that visually displays summary information and emotion-based feedback to the driver.
[1876] Vehicle ECU (Electronic Control Unit): A computing unit for real-time data processing.
[1877] software
[1878] Speech recognition engine (e.g., Google Speech Recognition API): Converts speech data into text data.
[1879] Natural language processing engines (e.g., the summarization model in the transformers library): Extract key points from text data.
[1880] Emotion recognition engine (e.g., the sentiment-analysis model in the transformers library): Recognizes emotions from the driver's statements.
[1881] Display control software: Software used to properly display information on a display.
[1882] Overall system processing
[1883] The server temporarily stores the audio data captured by the microphone locally. Then, it converts the audio data into text data using a speech recognition engine. This text data is analyzed by a natural language processing engine to extract key points. These extracted points, along with the driver's emotions, are analyzed by an emotion recognition engine, and feedback is provided to the driver based on this information. The feedback is displayed visually through the display.
[1884] Specific example
[1885] 1. Audio capture and text conversion:
[1886] When the driver says, "Turn right and then be careful at the intersection. A left turn follows," the voice is captured by the microphone. This voice data is then converted into text by a speech recognition engine, which reads, "Turn right and then be careful at the intersection. A left turn follows."
[1887] 2. Extracting key points and recognizing emotions:
[1888] Text data is input into a natural language processing engine (e.g., a summarization model), and key points such as "turn right, pay attention to the intersection, turn left" are extracted. Furthermore, an emotion recognition engine determines that "stress has been detected."
[1889] 3. Display of information:
[1890] Based on the extracted key points and emotion recognition results, information such as "Turn right, pay attention to intersection, turn left" is highlighted on the display. Furthermore, if stress is detected, messages and information encouraging relaxation are provided.
[1891] Example of a prompt:
[1892] "User comment: After turning right, be careful at the intersection. A left turn follows.
[1893] This allows drivers to receive necessary information and appropriate feedback in real time, enabling them to continue driving safely and effectively. This system contributes to improved traffic safety and reduced driver stress.
[1894] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[1895] Step 1:
[1896] The server captures the driver's speech and surrounding sounds using a microphone. The audio data is input and temporarily stored locally in digital format.
[1897] Step 2:
[1898] The server sends the captured audio data to a speech recognition engine, where it is converted into text data. The input is audio data, and the output is text data. The speech recognition engine (e.g., Google Speech Recognition API) analyzes the audio waveform and generates the corresponding text.
[1899] Step 3:
[1900] The server inputs text data into a natural language processing engine and extracts key points. The input data is text data, and the output data is key point information. The natural language processing engine (e.g., the summarization model in the transformers library) analyzes the content of the text and extracts the important information by summarizing it.
[1901] Step 4:
[1902] The server inputs the extracted text data into an emotion recognition engine to recognize the driver's emotions. The input data is the extracted text data, and the output data is emotion information. The emotion recognition engine (e.g., the sentiment-analysis model from the transformers library) analyzes the emotions contained in the text and identifies emotional states such as stress, excitement, and satisfaction.
[1903] Step 5:
[1904] The server integrates extracted key information and recognized sentiment information and transmits it to the display device in a visual format. Input data consists of key information and sentiment information, while output data is visual feedback. Display control software determines the display format based on this data and provides appropriate information to the driver.
[1905] As a concrete example, if the driver says, "Turn right, then be careful at the intersection. Then turn left," in step 1, the voice data is captured, and in step 2, the speech recognition engine converts it into text data. In step 3, the key points "turn right, be careful at the intersection, turn left" are extracted, and in step 4, "stress" is recognized. Finally, in step 5, the important points are highlighted on the display, and if stress is detected, a message encouraging relaxation is added.
[1906] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[1907] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1908] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[1909] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1910] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[1911] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[1912] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[1913] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[1914] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[1915] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[1916] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[1917] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[1918] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[1919] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1920] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[1921] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[1922] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[1923] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[1924] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[1925] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do n...
Claims
1. A means of capturing audio data, A means of converting captured audio data into text data, Methods for extracting key points from text data, A means of displaying the extracted key points as a summary, A system that includes this.
2. A means of capturing audio data, A means of converting captured audio data into text data, A means of analyzing text data to recognize use cases and business flows, A means of generating an image diagram based on recognized use cases and business flows, A means for displaying the generated image diagram, A system that includes this.
3. A means of capturing audio data, A means of converting captured audio data into text data, A method for extracting key points from MTG text data, A method for generating a summary as a single visual by combining extracted points and related illustrations, Means for displaying the generated summary, A system that includes this.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A