system

The system addresses the challenge of understanding technical terms in meetings by real-time transcription and automatic minute generation, improving efficiency and accessibility for all participants.

JP2026062198APending Publication Date: 2026-04-09SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-09-30
Publication Date
2026-04-09

AI Technical Summary

Technical Problem

Participants in meetings, especially those attending for the first time or with hearing impairments, struggle to understand technical terms and difficult words, leading to inefficient understanding and manual transcription of meeting minutes being time-consuming.

Method used

A system that transcribes meeting audio data in real time, saves the transcribed text data, automatically generates meeting minutes, detects and retrieves the meanings of technical terms and difficult words, and presents them to the user using a dictionary database or external API.

Benefits of technology

Enhances meeting efficiency by allowing users to understand meeting content in real time, saving time and effort in manual transcription, and providing an easily understandable environment for first-time participants or those with hearing impairments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026062198000001_ABST
    Figure 2026062198000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] A means for transcribing meeting audio data in real time, Means for saving the transcribed text data, A means for automatically generating meeting minutes based on the transcribed text data, A method for detecting technical terms and difficult words that appear during a meeting and obtaining their meanings, A system that includes means for presenting the meanings of the aforementioned technical terms and difficult words to the user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In recent years, there has been a demand for improving the efficiency of meetings and consultations. As part of this, automatic transcription of meeting content and automatic generation of minutes using speech recognition technology have been emphasized. However, for users who first participate in a specific organization or participants with hearing impairments, it is often difficult to understand technical terms and difficult words that appear during meetings. As a result, they may not be able to fully understand the content of the meeting and may lose opportunities for skill improvement. In addition, since it takes time and effort to create minutes manually, an efficient method is required.

[0005] There is a demand for providing a system that solves such problems and enables users who first participate in a specific organization or participants with hearing impairments to quickly and accurately understand the content of a meeting.

Means for Solving the Problems

[0006] The present invention provides a system for transcribing meeting audio data in real time. This system includes means for transcribing meeting audio data in real time, means for saving the transcribed text data, and means for automatically generating meeting minutes based on the transcribed text data.

[0007] Furthermore, the system has a means to detect and retrieve the meanings of technical terms and difficult words that appear during meetings, as well as a means to present the meanings of these terms to the user. The meanings of technical terms and difficult words can be obtained from a dictionary database or an external API. This allows users to understand the meaning of terms in real time through simple operations such as clicks and taps.

[0008] In addition, meeting audio data is sent to the server at regular intervals and transcribed in real time, allowing users to understand the meeting content in real time. This system improves meeting efficiency and speeds up understanding, providing an easy-to-understand environment for users joining a particular organization for the first time or for participants with hearing impairments.

[0009] "Audio data" refers to digital audio information that records speeches and conversations during meetings, discussions, and other similar events.

[0010] "Real-time" refers to the process of processing and responding to information as it unfolds in a meeting or event.

[0011] "Transcription" is the process of analyzing audio data and converting the spoken content into text data.

[0012] "Text data" refers to digital data in the form of a document, obtained as a result of transcription.

[0013] A "database" is a collection of information that stores and manages systematically structured data, making it easy to search and retrieve.

[0014] "Meeting minutes" are documents that record the content of meetings and discussions, and are created for later reference.

[0015] "Technical terms" are words or expressions used in a specific field or industry that have a particular meaning and are generally unfamiliar to the public.

[0016] "Difficult words" are words or expressions that are difficult for the average person to understand without specialized knowledge.

[0017] A "dictionary database" is an organized collection of data that provides information about the meanings and usages of words.

[0018] An "external API" is an interface for accessing functions and data provided by other systems or services.

[0019] "Users" refer to the people who use this system, and often specifically to the participants in a meeting.

[0020] A "server" is a computer system that processes and provides data in response to requests from client terminals.

[0021] A "device" is a device that a user directly operates (such as a computer, smartphone, or tablet).

[0022] A "dictionary API" is an API provided by a dictionary service, serving as an interface for obtaining information about the meaning and usage of words. [Brief explanation of the drawing]

[0023] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] It is a conceptual diagram showing an example of the main functions of a data processing device and a smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] Shows an emotion map to which multiple emotions are mapped. [Figure 10] Shows an emotion map to which multiple emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] [[ID=3C]]It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when the emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when the emotion engine is combined.

Mode for Carrying Out the Invention

[0024] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0025] First, let's explain the terminology used in the following explanation.

[0026] In the following embodiments, the signed processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Furthermore, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), and APU (Accelerated Processing Unit).

[0027] In the following embodiments, signed RAM (Random Access Memory) is a memory that temporarily stores information and is used as work memory by the processor.

[0028] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.

[0029] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0030] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0031] [First Embodiment]

[0032] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0033] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0034] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0035] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0036] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0037] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0038] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0039] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0040] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0041] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0042] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0043] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0044] The system for implementing the present invention has the following configuration. First, it transcribes the audio data of a meeting in real time and saves the text data. Next, it automatically generates meeting minutes based on the saved text data, and further detects technical terms and difficult words that appear during the meeting, obtains their meanings, and presents them to the user.

[0045] Functional Embodiments

[0046] 1. Collection and transcription of audio data

[0047] The terminal records audio data during the meeting in real time. The recorded audio data is sent to the server in increments of a certain buffer size. The server sends the received audio data to a speech recognition API and converts it into text data. This text data is temporarily stored in the server's memory.

[0048] 2. Saving text data

[0049] The server stores the transcribed text data obtained from the speech recognition API in a database. It is important to add a timestamp and meeting ID to the text data. This is helpful for later searching and extracting specific meeting content.

[0050] 3. Automatic generation of meeting minutes

[0051] After the meeting ends, or at an appropriate time, the server retrieves all text data from the database based on the meeting ID. The server sorts the retrieved text data chronologically and automatically generates meeting minutes according to a standard format. These minutes are stored in the database for later reference by the user.

[0052] 4. Acquisition and presentation of the meanings of technical terms and difficult words.

[0053] The server analyzes the transcribed text data to detect technical terms and difficult words. For detected words, the server retrieves their meaning using a dictionary database or an external API (dictionary API). The retrieved meanings are stored in the database.

[0054] The device displays a pop-up showing the meaning of a word when the user clicks on it in the text data displayed during the meeting. This display is provided in real time, allowing users to understand technical terms as the meeting progresses.

[0055] Specific example

[0056] At a meeting they are attending for the first time, the user hears the word "refactoring."

[0057] 1. The device records the meeting audio and sends the data to the server.

[0058] 2. The server receives "refactoring" as text data via the speech recognition API and stores it temporarily.

[0059] 3. The server saves the text data to the database and simultaneously detects the word "refactoring".

[0060] 4. The server uses the dictionary API to retrieve the meaning of "refactoring" and saves it to the database.

[0061] 5. When a user clicks the word "refactoring" during a meeting, the terminal will display a pop-up with the meaning retrieved from the server.

[0062] This system allows users to understand the progress of a meeting in real time and to instantly check the meaning of technical terms and difficult words. This deepens understanding of the meeting content and provides an environment where users participating in a particular organization for the first time, or participants with hearing impairments, can quickly obtain useful information.

[0063] The following describes the processing flow.

[0064] Step 1:

[0065] The device records the meeting audio in real time. Recording begins when the user launches the meeting recording application during the meeting and presses the start recording button. The recorded data is stored in a buffer.

[0066] Step 2:

[0067] The device sends the recorded data to the server at regular intervals. For example, it clears the buffer every 10 seconds and sends the current recorded data to the server.

[0068] Step 3:

[0069] The server sends the received audio data to a speech recognition API (e.g., Google® Speech-to-Text). The server creates an API request and passes the audio data to the API.

[0070] Step 4:

[0071] The speech recognition API converts the audio data into text and returns it to the server. The API parses the audio data and returns the result in string format.

[0072] Step 5:

[0073] The server receives the text data and temporarily stores it in memory. The text data is stored along with the meeting ID and timestamp.

[0074] Step 6:

[0075] The server saves the text data to the database. The saved data is tagged with a meeting ID and timestamp, making it easy to search later.

[0076] Step 7:

[0077] After the meeting ends, the server retrieves all text data for the corresponding meeting ID from the database. The server uses the meeting ID as a key to extract the relevant text data.

[0078] Step 8:

[0079] The server sorts the acquired text data chronologically and automatically generates meeting minutes according to a standard format. The server concatenates the text data and adds appropriate paragraphs and headings.

[0080] Step 9:

[0081] The server generates meeting minutes and saves them to a database. The saved minutes are then managed so that users can access them later.

[0082] Step 10:

[0083] During the meeting, the server analyzes the text data and detects technical terms and difficult words. The server then uses natural language processing to create a list of these technical terms and difficult words.

[0084] Step 11:

[0085] The server uses a dictionary API to retrieve the meaning of each word it detects. The server sends a request to the dictionary API and receives the returned result.

[0086] Step 12:

[0087] The server stores the meaning of the words it retrieves in a database. The meaning data is also stored in association with the conference ID.

[0088] Step 13:

[0089] The user clicks on a word in the text data displayed during the meeting. The device detects the click event and requests information about that word from the server.

[0090] Step 14:

[0091] The server receives the request and retrieves the meaning of the corresponding word from the database. The server queries the semantic data and returns the results to the terminal.

[0092] Step 15:

[0093] The device displays semantic data in a pop-up window. Users can see the detailed meaning of the clicked word in real time.

[0094] These detailed processing steps enable a system where meeting content is transcribed in real time and the meanings of technical terms are provided immediately.

[0095] (Example 1)

[0096] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0097] Meetings often involve a lot of jargon and complex terminology, making it difficult to understand them in real time. Furthermore, manually creating meeting minutes is time-consuming and laborious, and may lack accuracy. Moreover, real-time understanding is even more challenging for participants with hearing impairments or those attending a meeting for the first time. A system is needed to address these challenges.

[0098] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0099] In this invention, the server includes means for transcribing meeting audio data in real time, means for storing the transcribed text data, means for automatically generating meeting minutes based on the transcribed text data, means for detecting technical terms and difficult words that appear during the meeting, means for obtaining the meanings of the detected technical terms and difficult words using a dictionary database or an external API, means for presenting the meanings of the technical terms and difficult words to the user in real time, and means for transmitting the audio data to the server at regular intervals for transcription. As a result, the meanings of technical terms and difficult words can be checked in real time during the meeting, which facilitates understanding of technical terms and enables the automatic generation of meeting minutes, saving time and effort.

[0100] "Audio data" refers to information recorded in digital format using audio.

[0101] "Transcription" refers to the process of converting audio data into text data.

[0102] "Text data" refers to digital information stored as characters.

[0103] "Meeting minutes" refers to a document that records the content and progress of discussions at a meeting.

[0104] "Technical jargon" refers to specialized terms used in a particular field or industry that are generally difficult for the average person to understand.

[0105] "Difficult words" refer to words with special meanings that are not easily understood by the general public.

[0106] A "dictionary database" refers to a digital information source that stores the meanings and definitions of words.

[0107] An "external API" refers to a program that provides an interface for interacting with other systems and services.

[0108] A "server" refers to a computer system that processes and stores data over a network.

[0109] "Terminal" refers to a computer or smart device that a user directly operates.

[0110] "User" refers to a person who uses this system.

[0111] "Real-time" refers to processing that occurs instantly without delay.

[0112] A "timestamp" refers to information that indicates the time when data was recorded.

[0113] "Standard format" refers to the structure and format of a document that follows certain rules and conventions.

[0114] The system of the present invention transcribes meeting audio data in real time, saves the text data, automatically generates meeting minutes, and presents the meanings of technical terms and difficult words to the user. The system of the present invention uses the following specific hardware and software.

[0115] First, the terminal records the meeting audio data in real time using its built-in or external microphone. The recorded audio data is sent to the server in buffer sizes, such as every 5 seconds. The terminal uses an audio recording library, such as the AudioRecord class, for audio recording.

[0116] The server sends the received audio data to the Google Speech-to-Text API, where it is converted into text data. The converted text data is temporarily stored in the server's memory and then stored in a database. Here, each text data entry is assigned a timestamp and a meeting ID. The database used is a relational database such as MySQL® or PostgreSQL.

[0117] After the meeting ends, the server retrieves all text data from the database based on the meeting ID and sorts it chronologically based on the timestamp. It then automatically generates meeting minutes according to a standard format. These minutes are stored in the database by the server for later reference by the user.

[0118] Furthermore, the server analyzes the transcribed text data to detect technical terms and difficult words. For these words, the server retrieves their meanings using dictionary databases such as the Oxford Dictionaries API or external APIs. The retrieved meanings are then stored in the database. When a user clicks on a word in the text data displayed during a meeting, the terminal displays a pop-up with the meaning of the word retrieved from the server. This allows users to understand technical terms simultaneously with the progress of the meeting.

[0119] As a concrete example, let's consider a scenario where a user hears the word "refactoring" at a meeting they are attending for the first time.

[0120] 1. The device records the meeting audio and sends the data to the server.

[0121] 2. The server receives "refactoring" as text data via the Google Speech-to-Text API and stores it temporarily.

[0122] 3. The server saves the text data to the database and simultaneously detects the word "refactoring".

[0123] 4. The server uses the Oxford Dictionaries API to retrieve the meaning of "refactoring" and saves it to the database.

[0124] 5. When a user clicks the word "refactoring" during a meeting, the terminal will display a pop-up with the meaning retrieved from the server.

[0125] This system allows users to understand the progress of a meeting in real time and to instantly check the meaning of technical terms and difficult words. As a result, understanding of the meeting content is enhanced, and a valuable environment is provided where users, especially those joining a particular organization for the first time or participants with hearing impairments, can quickly obtain useful information.

[0126] Example of a prompt:

[0127] "A meeting audio about a new system has been recorded. Please generate a program that converts the recorded audio data into text data, detects technical terms and difficult words, and displays their meanings. Please specify the names of the speech recognition API and dictionary API to be used, and provide a detailed explanation of the processing steps."

[0128] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0129] Step 1:

[0130] The device records meeting audio data in real time using its built-in microphone or an external microphone. The recorded audio data is buffered at a fixed buffer size (e.g., every 5 seconds). The input is the meeting audio data, and the output is the buffered audio data. This buffered audio data is temporarily held within the device for subsequent processing.

[0131] Step 2:

[0132] The terminal sends buffered audio data to the server at pre-configured times. The input is the buffered audio data, and the output is the audio data sent to the server. This audio data is transmitted via an HTTP request.

[0133] Step 3:

[0134] The server sends the received audio data to the Google Speech-to-Text API, where it is converted into text data. The input is the audio data sent to the server, and the output is the text data returned by the API. Audio data is sent via an HTTP request, and text data is received as a JSON response.

[0135] Step 4:

[0136] The server adds a timestamp and meeting ID to the received text data and saves it to the database. The input is the text data returned from the API, along with the current timestamp and meeting ID, and the output is the text data saved in the database. Specifically, an SQL query is created to insert the data into the database with the timestamp and meeting ID attached.

[0137] Step 5:

[0138] The server retrieves all text data from the database based on the meeting ID after the meeting ends and sorts it chronologically. The input is the meeting ID, and the output is the chronologically sorted text data. A database query is executed to retrieve the text data and sort it based on the timestamp.

[0139] Step 6:

[0140] The server automatically generates meeting minutes according to a standard format using sorted text data. The input is text data sorted chronologically, and the output is automatically generated meeting minutes. A dedicated template engine is used to format the text data into meeting minutes.

[0141] Step 7:

[0142] The server analyzes the transcribed text data to detect technical terms and obscure words. The input is text data, and the output is a list of detected words. A natural language processing library (e.g., spaCy) is used to analyze the text and extract specific key words.

[0143] Step 8:

[0144] The server retrieves the meanings of detected technical terms and difficult words using a dictionary database or an external API. The input is a list of detected words, and the output is the meaning information of those words. An API request is created, and the response is received in JSON format.

[0145] Step 9:

[0146] The server stores the retrieved semantic information in a database. The input is the semantic information of a word, and the output is the semantic information stored in the database. An SQL query is created to insert the word and its meaning into the database.

[0147] Step 10:

[0148] The terminal retrieves the meaning of a word from the server and displays it in a pop-up window when the user clicks on a word in the text data displayed during a meeting. The input is the clicked word, and the output is the meaning of the word displayed in the pop-up window. An AJAX request is sent, and the retrieved data is processed and displayed using JavaScript (registered trademark).

[0149] These steps enable real-time transcription during meetings, automatic generation of meeting minutes, and retrieval and display of the meanings of technical terms.

[0150] (Application Example 1)

[0151] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0152] Meetings and work instructions within the factory often contain a lot of technical jargon and difficult-to-understand words, making it challenging for new or first-time workers to quickly respond to and understand them. Furthermore, there is a lack of systems to accurately transcribe meeting and instruction content in real time and present it in an easily understandable format.

[0153] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0154] In this invention, the server includes means for transcribing meeting audio data in real time, means for storing the transcribed text data, means for automatically generating meeting minutes based on the transcribed text data, means for detecting technical terms and difficult words that appear during the meeting and obtaining their meanings, means for presenting the meanings of the technical terms and difficult words to the user, and means for transcribing work instructions and meeting audio in the factory in real time and providing workers with the meanings of the detected technical terms. This enables new workers and workers participating for the first time to quickly understand technical terms and difficult words and accurately grasp the content of work instructions and meetings in the factory.

[0155] "Meeting audio data" refers to data that digitally records the speech spoken during a meeting.

[0156] "Methods for real-time transcription" refers to technologies or devices that instantly convert audio from meetings, work instructions, etc., into text data.

[0157] "Transcribed text data" refers to text data converted from audio in real time.

[0158] "Means for saving text data" refers to the technology or device used to save converted text data to a storage device or similar.

[0159] "Means for automatically generating meeting minutes" refers to a technology or device that automatically generates a document recording the content of a meeting by rearranging saved text data in chronological order.

[0160] "Technical terms and difficult vocabulary" refers to specialized terms or words that are not commonly used or are difficult to understand.

[0161] "Means for detecting technical terms and difficult words and obtaining their meanings" refers to technologies or devices that discover technical terms and difficult words within text data and obtain their meanings from external dictionary databases or APIs.

[0162] "Means of presenting to the user" refers to a technology or device that displays the meaning of acquired words in a format that is easy for the user to understand.

[0163] "A means of transcribing audio of work instructions and meetings within a factory in real time" refers to a technology or device that instantly converts audio of work instructions and meetings conducted within a factory into text data.

[0164] "Means of providing workers with the meaning of detected technical terms" refers to a technology or device that displays the meaning of detected technical terms in real time in a format that is easy for workers to understand.

[0165] This invention is a system for improving the efficiency of meetings and work instructions within a factory. This system supports rapid understanding and response by transcribing audio data in real time, saving and analyzing text data, automatically generating meeting minutes, and immediately presenting the meanings of technical terms and difficult words to the user.

[0166] 1. Hardware and software to be used

[0167] The system uses the following hardware and software.

[0168] Hardware:

[0169] Microphone (for recording meetings and work instructions)

[0170] Server (for data processing and management)

[0171] Terminals in stores or factories (to display information to users)

[0172] software:

[0173] Speech recognition API (e.g., Google Speech Recognition API): Used to convert recorded audio into text data.

[0174] Dictionary API (external dictionary database): For obtaining the meanings of technical terms and difficult words.

[0175] Database management system: For storing and managing text data and semantic information.

[0176] Dedicated application: To display information to the user in real time.

[0177] 2. Overall System Functions and Processing

[0178] Audio data collection and transcription

[0179] The terminal records audio data of meetings and work instructions in real time and sends the data to the server in fixed buffer sizes. The server sends the received audio data to a speech recognition API and converts it into text data. This text data is temporarily stored in the server's memory.

[0180] Saving text data

[0181] The server stores the transcribed text data obtained from the speech recognition API into a database. A timestamp and meeting ID are added to the text data. This is useful for later searching and extracting specific meeting content.

[0182] Automatic generation of meeting minutes

[0183] After the meeting ends, or at an appropriate time, the server retrieves all text data from the database based on the meeting ID. The server sorts the retrieved text data chronologically and automatically generates meeting minutes according to a standard format. These minutes are stored in the database for later reference by the user.

[0184] Acquisition and presentation of the meanings of technical terms and difficult words.

[0185] The server analyzes the transcribed text data to detect technical terms and difficult words. For detected words, the server retrieves their meaning using a dictionary database or an external API (dictionary API). The retrieved meanings are stored in the database. When a user clicks on a word in the text data displayed in real time, the meaning of the corresponding word is displayed in a pop-up window.

[0186] 3. Specific examples and prompt messages

[0187] Specific example

[0188] Scenario: A new worker in a factory is instructed for the first time about the term "PMI (Preventive Maintenance)."

[0189] 1. Record audio in real time and send it to the server.

[0190] 2. Translate "PMI" into text using a speech recognition API.

[0191] 3. Save to the database.

[0192] 4. Use the dictionary API to retrieve and save the meaning of "PMI" (preventive maintenance).

[0193] 5. When a new employee clicks on "PMI," its meaning will be displayed.

[0194] Example of a prompt

[0195] Retrieve the meaning of the technical term "PMI" clicked by the user from the dictionary database and display it in real time as "preventive maintenance".

[0196] This system allows for quick understanding of work instructions and meeting content within the factory, enabling new workers and those participating for the first time to adapt smoothly.

[0197] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0198] Step 1:

[0199] The terminal records audio of meetings and work instructions within the factory in real time.

[0200] Input: Audio from meetings or work instructions

[0201] Output: Recorded audio data

[0202] Specific operation: The device's microphone collects audio and stores the data at regular intervals of a certain buffer size.

[0203] Step 2:

[0204] The device sends the recorded audio data to the server in increments of a certain buffer size.

[0205] Input: Recorded audio data

[0206] Output: Audio data sent to the server

[0207] Specific operation: Once the audio data buffer is full, data transfer to the server begins.

[0208] Step 3:

[0209] The server sends the received audio data to a speech recognition API, where it is converted into text data.

[0210] Input: Sent audio data

[0211] Output: Text data returned from the speech recognition API

[0212] Specific operation: The server sends the received audio data to a speech recognition service such as the Google Speech Recognition API.

[0213] Step 4:

[0214] The server temporarily stores the converted text data in memory.

[0215] Input: Text data returned from the speech recognition API

[0216] Output: Text data stored in memory

[0217] Specific operation: The converted text data is given a meeting ID and timestamp, and then saved to the server's memory.

[0218] Step 5:

[0219] The server saves text data to the database.

[0220] Input: Text data stored in memory

[0221] Output: Text data stored in the database

[0222] Specific operation: Store the text data saved in memory in the database along with the meeting ID and timestamp.

[0223] Step 6:

[0224] After the meeting ends, the server retrieves all text data from the database based on the meeting ID.

[0225] Input: Meeting ID

[0226] Output: All text data related to the meeting

[0227] Specific operation: Use a database query to extract text data corresponding to a specified meeting ID.

[0228] Step 7:

[0229] The server sorts the acquired text data chronologically and automatically generates meeting minutes.

[0230] Input: Acquired text data

[0231] Output: Auto-generated meeting minutes

[0232] Specific operation: Sort text data based on timestamps and generate meeting minutes according to a standard format.

[0233] Step 8:

[0234] The server saves meeting minutes to a database so users can access them later.

[0235] Input: Generated meeting minutes

[0236] Output: Meeting minutes saved in the database

[0237] Specific action: Store the generated meeting minutes in the database.

[0238] Step 9:

[0239] The server analyzes the transcribed text data to detect technical terms and difficult words.

[0240] Input: Transcribed text data

[0241] Output: Detected technical terms and difficult words

[0242] Specific operation: Use a text analysis algorithm to identify technical terms and difficult words within the text data.

[0243] Step 10:

[0244] The server retrieves the meaning of a word detected using a dictionary API or dictionary database and stores it in the database.

[0245] Input: Detected technical terms or difficult words

[0246] Output: Meaning of the retrieved word

[0247] Specific operation: Access the dictionary API, retrieve the meaning of the detected word, and store it in the database.

[0248] Step 11:

[0249] When a user clicks on a technical term during a meeting or while working, the device retrieves the meaning of the word from the server and displays it in a pop-up window.

[0250] Input: Technical terms clicked by the user

[0251] Output: Meaning of the word displayed in the pop-up window

[0252] Specific operation: When the user clicks, the meaning of the corresponding word is retrieved from the server and displayed on the device.

[0253] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0254] The system implementing this invention transcribes meeting audio data in real time and saves the text data. Furthermore, it has a function to automatically generate meeting minutes based on the saved text data, and also detects technical terms and difficult words that appear during the meeting, retrieves their meanings, and presents them to the user. In addition to this basic configuration, the present invention incorporates an emotion engine that recognizes the user's emotional state and appropriately adjusts the meanings of technical terms and difficult words based on the results.

[0255] Functional Embodiments

[0256] 1. Collection and transcription of audio data

[0257] The terminal records audio data in real time during the meeting. The recorded data is stored in a buffer and sent to the server at regular intervals. The server sends the received audio data to a speech recognition API and converts it into text data. The converted text data is stored in a database along with a timestamp and meeting ID.

[0258] 2. Automatic generation of meeting minutes

[0259] After the meeting ends, or at an appropriate time, the server retrieves all text data from the database based on the meeting ID. The retrieved text data is sorted chronologically, and meeting minutes are automatically generated according to a standard format. The automatically generated minutes are saved in the database for later reference by the user.

[0260] 3. Acquisition and presentation of the meanings of technical terms and difficult words.

[0261] The server analyzes the transcribed text data to detect technical terms and difficult words. For detected words, the server uses a dictionary database or dictionary API to retrieve their meanings. The retrieved meanings are stored in the database. When the user clicks on a word, the terminal displays its meaning in a pop-up window.

[0262] 4. Introduction of an emotional engine

[0263] The terminal or server recognizes the user's emotional state in real time during the meeting. The emotion engine analyzes the user's voice tone, facial expressions, or biosignals to identify their emotional state. Specifically, it can distinguish emotions such as anger, joy, surprise, and sadness.

[0264] 5. Adjusting information presentation based on emotions

[0265] Based on the user's emotional state, the server appropriately adjusts the meaning of technical terms and difficult words. For example, if the user is confused, it can add more detailed explanations and examples. On the other hand, if the server determines that the user understands, it will only provide a concise explanation.

[0266] Specific example

[0267] 1. At a meeting they are attending for the first time, the user hears the word "refactoring."

[0268] The device records the meeting audio and sends the data to the server.

[0269] The server receives "refactoring" as text data via the speech recognition API and stores it temporarily.

[0270] The server saves that text data to the database and simultaneously detects the word "refactoring".

[0271] The server uses a dictionary API to retrieve the meaning of "refactoring" and saves it to the database.

[0272] 2. The user clicks the word "refactoring".

[0273] The device detects a click event and requests information about that word from the server.

[0274] The server receives the request, retrieves the meaning of the corresponding word from the database, and returns it to the terminal.

[0275] The device displays semantic data as a pop-up.

[0276] 3. Simultaneously, the emotion engine analyzes the user's emotional state and determines that they are confused.

[0277] Based on the sentiment information acquired by the server, more detailed explanations and additional information will be included in the pop-up.

[0278] This system allows users to understand the progress of meetings in real time, comprehend technical terms, and receive appropriate information tailored to their emotional state. In this way, it provides an easily understandable meeting environment, even for users joining a particular organization for the first time or participants with hearing impairments.

[0279] The following describes the processing flow.

[0280] Step 1:

[0281] The device records the meeting audio in real time. Recording begins when the user launches the meeting recording application and presses the start recording button. The recorded data is stored in the audio buffer.

[0282] Step 2:

[0283] The terminal sends data stored in the audio buffer to the server at regular intervals. For example, it sends all the audio data accumulated in the buffer to the server every 10 seconds.

[0284] Step 3:

[0285] The server sends the received voice data to the speech recognition API. The server creates an API request and passes the voice data to the API. APIs such as Google Speech-to-Text and IBM Watson (registered trademark) are used.

[0286] Step 4:

[0287] The speech recognition API converts the voice data into text and returns the result to the server. The API analyzes the voice data and generates text data in string format.

[0288] Step 5:

[0289] The server receives the text data returned from the speech recognition API and temporarily stores it in memory. At this time, the text data is assigned a timestamp and a meeting ID.

[0290] Step 6:

[0291] The server stores the text data in the database. When storing, a meeting ID and a timestamp are assigned, so it can be used for later search and minutes generation.

[0292] Step 7:

[0293] After the meeting ends, the server retrieves all the text data from the database based on the corresponding meeting ID. After retrieval, the server sorts the text data in chronological order.

[0294] Step 8:

[0295] Based on the sorted text data, the server automatically generates minutes according to a fixed format. The automatically generated minutes are organized into paragraphs and headings and formatted into an easy-to-read form.

[0296] Step 9:

[0297] The server saves the minutes generated in the database. Also, an index is added so that the user can access it later.

[0298] Step 10:

[0299] During the meeting, the server analyzes the text data sent in real time and detects technical terms and difficult words. Using natural language processing (NLP) technology, words that match specific conditions are listed.

[0300] Step 11:

[0301] For each word detected by the server, a dictionary API is used to obtain its meaning. The data returned from the dictionary API includes the detailed meaning and examples of the word.

[0302] Step 12:

[0303] The server saves the meaning of the words obtained in the database. Since the meaning data is also saved in association with the meeting ID, it can be easily accessed by the user later.

[0304] Step 13:

[0305] The terminal prepares sensors for recognizing the user's emotional state in real time and collects voice tones, facial expression analysis, or biometric signals. The emotion engine analyzes emotions such as anger, joy, surprise, and sadness in real time.

[0306] Step 14:

[0307] The server receives the emotional state data from the emotion engine and adjusts the meaning of technical terms and difficult words. For example, if the user is judged to be confused, more detailed explanations are added.

[0308] Step 15:

[0309] The user clicks on a word in the text data during a meeting. The device detects the click event and sends a request to the server regarding that word.

[0310] Step 16:

[0311] The server receives the request, retrieves the meaning of the corresponding word from the database, makes adjustments according to the emotional state, and then returns it to the terminal.

[0312] Step 17:

[0313] The device displays semantic data it has acquired as a pop-up. Users can see the meaning of the word and additional information adjusted by the sentiment engine in real time.

[0314] This detailed processing step enables a system that not only transcribes meeting content in real time, but also provides appropriate meanings for technical terms based on the user's emotional state.

[0315] (Example 2)

[0316] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0317] Traditional conferencing systems made it difficult to transcribe meeting audio data in real time, resulting in significant time and effort required for meeting minutes creation. Furthermore, understanding the meaning of specialized terminology and complex vocabulary used during meetings was challenging, posing a significant barrier to comprehension, especially for new or non-expert users. Additionally, the inability to consider the emotional state of participants during meetings led to inappropriate information presentation, posing a significant problem.

[0318] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for transcribing the audio data of a meeting in real time, means for storing the transcribed text data, means for automatically generating meeting minutes based on the transcribed text data, means for detecting technical terms and difficult words that appear during the meeting and obtaining their meanings, means for presenting the meanings of the technical terms and difficult words to the user, and means for recognizing the user's emotions in real time and adjusting the meanings of the technical terms and difficult words. This enables real-time transcription of meeting content, automatic generation of meeting minutes, immediate understanding of technical terms, and presentation of appropriate information based on the user's emotional state.

[0319] "Meeting audio data" refers to all audio generated during a meeting, including discussions, debates, and the content of statements made.

[0320] "Real-time" refers to an event or process that occurs immediately, without any delays or waiting periods, and is reflected instantly.

[0321] "Transcription" is the process of converting audio data into text format, recording the content of the audio as text data.

[0322] "Text data" refers to a data format composed of characters and symbols, and is data generated as a result of transcribing audio data into text.

[0323] "Saving" refers to storing data in a way that makes it usable for a long period of time, and is the act of recording it in a database or storage device.

[0324] "Meeting minutes" are documents that record what was said and decided during a meeting, and are official records intended for later reference.

[0325] "Automatic generation" means creating data or documents using specific algorithms or programs without human intervention.

[0326] "Detecting" refers to the act of finding specific elements or patterns within data or information, and this typically involves using analytical techniques.

[0327] "Acquiring" refers to the act of extracting necessary information from external sources or databases, and this is often done through APIs.

[0328] A "user" refers to a person or organization that uses a system or service, and is the entity that performs operations through the system interface.

[0329] "Recognizing emotions in real time" means instantly understanding a user's emotional state using specific technologies (such as voice analysis or facial expression analysis).

[0330] "Adjusting" refers to the act of changing the content or settings based on specific criteria or conditions to achieve an optimal state.

[0331] In an embodiment of the present invention, the system is capable of transcribing meeting audio data in real time and saving the text data. Furthermore, it can automatically generate meeting minutes based on the saved text data, and in addition, it can detect technical terms and difficult words that appear during the meeting, obtain their meanings, and present them to the user. Moreover, the present invention provides a system that recognizes the user's emotional state in real time and appropriately adjusts the meanings of technical terms and difficult words.

[0332] 1. Hardware and software to be used

[0333] The hardware requires devices (PCs, smartphones, tablets, etc.) to be used by meeting participants, and these devices must include high-sensitivity microphones. The server will be either a cloud environment or an on-premises server.

[0334] The software used includes a speech recognition API (e.g., Google Cloud Speech-to-Text API), a database (e.g., PostgreSQL), and a natural language processing library (e.g., spaCy). External dictionary APIs (e.g., Wiktionary API) are used to obtain the meanings of technical terms and complex words. Speech tone analysis (e.g., Microsoft® Azure® Emotion API) and facial expression analysis (e.g., OpenCV and Dlib) are used to recognize emotional states. Furthermore, generative AI models (e.g., OpenAI® GPT-3®) are used to refine the meanings of technical terms.

[0335] 2. Processing flow and specific data processing / calculations

[0336] The terminal records audio data during the meeting in real time, and the recorded data is temporarily stored in a buffer. At regular intervals, the terminal divides the data in the buffer into packets and sends them to the server via the internet. The HTTPS protocol is used for this.

[0337] The server sends the received audio data to a speech recognition API for transcription. The converted audio data is stored in a database along with a timestamp and meeting ID.

[0338] After the meeting ends, or at an appropriate time, the server retrieves all text data from the database based on the meeting ID. The retrieved text data is sorted chronologically, and meeting minutes are automatically generated using a natural language processing library. The automatically generated minutes are saved in the database for later reference by the user.

[0339] The server analyzes stored text data to detect technical terms and obscure words. This uses techniques such as TF-IDF and morphological analysis. The meaning of detected words is retrieved from a dictionary database or via an external API and stored in the database.

[0340] When a user clicks on a word, the device detects the click event and requests information about that word from the server. The server receives the request, retrieves the meaning data for the corresponding word from its database, and returns it to the device. The device then displays the meaning data in a pop-up window.

[0341] During the meeting, terminals and servers analyze voice tone and facial expressions in real time to recognize emotional states. An emotion engine evaluates this data to determine the user's specific emotional state (e.g., confused, understanding). If confused is detected, the server uses a generative AI model to further refine the meaning of technical terms and complex words, including detailed explanations and additional information.

[0342] 3. Specific Examples and Examples of Prompt Statements

[0343] As a concrete example, if a user hears the word "refactoring" in a meeting they are attending for the first time and does not understand it, it will behave as follows:

[0344] The device records the meeting audio and sends the data to the server.

[0345] The server receives "refactoring" as text data via the speech recognition API and stores it temporarily.

[0346] The server saves that text data to a database and detects the word "refactoring".

[0347] The server uses a dictionary API to retrieve the meaning of "refactoring" and saves it to the database.

[0348] When a user clicks the word "refactoring," the device detects the click event and requests information about that word from the server.

[0349] The server receives the request, retrieves the meaning of the corresponding word from the database, and returns it to the terminal.

[0350] The device displays semantic data as a pop-up.

[0351] At the same time, the emotion engine analyzes the user's emotional state and determines that they are confused.

[0352] Based on the sentiment information acquired by the server, more detailed explanations and additional information will be included in the pop-up.

[0353] Example of a prompt:

[0354] "Could you explain the meaning of refactoring?"

[0355] "Please explain refactoring with an example."

[0356] "Please explain how refactoring can benefit a project."

[0357] This invention makes it easier for participants to understand the progress of a meeting in real time, to immediately grasp the meaning of technical terms and difficult words, and to create an environment where appropriate information is provided according to their emotional state.

[0358] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0359] Step 1: Collect audio data

[0360] The device records audio data during the meeting in real time. Audio acquired through a high-sensitivity microphone is temporarily stored in a buffer. The input is the meeting audio data, and the output is the audio data stored in the buffer. Recording software runs on the device, sampling the audio data second by second and saving it to the buffer.

[0361] Step 2: Sending the audio data

[0362] The terminal periodically divides the audio data in its buffer into packets and sends them to the server via the internet. The input is the audio data in the buffer, and the output is the audio data packets sent to the server. This process is performed securely using the HTTPS protocol.

[0363] Step 3: Transcribing the audio data

[0364] The server receives audio data and sends it to a speech recognition API (e.g., Google Cloud Speech-to-Text API) to convert it into text data. The input is audio data, and the output is transcribed text data. The converted text data is temporarily stored in server memory.

[0365] Step 4: Save text data

[0366] The server saves the transcribed text data, along with the timestamp and meeting ID, to a database. The input is the text data, timestamp, and meeting ID, and the output is the text data saved in the database. The saving process is performed using an SQL query.

[0367] Step 5: Automatic generation of meeting minutes

[0368] After the meeting ends, the server retrieves text data from the database based on the meeting ID. The input is the meeting ID, and the output is the retrieved text data. The server sorts the retrieved text data chronologically and automatically generates meeting minutes using a natural language processing library (e.g., spaCy). The automatically generated meeting minutes are then saved back to the database.

[0369] Step 6: Detection of technical terms and difficult words

[0370] The server analyzes stored text data and uses TF-IDF and morphological analysis to detect technical terms and difficult words. The input is stored text data, and the output is a list of detected words. Natural language processing techniques are used for analysis, and words of high importance are extracted.

[0371] Step 7: Acquire the meaning of technical terms and difficult words.

[0372] The server queries a dictionary database (e.g., the Wiktionary API) for the detected words and retrieves their meanings. The input is a list of detected words, and the output is a list of the retrieved meanings. The retrieved word meanings are then saved back into the database.

[0373] Step 8: Presentation of technical terms and difficult vocabulary.

[0374] The device detects an event when the user clicks on a word and requests that information from the server. The input is the click event, and the output is the requested word information. The server receives the request, retrieves the meaning data for the corresponding word from the database, and returns it to the device. The device then displays the meaning data in a pop-up window.

[0375] Step 9: Recognizing your emotional state

[0376] The terminal or server analyzes the user's voice tone and facial expressions in real time during a meeting to recognize their emotional state. Input is voice tone data and facial expression data, and output is the analyzed emotional state data. This utilizes the Microsoft Azure Emotion API, OpenCV, and Dlib.

[0377] Step 10: Adjusting the Information Presentation

[0378] The server adjusts the meaning of technical terms and difficult words appropriately based on the user's emotional state. The input is emotional state data and the meaning of technical terms, while the output is the adjusted semantic data. If the user is confused, a generative AI model (e.g., OpenAI GPT-3) is used to generate a detailed explanation, which is returned to the device. Based on this returned data, the device displays more detailed information in a pop-up window.

[0379] (Application Example 2)

[0380] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0381] Conventional meeting recording systems often simply convert meeting audio data into text and automatically generate meeting minutes. However, these systems often make it difficult to understand the meaning of technical terms and complex vocabulary, and they fail to provide appropriate information based on the user's emotional state, making it difficult to understand the progress of the meeting and its subsequent analysis. Furthermore, they lack the functionality to respond in real time when security risks occur, requiring immediate countermeasures.

[0382] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for transcribing the audio data of a meeting in real time, means for storing the transcribed text data, means for automatically generating meeting minutes based on the transcribed text data, means for detecting technical terms and difficult words that appear during the meeting and obtaining their meanings, means for presenting the meanings of the technical terms and difficult words to the user, means for analyzing the user's emotional state and adjusting the content of the terms presented based on that emotional state, and means for sending emergency alerts based on specific terms detected during the conversation and the user's emotional state. As a result, the user can grasp the progress of the meeting in real time, easily understand the meanings of technical terms and difficult words, be provided with appropriate information according to their emotional state, and respond immediately to security risks.

[0383] "Audio data" refers to digital recordings of voices generated during meetings, conversations, and other similar events.

[0384] "Transcription" is the process of analyzing audio data and converting it into text data.

[0385] "Text data" refers to data expressed as characters, specifically a string of characters representing the content of a conversation generated through transcription.

[0386] "Meeting minutes" are documents that summarize the content of a meeting, recording what was said and decided during the meeting.

[0387] "Technical terms" are specific terms used in a particular field or industry, and are difficult to understand with general knowledge.

[0388] "Difficult words" are words that are hard to understand or words that are not commonly used.

[0389] "Emotional state" refers to the user's emotional condition and includes a variety of emotions such as joy, anger, sadness, and surprise.

[0390] An "emergency alert" is a notification sent when an emergency occurs or specific conditions are met, and it is a warning that requires immediate action.

[0391] A "user" refers to a person who uses a system or application.

[0392] A "server" is a computer system that processes information over a network and provides services to clients.

[0393] This invention provides a system that supports the progress of meetings in real time, enabling the understanding of important information and the management of security risks. A system implementing this invention includes the following hardware and software:

[0394] 1. Collection and transcription of audio data:

[0395] The terminal records audio data in real time during the meeting. The recorded data is stored in a buffer and sent to the server at regular intervals. The server converts the received audio data into text data using the Google Web Speech API. The converted text data is stored in a database along with a timestamp and meeting ID.

[0396] 2. Automatic generation of meeting minutes:

[0397] After the meeting ends, or at an appropriate time, the server retrieves all text data from the database based on the meeting ID. The retrieved text data is sorted chronologically, and meeting minutes are automatically generated according to a standard format. The automatically generated minutes are saved in the database for later reference by the user.

[0398] 3. Acquisition and presentation of the meanings of technical terms and difficult words:

[0399] The server analyzes the transcribed text data to detect technical terms and difficult words. For detected words, the server uses a dictionary database or dictionary API to retrieve their meanings. The retrieved meanings are stored in the database. When the user clicks on a word, the terminal displays its meaning in a pop-up window.

[0400] 4. Introduction of the Emotion Engine:

[0401] The terminal or server recognizes the user's emotional state in real time during the meeting. The emotion engine analyzes the user's voice tone, facial expressions, or biosignals to identify their emotional state. Specifically, it can distinguish emotions such as anger, joy, surprise, and sadness.

[0402] 5. Adjusting information presentation based on emotions:

[0403] Based on the user's emotional state, the server appropriately adjusts the meaning of technical terms and difficult words. For example, if the user is confused, it can add more detailed explanations and examples. On the other hand, if the server determines that the user understands, it will only provide a concise explanation.

[0404] 6. Sending an emergency alert:

[0405] The server sends emergency alerts based on specific security terms detected during the conversation (e.g., "confidential" or "threat") and the user's emotional state. If the emotion engine detects high levels of tension or anxiety in the user, a notification is sent to the administrator.

[0406] Specific example

[0407] 1. An emergency occurs during the conversation:

[0408] "We have received a report that an intruder entered the office during our conversation. We need to take immediate action."

[0409] This allows users to understand the progress of meetings in real time, comprehend technical terms and complex vocabulary more easily, receive appropriate information tailored to their emotional state, and respond immediately to security risks.

[0410] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0411] Step 1:

[0412] The terminal records the audio data of the meeting in real time. The input is the audio data of the meeting, which the terminal stores in a buffer in digital audio format. The output is digital audio data that is sent to the server at regular intervals.

[0413] Step 2:

[0414] The server sends the received audio data to the Google Web Speech API, where it is converted into text data. The input is digital audio data, and the server uses a speech recognition API to transcribe it. The output is text data stored on the server along with a timestamp and meeting ID.

[0415] Step 3:

[0416] The server analyzes stored text data to detect technical terms and obscure words. The input is text data, and the server uses a natural language processing engine to identify these terms. The output is a list of detected words.

[0417] Step 4:

[0418] The server retrieves the meaning of detected words using a dictionary database or dictionary API. The input is a list of detected words, and the server generates an API request to retrieve the meaning of each word. The output is the meaning information and explanation for each word.

[0419] Step 5:

[0420] The server stores the semantic information of words it has retrieved in a database, and the terminal displays that semantic information upon user request. Input consists of user click events and word requests; the terminal retrieves the word's meaning from the server and displays it in a pop-up. Output is the semantic information presented to the user.

[0421] Step 6:

[0422] The terminal or server recognizes the user's emotional state in real time during a meeting. Inputs include the user's voice tone, facial expressions, or biosignals, which an emotion engine analyzes to identify the emotional state. Output is data on the recognized emotional state.

[0423] Step 7:

[0424] Based on the user's emotional state, the server appropriately adjusts the meaning of technical terms and difficult words. The input consists of emotional state data and word semantic information; the server includes detailed explanations and additional information if the user is confused. The output is the adjusted semantic information.

[0425] Step 8:

[0426] The server sends an emergency alert based on specific security terms and the user's emotional state detected during the conversation. The input is the detected security terms and the user's emotional state data, which the server uses to generate an alert notification and send it to the administrator. The output is the emergency alert sent to the administrator.

[0427] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0428] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0429] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0430] [Second Embodiment]

[0431] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0432] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0433] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0434] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0435] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0436] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0437] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0438] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0439] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0440] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0441] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0442] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0443] The system for implementing the present invention has the following configuration. First, it transcribes the audio data of a meeting in real time and saves the text data. Next, it automatically generates meeting minutes based on the saved text data, and further detects technical terms and difficult words that appear during the meeting, obtains their meanings, and presents them to the user.

[0444] Functional Embodiments

[0445] 1. Collection and transcription of audio data

[0446] The terminal records audio data during the meeting in real time. The recorded audio data is sent to the server in increments of a certain buffer size. The server sends the received audio data to a speech recognition API and converts it into text data. This text data is temporarily stored in the server's memory.

[0447] 2. Saving text data

[0448] The server stores the transcribed text data obtained from the speech recognition API in a database. It is important to add a timestamp and meeting ID to the text data. This is helpful for later searching and extracting specific meeting content.

[0449] 3. Automatic generation of meeting minutes

[0450] After the meeting ends, or at an appropriate time, the server retrieves all text data from the database based on the meeting ID. The server sorts the retrieved text data chronologically and automatically generates meeting minutes according to a standard format. These minutes are stored in the database for later reference by the user.

[0451] 4. Acquisition and presentation of the meanings of technical terms and difficult words.

[0452] The server analyzes the transcribed text data to detect technical terms and difficult words. For detected words, the server retrieves their meaning using a dictionary database or an external API (dictionary API). The retrieved meanings are stored in the database.

[0453] The device displays a pop-up showing the meaning of a word when the user clicks on it in the text data displayed during the meeting. This display is provided in real time, allowing users to understand technical terms as the meeting progresses.

[0454] Specific example

[0455] At a meeting they are attending for the first time, the user hears the word "refactoring."

[0456] 1. The device records the meeting audio and sends the data to the server.

[0457] 2. The server receives "refactoring" as text data via the speech recognition API and stores it temporarily.

[0458] 3. The server saves the text data to the database and simultaneously detects the word "refactoring".

[0459] 4. The server uses the dictionary API to retrieve the meaning of "refactoring" and saves it to the database.

[0460] 5. When a user clicks the word "refactoring" during a meeting, the terminal will display a pop-up with the meaning retrieved from the server.

[0461] This system allows users to understand the progress of a meeting in real time and to instantly check the meaning of technical terms and difficult words. This deepens understanding of the meeting content and provides an environment where users participating in a particular organization for the first time, or participants with hearing impairments, can quickly obtain useful information.

[0462] The following describes the processing flow.

[0463] Step 1:

[0464] The device records the meeting audio in real time. Recording begins when the user launches the meeting recording application during the meeting and presses the start recording button. The recorded data is stored in a buffer.

[0465] Step 2:

[0466] The device sends the recorded data to the server at regular intervals. For example, it clears the buffer every 10 seconds and sends the current recorded data to the server.

[0467] Step 3:

[0468] The server sends the received audio data to a speech recognition API (e.g., Google Speech-to-Text). The server creates an API request and passes the audio data to the API.

[0469] Step 4:

[0470] The speech recognition API converts the audio data into text and returns it to the server. The API parses the audio data and returns the result in string format.

[0471] Step 5:

[0472] The server receives the text data and temporarily stores it in memory. The text data is stored along with the meeting ID and timestamp.

[0473] Step 6:

[0474] The server saves the text data to the database. The saved data is tagged with a meeting ID and timestamp, making it easy to search later.

[0475] Step 7:

[0476] After the meeting ends, the server retrieves all text data for the corresponding meeting ID from the database. The server uses the meeting ID as a key to extract the relevant text data.

[0477] Step 8:

[0478] The server sorts the acquired text data chronologically and automatically generates meeting minutes according to a standard format. The server concatenates the text data and adds appropriate paragraphs and headings.

[0479] Step 9:

[0480] The server generates meeting minutes and saves them to a database. The saved minutes are then managed so that users can access them later.

[0481] Step 10:

[0482] During the meeting, the server analyzes the text data and detects technical terms and difficult words. The server then uses natural language processing to create a list of these technical terms and difficult words.

[0483] Step 11:

[0484] The server uses a dictionary API to retrieve the meaning of each word it detects. The server sends a request to the dictionary API and receives the returned result.

[0485] Step 12:

[0486] The server stores the meaning of the words it retrieves in a database. The meaning data is also stored in association with the conference ID.

[0487] Step 13:

[0488] The user clicks on a word in the text data displayed during the meeting. The device detects the click event and requests information about that word from the server.

[0489] Step 14:

[0490] The server receives the request and retrieves the meaning of the corresponding word from the database. The server queries the semantic data and returns the results to the terminal.

[0491] Step 15:

[0492] The device displays semantic data in a pop-up window. Users can see the detailed meaning of the clicked word in real time.

[0493] These detailed processing steps enable a system where meeting content is transcribed in real time and the meanings of technical terms are provided immediately.

[0494] (Example 1)

[0495] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0496] Meetings often involve a lot of jargon and complex terminology, making it difficult to understand them in real time. Furthermore, manually creating meeting minutes is time-consuming and laborious, and may lack accuracy. Moreover, real-time understanding is even more challenging for participants with hearing impairments or those attending a meeting for the first time. A system is needed to address these challenges.

[0497] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0498] In this invention, the server includes means for transcribing meeting audio data in real time, means for storing the transcribed text data, means for automatically generating meeting minutes based on the transcribed text data, means for detecting technical terms and difficult words that appear during the meeting, means for obtaining the meanings of the detected technical terms and difficult words using a dictionary database or an external API, means for presenting the meanings of the technical terms and difficult words to the user in real time, and means for transmitting the audio data to the server at regular intervals for transcription. As a result, the meanings of technical terms and difficult words can be checked in real time during the meeting, which facilitates understanding of technical terms and enables the automatic generation of meeting minutes, saving time and effort.

[0499] "Audio data" refers to information recorded in digital format using audio.

[0500] "Transcription" refers to the process of converting audio data into text data.

[0501] "Text data" refers to digital information stored as characters.

[0502] "Meeting minutes" refers to a document that records the content and progress of discussions at a meeting.

[0503] "Technical jargon" refers to specialized terms used in a particular field or industry that are generally difficult for the average person to understand.

[0504] "Difficult words" refer to words with special meanings that are not easily understood by the general public.

[0505] A "dictionary database" refers to a digital information source that stores the meanings and definitions of words.

[0506] An "external API" refers to a program that provides an interface for interacting with other systems and services.

[0507] A "server" refers to a computer system that processes and stores data over a network.

[0508] "Terminal" refers to a computer or smart device that a user directly operates.

[0509] "User" refers to a person who uses this system.

[0510] "Real-time" refers to processing that occurs instantly without delay.

[0511] A "timestamp" refers to information that indicates the time when data was recorded.

[0512] "Standard format" refers to the structure and format of a document that follows certain rules and conventions.

[0513] The system of the present invention transcribes meeting audio data in real time, saves the text data, automatically generates meeting minutes, and presents the meanings of technical terms and difficult words to the user. The system of the present invention uses the following specific hardware and software.

[0514] First, the terminal records the meeting audio data in real time using its built-in or external microphone. The recorded audio data is sent to the server in buffer sizes, such as every 5 seconds. The terminal uses an audio recording library, such as the AudioRecord class, for audio recording.

[0515] The server sends the received audio data to the Google Speech-to-Text API, where it is converted into text data. The converted text data is temporarily stored in the server's memory and then stored in a database. Here, each text data entry is assigned a timestamp and a meeting ID. A relational database such as MySQL or PostgreSQL is used for the database.

[0516] After the meeting ends, the server retrieves all text data from the database based on the meeting ID and sorts it chronologically based on the timestamp. It then automatically generates meeting minutes according to a standard format. These minutes are stored in the database by the server for later reference by the user.

[0517] Furthermore, the server analyzes the transcribed text data to detect technical terms and difficult words. For these words, the server retrieves their meanings using dictionary databases such as the Oxford Dictionaries API or external APIs. The retrieved meanings are then stored in the database. When a user clicks on a word in the text data displayed during a meeting, the terminal displays a pop-up with the meaning of the word retrieved from the server. This allows users to understand technical terms simultaneously with the progress of the meeting.

[0518] As a concrete example, let's consider a scenario where a user hears the word "refactoring" at a meeting they are attending for the first time.

[0519] 1. The device records the meeting audio and sends the data to the server.

[0520] 2. The server receives "refactoring" as text data via the Google Speech-to-Text API and stores it temporarily.

[0521] 3. The server saves the text data to the database and simultaneously detects the word "refactoring".

[0522] 4. The server uses the Oxford Dictionaries API to retrieve the meaning of "refactoring" and saves it to the database.

[0523] 5. When a user clicks the word "refactoring" during a meeting, the terminal will display a pop-up with the meaning retrieved from the server.

[0524] This system allows users to understand the progress of a meeting in real time and to instantly check the meaning of technical terms and difficult words. As a result, understanding of the meeting content is enhanced, and a valuable environment is provided where users, especially those joining a particular organization for the first time or participants with hearing impairments, can quickly obtain useful information.

[0525] Example of a prompt:

[0526] "A meeting audio about a new system has been recorded. Please generate a program that converts the recorded audio data into text data, detects technical terms and difficult words, and displays their meanings. Please specify the names of the speech recognition API and dictionary API to be used, and provide a detailed explanation of the processing steps."

[0527] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0528] Step 1:

[0529] The device records meeting audio data in real time using its built-in microphone or an external microphone. The recorded audio data is buffered at a fixed buffer size (e.g., every 5 seconds). The input is the meeting audio data, and the output is the buffered audio data. This buffered audio data is temporarily held within the device for subsequent processing.

[0530] Step 2:

[0531] The terminal sends buffered audio data to the server at pre-configured times. The input is the buffered audio data, and the output is the audio data sent to the server. This audio data is transmitted via an HTTP request.

[0532] Step 3:

[0533] The server sends the received audio data to the Google Speech-to-Text API, where it is converted into text data. The input is the audio data sent to the server, and the output is the text data returned by the API. Audio data is sent via an HTTP request, and text data is received as a JSON response.

[0534] Step 4:

[0535] The server adds a timestamp and meeting ID to the received text data and saves it to the database. The input is the text data returned from the API, along with the current timestamp and meeting ID, and the output is the text data saved in the database. Specifically, an SQL query is created to insert the data into the database with the timestamp and meeting ID attached.

[0536] Step 5:

[0537] The server retrieves all text data from the database based on the meeting ID after the meeting ends and sorts it chronologically. The input is the meeting ID, and the output is the chronologically sorted text data. A database query is executed to retrieve the text data and sort it based on the timestamp.

[0538] Step 6:

[0539] The server automatically generates meeting minutes according to a standard format using sorted text data. The input is text data sorted chronologically, and the output is automatically generated meeting minutes. A dedicated template engine is used to format the text data into meeting minutes.

[0540] Step 7:

[0541] The server analyzes the transcribed text data to detect technical terms and obscure words. The input is text data, and the output is a list of detected words. A natural language processing library (e.g., spaCy) is used to analyze the text and extract specific key words.

[0542] Step 8:

[0543] The server retrieves the meanings of detected technical terms and difficult words using a dictionary database or an external API. The input is a list of detected words, and the output is the meaning information of those words. An API request is created, and the response is received in JSON format.

[0544] Step 9:

[0545] The server stores the retrieved semantic information in a database. The input is the semantic information of a word, and the output is the semantic information stored in the database. An SQL query is created to insert the word and its meaning into the database.

[0546] Step 10:

[0547] The terminal retrieves the meaning of a word from the server and displays it in a pop-up window when the user clicks on a word in the text data displayed during a meeting. The input is the clicked word, and the output is the meaning of the word displayed in the pop-up window. An AJAX request is sent, and the retrieved data is processed and displayed using JavaScript.

[0548] These steps enable real-time transcription during meetings, automatic generation of meeting minutes, and retrieval and display of the meanings of technical terms.

[0549] (Application Example 1)

[0550] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0551] Meetings and work instructions within the factory often contain a lot of technical jargon and difficult-to-understand words, making it challenging for new or first-time workers to quickly respond to and understand them. Furthermore, there is a lack of systems to accurately transcribe meeting and instruction content in real time and present it in an easily understandable format.

[0552] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0553] In this invention, the server includes means for transcribing meeting audio data in real time, means for storing the transcribed text data, means for automatically generating meeting minutes based on the transcribed text data, means for detecting technical terms and difficult words that appear during the meeting and obtaining their meanings, means for presenting the meanings of the technical terms and difficult words to the user, and means for transcribing work instructions and meeting audio in the factory in real time and providing workers with the meanings of the detected technical terms. This enables new workers and workers participating for the first time to quickly understand technical terms and difficult words and accurately grasp the content of work instructions and meetings in the factory.

[0554] "Meeting audio data" refers to data that digitally records the speech spoken during a meeting.

[0555] "Methods for real-time transcription" refers to technologies or devices that instantly convert audio from meetings, work instructions, etc., into text data.

[0556] "Transcribed text data" refers to text data converted from audio in real time.

[0557] "Means for saving text data" refers to the technology or device used to save converted text data to a storage device or similar.

[0558] "Means for automatically generating meeting minutes" refers to a technology or device that automatically generates a document recording the content of a meeting by rearranging saved text data in chronological order.

[0559] "Technical terms and difficult vocabulary" refers to specialized terms or words that are not commonly used or are difficult to understand.

[0560] "Means for detecting technical terms and difficult words and obtaining their meanings" refers to technologies or devices that discover technical terms and difficult words within text data and obtain their meanings from external dictionary databases or APIs.

[0561] "Means of presenting to the user" refers to a technology or device that displays the meaning of acquired words in a format that is easy for the user to understand.

[0562] "A means of transcribing audio of work instructions and meetings within a factory in real time" refers to a technology or device that instantly converts audio of work instructions and meetings conducted within a factory into text data.

[0563] "Means of providing workers with the meaning of detected technical terms" refers to a technology or device that displays the meaning of detected technical terms in real time in a format that is easy for workers to understand.

[0564] This invention is a system for improving the efficiency of meetings and work instructions within a factory. This system supports rapid understanding and response by transcribing audio data in real time, saving and analyzing text data, automatically generating meeting minutes, and immediately presenting the meanings of technical terms and difficult words to the user.

[0565] 1. Hardware and software to be used

[0566] The system uses the following hardware and software.

[0567] Hardware:

[0568] Microphone (for recording meetings and work instructions)

[0569] Server (for data processing and management)

[0570] Terminals in stores or factories (to display information to users)

[0571] software:

[0572] Speech recognition API (e.g., Google Speech Recognition API): Used to convert recorded audio into text data.

[0573] Dictionary API (external dictionary database): For obtaining the meanings of technical terms and difficult words.

[0574] Database management system: For storing and managing text data and semantic information.

[0575] Dedicated application: To display information to the user in real time.

[0576] 2. Overall System Functions and Processing

[0577] Audio data collection and transcription

[0578] The terminal records audio data of meetings and work instructions in real time and sends the data to the server in fixed buffer sizes. The server sends the received audio data to a speech recognition API and converts it into text data. This text data is temporarily stored in the server's memory.

[0579] Saving text data

[0580] The server stores the transcribed text data obtained from the speech recognition API into a database. A timestamp and meeting ID are added to the text data. This is useful for later searching and extracting specific meeting content.

[0581] Automatic generation of meeting minutes

[0582] After the meeting ends, or at an appropriate time, the server retrieves all text data from the database based on the meeting ID. The server sorts the retrieved text data chronologically and automatically generates meeting minutes according to a standard format. These minutes are stored in the database for later reference by the user.

[0583] Acquisition and presentation of the meanings of technical terms and difficult words.

[0584] The server analyzes the transcribed text data to detect technical terms and difficult words. For detected words, the server retrieves their meaning using a dictionary database or an external API (dictionary API). The retrieved meanings are stored in the database. When a user clicks on a word in the text data displayed in real time, the meaning of the corresponding word is displayed in a pop-up window.

[0585] 3. Specific examples and prompt messages

[0586] Specific example

[0587] Scenario: A new worker in a factory is instructed for the first time about the term "PMI (Preventive Maintenance)."

[0588] 1. Record audio in real time and send it to the server.

[0589] 2. Translate "PMI" into text using a speech recognition API.

[0590] 3. Save to the database.

[0591] 4. Use the dictionary API to retrieve and save the meaning of "PMI" (preventive maintenance).

[0592] 5. When a new employee clicks on "PMI," its meaning will be displayed.

[0593] Example of a prompt

[0594] Retrieve the meaning of the technical term "PMI" clicked by the user from the dictionary database and display it in real time as "preventive maintenance".

[0595] This system allows for quick understanding of work instructions and meeting content within the factory, enabling new workers and those participating for the first time to adapt smoothly.

[0596] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0597] Step 1:

[0598] The terminal records audio of meetings and work instructions within the factory in real time.

[0599] Input: Audio from meetings or work instructions

[0600] Output: Recorded audio data

[0601] Specific operation: The device's microphone collects audio and stores the data at regular intervals of a certain buffer size.

[0602] Step 2:

[0603] The device sends the recorded audio data to the server in increments of a certain buffer size.

[0604] Input: Recorded audio data

[0605] Output: Audio data sent to the server

[0606] Specific operation: Once the audio data buffer is full, data transfer to the server begins.

[0607] Step 3:

[0608] The server sends the received audio data to a speech recognition API, where it is converted into text data.

[0609] Input: Sent audio data

[0610] Output: Text data returned from the speech recognition API

[0611] Specific operation: The server sends the received audio data to a speech recognition service such as the Google Speech Recognition API.

[0612] Step 4:

[0613] The server temporarily stores the converted text data in memory.

[0614] Input: Text data returned from the speech recognition API

[0615] Output: Text data stored in memory

[0616] Specific operation: The converted text data is given a meeting ID and timestamp, and then saved to the server's memory.

[0617] Step 5:

[0618] The server saves text data to the database.

[0619] Input: Text data stored in memory

[0620] Output: Text data stored in the database

[0621] Specific operation: Store the text data saved in memory in the database along with the meeting ID and timestamp.

[0622] Step 6:

[0623] After the meeting ends, the server retrieves all text data from the database based on the meeting ID.

[0624] Input: Meeting ID

[0625] Output: All text data related to the meeting

[0626] Specific operation: Use a database query to extract text data corresponding to a specified meeting ID.

[0627] Step 7:

[0628] The server sorts the acquired text data chronologically and automatically generates meeting minutes.

[0629] Input: Acquired text data

[0630] Output: Auto-generated meeting minutes

[0631] Specific operation: Sort text data based on timestamps and generate meeting minutes according to a standard format.

[0632] Step 8:

[0633] The server saves meeting minutes to a database so users can access them later.

[0634] Input: Generated meeting minutes

[0635] Output: Meeting minutes saved in the database

[0636] Specific action: Store the generated meeting minutes in the database.

[0637] Step 9:

[0638] The server analyzes the transcribed text data to detect technical terms and difficult words.

[0639] Input: Transcribed text data

[0640] Output: Detected technical terms and difficult words

[0641] Specific operation: Use a text analysis algorithm to identify technical terms and difficult words within the text data.

[0642] Step 10:

[0643] The server retrieves the meaning of a word detected using a dictionary API or dictionary database and stores it in the database.

[0644] Input: Detected technical terms or difficult words

[0645] Output: Meaning of the retrieved word

[0646] Specific operation: Access the dictionary API, retrieve the meaning of the detected word, and store it in the database.

[0647] Step 11:

[0648] When a user clicks on a technical term during a meeting or while working, the device retrieves the meaning of the word from the server and displays it in a pop-up window.

[0649] Input: Technical terms clicked by the user

[0650] Output: Meaning of the word displayed in the pop-up window

[0651] Specific operation: When the user clicks, the meaning of the corresponding word is retrieved from the server and displayed on the device.

[0652] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0653] The system implementing this invention transcribes meeting audio data in real time and saves the text data. Furthermore, it has a function to automatically generate meeting minutes based on the saved text data, and also detects technical terms and difficult words that appear during the meeting, retrieves their meanings, and presents them to the user. In addition to this basic configuration, the present invention incorporates an emotion engine that recognizes the user's emotional state and appropriately adjusts the meanings of technical terms and difficult words based on the results.

[0654] Functional Embodiments

[0655] 1. Collection and transcription of audio data

[0656] The terminal records audio data in real time during the meeting. The recorded data is stored in a buffer and sent to the server at regular intervals. The server sends the received audio data to a speech recognition API and converts it into text data. The converted text data is stored in a database along with a timestamp and meeting ID.

[0657] 2. Automatic generation of meeting minutes

[0658] After the meeting ends, or at an appropriate time, the server retrieves all text data from the database based on the meeting ID. The retrieved text data is sorted chronologically, and meeting minutes are automatically generated according to a standard format. The automatically generated minutes are saved in the database for later reference by the user.

[0659] 3. Acquisition and presentation of the meanings of technical terms and difficult words.

[0660] The server analyzes the transcribed text data to detect technical terms and difficult words. For detected words, the server uses a dictionary database or dictionary API to retrieve their meanings. The retrieved meanings are stored in the database. When the user clicks on a word, the terminal displays its meaning in a pop-up window.

[0661] 4. Introduction of an emotional engine

[0662] The terminal or server recognizes the user's emotional state in real time during the meeting. The emotion engine analyzes the user's voice tone, facial expressions, or biosignals to identify their emotional state. Specifically, it can distinguish emotions such as anger, joy, surprise, and sadness.

[0663] 5. Adjusting information presentation based on emotions

[0664] Based on the user's emotional state, the server appropriately adjusts the meaning of technical terms and difficult words. For example, if the user is confused, it can add more detailed explanations and examples. On the other hand, if the server determines that the user understands, it will only provide a concise explanation.

[0665] Specific example

[0666] 1. At a meeting they are attending for the first time, the user hears the word "refactoring."

[0667] The device records the meeting audio and sends the data to the server.

[0668] The server receives "refactoring" as text data via the speech recognition API and stores it temporarily.

[0669] The server saves that text data to the database and simultaneously detects the word "refactoring".

[0670] The server uses a dictionary API to retrieve the meaning of "refactoring" and saves it to the database.

[0671] 2. The user clicks the word "refactoring".

[0672] The device detects a click event and requests information about that word from the server.

[0673] The server receives the request, retrieves the meaning of the corresponding word from the database, and returns it to the terminal.

[0674] The device displays semantic data as a pop-up.

[0675] 3. Simultaneously, the emotion engine analyzes the user's emotional state and determines that they are confused.

[0676] Based on the sentiment information acquired by the server, more detailed explanations and additional information will be included in the pop-up.

[0677] This system allows users to understand the progress of meetings in real time, comprehend technical terms, and receive appropriate information tailored to their emotional state. In this way, it provides an easily understandable meeting environment, even for users joining a particular organization for the first time or participants with hearing impairments.

[0678] The following describes the processing flow.

[0679] Step 1:

[0680] The device records the meeting audio in real time. Recording begins when the user launches the meeting recording application and presses the start recording button. The recorded data is stored in the audio buffer.

[0681] Step 2:

[0682] The terminal sends data stored in the audio buffer to the server at regular intervals. For example, it sends all the audio data accumulated in the buffer to the server every 10 seconds.

[0683] Step 3:

[0684] The server sends the received audio data to a speech recognition API. The server creates an API request and passes the audio data to the API. Examples of APIs used include Google Speech-to-Text and IBM Watson.

[0685] Step 4:

[0686] The speech recognition API converts the audio data into text and returns the result to the server. The API analyzes the audio data and generates text data in string format.

[0687] Step 5:

[0688] The server receives the text data returned from the speech recognition API and temporarily stores it in memory. At this time, a timestamp and meeting ID are added to the text data.

[0689] Step 6:

[0690] The server saves the text data to the database. A meeting ID and timestamp are added during saving, making it useful for later searching and generating meeting minutes.

[0691] Step 7:

[0692] After the meeting ends, the server retrieves all text data from the database based on the corresponding meeting ID. After retrieval, the server sorts the text data chronologically.

[0693] Step 8:

[0694] The server automatically generates meeting minutes based on sorted text data, following a standard format. The automatically generated minutes are formatted for readability, with paragraphs and headings neatly organized.

[0695] Step 9:

[0696] The server saves the generated meeting minutes to a database. It also adds an index to allow users to access them later.

[0697] Step 10:

[0698] During the meeting, the server analyzes the text data transmitted in real time to detect technical terms and difficult words. Natural language processing (NLP) techniques are used to list words that match specific criteria.

[0699] Step 11:

[0700] The server uses a dictionary API to retrieve the meaning of each word it detects. The data returned from the dictionary API includes a detailed definition of the word and example sentences.

[0701] Step 12:

[0702] The server stores the meaning of the words it retrieves in a database. Since the meaning data is also stored in association with the meeting ID, users can easily access it later.

[0703] Step 13:

[0704] The device is equipped with sensors to recognize the user's emotional state in real time, collecting voice tone, facial expression analysis, or biosignals. The emotion engine analyzes emotions such as anger, joy, surprise, and sadness in real time.

[0705] Step 14:

[0706] The server receives emotional state data from the emotion engine and adjusts the meaning of technical terms and complex words. For example, if it determines that the user is confused, it adds a more detailed explanation.

[0707] Step 15:

[0708] The user clicks on a word in the text data during a meeting. The device detects the click event and sends a request to the server regarding that word.

[0709] Step 16:

[0710] The server receives the request, retrieves the meaning of the corresponding word from the database, makes adjustments according to the emotional state, and then returns it to the terminal.

[0711] Step 17:

[0712] The device displays semantic data it has acquired as a pop-up. Users can see the meaning of the word and additional information adjusted by the sentiment engine in real time.

[0713] This detailed processing step enables a system that not only transcribes meeting content in real time, but also provides appropriate meanings for technical terms based on the user's emotional state.

[0714] (Example 2)

[0715] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0716] Traditional conferencing systems made it difficult to transcribe meeting audio data in real time, resulting in significant time and effort required for creating meeting minutes. Furthermore, understanding the meaning of specialized terminology and complex vocabulary used during meetings was challenging, posing a significant barrier to comprehension, especially for new or non-expert users. Additionally, the inability to consider the emotional state of participants during meetings led to inappropriate information presentation, posing a significant problem.

[0717] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for transcribing the audio data of a meeting in real time, means for storing the transcribed text data, means for automatically generating meeting minutes based on the transcribed text data, means for detecting technical terms and difficult words that appear during the meeting and obtaining their meanings, means for presenting the meanings of the technical terms and difficult words to the user, and means for recognizing the user's emotions in real time and adjusting the meanings of the technical terms and difficult words. This enables real-time transcription of meeting content, automatic generation of meeting minutes, immediate understanding of technical terms, and presentation of appropriate information based on the user's emotional state.

[0718] "Meeting audio data" refers to all audio generated during a meeting, including discussions, debates, and the content of statements made.

[0719] "Real-time" refers to an event or process that occurs immediately, without any delays or waiting periods, and is reflected instantly.

[0720] "Transcription" is the process of converting audio data into text format, recording the content of the audio as text data.

[0721] "Text data" refers to a data format composed of characters and symbols, and is data generated as a result of transcribing audio data into text.

[0722] "Saving" refers to storing data in a way that makes it usable for a long period of time, and is the act of recording it in a database or storage device.

[0723] "Meeting minutes" are documents that record what was said and decided during a meeting, and are official records intended for later reference.

[0724] "Automatic generation" means creating data or documents using specific algorithms or programs without human intervention.

[0725] "Detecting" refers to the act of finding specific elements or patterns within data or information, and this typically involves using analytical techniques.

[0726] "Acquiring" refers to the act of extracting necessary information from external sources or databases, and this is often done through APIs.

[0727] A "user" refers to a person or organization that uses a system or service, and is the entity that performs operations through the system interface.

[0728] "Recognizing emotions in real time" means instantly understanding a user's emotional state using specific technologies (such as voice analysis or facial expression analysis).

[0729] "Adjusting" refers to the act of changing the content or settings based on specific criteria or conditions to achieve an optimal state.

[0730] In an embodiment of the present invention, the system is capable of transcribing meeting audio data in real time and saving the text data. Furthermore, it can automatically generate meeting minutes based on the saved text data, and in addition, it can detect technical terms and difficult words that appear during the meeting, obtain their meanings, and present them to the user. Moreover, the present invention provides a system that recognizes the user's emotional state in real time and appropriately adjusts the meanings of technical terms and difficult words.

[0731] 1. Hardware and software to be used

[0732] The hardware requires devices (PCs, smartphones, tablets, etc.) to be used by meeting participants, and these devices must include high-sensitivity microphones. The server will be either a cloud environment or an on-premises server.

[0733] The software used includes a speech recognition API (e.g., Google Cloud Speech-to-Text API), a database (e.g., PostgreSQL), and a natural language processing library (e.g., spaCy). External dictionary APIs (e.g., Wiktionary API) are used to obtain the meanings of technical terms and complex words. Speech tone analysis (e.g., Microsoft Azure Emotion API) and facial expression analysis (e.g., OpenCV and Dlib) are used to recognize emotional states. Furthermore, generative AI models (e.g., OpenAI GPT-3) are used to refine the meanings of technical terms.

[0734] 2. Processing flow and specific data processing / calculations

[0735] The terminal records audio data during the meeting in real time, and the recorded data is temporarily stored in a buffer. At regular intervals, the terminal divides the data in the buffer into packets and sends them to the server via the internet. The HTTPS protocol is used for this.

[0736] The server sends the received audio data to a speech recognition API for transcription. The converted audio data is stored in a database along with a timestamp and meeting ID.

[0737] After the meeting ends, or at an appropriate time, the server retrieves all text data from the database based on the meeting ID. The retrieved text data is sorted chronologically, and meeting minutes are automatically generated using a natural language processing library. The automatically generated minutes are saved in the database for later reference by the user.

[0738] The server analyzes stored text data to detect technical terms and obscure words. This uses techniques such as TF-IDF and morphological analysis. The meaning of detected words is retrieved from a dictionary database or via an external API and stored in the database.

[0739] When a user clicks on a word, the device detects the click event and requests information about that word from the server. The server receives the request, retrieves the meaning data for the corresponding word from its database, and returns it to the device. The device then displays the meaning data in a pop-up window.

[0740] During the meeting, terminals and servers analyze voice tone and facial expressions in real time to recognize emotional states. An emotion engine evaluates this data to determine the user's specific emotional state (e.g., confused, understanding). If confused is detected, the server uses a generative AI model to further refine the meaning of technical terms and complex words, including detailed explanations and additional information.

[0741] 3. Specific Examples and Examples of Prompt Statements

[0742] As a concrete example, if a user hears the word "refactoring" in a meeting they are attending for the first time and does not understand it, it will behave as follows:

[0743] The device records the meeting audio and sends the data to the server.

[0744] The server receives "refactoring" as text data via the speech recognition API and stores it temporarily.

[0745] The server saves that text data to a database and detects the word "refactoring".

[0746] The server uses a dictionary API to retrieve the meaning of "refactoring" and saves it to the database.

[0747] When a user clicks the word "refactoring," the device detects the click event and requests information about that word from the server.

[0748] The server receives the request, retrieves the meaning of the corresponding word from the database, and returns it to the terminal.

[0749] The device displays semantic data as a pop-up.

[0750] At the same time, the emotion engine analyzes the user's emotional state and determines that they are confused.

[0751] Based on the sentiment information acquired by the server, more detailed explanations and additional information will be included in the pop-up.

[0752] Example of a prompt:

[0753] "Could you explain the meaning of refactoring?"

[0754] "Please explain refactoring with an example."

[0755] "Please explain how refactoring can benefit a project."

[0756] This invention makes it easier for participants to understand the progress of a meeting in real time, to immediately grasp the meaning of technical terms and difficult words, and to create an environment where appropriate information is provided according to their emotional state.

[0757] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0758] Step 1: Collect audio data

[0759] The device records audio data during the meeting in real time. Audio acquired through a high-sensitivity microphone is temporarily stored in a buffer. The input is the meeting audio data, and the output is the audio data stored in the buffer. Recording software runs on the device, sampling the audio data second by second and saving it to the buffer.

[0760] Step 2: Sending the audio data

[0761] The terminal periodically divides the audio data in its buffer into packets and sends them to the server via the internet. The input is the audio data in the buffer, and the output is the audio data packets sent to the server. This process is performed securely using the HTTPS protocol.

[0762] Step 3: Transcribing the audio data

[0763] The server receives audio data and sends it to a speech recognition API (e.g., Google Cloud Speech-to-Text API) to convert it into text data. The input is audio data, and the output is transcribed text data. The converted text data is temporarily stored in server memory.

[0764] Step 4: Save text data

[0765] The server saves the transcribed text data, along with the timestamp and meeting ID, to a database. The input is the text data, timestamp, and meeting ID, and the output is the text data saved in the database. The saving process is performed using an SQL query.

[0766] Step 5: Automatic generation of meeting minutes

[0767] After the meeting ends, the server retrieves text data from the database based on the meeting ID. The input is the meeting ID, and the output is the retrieved text data. The server sorts the retrieved text data chronologically and automatically generates meeting minutes using a natural language processing library (e.g., spaCy). The automatically generated meeting minutes are then saved back to the database.

[0768] Step 6: Detection of technical terms and difficult words

[0769] The server analyzes stored text data and uses TF-IDF and morphological analysis to detect technical terms and difficult words. The input is stored text data, and the output is a list of detected words. Natural language processing techniques are used for analysis, and words of high importance are extracted.

[0770] Step 7: Acquire the meaning of technical terms and difficult words.

[0771] The server queries a dictionary database (e.g., the Wiktionary API) for the detected words and retrieves their meanings. The input is a list of detected words, and the output is a list of retrieved meanings. The retrieved word meanings are then saved back into the database.

[0772] Step 8: Presentation of technical terms and difficult vocabulary.

[0773] The device detects an event when the user clicks on a word and requests that information from the server. The input is the click event, and the output is the requested word information. The server receives the request, retrieves the meaning data for the corresponding word from the database, and returns it to the device. The device then displays the meaning data in a pop-up window.

[0774] Step 9: Recognizing your emotional state

[0775] The terminal or server analyzes the user's voice tone and facial expressions in real time during a meeting to recognize their emotional state. Input is voice tone data and facial expression data, and output is the analyzed emotional state data. This utilizes the Microsoft Azure Emotion API, OpenCV, and Dlib.

[0776] Step 10: Adjusting the Information Presentation

[0777] The server adjusts the meaning of technical terms and difficult words appropriately based on the user's emotional state. The input is emotional state data and the meaning of technical terms, while the output is the adjusted semantic data. If the user is confused, a generative AI model (e.g., OpenAI GPT-3) is used to generate a detailed explanation, which is returned to the device. Based on this returned data, the device displays more detailed information in a pop-up window.

[0778] (Application Example 2)

[0779] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0780] Conventional meeting recording systems often simply convert meeting audio data into text and automatically generate meeting minutes. However, these systems often make it difficult to understand the meaning of technical terms and complex vocabulary, and they fail to provide appropriate information based on the user's emotional state, making it difficult to understand the progress of the meeting and its subsequent analysis. Furthermore, they lack the functionality to respond in real time when security risks occur, requiring immediate countermeasures.

[0781] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for transcribing the audio data of a meeting in real time, means for storing the transcribed text data, means for automatically generating meeting minutes based on the transcribed text data, means for detecting technical terms and difficult words that appear during the meeting and obtaining their meanings, means for presenting the meanings of the technical terms and difficult words to the user, means for analyzing the user's emotional state and adjusting the content of the terms presented based on that emotional state, and means for sending emergency alerts based on specific terms detected during the conversation and the user's emotional state. As a result, the user can grasp the progress of the meeting in real time, easily understand the meanings of technical terms and difficult words, be provided with appropriate information according to their emotional state, and respond immediately to security risks.

[0782] "Audio data" refers to digital recordings of voices generated during meetings, conversations, and other similar events.

[0783] "Transcription" is the process of analyzing audio data and converting it into text data.

[0784] "Text data" refers to data expressed as characters, specifically a string of characters representing the content of a conversation generated through transcription.

[0785] "Meeting minutes" are documents that summarize the content of a meeting, recording the statements made and decisions made during the meeting.

[0786] "Technical terms" are specific terms used in a particular field or industry, and are difficult to understand with general knowledge.

[0787] "Difficult words" are words that are hard to understand or words that are not commonly used.

[0788] "Emotional state" refers to the user's emotional condition and includes a variety of emotions such as joy, anger, sadness, and surprise.

[0789] An "emergency alert" is a notification sent when an emergency occurs or specific conditions are met, and it is a warning that requires immediate action.

[0790] "User" refers to a person who uses a system or application.

[0791] A "server" is a computer system that processes information over a network and provides services to clients.

[0792] This invention provides a system that supports the progress of meetings in real time, enabling the understanding of important information and the management of security risks. A system implementing this invention includes the following hardware and software:

[0793] 1. Collection and transcription of audio data:

[0794] The terminal records audio data in real time during the meeting. The recorded data is stored in a buffer and sent to the server at regular intervals. The server converts the received audio data into text data using the Google Web Speech API. The converted text data is stored in a database along with a timestamp and meeting ID.

[0795] 2. Automatic generation of meeting minutes:

[0796] After the meeting ends, or at an appropriate time, the server retrieves all text data from the database based on the meeting ID. The retrieved text data is sorted chronologically, and meeting minutes are automatically generated according to a standard format. The automatically generated minutes are saved in the database for later reference by the user.

[0797] 3. Acquisition and presentation of the meanings of technical terms and difficult words:

[0798] The server analyzes the transcribed text data to detect technical terms and difficult words. For detected words, the server uses a dictionary database or dictionary API to retrieve their meanings. The retrieved meanings are stored in the database. When the user clicks on a word, the terminal displays its meaning in a pop-up window.

[0799] 4. Introduction of the Emotion Engine:

[0800] The terminal or server recognizes the user's emotional state in real time during the meeting. The emotion engine analyzes the user's voice tone, facial expressions, or biosignals to identify their emotional state. Specifically, it can distinguish emotions such as anger, joy, surprise, and sadness.

[0801] 5. Adjusting information presentation based on emotions:

[0802] Based on the user's emotional state, the server appropriately adjusts the meaning of technical terms and difficult words. For example, if the user is confused, it can add more detailed explanations and examples. On the other hand, if the server determines that the user understands, it will only provide a concise explanation.

[0803] 6. Sending an emergency alert:

[0804] The server sends emergency alerts based on specific security terms detected during the conversation (e.g., "confidential" or "threat") and the user's emotional state. If the emotion engine detects high levels of tension or anxiety in the user, a notification is sent to the administrator.

[0805] Specific example

[0806] 1. An emergency occurs during the conversation:

[0807] "We have received a report that an intruder entered the office during our conversation. We need to take immediate action."

[0808] This allows users to understand the progress of meetings in real time, comprehend technical terms and complex vocabulary more easily, receive appropriate information tailored to their emotional state, and respond immediately to security risks.

[0809] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0810] Step 1:

[0811] The terminal records the audio data of the meeting in real time. The input is the audio data of the meeting, which the terminal stores in a buffer in digital audio format. The output is digital audio data that is sent to the server at regular intervals.

[0812] Step 2:

[0813] The server sends the received audio data to the Google Web Speech API, where it is converted into text data. The input is digital audio data, and the server uses a speech recognition API to transcribe it. The output is text data stored on the server along with a timestamp and meeting ID.

[0814] Step 3:

[0815] The server analyzes stored text data to detect technical terms and obscure words. The input is text data, and the server uses a natural language processing engine to identify these terms. The output is a list of detected words.

[0816] Step 4:

[0817] The server retrieves the meaning of detected words using a dictionary database or dictionary API. The input is a list of detected words, and the server generates an API request to retrieve the meaning of each word. The output is the meaning information and explanation for each word.

[0818] Step 5:

[0819] The server stores the semantic information of words it has retrieved in a database, and the terminal displays that semantic information upon user request. Input consists of user click events and word requests; the terminal retrieves the word's meaning from the server and displays it in a pop-up. Output is the semantic information presented to the user.

[0820] Step 6:

[0821] The terminal or server recognizes the user's emotional state in real time during a meeting. Inputs include the user's voice tone, facial expressions, or biosignals, which the emotion engine analyzes to identify the emotional state. Output is data on the recognized emotional state.

[0822] Step 7:

[0823] Based on the user's emotional state, the server appropriately adjusts the meaning of technical terms and difficult words. Input consists of emotional state data and word semantic information; the server includes detailed explanations and additional information if the user is confused. Output is the adjusted semantic information.

[0824] Step 8:

[0825] The server sends an emergency alert based on specific security terms and the user's emotional state detected during the conversation. The input is the detected security terms and user emotional state data, which the server uses to generate an alert notification and send it to the administrator. The output is the emergency alert sent to the administrator.

[0826] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0827] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0828] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0829] [Third Embodiment]

[0830] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0831] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0832] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0833] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0834] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0835] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0836] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0837] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0838] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0839] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0840] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0841] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0842] The system for implementing the present invention has the following configuration. First, it transcribes the audio data of a meeting in real time and saves the text data. Next, it automatically generates meeting minutes based on the saved text data, and further detects technical terms and difficult words that appear during the meeting, obtains their meanings, and presents them to the user.

[0843] Functional Embodiments

[0844] 1. Collection and transcription of audio data

[0845] The terminal records audio data during the meeting in real time. The recorded audio data is sent to the server in increments of a certain buffer size. The server sends the received audio data to a speech recognition API and converts it into text data. This text data is temporarily stored in the server's memory.

[0846] 2. Saving text data

[0847] The server stores the transcribed text data obtained from the speech recognition API in a database. It is important to add a timestamp and meeting ID to the text data. This is helpful for later searching and extracting specific meeting content.

[0848] 3. Automatic generation of meeting minutes

[0849] After the meeting ends, or at an appropriate time, the server retrieves all text data from the database based on the meeting ID. The server sorts the retrieved text data chronologically and automatically generates meeting minutes according to a standard format. These minutes are stored in the database for later reference by the user.

[0850] 4. Acquisition and presentation of the meanings of technical terms and difficult words.

[0851] The server analyzes the transcribed text data to detect technical terms and difficult words. For detected words, the server retrieves their meaning using a dictionary database or an external API (dictionary API). The retrieved meanings are stored in the database.

[0852] The device displays a pop-up showing the meaning of a word when the user clicks on it in the text data displayed during the meeting. This display is provided in real time, allowing users to understand technical terms as the meeting progresses.

[0853] Specific example

[0854] At a meeting they are attending for the first time, the user hears the word "refactoring."

[0855] 1. The device records the meeting audio and sends the data to the server.

[0856] 2. The server receives "refactoring" as text data via the speech recognition API and stores it temporarily.

[0857] 3. The server saves the text data to the database and simultaneously detects the word "refactoring".

[0858] 4. The server uses the dictionary API to retrieve the meaning of "refactoring" and saves it to the database.

[0859] 5. When a user clicks the word "refactoring" during a meeting, the terminal will display a pop-up with the meaning retrieved from the server.

[0860] This system allows users to understand the progress of a meeting in real time and to instantly check the meaning of technical terms and difficult words. This deepens understanding of the meeting content and provides an environment where users participating in a particular organization for the first time, or participants with hearing impairments, can quickly obtain useful information.

[0861] The following describes the processing flow.

[0862] Step 1:

[0863] The device records the meeting audio in real time. Recording begins when the user launches the meeting recording application during the meeting and presses the start recording button. The recorded data is stored in a buffer.

[0864] Step 2:

[0865] The device sends the recorded data to the server at regular intervals. For example, it clears the buffer every 10 seconds and sends the current recorded data to the server.

[0866] Step 3:

[0867] The server sends the received audio data to a speech recognition API (e.g., Google Speech-to-Text). The server creates an API request and passes the audio data to the API.

[0868] Step 4:

[0869] The speech recognition API converts the audio data into text and returns it to the server. The API parses the audio data and returns the result in string format.

[0870] Step 5:

[0871] The server receives the text data and temporarily stores it in memory. The text data is stored along with the meeting ID and timestamp.

[0872] Step 6:

[0873] The server saves the text data to the database. The saved data is tagged with a meeting ID and timestamp, making it easy to search later.

[0874] Step 7:

[0875] After the meeting ends, the server retrieves all text data for the corresponding meeting ID from the database. The server uses the meeting ID as a key to extract the relevant text data.

[0876] Step 8:

[0877] The server sorts the acquired text data chronologically and automatically generates meeting minutes according to a standard format. The server concatenates the text data and adds appropriate paragraphs and headings.

[0878] Step 9:

[0879] The server generates meeting minutes and saves them to a database. The saved minutes are then managed so that users can access them later.

[0880] Step 10:

[0881] During the meeting, the server analyzes the text data and detects technical terms and difficult words. The server then uses natural language processing to create a list of these technical terms and difficult words.

[0882] Step 11:

[0883] The server uses a dictionary API to retrieve the meaning of each word it detects. The server sends a request to the dictionary API and receives the returned result.

[0884] Step 12:

[0885] The server stores the meaning of the words it retrieves in a database. The meaning data is also stored in association with the conference ID.

[0886] Step 13:

[0887] The user clicks on a word in the text data displayed during the meeting. The device detects the click event and requests information about that word from the server.

[0888] Step 14:

[0889] The server receives the request and retrieves the meaning of the corresponding word from the database. The server queries the semantic data and returns the results to the terminal.

[0890] Step 15:

[0891] The device displays semantic data in a pop-up window. Users can see the detailed meaning of the clicked word in real time.

[0892] These detailed processing steps enable a system where meeting content is transcribed in real time and the meanings of technical terms are provided immediately.

[0893] (Example 1)

[0894] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0895] Meetings often involve a lot of jargon and complex terminology, making it difficult to understand them in real time. Furthermore, manually creating meeting minutes is time-consuming and laborious, and may lack accuracy. Moreover, real-time understanding is even more challenging for participants with hearing impairments or those attending a meeting for the first time. A system is needed to address these challenges.

[0896] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0897] In this invention, the server includes means for transcribing meeting audio data in real time, means for storing the transcribed text data, means for automatically generating meeting minutes based on the transcribed text data, means for detecting technical terms and difficult words that appear during the meeting, means for obtaining the meanings of the detected technical terms and difficult words using a dictionary database or an external API, means for presenting the meanings of the technical terms and difficult words to the user in real time, and means for transmitting the audio data to the server at regular intervals for transcription. As a result, the meanings of technical terms and difficult words can be checked in real time during the meeting, which facilitates understanding of technical terms and enables the automatic generation of meeting minutes, saving time and effort.

[0898] "Audio data" refers to information recorded in digital format using audio.

[0899] "Transcription" refers to the process of converting audio data into text data.

[0900] "Text data" refers to digital information stored as characters.

[0901] "Meeting minutes" refers to a document that records the content and progress of discussions at a meeting.

[0902] "Technical jargon" refers to specialized terms used in a particular field or industry that are generally difficult for the average person to understand.

[0903] "Difficult words" refer to words with special meanings that are not easily understood by the general public.

[0904] A "dictionary database" refers to a digital information source that stores the meanings and definitions of words.

[0905] An "external API" refers to a program that provides an interface for interacting with other systems and services.

[0906] A "server" refers to a computer system that processes and stores data over a network.

[0907] "Terminal" refers to a computer or smart device that a user directly operates.

[0908] "User" refers to a person who uses this system.

[0909] "Real-time" refers to processing that occurs instantly without delay.

[0910] A "timestamp" refers to information that indicates the time when data was recorded.

[0911] "Standard format" refers to the structure and format of a document that follows certain rules and conventions.

[0912] The system of the present invention transcribes meeting audio data in real time, saves the text data, automatically generates meeting minutes, and presents the meanings of technical terms and difficult words to the user. The system of the present invention uses the following specific hardware and software.

[0913] First, the terminal records the meeting audio data in real time using its built-in or external microphone. The recorded audio data is sent to the server in buffer sizes, such as every 5 seconds. The terminal uses an audio recording library, such as the AudioRecord class, for audio recording.

[0914] The server sends the received audio data to the Google Speech-to-Text API, where it is converted into text data. The converted text data is temporarily stored in the server's memory and then stored in a database. Here, each text data entry is assigned a timestamp and a meeting ID. A relational database such as MySQL or PostgreSQL is used for the database.

[0915] After the meeting ends, the server retrieves all text data from the database based on the meeting ID and sorts it chronologically based on the timestamp. It then automatically generates meeting minutes according to a standard format. These minutes are stored in the database by the server for later reference by the user.

[0916] Furthermore, the server analyzes the transcribed text data to detect technical terms and difficult words. For these words, the server retrieves their meanings using dictionary databases such as the Oxford Dictionaries API or external APIs. The retrieved meanings are then stored in the database. When a user clicks on a word in the text data displayed during a meeting, the terminal displays a pop-up with the meaning of the word retrieved from the server. This allows users to understand technical terms simultaneously with the progress of the meeting.

[0917] As a concrete example, let's consider a scenario where a user hears the word "refactoring" at a meeting they are attending for the first time.

[0918] 1. The device records the meeting audio and sends the data to the server.

[0919] 2. The server receives "refactoring" as text data via the Google Speech-to-Text API and stores it temporarily.

[0920] 3. The server saves the text data to the database and simultaneously detects the word "refactoring".

[0921] 4. The server uses the Oxford Dictionaries API to retrieve the meaning of "refactoring" and saves it to the database.

[0922] 5. When a user clicks the word "refactoring" during a meeting, the terminal will display a pop-up with the meaning retrieved from the server.

[0923] This system allows users to understand the progress of a meeting in real time and to instantly check the meaning of technical terms and difficult words. As a result, understanding of the meeting content is enhanced, and a valuable environment is provided where users, especially those joining a particular organization for the first time or participants with hearing impairments, can quickly obtain useful information.

[0924] Example of a prompt:

[0925] "A meeting audio about a new system has been recorded. Please generate a program that converts the recorded audio data into text data, detects technical terms and difficult words, and displays their meanings. Please specify the names of the speech recognition API and dictionary API to be used, and provide a detailed explanation of the processing steps."

[0926] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0927] Step 1:

[0928] The device records meeting audio data in real time using its built-in microphone or an external microphone. The recorded audio data is buffered at a fixed buffer size (e.g., every 5 seconds). The input is the meeting audio data, and the output is the buffered audio data. This buffered audio data is temporarily held within the device for subsequent processing.

[0929] Step 2:

[0930] The terminal sends buffered audio data to the server at pre-configured times. The input is the buffered audio data, and the output is the audio data sent to the server. This audio data is transmitted via an HTTP request.

[0931] Step 3:

[0932] The server sends the received audio data to the Google Speech-to-Text API, where it is converted into text data. The input is the audio data sent to the server, and the output is the text data returned by the API. Audio data is sent via an HTTP request, and text data is received as a JSON response.

[0933] Step 4:

[0934] The server adds a timestamp and meeting ID to the received text data and saves it to the database. The input is the text data returned from the API, along with the current timestamp and meeting ID, and the output is the text data saved in the database. Specifically, an SQL query is created to insert the data into the database with the timestamp and meeting ID attached.

[0935] Step 5:

[0936] The server retrieves all text data from the database based on the meeting ID after the meeting ends and sorts it chronologically. The input is the meeting ID, and the output is the chronologically sorted text data. A database query is executed to retrieve the text data and sort it based on the timestamp.

[0937] Step 6:

[0938] The server automatically generates meeting minutes according to a standard format using sorted text data. The input is text data sorted chronologically, and the output is automatically generated meeting minutes. A dedicated template engine is used to format the text data into meeting minutes.

[0939] Step 7:

[0940] The server analyzes the transcribed text data to detect technical terms and obscure words. The input is text data, and the output is a list of detected words. A natural language processing library (e.g., spaCy) is used to analyze the text and extract specific key words.

[0941] Step 8:

[0942] The server retrieves the meanings of detected technical terms and difficult words using a dictionary database or an external API. The input is a list of detected words, and the output is the meaning information of those words. An API request is created, and the response is received in JSON format.

[0943] Step 9:

[0944] The server stores the retrieved semantic information in a database. The input is the semantic information of a word, and the output is the semantic information stored in the database. An SQL query is created to insert the word and its meaning into the database.

[0945] Step 10:

[0946] The terminal retrieves the meaning of a word from the server and displays it in a pop-up window when the user clicks on a word in the text data displayed during a meeting. The input is the clicked word, and the output is the meaning of the word displayed in the pop-up window. An AJAX request is sent, and the retrieved data is processed and displayed using JavaScript.

[0947] These steps enable real-time transcription during meetings, automatic generation of meeting minutes, and retrieval and display of the meanings of technical terms.

[0948] (Application Example 1)

[0949] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0950] Meetings and work instructions within the factory often contain a lot of technical jargon and difficult-to-understand words, making it challenging for new or first-time workers to quickly respond to and understand them. Furthermore, there is a lack of systems to accurately transcribe meeting and instruction content in real time and present it in an easily understandable format.

[0951] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0952] In this invention, the server includes means for transcribing meeting audio data in real time, means for storing the transcribed text data, means for automatically generating meeting minutes based on the transcribed text data, means for detecting technical terms and difficult words that appear during the meeting and obtaining their meanings, means for presenting the meanings of the technical terms and difficult words to the user, and means for transcribing work instructions and meeting audio in the factory in real time and providing workers with the meanings of the detected technical terms. This enables new workers and workers participating for the first time to quickly understand technical terms and difficult words and accurately grasp the content of work instructions and meetings in the factory.

[0953] "Meeting audio data" refers to data that digitally records the speech spoken during a meeting.

[0954] "Methods for real-time transcription" refers to technologies or devices that instantly convert audio from meetings, work instructions, etc., into text data.

[0955] "Transcribed text data" refers to text data converted from audio in real time.

[0956] "Means for saving text data" refers to the technology or device used to save converted text data to a storage device or similar.

[0957] "Means for automatically generating meeting minutes" refers to a technology or device that automatically generates a document recording the content of a meeting by rearranging saved text data in chronological order.

[0958] "Technical terms and difficult vocabulary" refers to specialized terms or words that are not commonly used or are difficult to understand.

[0959] "Means for detecting technical terms and difficult words and obtaining their meanings" refers to technologies or devices that discover technical terms and difficult words within text data and obtain their meanings from external dictionary databases or APIs.

[0960] "Means of presenting to the user" refers to a technology or device that displays the meaning of acquired words in a format that is easy for the user to understand.

[0961] "A means of transcribing audio of work instructions and meetings within a factory in real time" refers to a technology or device that instantly converts audio of work instructions and meetings conducted within a factory into text data.

[0962] "Means of providing workers with the meaning of detected technical terms" refers to a technology or device that displays the meaning of detected technical terms in real time in a format that is easy for workers to understand.

[0963] This invention is a system for improving the efficiency of meetings and work instructions within a factory. This system supports rapid understanding and response by transcribing audio data in real time, saving and analyzing text data, automatically generating meeting minutes, and immediately presenting the meanings of technical terms and difficult words to the user.

[0964] 1. Hardware and software to be used

[0965] The system uses the following hardware and software.

[0966] Hardware:

[0967] Microphone (for recording meetings and work instructions)

[0968] Server (for data processing and management)

[0969] Terminals in stores or factories (to display information to users)

[0970] software:

[0971] Speech recognition API (e.g., Google Speech Recognition API): Used to convert recorded audio into text data.

[0972] Dictionary API (external dictionary database): For obtaining the meanings of technical terms and difficult words.

[0973] Database management system: For storing and managing text data and semantic information.

[0974] Dedicated application: To display information to the user in real time.

[0975] 2. Overall System Functions and Processing

[0976] Audio data collection and transcription

[0977] The terminal records audio data of meetings and work instructions in real time and sends the data to the server in fixed buffer sizes. The server sends the received audio data to a speech recognition API and converts it into text data. This text data is temporarily stored in the server's memory.

[0978] Saving text data

[0979] The server stores the transcribed text data obtained from the speech recognition API into a database. A timestamp and meeting ID are added to the text data. This is useful for later searching and extracting specific meeting content.

[0980] Automatic generation of meeting minutes

[0981] After the meeting ends, or at an appropriate time, the server retrieves all text data from the database based on the meeting ID. The server sorts the retrieved text data chronologically and automatically generates meeting minutes according to a standard format. These minutes are stored in the database for later reference by the user.

[0982] Acquisition and presentation of the meanings of technical terms and difficult words.

[0983] The server analyzes the transcribed text data to detect technical terms and difficult words. For detected words, the server retrieves their meaning using a dictionary database or an external API (dictionary API). The retrieved meanings are stored in the database. When a user clicks on a word in the text data displayed in real time, the meaning of the corresponding word is displayed in a pop-up window.

[0984] 3. Specific examples and prompt messages

[0985] Specific example

[0986] Scenario: A new worker in a factory is instructed for the first time about the term "PMI (Preventive Maintenance)."

[0987] 1. Record audio in real time and send it to the server.

[0988] 2. Translate "PMI" into text using a speech recognition API.

[0989] 3. Save to the database.

[0990] 4. Use the dictionary API to retrieve and save the meaning of "PMI" (preventive maintenance).

[0991] 5. When a new employee clicks on "PMI," its meaning will be displayed.

[0992] Example of a prompt

[0993] Retrieve the meaning of the technical term "PMI" clicked by the user from the dictionary database and display it in real time as "preventive maintenance".

[0994] This system allows for quick understanding of work instructions and meeting content within the factory, enabling new workers and those participating for the first time to adapt smoothly.

[0995] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0996] Step 1:

[0997] The terminal records audio of meetings and work instructions within the factory in real time.

[0998] Input: Audio from meetings or work instructions

[0999] Output: Recorded audio data

[1000] Specific operation: The device's microphone collects audio and stores the data at regular intervals of a certain buffer size.

[1001] Step 2:

[1002] The device sends the recorded audio data to the server in increments of a certain buffer size.

[1003] Input: Recorded audio data

[1004] Output: Audio data sent to the server

[1005] Specific operation: Once the audio data buffer is full, data transfer to the server begins.

[1006] Step 3:

[1007] The server sends the received audio data to a speech recognition API, where it is converted into text data.

[1008] Input: Sent audio data

[1009] Output: Text data returned from the speech recognition API

[1010] Specific operation: The server sends the received audio data to a speech recognition service such as the Google Speech Recognition API.

[1011] Step 4:

[1012] The server temporarily stores the converted text data in memory.

[1013] Input: Text data returned from the speech recognition API

[1014] Output: Text data stored in memory

[1015] Specific operation: The converted text data is given a meeting ID and timestamp, and then saved to the server's memory.

[1016] Step 5:

[1017] The server saves text data to the database.

[1018] Input: Text data stored in memory

[1019] Output: Text data stored in the database

[1020] Specific operation: Store the text data saved in memory in the database along with the meeting ID and timestamp.

[1021] Step 6:

[1022] After the meeting ends, the server retrieves all text data from the database based on the meeting ID.

[1023] Input: Meeting ID

[1024] Output: All text data related to the meeting

[1025] Specific operation: Use a database query to extract text data corresponding to a specified meeting ID.

[1026] Step 7:

[1027] The server sorts the acquired text data chronologically and automatically generates meeting minutes.

[1028] Input: Acquired text data

[1029] Output: Auto-generated meeting minutes

[1030] Specific operation: Sort text data based on timestamps and generate meeting minutes according to a standard format.

[1031] Step 8:

[1032] The server saves meeting minutes to a database so users can access them later.

[1033] Input: Generated meeting minutes

[1034] Output: Meeting minutes saved in the database

[1035] Specific action: Store the generated meeting minutes in the database.

[1036] Step 9:

[1037] The server analyzes the transcribed text data to detect technical terms and difficult words.

[1038] Input: Transcribed text data

[1039] Output: Detected technical terms and difficult words

[1040] Specific operation: Use a text analysis algorithm to identify technical terms and difficult words within the text data.

[1041] Step 10:

[1042] The server retrieves the meaning of a word detected using a dictionary API or dictionary database and stores it in the database.

[1043] Input: Detected technical terms or difficult words

[1044] Output: Meaning of the retrieved word

[1045] Specific operation: Access the dictionary API, retrieve the meaning of the detected word, and store it in the database.

[1046] Step 11:

[1047] When a user clicks on a technical term during a meeting or while working, the device retrieves the meaning of the word from the server and displays it in a pop-up window.

[1048] Input: Technical terms clicked by the user

[1049] Output: Meaning of the word displayed in the pop-up window

[1050] Specific operation: When the user clicks, the meaning of the corresponding word is retrieved from the server and displayed on the device.

[1051] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[1052] The system implementing this invention transcribes meeting audio data in real time and saves the text data. Furthermore, it has a function to automatically generate meeting minutes based on the saved text data, and also detects technical terms and difficult words that appear during the meeting, retrieves their meanings, and presents them to the user. In addition to this basic configuration, the present invention incorporates an emotion engine that recognizes the user's emotional state and appropriately adjusts the meanings of technical terms and difficult words based on the results.

[1053] Functional Embodiments

[1054] 1. Collection and transcription of audio data

[1055] The terminal records audio data in real time during the meeting. The recorded data is stored in a buffer and sent to the server at regular intervals. The server sends the received audio data to a speech recognition API and converts it into text data. The converted text data is stored in a database along with a timestamp and meeting ID.

[1056] 2. Automatic generation of meeting minutes

[1057] After the meeting ends, or at an appropriate time, the server retrieves all text data from the database based on the meeting ID. The retrieved text data is sorted chronologically, and meeting minutes are automatically generated according to a standard format. The automatically generated minutes are saved in the database for later reference by the user.

[1058] 3. Acquisition and presentation of the meanings of technical terms and difficult words.

[1059] The server analyzes the transcribed text data to detect technical terms and difficult words. For detected words, the server uses a dictionary database or dictionary API to retrieve their meanings. The retrieved meanings are stored in the database. When the user clicks on a word, the terminal displays its meaning in a pop-up window.

[1060] 4. Introduction of an emotional engine

[1061] The terminal or server recognizes the user's emotional state in real time during the meeting. The emotion engine analyzes the user's voice tone, facial expressions, or biosignals to identify their emotional state. Specifically, it can distinguish emotions such as anger, joy, surprise, and sadness.

[1062] 5. Adjusting information presentation based on emotions

[1063] Based on the user's emotional state, the server appropriately adjusts the meaning of technical terms and difficult words. For example, if the user is confused, it can add more detailed explanations and examples. On the other hand, if the server determines that the user understands, it will only provide a concise explanation.

[1064] Specific example

[1065] 1. At a meeting they are attending for the first time, the user hears the word "refactoring."

[1066] The device records the meeting audio and sends the data to the server.

[1067] The server receives "refactoring" as text data via the speech recognition API and stores it temporarily.

[1068] The server saves that text data to the database and simultaneously detects the word "refactoring".

[1069] The server uses a dictionary API to retrieve the meaning of "refactoring" and saves it to the database.

[1070] 2. The user clicks the word "refactoring".

[1071] The device detects a click event and requests information about that word from the server.

[1072] The server receives the request, retrieves the meaning of the corresponding word from the database, and returns it to the terminal.

[1073] The device displays semantic data as a pop-up.

[1074] 3. Simultaneously, the emotion engine analyzes the user's emotional state and determines that they are confused.

[1075] Based on the sentiment information acquired by the server, more detailed explanations and additional information will be included in the pop-up.

[1076] This system allows users to understand the progress of meetings in real time, comprehend technical terms, and receive appropriate information tailored to their emotional state. In this way, it provides an easily understandable meeting environment, even for users joining a particular organization for the first time or participants with hearing impairments.

[1077] The following describes the processing flow.

[1078] Step 1:

[1079] The device records the meeting audio in real time. Recording begins when the user launches the meeting recording application and presses the start recording button. The recorded data is stored in the audio buffer.

[1080] Step 2:

[1081] The terminal sends data stored in the audio buffer to the server at regular intervals. For example, it sends all the audio data accumulated in the buffer to the server every 10 seconds.

[1082] Step 3:

[1083] The server sends the received audio data to a speech recognition API. The server creates an API request and passes the audio data to the API. Examples of APIs used include Google Speech-to-Text and IBM Watson.

[1084] Step 4:

[1085] The speech recognition API converts the audio data into text and returns the result to the server. The API analyzes the audio data and generates text data in string format.

[1086] Step 5:

[1087] The server receives the text data returned from the speech recognition API and temporarily stores it in memory. At this time, a timestamp and meeting ID are added to the text data.

[1088] Step 6:

[1089] The server saves the text data to the database. A meeting ID and timestamp are added during saving, making it useful for later searching and generating meeting minutes.

[1090] Step 7:

[1091] After the meeting ends, the server retrieves all text data from the database based on the corresponding meeting ID. After retrieval, the server sorts the text data chronologically.

[1092] Step 8:

[1093] The server automatically generates meeting minutes based on sorted text data, following a standard format. The automatically generated minutes are formatted for readability, with paragraphs and headings neatly organized.

[1094] Step 9:

[1095] The server saves the generated meeting minutes to a database. It also adds an index to allow users to access them later.

[1096] Step 10:

[1097] During the meeting, the server analyzes the text data transmitted in real time to detect technical terms and difficult words. Natural language processing (NLP) techniques are used to list words that match specific criteria.

[1098] Step 11:

[1099] The server uses a dictionary API to retrieve the meaning of each word it detects. The data returned from the dictionary API includes a detailed definition of the word and example sentences.

[1100] Step 12:

[1101] The server stores the meaning of the words it retrieves in a database. Since the meaning data is also stored in association with the meeting ID, users can easily access it later.

[1102] Step 13:

[1103] The device is equipped with sensors to recognize the user's emotional state in real time, collecting voice tone, facial expression analysis, or biosignals. The emotion engine analyzes emotions such as anger, joy, surprise, and sadness in real time.

[1104] Step 14:

[1105] The server receives emotional state data from the emotion engine and adjusts the meaning of technical terms and complex words. For example, if it determines that the user is confused, it adds a more detailed explanation.

[1106] Step 15:

[1107] The user clicks on a word in the text data during a meeting. The device detects the click event and sends a request to the server regarding that word.

[1108] Step 16:

[1109] The server receives the request, retrieves the meaning of the corresponding word from the database, makes adjustments according to the emotional state, and then returns it to the terminal.

[1110] Step 17:

[1111] The device displays semantic data it has acquired as a pop-up. Users can see the meaning of the word and additional information adjusted by the sentiment engine in real time.

[1112] This detailed processing step enables a system that not only transcribes meeting content in real time, but also provides appropriate meanings for technical terms based on the user's emotional state.

[1113] (Example 2)

[1114] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[1115] Traditional conferencing systems made it difficult to transcribe meeting audio data in real time, resulting in significant time and effort required for creating meeting minutes. Furthermore, understanding the meaning of specialized terminology and complex vocabulary used during meetings was challenging, posing a significant barrier to comprehension, especially for new or non-expert users. Additionally, the inability to consider the emotional state of participants during meetings led to inappropriate information presentation, posing a significant problem.

[1116] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for transcribing the audio data of a meeting in real time, means for storing the transcribed text data, means for automatically generating meeting minutes based on the transcribed text data, means for detecting technical terms and difficult words that appear during the meeting and obtaining their meanings, means for presenting the meanings of the technical terms and difficult words to the user, and means for recognizing the user's emotions in real time and adjusting the meanings of the technical terms and difficult words. This enables real-time transcription of meeting content, automatic generation of meeting minutes, immediate understanding of technical terms, and presentation of appropriate information based on the user's emotional state.

[1117] "Meeting audio data" refers to all audio generated during a meeting, including discussions, debates, and the content of statements made.

[1118] "Real-time" refers to an event or process that occurs immediately, without any delays or waiting periods, and is reflected instantly.

[1119] "Transcription" is the process of converting audio data into text format, recording the content of the audio as text data.

[1120] "Text data" refers to a data format composed of characters and symbols, and is data generated as a result of transcribing audio data into text.

[1121] "Saving" refers to storing data in a way that makes it usable for a long period of time, and is the act of recording it in a database or storage device.

[1122] "Meeting minutes" are documents that record what was said and decided during a meeting, and are official records intended for later reference.

[1123] "Automatic generation" means creating data or documents using specific algorithms or programs without human intervention.

[1124] "Detecting" refers to the act of finding specific elements or patterns within data or information, and this typically involves using analytical techniques.

[1125] "Acquiring" refers to the act of extracting necessary information from external sources or databases, and this is often done through APIs.

[1126] A "user" refers to a person or organization that uses a system or service, and is the entity that performs operations through the system interface.

[1127] "Recognizing emotions in real time" means instantly understanding a user's emotional state using specific technologies (such as voice analysis or facial expression analysis).

[1128] "Adjusting" refers to the act of changing the content or settings based on specific criteria or conditions to achieve an optimal state.

[1129] In an embodiment of the present invention, the system is capable of transcribing meeting audio data in real time and saving the text data. Furthermore, it can automatically generate meeting minutes based on the saved text data, and in addition, it can detect technical terms and difficult words that appear during the meeting, obtain their meanings, and present them to the user. Moreover, the present invention provides a system that recognizes the user's emotional state in real time and appropriately adjusts the meanings of technical terms and difficult words.

[1130] 1. Hardware and software to be used

[1131] The hardware requires devices (PCs, smartphones, tablets, etc.) to be used by meeting participants, and these devices must include high-sensitivity microphones. The server will be either a cloud environment or an on-premises server.

[1132] The software used includes a speech recognition API (e.g., Google Cloud Speech-to-Text API), a database (e.g., PostgreSQL), and a natural language processing library (e.g., spaCy). External dictionary APIs (e.g., Wiktionary API) are used to obtain the meanings of technical terms and complex words. Speech tone analysis (e.g., Microsoft Azure Emotion API) and facial expression analysis (e.g., OpenCV and Dlib) are used to recognize emotional states. Furthermore, generative AI models (e.g., OpenAI GPT-3) are used to refine the meanings of technical terms.

[1133] 2. Processing flow and specific data processing / calculations

[1134] The terminal records audio data during the meeting in real time, and the recorded data is temporarily stored in a buffer. At regular intervals, the terminal divides the data in the buffer into packets and sends them to the server via the internet. The HTTPS protocol is used for this.

[1135] The server sends the received audio data to a speech recognition API for transcription. The converted audio data is stored in a database along with a timestamp and meeting ID.

[1136] After the meeting ends, or at an appropriate time, the server retrieves all text data from the database based on the meeting ID. The retrieved text data is sorted chronologically, and meeting minutes are automatically generated using a natural language processing library. The automatically generated minutes are saved in the database for later reference by the user.

[1137] The server analyzes stored text data to detect technical terms and obscure words. This uses techniques such as TF-IDF and morphological analysis. The meaning of detected words is retrieved from a dictionary database or via an external API and stored in the database.

[1138] When a user clicks on a word, the device detects the click event and requests information about that word from the server. The server receives the request, retrieves the meaning data for the corresponding word from its database, and returns it to the device. The device then displays the meaning data in a pop-up window.

[1139] During the meeting, terminals and servers analyze voice tone and facial expressions in real time to recognize emotional states. An emotion engine evaluates this data to determine the user's specific emotional state (e.g., confused, understanding). If confused is detected, the server uses a generative AI model to further refine the meaning of technical terms and complex words, including detailed explanations and additional information.

[1140] 3. Specific Examples and Examples of Prompt Statements

[1141] As a concrete example, if a user hears the word "refactoring" in a meeting they are attending for the first time and does not understand it, it will behave as follows:

[1142] The device records the meeting audio and sends the data to the server.

[1143] The server receives "refactoring" as text data via the speech recognition API and stores it temporarily.

[1144] The server saves that text data to a database and detects the word "refactoring".

[1145] The server uses a dictionary API to retrieve the meaning of "refactoring" and saves it to the database.

[1146] When a user clicks the word "refactoring," the device detects the click event and requests information about that word from the server.

[1147] The server receives the request, retrieves the meaning of the corresponding word from the database, and returns it to the terminal.

[1148] The device displays semantic data as a pop-up.

[1149] At the same time, the emotion engine analyzes the user's emotional state and determines that they are confused.

[1150] Based on the sentiment information acquired by the server, more detailed explanations and additional information will be included in the pop-up.

[1151] Example of a prompt:

[1152] "Could you explain the meaning of refactoring?"

[1153] "Please explain refactoring with an example."

[1154] "Please explain how refactoring can benefit a project."

[1155] This invention makes it easier for participants to understand the progress of a meeting in real time, to immediately grasp the meaning of technical terms and difficult words, and to create an environment where appropriate information is provided according to their emotional state.

[1156] The flow of the specific processing in Example 2 will be explained using Figure 13.

[1157] Step 1: Collect audio data

[1158] The device records audio data during the meeting in real time. Audio acquired through a high-sensitivity microphone is temporarily stored in a buffer. The input is the meeting audio data, and the output is the audio data stored in the buffer. Recording software runs on the device, sampling the audio data second by second and saving it to the buffer.

[1159] Step 2: Sending the audio data

[1160] The terminal periodically divides the audio data in its buffer into packets and sends them to the server via the internet. The input is the audio data in the buffer, and the output is the audio data packets sent to the server. This process is performed securely using the HTTPS protocol.

[1161] Step 3: Transcribing the audio data

[1162] The server receives audio data and sends it to a speech recognition API (e.g., Google Cloud Speech-to-Text API) to convert it into text data. The input is audio data, and the output is transcribed text data. The converted text data is temporarily stored in server memory.

[1163] Step 4: Save text data

[1164] The server saves the transcribed text data, along with the timestamp and meeting ID, to a database. The input is the text data, timestamp, and meeting ID, and the output is the text data saved in the database. The saving process is performed using an SQL query.

[1165] Step 5: Automatic generation of meeting minutes

[1166] After the meeting ends, the server retrieves text data from the database based on the meeting ID. The input is the meeting ID, and the output is the retrieved text data. The server sorts the retrieved text data chronologically and automatically generates meeting minutes using a natural language processing library (e.g., spaCy). The automatically generated meeting minutes are then saved back to the database.

[1167] Step 6: Detection of technical terms and difficult words

[1168] The server analyzes stored text data and uses TF-IDF and morphological analysis to detect technical terms and difficult words. The input is stored text data, and the output is a list of detected words. Natural language processing techniques are used for analysis, and words of high importance are extracted.

[1169] Step 7: Acquire the meaning of technical terms and difficult words.

[1170] The server queries a dictionary database (e.g., the Wiktionary API) for the detected words and retrieves their meanings. The input is a list of detected words, and the output is a list of retrieved meanings. The retrieved word meanings are then saved back into the database.

[1171] Step 8: Presentation of technical terms and difficult vocabulary.

[1172] The device detects an event when the user clicks on a word and requests that information from the server. The input is the click event, and the output is the requested word information. The server receives the request, retrieves the meaning data for the corresponding word from the database, and returns it to the device. The device then displays the meaning data in a pop-up window.

[1173] Step 9: Recognizing your emotional state

[1174] The terminal or server analyzes the user's voice tone and facial expressions in real time during a meeting to recognize their emotional state. Input is voice tone data and facial expression data, and output is the analyzed emotional state data. This utilizes the Microsoft Azure Emotion API, OpenCV, and Dlib.

[1175] Step 10: Adjusting the Information Presentation

[1176] The server adjusts the meaning of technical terms and difficult words appropriately based on the user's emotional state. The input is emotional state data and the meaning of technical terms, while the output is the adjusted semantic data. If the user is confused, a generative AI model (e.g., OpenAI GPT-3) is used to generate a detailed explanation, which is returned to the device. Based on this returned data, the device displays more detailed information in a pop-up window.

[1177] (Application Example 2)

[1178] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[1179] Conventional meeting recording systems often simply convert meeting audio data into text and automatically generate meeting minutes. However, these systems often make it difficult to understand the meaning of technical terms and complex vocabulary, and they fail to provide appropriate information based on the user's emotional state, making it difficult to understand the progress of the meeting and its subsequent analysis. Furthermore, they lack the functionality to respond in real time when security risks occur, requiring immediate countermeasures.

[1180] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for transcribing the audio data of a meeting in real time, means for storing the transcribed text data, means for automatically generating meeting minutes based on the transcribed text data, means for detecting technical terms and difficult words that appear during the meeting and obtaining their meanings, means for presenting the meanings of the technical terms and difficult words to the user, means for analyzing the user's emotional state and adjusting the content of the terms presented based on that emotional state, and means for sending emergency alerts based on specific terms detected during the conversation and the user's emotional state. As a result, the user can grasp the progress of the meeting in real time, easily understand the meanings of technical terms and difficult words, be provided with appropriate information according to their emotional state, and respond immediately to security risks.

[1181] "Audio data" refers to digital recordings of voices generated during meetings, conversations, and other similar events.

[1182] "Transcription" is the process of analyzing audio data and converting it into text data.

[1183] "Text data" refers to data expressed as characters, specifically a string of characters representing the content of a conversation generated through transcription.

[1184] "Meeting minutes" are documents that summarize the content of a meeting, recording the statements made and decisions made during the meeting.

[1185] "Technical terms" are specific terms used in a particular field or industry, and are difficult to understand with general knowledge.

[1186] "Difficult words" are words that are hard to understand or words that are not commonly used.

[1187] "Emotional state" refers to the user's emotional condition and includes a variety of emotions such as joy, anger, sadness, and surprise.

[1188] An "emergency alert" is a notification sent when an emergency occurs or specific conditions are met, and it is a warning that requires immediate action.

[1189] "User" refers to a person who uses a system or application.

[1190] A "server" is a computer system that processes information over a network and provides services to clients.

[1191] This invention provides a system that supports the progress of meetings in real time, enabling the understanding of important information and the management of security risks. A system implementing this invention includes the following hardware and software:

[1192] 1. Collection and transcription of audio data:

[1193] The terminal records audio data in real time during the meeting. The recorded data is stored in a buffer and sent to the server at regular intervals. The server converts the received audio data into text data using the Google Web Speech API. The converted text data is stored in a database along with a timestamp and meeting ID.

[1194] 2. Automatic generation of meeting minutes:

[1195] After the meeting ends, or at an appropriate time, the server retrieves all text data from the database based on the meeting ID. The retrieved text data is sorted chronologically, and meeting minutes are automatically generated according to a standard format. The automatically generated minutes are saved in the database for later reference by the user.

[1196] 3. Acquisition and presentation of the meanings of technical terms and difficult words:

[1197] The server analyzes the transcribed text data to detect technical terms and difficult words. For detected words, the server uses a dictionary database or dictionary API to retrieve their meanings. The retrieved meanings are stored in the database. When the user clicks on a word, the terminal displays its meaning in a pop-up window.

[1198] 4. Introduction of the Emotion Engine:

[1199] The terminal or server recognizes the user's emotional state in real time during the meeting. The emotion engine analyzes the user's voice tone, facial expressions, or biosignals to identify their emotional state. Specifically, it can distinguish emotions such as anger, joy, surprise, and sadness.

[1200] 5. Adjusting information presentation based on emotions:

[1201] Based on the user's emotional state, the server appropriately adjusts the meaning of technical terms and difficult words. For example, if the user is confused, it can add more detailed explanations and examples. On the other hand, if the server determines that the user understands, it will only provide a concise explanation.

[1202] 6. Sending an emergency alert:

[1203] The server sends emergency alerts based on specific security terms detected during the conversation (e.g., "confidential" or "threat") and the user's emotional state. If the emotion engine detects high levels of tension or anxiety in the user, a notification is sent to the administrator.

[1204] Specific example

[1205] 1. An emergency occurs during the conversation:

[1206] "We have received a report that an intruder entered the office during our conversation. We need to take immediate action."

[1207] This allows users to understand the progress of meetings in real time, comprehend technical terms and complex vocabulary more easily, receive appropriate information tailored to their emotional state, and respond immediately to security risks.

[1208] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[1209] Step 1:

[1210] The terminal records the audio data of the meeting in real time. The input is the audio data of the meeting, which the terminal stores in a buffer in digital audio format. The output is digital audio data that is sent to the server at regular intervals.

[1211] Step 2:

[1212] The server sends the received audio data to the Google Web Speech API, where it is converted into text data. The input is digital audio data, and the server uses a speech recognition API to transcribe it. The output is text data stored on the server along with a timestamp and meeting ID.

[1213] Step 3:

[1214] The server analyzes stored text data to detect technical terms and obscure words. The input is text data, and the server uses a natural language processing engine to identify these terms. The output is a list of detected words.

[1215] Step 4:

[1216] The server retrieves the meaning of detected words using a dictionary database or dictionary API. The input is a list of detected words, and the server generates an API request to retrieve the meaning of each word. The output is the meaning information and explanation for each word.

[1217] Step 5:

[1218] The server stores the semantic information of words it has retrieved in a database, and the terminal displays that semantic information upon user request. Input consists of user click events and word requests; the terminal retrieves the word's meaning from the server and displays it in a pop-up. Output is the semantic information presented to the user.

[1219] Step 6:

[1220] The terminal or server recognizes the user's emotional state in real time during a meeting. Inputs include the user's voice tone, facial expressions, or biosignals, which the emotion engine analyzes to identify the emotional state. Output is data on the recognized emotional state.

[1221] Step 7:

[1222] Based on the user's emotional state, the server appropriately adjusts the meaning of technical terms and difficult words. Input consists of emotional state data and word semantic information; the server includes detailed explanations and additional information if the user is confused. Output is the adjusted semantic information.

[1223] Step 8:

[1224] The server sends an emergency alert based on specific security terms and the user's emotional state detected during the conversation. The input is the detected security terms and user emotional state data, which the server uses to generate an alert notification and send it to the administrator. The output is the emergency alert sent to the administrator.

[1225] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[1226] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1227] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[1228] [Fourth Embodiment]

[1229] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[1230] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1231] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1232] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[1233] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[1234] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[1235] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[1236] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[1237] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[1238] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1239] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1240] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[1241] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1242] The system for implementing the present invention has the following configuration. First, it transcribes the audio data of a meeting in real time and saves the text data. Next, it automatically generates meeting minutes based on the saved text data, and further detects technical terms and difficult words that appear during the meeting, obtains their meanings, and presents them to the user.

[1243] Functional Embodiments

[1244] 1. Collection and transcription of audio data

[1245] The terminal records audio data during the meeting in real time. The recorded audio data is sent to the server in increments of a certain buffer size. The server sends the received audio data to a speech recognition API and converts it into text data. This text data is temporarily stored in the server's memory.

[1246] 2. Saving text data

[1247] The server stores the transcribed text data obtained from the speech recognition API in a database. It is important to add a timestamp and meeting ID to the text data. This is helpful for later searching and extracting specific meeting content.

[1248] 3. Automatic generation of meeting minutes

[1249] After the meeting ends, or at an appropriate time, the server retrieves all text data from the database based on the meeting ID. The server sorts the retrieved text data chronologically and automatically generates meeting minutes according to a standard format. These minutes are stored in the database for later reference by the user.

[1250] 4. Acquisition and presentation of the meanings of technical terms and difficult words.

[1251] The server analyzes the transcribed text data to detect technical terms and difficult words. For detected words, the server retrieves their meaning using a dictionary database or an external API (dictionary API). The retrieved meanings are stored in the database.

[1252] The device displays a pop-up showing the meaning of a word when the user clicks on it in the text data displayed during the meeting. This display is provided in real time, allowing users to understand technical terms as the meeting progresses.

[1253] Specific example

[1254] At a meeting they are attending for the first time, the user hears the word "refactoring."

[1255] 1. The device records the meeting audio and sends the data to the server.

[1256] 2. The server receives "refactoring" as text data via the speech recognition API and stores it temporarily.

[1257] 3. The server saves the text data to the database and simultaneously detects the word "refactoring".

[1258] 4. The server uses the dictionary API to retrieve the meaning of "refactoring" and saves it to the database.

[1259] 5. When a user clicks the word "refactoring" during a meeting, the terminal will display a pop-up with the meaning retrieved from the server.

[1260] This system allows users to understand the progress of a meeting in real time and to instantly check the meaning of technical terms and difficult words. This deepens understanding of the meeting content and provides an environment where users participating in a particular organization for the first time, or participants with hearing impairments, can quickly obtain useful information.

[1261] The following describes the processing flow.

[1262] Step 1:

[1263] The device records the meeting audio in real time. Recording begins when the user launches the meeting recording application during the meeting and presses the start recording button. The recorded data is stored in a buffer.

[1264] Step 2:

[1265] The device sends the recorded data to the server at regular intervals. For example, it clears the buffer every 10 seconds and sends the current recorded data to the server.

[1266] Step 3:

[1267] The server sends the received audio data to a speech recognition API (e.g., Google Speech-to-Text). The server creates an API request and passes the audio data to the API.

[1268] Step 4:

[1269] The speech recognition API converts the audio data into text and returns it to the server. The API parses the audio data and returns the result in string format.

[1270] Step 5:

[1271] The server receives the text data and temporarily stores it in memory. The text data is stored along with the meeting ID and timestamp.

[1272] Step 6:

[1273] The server saves the text data to the database. The saved data is tagged with a meeting ID and timestamp, making it easy to search later.

[1274] Step 7:

[1275] After the meeting ends, the server retrieves all text data for the corresponding meeting ID from the database. The server uses the meeting ID as a key to extract the relevant text data.

[1276] Step 8:

[1277] The server sorts the acquired text data chronologically and automatically generates meeting minutes according to a standard format. The server concatenates the text data and adds appropriate paragraphs and headings.

[1278] Step 9:

[1279] The server generates meeting minutes and saves them to a database. The saved minutes are then managed so that users can access them later.

[1280] Step 10:

[1281] During the meeting, the server analyzes the text data and detects technical terms and difficult words. The server then uses natural language processing to create a list of these technical terms and difficult words.

[1282] Step 11:

[1283] The server uses a dictionary API to retrieve the meaning of each word it detects. The server sends a request to the dictionary API and receives the returned result.

[1284] Step 12:

[1285] The server stores the meaning of the words it retrieves in a database. The meaning data is also stored in association with the conference ID.

[1286] Step 13:

[1287] The user clicks on a word in the text data displayed during the meeting. The device detects the click event and requests information about that word from the server.

[1288] Step 14:

[1289] The server receives the request and retrieves the meaning of the corresponding word from the database. The server queries the semantic data and returns the results to the terminal.

[1290] Step 15:

[1291] The device displays semantic data in a pop-up window. Users can see the detailed meaning of the clicked word in real time.

[1292] These detailed processing steps enable a system where meeting content is transcribed in real time and the meanings of technical terms are provided immediately.

[1293] (Example 1)

[1294] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1295] Meetings often involve a lot of jargon and complex terminology, making it difficult to understand them in real time. Furthermore, manually creating meeting minutes is time-consuming and laborious, and may lack accuracy. Moreover, real-time understanding is even more challenging for participants with hearing impairments or those attending a meeting for the first time. A system is needed to address these challenges.

[1296] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[1297] In this invention, the server includes means for transcribing meeting audio data in real time, means for storing the transcribed text data, means for automatically generating meeting minutes based on the transcribed text data, means for detecting technical terms and difficult words that appear during the meeting, means for obtaining the meanings of the detected technical terms and difficult words using a dictionary database or an external API, means for presenting the meanings of the technical terms and difficult words to the user in real time, and means for transmitting the audio data to the server at regular intervals for transcription. As a result, the meanings of technical terms and difficult words can be checked in real time during the meeting, which facilitates understanding of technical terms and enables the automatic generation of meeting minutes, saving time and effort.

[1298] "Audio data" refers to information recorded in digital format using audio.

[1299] "Transcription" refers to the process of converting audio data into text data.

[1300] "Text data" refers to digital information stored as characters.

[1301] "Meeting minutes" refers to a document that records the content and progress of discussions at a meeting.

[1302] "Technical jargon" refers to specialized terms used in a particular field or industry that are generally difficult for the average person to understand.

[1303] "Difficult words" refer to words with special meanings that are not easily understood by the general public.

[1304] A "dictionary database" refers to a digital information source that stores the meanings and definitions of words.

[1305] An "external API" refers to a program that provides an interface for interacting with other systems and services.

[1306] A "server" refers to a computer system that processes and stores data over a network.

[1307] "Terminal" refers to a computer or smart device that a user directly operates.

[1308] "User" refers to a person who uses this system.

[1309] "Real-time" refers to processing that occurs instantly without delay.

[1310] A "timestamp" refers to information that indicates the time when data was recorded.

[1311] "Standard format" refers to the structure and format of a document that follows certain rules and conventions.

[1312] The system of the present invention transcribes meeting audio data in real time, saves the text data, automatically generates meeting minutes, and presents the meanings of technical terms and difficult words to the user. The system of the present invention uses the following specific hardware and software.

[1313] First, the terminal records the meeting audio data in real time using its built-in or external microphone. The recorded audio data is sent to the server in buffer sizes, such as every 5 seconds. The terminal uses an audio recording library, such as the AudioRecord class, for audio recording.

[1314] The server sends the received audio data to the Google Speech-to-Text API, where it is converted into text data. The converted text data is temporarily stored in the server's memory and then stored in a database. Here, each text data entry is assigned a timestamp and a meeting ID. A relational database such as MySQL or PostgreSQL is used for the database.

[1315] After the meeting ends, the server retrieves all text data from the database based on the meeting ID and sorts it chronologically based on the timestamp. It then automatically generates meeting minutes according to a standard format. These minutes are stored in the database by the server for later reference by the user.

[1316] Furthermore, the server analyzes the transcribed text data to detect technical terms and difficult words. For these words, the server retrieves their meanings using dictionary databases such as the Oxford Dictionaries API or external APIs. The retrieved meanings are then stored in the database. When a user clicks on a word in the text data displayed during a meeting, the terminal displays a pop-up with the meaning of the word retrieved from the server. This allows users to understand technical terms simultaneously with the progress of the meeting.

[1317] As a concrete example, let's consider a scenario where a user hears the word "refactoring" at a meeting they are attending for the first time.

[1318] 1. The device records the meeting audio and sends the data to the server.

[1319] 2. The server receives "refactoring" as text data via the Google Speech-to-Text API and stores it temporarily.

[1320] 3. The server saves the text data to the database and simultaneously detects the word "refactoring".

[1321] 4. The server uses the Oxford Dictionaries API to retrieve the meaning of "refactoring" and saves it to the database.

[1322] 5. When a user clicks the word "refactoring" during a meeting, the terminal will display a pop-up with the meaning retrieved from the server.

[1323] This system allows users to understand the progress of a meeting in real time and to instantly check the meaning of technical terms and difficult words. As a result, understanding of the meeting content is enhanced, and a valuable environment is provided where users, especially those joining a particular organization for the first time or participants with hearing impairments, can quickly obtain useful information.

[1324] Example of a prompt:

[1325] "A meeting audio about a new system has been recorded. Please generate a program that converts the recorded audio data into text data, detects technical terms and difficult words, and displays their meanings. Please specify the names of the speech recognition API and dictionary API to be used, and provide a detailed explanation of the processing steps."

[1326] The flow of the specific processing in Example 1 will be explained using Figure 11.

[1327] Step 1:

[1328] The device records meeting audio data in real time using its built-in microphone or an external microphone. The recorded audio data is buffered at a fixed buffer size (e.g., every 5 seconds). The input is the meeting audio data, and the output is the buffered audio data. This buffered audio data is temporarily held within the device for subsequent processing.

[1329] Step 2:

[1330] The terminal sends buffered audio data to the server at pre-configured times. The input is the buffered audio data, and the output is the audio data sent to the server. This audio data is transmitted via an HTTP request.

[1331] Step 3:

[1332] The server sends the received audio data to the Google Speech-to-Text API, where it is converted into text data. The input is the audio data sent to the server, and the output is the text data returned by the API. Audio data is sent via an HTTP request, and text data is received as a JSON response.

[1333] Step 4:

[1334] The server adds a timestamp and meeting ID to the received text data and saves it to the database. The input is the text data returned from the API, along with the current timestamp and meeting ID, and the output is the text data saved in the database. Specifically, an SQL query is created to insert the data into the database with the timestamp and meeting ID attached.

[1335] Step 5:

[1336] The server retrieves all text data from the database based on the meeting ID after the meeting ends and sorts it chronologically. The input is the meeting ID, and the output is the chronologically sorted text data. A database query is executed to retrieve the text data and sort it based on the timestamp.

[1337] Step 6:

[1338] The server automatically generates meeting minutes according to a standard format using sorted text data. The input is text data sorted chronologically, and the output is automatically generated meeting minutes. A dedicated template engine is used to format the text data into meeting minutes.

[1339] Step 7:

[1340] The server analyzes the transcribed text data to detect technical terms and obscure words. The input is text data, and the output is a list of detected words. A natural language processing library (e.g., spaCy) is used to analyze the text and extract specific key words.

[1341] Step 8:

[1342] The server retrieves the meanings of detected technical terms and difficult words using a dictionary database or an external API. The input is a list of detected words, and the output is the meaning information of those words. An API request is created, and the response is received in JSON format.

[1343] Step 9:

[1344] The server stores the retrieved semantic information in a database. The input is the semantic information of a word, and the output is the semantic information stored in the database. An SQL query is created to insert the word and its meaning into the database.

[1345] Step 10:

[1346] The terminal retrieves the meaning of a word from the server and displays it in a pop-up window when the user clicks on a word in the text data displayed during a meeting. The input is the clicked word, and the output is the meaning of the word displayed in the pop-up window. An AJAX request is sent, and the retrieved data is processed and displayed using JavaScript.

[1347] These steps enable real-time transcription during meetings, automatic generation of meeting minutes, and retrieval and display of the meanings of technical terms.

[1348] (Application Example 1)

[1349] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1350] Meetings and work instructions within the factory often contain a lot of technical jargon and difficult-to-understand words, making it challenging for new or first-time workers to quickly respond to and understand them. Furthermore, there is a lack of systems to accurately transcribe meeting and instruction content in real time and present it in an easily understandable format.

[1351] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[1352] In this invention, the server includes means for transcribing meeting audio data in real time, means for storing the transcribed text data, means for automatically generating meeting minutes based on the transcribed text data, means for detecting technical terms and difficult words that appear during the meeting and obtaining their meanings, means for presenting the meanings of the technical terms and difficult words to the user, and means for transcribing work instructions and meeting audio in the factory in real time and providing workers with the meanings of the detected technical terms. This enables new workers and workers participating for the first time to quickly understand technical terms and difficult words and accurately grasp the content of work instructions and meetings in the factory.

[1353] "Meeting audio data" refers to data that digitally records the speech spoken during a meeting.

[1354] "Methods for real-time transcription" refers to technologies or devices that instantly convert audio from meetings, work instructions, etc., into text data.

[1355] "Transcribed text data" refers to text data converted from audio in real time.

[1356] "Means for saving text data" refers to the technology or device used to save converted text data to a storage device or similar.

[1357] "Means for automatically generating meeting minutes" refers to a technology or device that automatically generates a document recording the content of a meeting by rearranging saved text data in chronological order.

[1358] "Technical terms and difficult vocabulary" refers to specialized terms or words that are not commonly used or are difficult to understand.

[1359] "Means for detecting technical terms and difficult words and obtaining their meanings" refers to technologies or devices that discover technical terms and difficult words within text data and obtain their meanings from external dictionary databases or APIs.

[1360] "Means of presenting to the user" refers to a technology or device that displays the meaning of acquired words in a format that is easy for the user to understand.

[1361] "A means of transcribing audio of work instructions and meetings within a factory in real time" refers to a technology or device that instantly converts audio of work instructions and meetings conducted within a factory into text data.

[1362] "Means of providing workers with the meaning of detected technical terms" refers to a technology or device that displays the meaning of detected technical terms in real time in a format that is easy for workers to understand.

[1363] This invention is a system for improving the efficiency of meetings and work instructions within a factory. This system supports rapid understanding and response by transcribing audio data in real time, saving and analyzing text data, automatically generating meeting minutes, and immediately presenting the meanings of technical terms and difficult words to the user.

[1364] 1. Hardware and software to be used

[1365] The system uses the following hardware and software.

[1366] Hardware:

[1367] Microphone (for recording meetings and work instructions)

[1368] Server (for data processing and management)

[1369] Terminals in stores or factories (to display information to users)

[1370] software:

[1371] Speech recognition API (e.g., Google Speech Recognition API): Used to convert recorded audio into text data.

[1372] Dictionary API (external dictionary database): For obtaining the meanings of technical terms and difficult words.

[1373] Database management system: For storing and managing text data and semantic information.

[1374] Dedicated application: To display information to the user in real time.

[1375] 2. Overall System Functions and Processing

[1376] Audio data collection and transcription

[1377] The terminal records audio data of meetings and work instructions in real time and sends the data to the server in fixed buffer sizes. The server sends the received audio data to a speech recognition API and converts it into text data. This text data is temporarily stored in the server's memory.

[1378] Saving text data

[1379] The server stores the transcribed text data obtained from the speech recognition API into a database. A timestamp and meeting ID are added to the text data. This is useful for later searching and extracting specific meeting content.

[1380] Automatic generation of meeting minutes

[1381] After the meeting ends, or at an appropriate time, the server retrieves all text data from the database based on the meeting ID. The server sorts the retrieved text data chronologically and automatically generates meeting minutes according to a standard format. These minutes are stored in the database for later reference by the user.

[1382] Acquisition and presentation of the meanings of technical terms and difficult words.

[1383] The server analyzes the transcribed text data to detect technical terms and difficult words. For detected words, the server retrieves their meaning using a dictionary database or an external API (dictionary API). The retrieved meanings are stored in the database. When a user clicks on a word in the text data displayed in real time, the meaning of the corresponding word is displayed in a pop-up window.

[1384] 3. Specific examples and prompt messages

[1385] Specific example

[1386] Scenario: A new worker in a factory is instructed for the first time about the term "PMI (Preventive Maintenance)."

[1387] 1. Record audio in real time and send it to the server.

[1388] 2. Translate "PMI" into text using a speech recognition API.

[1389] 3. Save to the database.

[1390] 4. Use the dictionary API to retrieve and save the meaning of "PMI" (preventive maintenance).

[1391] 5. When a new employee clicks on "PMI," its meaning will be displayed.

[1392] Example of a prompt

[1393] Retrieve the meaning of the technical term "PMI" clicked by the user from the dictionary database and display it in real time as "preventive maintenance".

[1394] This system allows for quick understanding of work instructions and meeting content within the factory, enabling new workers and those participating for the first time to adapt smoothly.

[1395] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[1396] Step 1:

[1397] The terminal records audio of meetings and work instructions within the factory in real time.

[1398] Input: Audio from meetings or work instructions

[1399] Output: Recorded audio data

[1400] Specific operation: The device's microphone collects audio and stores the data at regular intervals of a certain buffer size.

[1401] Step 2:

[1402] The device sends the recorded audio data to the server in increments of a certain buffer size.

[1403] Input: Recorded audio data

[1404] Output: Audio data sent to the server

[1405] Specific operation: Once the audio data buffer is full, data transfer to the server begins.

[1406] Step 3:

[1407] The server sends the received audio data to a speech recognition API, where it is converted into text data.

[1408] Input: Sent audio data

[1409] Output: Text data returned from the speech recognition API

[1410] Specific operation: The server sends the received audio data to a speech recognition service such as the Google Speech Recognition API.

[1411] Step 4:

[1412] The server temporarily stores the converted text data in memory.

[1413] Input: Text data returned from the speech recognition API

[1414] Output: Text data stored in memory

[1415] Specific operation: The converted text data is given a meeting ID and timestamp, and then saved to the server's memory.

[1416] Step 5:

[1417] The server saves text data to the database.

[1418] Input: Text data stored in memory

[1419] Output: Text data stored in the database

[1420] Specific operation: Store the text data saved in memory in the database along with the meeting ID and timestamp.

[1421] Step 6:

[1422] After the meeting ends, the server retrieves all text data from the database based on the meeting ID.

[1423] Input: Meeting ID

[1424] Output: All text data related to the meeting

[1425] Specific operation: Use a database query to extract text data corresponding to a specified meeting ID.

[1426] Step 7:

[1427] The server sorts the acquired text data chronologically and automatically generates meeting minutes.

[1428] Input: Acquired text data

[1429] Output: Auto-generated meeting minutes

[1430] Specific operation: Sort text data based on timestamps and generate meeting minutes according to a standard format.

[1431] Step 8:

[1432] The server saves meeting minutes to a database so users can access them later.

[1433] Input: Generated meeting minutes

[1434] Output: Meeting minutes saved in the database

[1435] Specific action: Store the generated meeting minutes in the database.

[1436] Step 9:

[1437] The server analyzes the transcribed text data to detect technical terms and difficult words.

[1438] Input: Transcribed text data

[1439] Output: Detected technical terms and difficult words

[1440] Specific operation: Use a text analysis algorithm to identify technical terms and difficult words within the text data.

[1441] Step 10:

[1442] The server retrieves the meaning of a word detected using a dictionary API or dictionary database and stores it in the database.

[1443] Input: Detected technical terms or difficult words

[1444] Output: Meaning of the retrieved word

[1445] Specific operation: Access the dictionary API, retrieve the meaning of the detected word, and store it in the database.

[1446] Step 11:

[1447] When a user clicks on a technical term during a meeting or while working, the device retrieves the meaning of the word from the server and displays it in a pop-up window.

[1448] Input: Technical terms clicked by the user

[1449] Output: Meaning of the word displayed in the pop-up window

[1450] Specific operation: When the user clicks, the meaning of the corresponding word is retrieved from the server and displayed on the device.

[1451] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[1452] The system implementing this invention transcribes meeting audio data in real time and saves the text data. Furthermore, it has a function to automatically generate meeting minutes based on the saved text data, and also detects technical terms and difficult words that appear during the meeting, retrieves their meanings, and presents them to the user. In addition to this basic configuration, the present invention incorporates an emotion engine that recognizes the user's emotional state and appropriately adjusts the meanings of technical terms and difficult words based on the results.

[1453] Functional Embodiments

[1454] 1. Collection and transcription of audio data

[1455] The terminal records audio data in real time during the meeting. The recorded data is stored in a buffer and sent to the server at regular intervals. The server sends the received audio data to a speech recognition API and converts it into text data. The converted text data is stored in a database along with a timestamp and meeting ID.

[1456] 2. Automatic generation of meeting minutes

[1457] After the meeting ends, or at an appropriate time, the server retrieves all text data from the database based on the meeting ID. The retrieved text data is sorted chronologically, and meeting minutes are automatically generated according to a standard format. The automatically generated minutes are saved in the database for later reference by the user.

[1458] 3. Acquisition and presentation of the meanings of technical terms and difficult words.

[1459] The server analyzes the transcribed text data to detect technical terms and difficult words. For detected words, the server uses a dictionary database or dictionary API to retrieve their meanings. The retrieved meanings are stored in the database. When the user clicks on a word, the terminal displays its meaning in a pop-up window.

[1460] 4. Introduction of an emotional engine

[1461] The terminal or server recognizes the user's emotional state in real time during the meeting. The emotion engine analyzes the user's voice tone, facial expressions, or biosignals to identify their emotional state. Specifically, it can distinguish emotions such as anger, joy, surprise, and sadness.

[1462] 5. Adjusting information presentation based on emotions

[1463] Based on the user's emotional state, the server appropriately adjusts the meaning of technical terms and difficult words. For example, if the user is confused, it can add more detailed explanations and examples. On the other hand, if the server determines that the user understands, it will only provide a concise explanation.

[1464] Specific example

[1465] 1. At a meeting they are attending for the first time, the user hears the word "refactoring."

[1466] The device records the meeting audio and sends the data to the server.

[1467] The server receives "refactoring" as text data via the speech recognition API and stores it temporarily.

[1468] The server saves that text data to the database and simultaneously detects the word "refactoring".

[1469] The server uses a dictionary API to retrieve the meaning of "refactoring" and saves it to the database.

[1470] 2. The user clicks the word "refactoring".

[1471] The device detects a click event and requests information about that word from the server.

[1472] The server receives the request, retrieves the meaning of the corresponding word from the database, and returns it to the terminal.

[1473] The device displays semantic data as a pop-up.

[1474] 3. Simultaneously, the emotion engine analyzes the user's emotional state and determines that they are confused.

[1475] Based on the sentiment information acquired by the server, more detailed explanations and additional information will be included in the pop-up.

[1476] This system allows users to understand the progress of meetings in real time, comprehend technical terms, and receive appropriate information tailored to their emotional state. In this way, it provides an easily understandable meeting environment, even for users joining a particular organization for the first time or participants with hearing impairments.

[1477] The following describes the processing flow.

[1478] Step 1:

[1479] The device records the meeting audio in real time. Recording begins when the user launches the meeting recording application and presses the start recording button. The recorded data is stored in the audio buffer.

[1480] Step 2:

[1481] The terminal sends data stored in the audio buffer to the server at regular intervals. For example, it sends all the audio data accumulated in the buffer to the server every 10 seconds.

[1482] Step 3:

[1483] The server sends the received audio data to a speech recognition API. The server creates an API request and passes the audio data to the API. Examples of APIs used include Google Speech-to-Text and IBM Watson.

[1484] Step 4:

[1485] The speech recognition API converts the audio data into text and returns the result to the server. The API analyzes the audio data and generates text data in string format.

[1486] Step 5:

[1487] The server receives the text data returned from the speech recognition API and temporarily stores it in memory. At this time, a timestamp and meeting ID are added to the text data.

[1488] Step 6:

[1489] The server saves the text data to the database. A meeting ID and timestamp are added during saving, making it useful for later searching and generating meeting minutes.

[1490] Step 7:

[1491] After the meeting ends, the server retrieves all text data from the database based on the corresponding meeting ID. After retrieval, the server sorts the text data chronologically.

[1492] Step 8:

[1493] The server automatically generates meeting minutes based on sorted text data, following a standard format. The automatically generated minutes are formatted for readability, with paragraphs and headings neatly organized.

[1494] Step 9:

[1495] The server saves the generated meeting minutes to a database. It also adds an index to allow users to access them later.

[1496] Step 10:

[1497] During the meeting, the server analyzes the text data transmitted in real time to detect technical terms and difficult words. Natural language processing (NLP) techniques are used to list words that match specific criteria.

[1498] Step 11:

[1499] The server uses a dictionary API to retrieve the meaning of each word it detects. The data returned from the dictionary API includes a detailed definition of the word and example sentences.

[1500] Step 12:

[1501] The server stores the meaning of the words it retrieves in a database. Since the meaning data is also stored in association with the meeting ID, users can easily access it later.

[1502] Step 13:

[1503] The device is equipped with sensors to recognize the user's emotional state in real time, collecting voice tone, facial expression analysis, or biosignals. The emotion engine analyzes emotions such as anger, joy, surprise, and sadness in real time.

[1504] Step 14:

[1505] The server receives emotional state data from the emotion engine and adjusts the meaning of technical terms and complex words. For example, if it determines that the user is confused, it adds a more detailed explanation.

[1506] Step 15:

[1507] The user clicks on a word in the text data during a meeting. The device detects the click event and sends a request to the server regarding that word.

[1508] Step 16:

[1509] The server receives the request, retrieves the meaning of the corresponding word from the database, makes adjustments according to the emotional state, and then returns it to the terminal.

[1510] Step 17:

[1511] The device displays semantic data it has acquired as a pop-up. Users can see the meaning of the word and additional information adjusted by the sentiment engine in real time.

[1512] This detailed processing step enables a system that not only transcribes meeting content in real time, but also provides appropriate meanings for technical terms based on the user's emotional state.

[1513] (Example 2)

[1514] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1515] Traditional conferencing systems made it difficult to transcribe meeting audio data in real time, resulting in significant time and effort required for creating meeting minutes. Furthermore, understanding the meaning of specialized terminology and complex vocabulary used during meetings was challenging, posing a significant barrier to comprehension, especially for new or non-expert users. Additionally, the inability to consider the emotional state of participants during meetings led to inappropriate information presentation, posing a significant problem.

[1516] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for transcribing the audio data of a meeting in real time, means for storing the transcribed text data, means for automatically generating meeting minutes based on the transcribed text data, means for detecting technical terms and difficult words that appear during the meeting and obtaining their meanings, means for presenting the meanings of the technical terms and difficult words to the user, and means for recognizing the user's emotions in real time and adjusting the meanings of the technical terms and difficult words. This enables real-time transcription of meeting content, automatic generation of meeting minutes, immediate understanding of technical terms, and presentation of appropriate information based on the user's emotional state.

[1517] "Meeting audio data" refers to all audio generated during a meeting, including discussions, debates, and the content of statements made.

[1518] "Real-time" refers to an event or process that occurs immediately, without any delays or waiting periods, and is reflected instantly.

[1519] "Transcription" is the process of converting audio data into text format, recording the content of the audio as text data.

[1520] "Text data" refers to a data format composed of characters and symbols, and is data generated as a result of transcribing audio data into text.

[1521] "Saving" refers to storing data in a way that makes it usable for a long period of time, and is the act of recording it in a database or storage device.

[1522] "Meeting minutes" are documents that record what was said and decided during a meeting, and are official records intended for later reference.

[1523] "Automatic generation" means creating data or documents using specific algorithms or programs without human intervention.

[1524] "Detecting" refers to the act of finding specific elements or patterns within data or information, and this typically involves using analytical techniques.

[1525] "Acquiring" refers to the act of extracting necessary information from external sources or databases, and this is often done through APIs.

[1526] A "user" refers to a person or organization that uses a system or service, and is the entity that performs operations through the system interface.

[1527] "Recognizing emotions in real time" means instantly understanding a user's emotional state using specific technologies (such as voice analysis or facial expression analysis).

[1528] "Adjusting" refers to the act of changing the content or settings based on specific criteria or conditions to achieve an optimal state.

[1529] In an embodiment of the present invention, the system is capable of transcribing meeting audio data in real time and saving the text data. Furthermore, it can automatically generate meeting minutes based on the saved text data, and in addition, it can detect technical terms and difficult words that appear during the meeting, obtain their meanings, and present them to the user. Moreover, the present invention provides a system that recognizes the user's emotional state in real time and appropriately adjusts the meanings of technical terms and difficult words.

[1530] 1. Hardware and software to be used

[1531] The hardware requires devices (PCs, smartphones, tablets, etc.) to be used by meeting participants, and these devices must include high-sensitivity microphones. The server will be either a cloud environment or an on-premises server.

[1532] The software used includes a speech recognition API (e.g., Google Cloud Speech-to-Text API), a database (e.g., PostgreSQL), and a natural language processing library (e.g., spaCy). External dictionary APIs (e.g., Wiktionary API) are used to obtain the meanings of technical terms and complex words. Speech tone analysis (e.g., Microsoft Azure Emotion API) and facial expression analysis (e.g., OpenCV and Dlib) are used to recognize emotional states. Furthermore, generative AI models (e.g., OpenAI GPT-3) are used to refine the meanings of technical terms.

[1533] 2. Processing flow and specific data processing / calculations

[1534] The terminal records audio data during the meeting in real time, and the recorded data is temporarily stored in a buffer. At regular intervals, the terminal divides the data in the buffer into packets and sends them to the server via the internet. The HTTPS protocol is used for this.

[1535] The server sends the received audio data to a speech recognition API for transcription. The converted audio data is stored in a database along with a timestamp and meeting ID.

[1536] After the meeting ends, or at an appropriate time, the server retrieves all text data from the database based on the meeting ID. The retrieved text data is sorted chronologically, and meeting minutes are automatically generated using a natural language processing library. The automatically generated minutes are saved in the database for later reference by the user.

[1537] The server analyzes stored text data to detect technical terms and obscure words. This uses techniques such as TF-IDF and morphological analysis. The meaning of detected words is retrieved from a dictionary database or via an external API and stored in the database.

[1538] When a user clicks on a word, the device detects the click event and requests information about that word from the server. The server receives the request, retrieves the meaning data for the corresponding word from its database, and returns it to the device. The device then displays the meaning data in a pop-up window.

[1539] During the meeting, terminals and servers analyze voice tone and facial expressions in real time to recognize emotional states. An emotion engine evaluates this data to determine the user's specific emotional state (e.g., confused, understanding). If confused is detected, the server uses a generative AI model to further refine the meaning of technical terms and complex words, including detailed explanations and additional information.

[1540] 3. Specific Examples and Examples of Prompt Statements

[1541] As a concrete example, if a user hears the word "refactoring" in a meeting they are attending for the first time and does not understand it, it will behave as follows:

[1542] The device records the meeting audio and sends the data to the server.

[1543] The server receives "refactoring" as text data via the speech recognition API and stores it temporarily.

[1544] The server saves that text data to a database and detects the word "refactoring".

[1545] The server uses a dictionary API to retrieve the meaning of "refactoring" and saves it to the database.

[1546] When a user clicks the word "refactoring," the device detects the click event and requests information about that word from the server.

[1547] The server receives the request, retrieves the meaning of the corresponding word from the database, and returns it to the terminal.

[1548] The device displays semantic data as a pop-up.

[1549] At the same time, the emotion engine analyzes the user's emotional state and determines that they are confused.

[1550] Based on the sentiment information acquired by the server, more detailed explanations and additional information will be included in the pop-up.

[1551] Example of a prompt:

[1552] "Could you explain the meaning of refactoring?"

[1553] "Please explain refactoring with an example."

[1554] "Please explain how refactoring can benefit a project."

[1555] This invention makes it easier for participants to understand the progress of a meeting in real time, to immediately grasp the meaning of technical terms and difficult words, and to create an environment where appropriate information is provided according to their emotional state.

[1556] The flow of the specific processing in Example 2 will be explained using Figure 13.

[1557] Step 1: Collect audio data

[1558] The device records audio data during the meeting in real time. Audio acquired through a high-sensitivity microphone is temporarily stored in a buffer. The input is the meeting audio data, and the output is the audio data stored in the buffer. Recording software runs on the device, sampling the audio data second by second and saving it to the buffer.

[1559] Step 2: Sending the audio data

[1560] The terminal periodically divides the audio data in its buffer into packets and sends them to the server via the internet. The input is the audio data in the buffer, and the output is the audio data packets sent to the server. This process is performed securely using the HTTPS protocol.

[1561] Step 3: Transcribing the audio data

[1562] The server receives audio data and sends it to a speech recognition API (e.g., Google Cloud Speech-to-Text API) to convert it into text data. The input is audio data, and the output is transcribed text data. The converted text data is temporarily stored in server memory.

[1563] Step 4: Save text data

[1564] The server saves the transcribed text data, along with the timestamp and meeting ID, to a database. The input is the text data, timestamp, and meeting ID, and the output is the text data saved in the database. The saving process is performed using an SQL query.

[1565] Step 5: Automatic generation of meeting minutes

[1566] After the meeting ends, the server retrieves text data from the database based on the meeting ID. The input is the meeting ID, and the output is the retrieved text data. The server sorts the retrieved text data chronologically and automatically generates meeting minutes using a natural language processing library (e.g., spaCy). The automatically generated meeting minutes are then saved back to the database.

[1567] Step 6: Detection of technical terms and difficult words

[1568] The server analyzes stored text data and uses TF-IDF and morphological analysis to detect technical terms and difficult words. The input is stored text data, and the output is a list of detected words. Natural language processing techniques are used for analysis, and words of high importance are extracted.

[1569] Step 7: Acquire the meaning of technical terms and difficult words.

[1570] The server queries a dictionary database (e.g., the Wiktionary API) for the detected words and retrieves their meanings. The input is a list of detected words, and the output is a list of retrieved meanings. The retrieved word meanings are then saved back into the database.

[1571] Step 8: Presentation of technical terms and difficult vocabulary.

[1572] The device detects an event when the user clicks on a word and requests that information from the server. The input is the click event, and the output is the requested word information. The server receives the request, retrieves the meaning data for the corresponding word from the database, and returns it to the device. The device then displays the meaning data in a pop-up window.

[1573] Step 9: Recognizing your emotional state

[1574] The terminal or server analyzes the user's voice tone and facial expressions in real time during a meeting to recognize their emotional state. Input is voice tone data and facial expression data, and output is the analyzed emotional state data. This utilizes the Microsoft Azure Emotion API, OpenCV, and Dlib.

[1575] Step 10: Adjusting the Information Presentation

[1576] The server adjusts the meaning of technical terms and difficult words appropriately based on the user's emotional state. The input is emotional state data and the meaning of technical terms, while the output is the adjusted semantic data. If the user is confused, a generative AI model (e.g., OpenAI GPT-3) is used to generate a detailed explanation, which is returned to the device. Based on this returned data, the device displays more detailed information in a pop-up window.

[1577] (Application Example 2)

[1578] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1579] Conventional meeting recording systems often simply convert meeting audio data into text and automatically generate meeting minutes. However, these systems often make it difficult to understand the meaning of technical terms and complex vocabulary, and they fail to provide appropriate information based on the user's emotional state, making it difficult to understand the progress of the meeting and its subsequent analysis. Furthermore, they lack the functionality to respond in real time when security risks occur, requiring immediate countermeasures.

[1580] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for transcribing the audio data of a meeting in real time, means for storing the transcribed text data, means for automatically generating meeting minutes based on the transcribed text data, means for detecting technical terms and difficult words that appear during the meeting and obtaining their meanings, means for presenting the meanings of the technical terms and difficult words to the user, means for analyzing the user's emotional state and adjusting the content of the terms presented based on that emotional state, and means for sending emergency alerts based on specific terms detected during the conversation and the user's emotional state. As a result, the user can grasp the progress of the meeting in real time, easily understand the meanings of technical terms and difficult words, be provided with appropriate information according to their emotional state, and respond immediately to security risks.

[1581] "Audio data" refers to digital recordings of voices generated during meetings, conversations, and other similar events.

[1582] "Transcription" is the process of analyzing audio data and converting it into text data.

[1583] "Text data" refers to data expressed as characters, specifically a string of characters representing the content of a conversation generated through transcription.

[1584] "Meeting minutes" are documents that summarize the content of a meeting, recording the statements made and decisions made during the meeting.

[1585] "Technical terms" are specific terms used in a particular field or industry, and are difficult to understand with general knowledge.

[1586] "Difficult words" are words that are hard to understand or words that are not commonly used.

[1587] "Emotional state" refers to the user's emotional condition and includes a variety of emotions such as joy, anger, sadness, and surprise.

[1588] An "emergency alert" is a notification sent when an emergency occurs or specific conditions are met, and it is a warning that requires immediate action.

[1589] "User" refers to a person who uses a system or application.

[1590] A "server" is a computer system that processes information over a network and provides services to clients.

[1591] This invention provides a system that supports the progress of meetings in real time, enabling the understanding of important information and the management of security risks. A system implementing this invention includes the following hardware and software:

[1592] 1. Collection and transcription of audio data:

[1593] The terminal records audio data in real time during the meeting. The recorded data is stored in a buffer and sent to the server at regular intervals. The server converts the received audio data into text data using the Google Web Speech API. The converted text data is stored in a database along with a timestamp and meeting ID.

[1594] 2. Automatic generation of meeting minutes:

[1595] After the meeting ends, or at an appropriate time, the server retrieves all text data from the database based on the meeting ID. The retrieved text data is sorted chronologically, and meeting minutes are automatically generated according to a standard format. The automatically generated minutes are saved in the database for later reference by the user.

[1596] 3. Acquisition and presentation of the meanings of technical terms and difficult words:

[1597] The server analyzes the transcribed text data to detect technical terms and difficult words. For detected words, the server uses a dictionary database or dictionary API to retrieve their meanings. The retrieved meanings are stored in the database. When the user clicks on a word, the terminal displays its meaning in a pop-up window.

[1598] 4. Introduction of the Emotion Engine:

[1599] The terminal or server recognizes the user's emotional state in real time during the meeting. The emotion engine analyzes the user's voice tone, facial expressions, or biosignals to identify their emotional state. Specifically, it can distinguish emotions such as anger, joy, surprise, and sadness.

[1600] 5. Adjusting information presentation based on emotions:

[1601] Based on the user's emotional state, the server appropriately adjusts the meaning of technical terms and difficult words. For example, if the user is confused, it can add more detailed explanations and examples. On the other hand, if the server determines that the user understands, it will only provide a concise explanation.

[1602] 6. Sending an emergency alert:

[1603] The server sends emergency alerts based on specific security terms detected during the conversation (e.g., "confidential" or "threat") and the user's emotional state. If the emotion engine detects high levels of tension or anxiety in the user, a notification is sent to the administrator.

[1604] Specific example

[1605] 1. An emergency occurs during the conversation:

[1606] "We have received a report that an intruder entered the office during our conversation. We need to take immediate action."

[1607] This allows users to understand the progress of meetings in real time, comprehend technical terms and complex vocabulary more easily, receive appropriate information tailored to their emotional state, and respond immediately to security risks.

[1608] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[1609] Step 1:

[1610] The terminal records the audio data of the meeting in real time. The input is the audio data of the meeting, which the terminal stores in a buffer in digital audio format. The output is digital audio data that is sent to the server at regular intervals.

[1611] Step 2:

[1612] The server sends the received audio data to the Google Web Speech API, where it is converted into text data. The input is digital audio data, and the server uses a speech recognition API to transcribe it. The output is text data stored on the server along with a timestamp and meeting ID.

[1613] Step 3:

[1614] The server analyzes stored text data to detect technical terms and obscure words. The input is text data, and the server uses a natural language processing engine to identify these terms. The output is a list of detected words.

[1615] Step 4:

[1616] The server retrieves the meaning of detected words using a dictionary database or dictionary API. The input is a list of detected words, and the server generates an API request to retrieve the meaning of each word. The output is the meaning information and explanation for each word.

[1617] Step 5:

[1618] The server stores the semantic information of words it has retrieved in a database, and the terminal displays that semantic information upon user request. Input consists of user click events and word requests; the terminal retrieves the word's meaning from the server and displays it in a pop-up. Output is the semantic information presented to the user.

[1619] Step 6:

[1620] The terminal or server recognizes the user's emotional state in real time during a meeting. Inputs include the user's voice tone, facial expressions, or biosignals, which the emotion engine analyzes to identify the emotional state. Output is data on the recognized emotional state.

[1621] Step 7:

[1622] Based on the user's emotional state, the server appropriately adjusts the meaning of technical terms and difficult words. Input consists of emotional state data and word semantic information; the server includes detailed explanations and additional information if the user is confused. Output is the adjusted semantic information.

[1623] Step 8:

[1624] The server sends an emergency alert based on specific security terms and the user's emotional state detected during the conversation. The input is the detected security terms and user emotional state data, which the server uses to generate an alert notification and send it to the administrator. The output is the emergency alert sent to the administrator.

[1625] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[1626] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1627] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[1628] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1629] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[1630] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[1631] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[1632] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[1633] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[1634] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[1635] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[1636] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[1637] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[1638] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1639] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[1640] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[1641] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[1642] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[1643] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[1644] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[1645] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted as being incorporated by reference.

[1646] The following is further disclosed regarding the embodiments described above.

[1647] (Claim 1)

[1648] A method for transcribing meeting audio data in real time,

[1649] Means for saving the transcribed text data,

[1650] A means for automatically generating meeting minutes based on the transcribed text data,

[1651] A method for detecting technical terms and difficult words that appear during a meeting and obtaining their meanings,

[1652] A system that includes means for presenting the meanings of the aforementioned technical terms and difficult words to the user.

[1653] (Claim 2)

[1654] The system according to claim 1, characterized in that the means for obtaining the meaning of the aforementioned technical terms and difficult words obtains information from a dictionary database or an external API.

[1655] (Claim 3)

[1656] The system according to claim 1, characterized in that it includes means for transmitting the audio data of the aforementioned meeting to a server at regular intervals and transcribing it, and performs transcription in real time.

[1657] "Example 1"

[1658] (Claim 1)

[1659] A method for transcribing meeting audio data in real time,

[1660] Means for saving the transcribed text data,

[1661] A means for automatically generating meeting minutes based on the transcribed text data,

[1662] A method for detecting technical terms and difficult words that appear during meetings,

[1663] A means for obtaining the meanings of the detected technical terms and difficult words using a dictionary database or an external API,

[1664] A means of presenting the meanings of the aforementioned technical terms and difficult words to the user in real time,

[1665] A system including means for transmitting the aforementioned audio data to a server at regular intervals and performing transcription.

[1666] (Claim 2)

[1667] The system according to claim 1, characterized in that the means for obtaining the meaning of the aforementioned technical terms and difficult words obtains information from a dictionary database or an external API.

[1668] (Claim 3)

[1669] The system according to claim 1, characterized by means for rearranging the minutes of the aforementioned meetings in chronological order and automatically generating them according to a standard format.

[1670] "Application Example 1"

[1671] (Claim 1)

[1672] A method for transcribing meeting audio data in real time,

[1673] Means for saving the transcribed text data,

[1674] A means for automatically generating meeting minutes based on the transcribed text data,

[1675] A method for detecting technical terms and difficult words that appear during a meeting and obtaining their meanings,

[1676] A means of presenting the meanings of the aforementioned technical terms and difficult words to the user,

[1677] A method for transcribing audio of work instructions and meetings within a factory in real time, and providing workers with the meanings of detected technical terms,

[1678] A system that includes this.

[1679] (Claim 2)

[1680] The system according to claim 1, characterized in that the means for obtaining the meaning of the aforementioned technical terms and difficult words obtains information from a dictionary database or an external API.

[1681] (Claim 3)

[1682] The system according to claim 1, characterized in that it includes means for transmitting audio data of the aforementioned meeting and work instructions within the factory to a server at regular intervals and transcribing it, and performs transcription in real time.

[1683] "Example 2 of combining an emotion engine"

[1684] (Claim 1)

[1685] A method for transcribing meeting audio data in real time,

[1686] Means for saving the transcribed text data,

[1687] A means for automatically generating meeting minutes based on the transcribed text data,

[1688] A method for detecting technical terms and difficult words that appear during a meeting and obtaining their meanings,

[1689] A means of presenting the meanings of the aforementioned technical terms and difficult words to the user,

[1690] A means to recognize user emotions in real time and adjust the meaning of technical terms and difficult words,

[1691] A system that includes this.

[1692] (Claim 2)

[1693] The system according to claim 1, characterized in that it obtains the meanings of the aforementioned technical terms and difficult words from a dictionary database or an external API.

[1694] (Claim 3)

[1695] The system according to claim 1, characterized in that it transmits the audio data of the aforementioned meeting to a server at regular intervals and performs transcription.

[1696] "Application example 2 when combining with an emotional engine"

[1697] (Claim 1)

[1698] A method for transcribing meeting audio data in real time,

[1699] Means for saving the transcribed text data,

[1700] A means for automatically generating meeting minutes based on the transcribed text data,

[1701] A method for detecting technical terms and difficult words that appear during a meeting and obtaining their meanings,

[1702] A means of presenting the meanings of the aforementioned technical terms and difficult words to the user,

[1703] A means for analyzing the user's emotional state and adjusting the content of the terms presented based on that emotional state,

[1704] A system that includes means for sending emergency alerts based on specific terms detected during a conversation and the user's emotional state.

[1705] (Claim 2)

[1706] The system according to claim 1, characterized in that the means for obtaining the meaning of the aforementioned technical terms and difficult words obtains information from a dictionary database or an external API.

[1707] (Claim 3)

[1708] The system according to claim 1, characterized in that it includes means for transmitting the audio data of the aforementioned meeting to a server at regular intervals and transcribing it, and performs transcription in real time. [Explanation of Symbols]

[1709] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. A method for transcribing meeting audio data in real time, Means for saving the transcribed text data, A means for automatically generating meeting minutes based on the transcribed text data, A method for detecting technical terms and difficult words that appear during a meeting and obtaining their meanings, A system that includes means for presenting the meanings of the aforementioned technical terms and difficult words to the user.

2. The system according to claim 1, characterized in that the means for obtaining the meaning of the aforementioned technical terms and difficult words obtains information from a dictionary database or an external API.

3. The system according to claim 1, characterized in that it includes means for transmitting the audio data of the aforementioned meeting to a server at regular intervals and transcribing it, and performs transcription in real time.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A