System
The system addresses inefficient meeting management by converting voice data to text, extracting keywords, and analyzing emotions, enhancing project management efficiency through real-time data analysis.
Patent Information
- Application Number
- JP2024122756
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-29
- Publication Date
- 2026-02-10
AI Technical Summary
Inefficient management of meeting and discussion information, manual recording and resource allocation, neglect of team members' emotional states, and slow decision-making processes hinder effective project management.
A system that receives voice data, converts it into text, extracts specific keywords, stores them in a database, and displays the data on a user interface, enabling real-time analysis and efficient management of meetings.
Facilitates rapid decision-making by automatically recording and analyzing meeting content and emotional states, improving project management efficiency and resource allocation.
Smart Images

Figure 2026021074000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] In modern companies, efficiently managing meeting and discussion information and making appropriate decisions is crucial. However, manual meeting recording, resource allocation, and risk management consume a significant amount of time and effort, and are inefficient. Furthermore, analysis of team members' emotional states and motivations is often neglected, often resulting in failure to take appropriate measures to boost morale. Furthermore, advanced analysis and rapid response are required to quickly identify key points and important risk factors from meetings and take appropriate measures. [Means for solving the problem]
[0005] The present invention provides a system including means for receiving voice data and converting it into text data, means for extracting specific keywords from the converted text data, means for saving the extracted keywords and their location information in a database, and means for displaying the saved data on a user interface. Furthermore, the system also includes means for analyzing voice data in real time and displaying text data containing specific keywords on the user interface in real time, and means for analyzing voice data of meetings or discussions and automatically extracting, saving, and displaying key points from the converted text data, thereby enabling efficient management of meetings and rapid decision-making.
[0006] "Audio data" refers to data in which audio signals are recorded in digital format, and includes the contents of meetings and conferences.
[0007] "Character data" is voice data converted into text format, and is used to record minutes and key points.
[0008] "Keywords" are specific important words or phrases extracted from the converted text data that indicate risks or important matters.
[0009] A "database" is a storage system that systematically stores converted character data, extracted keywords, and their location information, allowing for searching and management.
[0010] A "user interface" is a screen or operating means for exchanging information between the system and the user, and displays analysis results and important points.
[0011] "Real time" means that the analysis of the voice data and the display of the results are carried out immediately without delay.
[0012] "Analysis" is the process of recognizing voice data, converting it into text data, understanding its content, and extracting specific information.
[0013] "Extraction" is the process of identifying and extracting specific keywords or key points from text data.
[0014] "Storage" is the process of recording the analyzed data in a database so that it can be accessed later.
[0015] "Display" is the process of visually presenting the analysis results and key points to the user through a user interface. [Brief explanation of the drawings]
[0016] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0017] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0018] First, the terms used in the following description will be explained.
[0019] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0020] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0021] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0022] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0023] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0024] [First embodiment]
[0025] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0026] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0027] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0028] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0029] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0030] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0031] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0032] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0033] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0034] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0035] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0036] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0037] The present invention provides a system that receives voice data, converts it into text data, extracts specific keywords, stores them in a database, and displays them on a user interface in order to improve the efficiency of project management.
[0038] 1. Acceptance of audio data
[0039] First, the user uploads the recording of the meeting or conference from their own device to the server. Specifically, they click the "File Upload" button on the user interface and select the desired audio file from the file system. The server receives the audio file and saves it in the specified directory.
[0040] 2. Converting voice data to text data
[0041] Next, the server reads the received voice data as audio data, initializes a speech recognition library, and converts the voice data into text using a speech recognition service such as the Google Speech Recognition API. During this process, the server works to convert the voice content into text data with high accuracy.
[0042] 3. Keyword extraction
[0043] To extract specific keywords from the converted text data, the server uses a predefined list of important keywords. It checks whether these keywords exist in the text data and records their location information. For example, if keywords such as "important," "risk," and "deadline" are included, it records the words and their locations in a list.
[0044] 4. Saving to the database
[0045] The extracted keywords and their location information are stored in a database by the server. To maintain data integrity, the data is recorded with a current timestamp. This makes it clear what was discussed and when.
[0046] 5. Display in the user interface
[0047] Finally, the server displays the saved data on a user interface. Through the interface, users can check the text data and important keywords extracted from the recording data of meetings and discussions. For example, if a user uploads an audio file called "project_meeting.wav," the server analyzes the audio data and displays the text data "We need to discuss the risks and important deadlines of the next project" along with the keywords "risk" and "deadline" to the user.
[0048] In this way, the user can efficiently grasp the contents of the meeting, quickly confirm important points, and take necessary measures.
[0049] The processing flow will be explained below.
[0050] Step 1:
[0051] A user records an audio file of a meeting or conference and uses the system's "audio file upload" function on their device. They click the "file upload" button on the user interface and select the desired audio file from the file system.
[0052] Step 2:
[0053] The device uploads the selected audio file to the server via an HTTP request, and the server receives the request and saves the audio file in the specified directory.
[0054] Step 3:
[0055] The server reads the saved audio file and captures the audio data as audio data. Specifically, it initializes the speech recognition library and loads the audio file as an audio object.
[0056] Step 4:
[0057] The server calls a speech recognition service such as the Google Speech Recognition API to convert the audio data into text data. If the conversion process is successful, the content of the audio data is obtained in text format.
[0058] Step 5:
[0059] The server extracts specific keywords from the acquired text data, checks them against a predefined list of important keywords, identifies the keywords contained in the text data, and records their location information.
[0060] Step 6:
[0061] The server stores the converted text data, extracted keywords, and their location information in a database. When storing the data, it records the data with the current timestamp to maintain data integrity.
[0062] Step 7:
[0063] The server displays the stored data on a user interface. Users access the system from their devices and can view text data and important keywords in real time through the interface. The display includes specific discussion content, risks, and important points.
[0064] Step 8:
[0065] Users can check the interface and quickly take necessary measures and adjust resource allocation. The system records user interactions, contributing to the efficiency of future meetings.
[0066] Example 1
[0067] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0068] There is a demand for efficient recording of meeting and conference content so that important points can be quickly grasped. However, conventional systems require manual conversion of audio data into text data and manual extraction of important keywords, which is extremely time-consuming. In addition, because data is also saved and displayed manually, there are issues with information leaks and inconsistency.
[0069] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0070] In this invention, the server includes means for receiving voice data, means for converting the received voice data into text data, means for extracting specific keywords from the converted text data, means for saving the extracted keywords and their location information in a database, means for displaying the saved data on a user interface, means for converting voice data into text data using a voice recognition service, means for loading a predefined keyword list to extract specific keywords, means for recording a current timestamp when saving the text data including the location information of the extracted keywords in the database, and means for uploading voice data and checking the results of the analysis process via the user interface. This makes it possible to automatically and efficiently record the contents of meetings and discussions and easily grasp important points.
[0071] "Audio data" refers to data in which the contents of a meeting, conference, etc. are recorded in audio format.
[0072] "Character data" is data in text format that has been converted from voice data using voice recognition technology.
[0073] "Keywords" refer to specific words or phrases that are considered important in the content of a meeting or discussion.
[0074] A "database" is a digital system for efficiently storing and managing information, and in this case, it plays a role in storing text data based on voice data and extracted keywords.
[0075] A "user interface" is the part of a system that contains the screens and controls that allow a user to interact with the system.
[0076] A "means for receiving" is a technique or mechanism for capturing audio data from a user.
[0077] "Means for converting" refers to the technology or method for converting voice data into text data.
[0078] "Extraction means" refers to techniques or methods for extracting specific keywords from the converted character data.
[0079] "Storage means" refers to the technology or method for storing the extracted keywords and their location information in a database.
[0080] "Display means" refers to the techniques and methods for presenting the stored data in a user interface so that the user can check it.
[0081] "Speech Recognition Service" means a cloud-based or local service for analyzing and converting voice data into text data.
[0082] A "keyword list" is a collection of predefined important keywords that are used in text analysis.
[0083] A "timestamp" is time information that indicates that data was saved or processed at a specific time.
[0084] "Upload" refers to the operation or process by which a user sends audio data from their device to a server.
[0085] "Analysis" refers to a series of processes that converts voice data into text data and then extracts keywords.
[0086] The present invention aims to improve the efficiency of project management by providing a system that receives voice data, converts it into text data, extracts specific keywords, stores them in a database, and displays them on a user interface. Specifically, the invention can be implemented in the following forms.
[0087] First, the user uploads the recording of a meeting or conference from their own device to the server. The user clicks the "File Upload" button on the user interface and selects the desired audio file from the file system. The uploaded audio file is sent from the device to the server, and the server receives it and saves it in the specified directory. At this time, the file's metadata (e.g., file name, upload date and time, user ID, etc.) is also saved.
[0088] Next, the server reads the received voice data. It treats the voice data as audio data and initializes a voice recognition library. The system uses a voice recognition service such as the Google Speech Recognition API to convert the voice data into text. Through the conversion process, the content of the voice data is obtained as text data. For example, the resulting text data is, "We need to discuss the risks and important deadlines for the next project."
[0089] The server then extracts specific keywords from the converted text data using a predefined list of important keywords (e.g., "important," "risk," and "deadline"). The server uses a text analysis algorithm to identify the presence of keywords in the converted text data and record their location (e.g., which sentence or word). For example, if the words "risk" and "deadline" appear on the third line of the text, the server will add their location to the list.
[0090] The server then saves the extracted keywords and their location information, along with the text data, to a database. When saving, the server also records the current timestamp. This ensures data integrity. The saved data is stored in the "meeting_transcripts" table.
[0091] Finally, the server displays the stored data on a user interface. The server retrieves the stored text data and keywords from the database, formats them, and displays them to the user. Through the user interface, the user can view the text data and important keywords extracted from the recording of a meeting or discussion. For example, if a user uploads an audio file called "project_meeting.wav," the server analyzes the audio data and displays the text data, "We need to discuss the risks and important deadlines for the next project," along with the keywords "risk" and "deadline" to the user.
[0092] Specific examples
[0093] A user uploads an audio file called "project_meeting.wav." The server then converts it into text using the Google Speech Recognition API and extracts the keywords "risk" and "deadline." The extracted information is stored in a database and the user interface displays the message, "We need to discuss the risks and important deadlines for the next project."
[0094] Example of input prompt for generative AI model
[0095] An example of a prompt is, "Please upload the recording of the project meeting, 'project_meeting.wav'. The contents of this audio file will be converted into text data, and important keywords will be extracted and displayed. The service used is the Google Speech Recognition API."
[0096] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0097] Step 1:
[0098] The user uploads the audio data.
[0099] A user accesses the user interface of the project management system using a terminal. He clicks the "File Upload" button on the user interface and selects the desired audio file from the file system. The input is the audio file selected by the user (e.g., project_meeting.wav). The terminal sends the audio file to the server, and the audio data is sent to the server as the output.
[0100] Step 2:
[0101] The server receives and stores the audio data.
[0102] The server receives audio data sent from the device. It saves the received audio data in a specified directory (e.g., / uploaded_audio). At this time, it also saves metadata such as the file name, upload date and time, and user ID. The input is the audio data sent by the user. The output is an audio file saved in a specified directory on the server.
[0103] Step 3:
[0104] The server initializes the speech recognition library and reads the audio data.
[0105] The server initializes a speech recognition library (e.g., Google Speech Recognition API). It then reads the saved audio data and loads it into memory as audio data. The input is the saved audio data file. The output is the audio data loaded into memory.
[0106] Step 4:
[0107] The server converts the audio data into text data.
[0108] The server converts the audio data into text using an initialized speech recognition library. The speech recognition library analyzes the audio data and outputs the speech in text format. The input is the audio data loaded in memory. The output is text (e.g., "We need to discuss the risks and important deadlines for the next project").
[0109] Step 5:
[0110] The server loads the keyword list and extracts the keywords.
[0111] The server loads a predefined list of important keywords (e.g., "important," "risk," "deadline"), then checks for the presence of keywords in the converted text data and records location information. The input is the converted text data and the keyword list. The output is the extracted keywords (e.g., "risk," "deadline") and their location information.
[0112] Step 6:
[0113] The server stores the extracted data in a database.
[0114] The server stores the extracted keywords and text data, including their location information, in a database. The stored data also includes the current timestamp. The input is text data with keywords and location information added. The output is the information stored in the database.
[0115] Step 7:
[0116] The server displays the stored data in a user interface.
[0117] The server retrieves the stored text data and keywords from the database, formats them, and displays them on the user interface. Through the user interface, users can check the text data and important keywords extracted from the recording data of meetings and conferences. The input is the text data and keyword information retrieved from the database. The output is the text data and keywords displayed on the user interface.
[0118] (Application example 1)
[0119] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0120] Currently, in factory meetings and discussions, creating minutes and extracting important keywords is often done manually, which is extremely time-consuming. There is also the risk that important points may be overlooked, creating challenges for efficient project management and quality control. Large factories, in particular, where many meetings are held, require efficient and accurate management of meeting content. Therefore, a system is needed to streamline in-factory meeting records and quickly grasp important information.
[0121] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0122] In this invention, the server includes means for receiving voice data, means for converting the received voice data into text data, means for extracting specific keywords from the converted text data, means for storing the extracted keywords and their location information in a database, means for displaying the stored data on a user interface, means for collecting voice data from multifunction devices installed in an industrial environment, and means for displaying the text data and its keywords on an industrial user interface, thereby making it possible to efficiently record meetings in a factory and quickly grasp important information.
[0123] "Voice data" means data in digital form that records what a user says.
[0124] "Character data" refers to audio data converted into text format.
[0125] A "keyword" is a word or phrase that has an important meaning within the text data.
[0126] "Position information" is information indicating the position where a keyword appears within character data.
[0127] A "database" is a digital storage system for storing text data, keywords, location information, etc.
[0128] A "user interface" is a collection of screens and input devices that allow a user to interact with a system.
[0129] A "multifunction device" is a device for voice collection and data processing installed in an industrial environment.
[0130] An "industrial user interface" is an interface that displays data in a format suitable for business processes within a factory.
[0131] The system according to the present invention is for efficiently managing voice data from meetings and conferences in industrial environments and quickly grasping important information. Specifically, it includes the following components:
[0132] Hardware and software usage examples
[0133] Hardware:
[0134] Smart speaker: A device for collecting recordings of meetings and discussions within the factory.
[0135] Conference recording device: A device that records speech during a conference and provides the data to a server.
[0136] Robot: A machine installed in a factory that supports voice data collection and data management (e.g., Pepper).
[0137] software:
[0138] Python: A programming language for implementing server-side programs.
[0139] Google Speech Recognition API: A speech recognition service for converting voice data into text data.
[0140] SQLite: A database system for storing converted character data and keywords.
[0141] Flask: A web application framework for building user interfaces and displaying data.
[0142] System operation explanation
[0143] 1. Audio data collection:
[0144] When users hold meetings or conferences, smart speakers, conference recording devices, and robots collect audio data, which is then stored digitally and uploaded to a server.
[0145] 2. Audio to text conversion:
[0146] The server uses a Python program to read the received audio data and converts it into text using the Google Speech Recognition API. Through this process, the meeting contents are recorded in text format.
[0147] 3. Keyword extraction:
[0148] Predefined important keywords (e.g., "production line," "defective product," "maintenance") are automatically extracted from the converted text data, and the position in the text data where the keywords appear is also recorded.
[0149] 4. Save to database:
[0150] The extracted keywords, their location information, and the converted text data are stored in an SQLite database. A timestamp is added to the data when it is saved to ensure data integrity.
[0151] 5. Display in the user interface:
[0152] Through a web interface built with Flask, users can view saved meeting data and keywords. The displayed content includes text data and important keywords, allowing users to efficiently understand the content of the meeting.
[0153] Specific examples
[0154] Recorded data from quality control meetings held within the factory is collected using a smart speaker, and important keywords such as "defective products" and "maintenance" are automatically extracted and displayed on the server's web interface.
[0155] Prompt Sentence Examples
[0156] "Convert the meeting recording data into text data, extract specific keywords (e.g., "production line," "defective product," "maintenance"), and store them in a database. Create a web application that displays this in a user interface."
[0157] This will make it possible to efficiently record meetings within the factory and quickly grasp important information.
[0158] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0159] Step 1:
[0160] When a user starts a meeting, audio data is collected by a smart speaker, a meeting recording device, or a robot.
[0161] Input: Audio data during the meeting
[0162] Output: Digitally saved audio data file
[0163] What it does: A user starts a meeting, and the device collects audio data in real time and stores it in its internal storage.
[0164] Step 2:
[0165] The collected voice data is uploaded to a server, and the user uses a terminal to send the voice data to the server through a dedicated user interface.
[0166] Input: Audio data file saved on the device
[0167] Output: Audio data file uploaded to the server
[0168] Specific operation: The user clicks the "File Upload" button on the user interface to send the audio data file to the server.
[0169] Step 3:
[0170] The server reads the received audio data using a Python program.
[0171] Input: Uploaded audio data
[0172] Output: Audio data object
[0173] Specific operation: The server saves the audio file in a specific directory and reads the file using a Python program.
[0174] Step 4:
[0175] The server uses the Google Speech Recognition API to convert the audio data into text data.
[0176] Input: Audio data object
[0177] Output: Converted character data (text)
[0178] Specific operation: The server calls the Google Speech Recognition API and converts the voice data into text data.
[0179] Step 5:
[0180] The server extracts specific keywords from the converted character data.
[0181] Input: Character data
[0182] Output: Extracted keywords and their locations
[0183] Specific operation: The server matches the character data with a predefined keyword list and records the keywords and their location information.
[0184] Step 6:
[0185] The server stores the extracted keywords, their location information, and the converted text data in an SQLite database.
[0186] Input: Keywords, location information, text data
[0187] Output: Data stored in the database
[0188] Specific operation: The server connects to an SQLite database and stores keywords, location information, and text data along with timestamps.
[0189] Step 7:
[0190] The server displays the stored data in a web interface built with Flask.
[0191] Input: Data stored in a database
[0192] Output: Data displayed in the user interface
[0193] Specific operation: The server generates a user interface through a Flask application, allowing users to view the meeting content and keywords.
[0194] This allows users to efficiently manage meeting records within the factory and quickly grasp important information.
[0195] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0196] This invention combines a system that receives voice data, converts it into text data, extracts specific keywords, stores them in a database, and displays them on a user interface in order to improve the efficiency of project management, with an emotion engine that recognizes the user's emotions. This makes it possible to analyze not only the content of meetings and discussions, but also the emotional states of team members, to provide an optimal working environment.
[0197] Audio data reception
[0198] First, the user records the audio data of a meeting or conference and uses the system's "audio file upload" function on their device. They click the "File Upload" button on the user interface and select the desired audio file from the file system. The server receives the audio file and saves it in the specified directory.
[0199] Converting voice data to text
[0200] Next, the server reads the received voice data as audio data, initializes the voice recognition library, and converts it into text data using the Google Speech Recognition API, etc. This allows the voice content to be obtained in text format.
[0201] Keyword extraction
[0202] To extract specific keywords from the converted text data, the server uses a predefined list of important keywords. It checks whether these keywords exist in the text data and records their location information. For example, if keywords such as "important," "risk," and "deadline" are included, it records the words and their locations in a list.
[0203] Emotional state analysis using emotion engine
[0204] The server then analyzes the recorded voice data and identifies the user's emotional state using an emotion engine, which uses voice analysis technology to infer emotions from the tone, speed, and strength of the voice.
[0205] Saving to a database
[0206] The server stores the converted text data, extracted keywords and their location information, as well as the analyzed emotional state in a database. When storing the data, it records the current timestamp to maintain data integrity.
[0207] Display in the user interface
[0208] Finally, the server displays the saved data on a user interface. Users access the system from their devices and view text data, important keywords, and emotional states in real time through the interface. For example, the text data for a meeting might read, "We need to discuss the risks and important deadlines for the next project," along with the keywords "risk," "importance," and "deadline." Furthermore, the emotional states of the participants at the relevant points in the conversation might be displayed as "tense" or "excited."
[0209] In this way, users and teams can not only understand the content of the meeting, but also take into account emotional information to make appropriate responses and allocate resources.The system also records user interactions, contributing to the efficiency of future meetings.
[0210] The processing flow will be explained below.
[0211] Step 1:
[0212] A user records an audio file of a meeting or conference and uses the system's "audio file upload" function on their device. They click the "file upload" button on the user interface and select the desired audio file from the file system.
[0213] Step 2:
[0214] The device uploads the selected audio file to the server via an HTTP request, and the server receives the request and saves the audio file in the specified directory.
[0215] Step 3:
[0216] The server reads the saved audio file and captures the audio data as audio data. Specifically, it initializes the speech recognition library and loads the audio file as an audio object.
[0217] Step 4:
[0218] The server calls a speech recognition service such as the Google Speech Recognition API to convert the audio data into text data. If the conversion process is successful, the content of the audio data is obtained in text format.
[0219] Step 5:
[0220] The server extracts specific keywords from the acquired text data, checks them against a predefined list of important keywords, identifies the keywords contained in the text data, and records their location information.
[0221] Step 6:
[0222] The server inputs the voice data into the emotion engine to analyze the emotional state of the user or speaker. The emotion engine analyzes the tone, rate, intensity, etc. of the voice to determine emotions (e.g., tension, excitement, satisfaction, etc.).
[0223] Step 7:
[0224] The server stores the converted text data, extracted keywords, and analyzed emotional states in a database, adding a current timestamp to each piece of data to maintain data integrity.
[0225] Step 8:
[0226] The server displays the stored data on a user interface, and the user accesses the system from a terminal and checks the text data, important keywords, and emotional state in real time through the interface.
[0227] Step 9:
[0228] Users can review the displayed data to understand the content of the meeting and the emotional state of the speakers, which allows them to take appropriate action and adjust resource allocation.
[0229] Step 10:
[0230] The user can take necessary measures and allocate resources to improve the efficiency of future meetings. The system records these interactions and uses them for future analysis and improvement.
[0231] Example 2
[0232] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0233] Conventional project management systems make it difficult to effectively grasp the content of meetings and discussions. In particular, they do not support extracting information from voice data or analyzing users' emotional states, making it difficult to grasp the psychological state and important points of team members in real time.
[0234] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0235] In this invention, the server includes means for receiving voice data, means for converting the received voice data into character data, means for extracting specific keywords from the converted character data, means for saving the extracted keywords and their location information in a database, means for analyzing the emotional state of a user from the voice data using an emotion engine, means for saving the analyzed emotional state in a database, and means for displaying the saved data on a user interface. This makes it possible to grasp in detail not only the content of meetings and discussions but also the emotional state of the speakers, thereby making project management more efficient.
[0236] "Audio data" refers to digital data of audio recorded during a meeting or conference.
[0237] The "receiving means" is a device or software for capturing voice data sent from a user into the system.
[0238] "Means for converting into text data" refers to a device or program that includes an algorithm or voice recognition technology for analyzing voice data and converting it into text format data.
[0239] A "keyword extraction means" is a device or software that detects and extracts specified important words or phrases from character data.
[0240] An "emotion engine" is an algorithm or software that estimates emotional states through the analysis of voice data.
[0241] The "means for storing in a database" is a storage system for organizing and recording the converted character data, extracted keywords, and analyzed emotional states.
[0242] The "means for displaying on the user interface" refers to a screen or software for displaying the saved data in a form that can be viewed by the user.
[0243] The present invention aims to improve the efficiency of project management by providing a system that receives voice data, converts it into text data, extracts keywords, analyzes emotional states, stores the data in a database, and displays it on a user interface.
[0244] Hardware and software used
[0245] The main components of the system are the server, the terminal, and the user. The specific hardware and software used are specified below:
[0246] 1. Hardware
[0247] Server: A powerful computer for data processing and storage
[0248] Terminal: The device used by the user to access the site (PC, tablet, smartphone, etc.)
[0249] Microphone: A device for recording audio data.
[0250] 2. Software
[0251] Speech recognition library: Google Speech Recognition API
[0252] Database management system: MySQL or PostgreSQL
[0253] Emotion engine: a program with a voice analysis algorithm
[0254] User Interface: Web applications implemented using HTML, CSS, and JavaScript
[0255] Program processing
[0256] Audio data reception
[0257] Users record audio from meetings and conferences and upload it to the server via their devices. Specifically, they click the "File Upload" button on the user interface and select an audio file from the file system. The server receives the audio file and saves it in the specified directory.
[0258] Converting voice data to text
[0259] The server reads the received voice data as audio data and initializes a voice recognition library such as the Google Speech Recognition API. The voice data is converted into text data using the voice recognition library, and the voice content is obtained in text format.
[0260] Keyword extraction
[0261] The server extracts specific keywords from the text data. Using a predefined list of important keywords (e.g., "important," "risk," "deadline"), it searches for the presence of these keywords in the text data. It records the keywords found and their location information in a list.
[0262] Emotional state analysis using emotion engine
[0263] The server feeds the recorded voice data to the emotion engine, which analyzes factors such as tone, speed, and strength of the voice. Based on the analysis results, the user's emotional state is identified and the results are stored in a database.
[0264] Saving to a database
[0265] The server stores the converted text data, extracted keywords and their location information, and analyzed emotional state in a database, where the data is recorded with the current timestamp to maintain data integrity.
[0266] Display in the user interface
[0267] The server displays the stored data on a user interface, and users can access the system from their devices and check the text data, important keywords, and emotional state in real time through the interface.
[0268] Specific examples
[0269] For example, the text data for a meeting might read, "We need to discuss risks and important deadlines for the next project," along with keywords like "risk," "important," and "deadline." Emotional states are also displayed as "tension" or "excitement." Based on this information, users and teams can take appropriate action and allocate resources.
[0270] Prompt Sentence Examples
[0271] Prompt: "Transcribe what is said in the next meeting, extract key keywords, and analyze the emotional state of the conversation. For example, display keywords such as risk, importance, and deadline, along with the associated emotions."
[0272] In this way, the system analyzes the content of meetings and discussions in detail, aiming to improve the efficiency of project management.
[0273] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0274] Step 1: Record and upload audio data
[0275] The user records the audio of a meeting or conference and selects the audio file from the device's file system. Next, the user uses the system's "file upload" function from the device to send the audio file to the server. The input is the audio file, and the output is an audio file saved in a specified directory on the server. Specifically, the user presses the record button, and then clicks the upload button after recording is complete.
[0276] Step 2: Receive and save the audio file
[0277] The server receives audio files sent through the user interface and saves them in the specified directory. The received audio files are recorded in a log, along with metadata to maintain consistency. The input is the uploaded audio file, and the output is a file saved on the server. Specifically, the server receives a send request and processes the saving of the audio file.
[0278] Step 3: Loading the audio data
[0279] The server reads the saved audio file as audio data. During this process, the file format and encoding are checked. The input is the saved audio file, and the output is the read audio data. Specifically, the server uses Python or JavaScript libraries to open the audio file.
[0280] Step 4: Initializing the speech recognition library
[0281] The server initializes a speech recognition library such as the Google Speech Recognition API, sets an API key, and runs the authentication process. The input is the API key, and the output is the initialized speech recognition library. Specifically, the server sets the API key and initializes the library.
[0282] Step 5: Converting audio data to text
[0283] The server uses a speech recognition library to convert voice data into text data. The input is the read audio data, and the output is the converted text data. Specifically, the server sends the voice data to the speech recognition API and stores the returned text data in memory.
[0284] Step 6: Keyword extraction
[0285] The server loads a predefined list of important keywords (e.g., "important," "risk," "deadline") and extracts these keywords from the text data. The input is the converted text data and a keyword list, and the output is the extracted keywords and their location information. Specifically, it performs regular expression and pattern matching on the text data.
[0286] Step 7: Emotional state analysis by the emotion engine
[0287] The server supplies the recorded voice data to the emotion engine, which analyzes the tone, speed, and strength of the voice. The input is the voice data, and the output is information about the user's emotional state. Specifically, the engine uses a voice analysis algorithm to estimate the emotional state.
[0288] Step 8: Saving to the Database
[0289] The server stores the converted text data, extracted keywords and their location information, and analyzed emotional state in a database. During this process, the current timestamp is also recorded. The input is the text data, keywords, location information, and emotional state, and the output is the complete dataset stored in the database. Specifically, it generates an SQL query and inserts it into the database.
[0290] Step 9: Display in the User Interface
[0291] The server retrieves the necessary data from the database and displays it on the user interface. The input is the information stored in the database, and the output is the displayed text data, keywords, and emotional state. Specifically, the interface is constructed using HTML and JavaScript, and the data is displayed dynamically.
[0292] In this way, the system can sequentially perform specific tasks at each processing step, ultimately providing the information the user needs.
[0293] (Application example 2)
[0294] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0295] In today's factories, meetings and communication between workers are not managed efficiently, and improvements are needed, particularly in quality control and work efficiency. There is a lack of means to grasp the situation in real time using voice data, and there is no system that analyzes the emotional state of workers and provides appropriate feedback. This makes it difficult to detect problems early and take appropriate measures.
[0296] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0297] In this invention, the server includes means for receiving voice data, means for converting the received voice data into character data, means for extracting specific keywords from the converted character data, means for saving the extracted keywords and their location information in a database, means for displaying the saved data on a user interface, means for analyzing an emotional state from the voice data, means for saving the analyzed emotional state in a database, and means for displaying the saved data and the emotional state on a user interface. This enables real-time monitoring and efficient management of meetings and work status in a factory, and realizes quality control and improved work efficiency based on the emotional state of workers.
[0298] 1. "Voice data" refers to data that includes voice information from meetings or communications between workers.
[0299] 2. "Text data" means text-format data converted from voice data using voice recognition technology.
[0300] 3. "Keywords" refer to words or phrases that are particularly important in a meeting or work content.
[0301] 4. "Location information" is information that indicates the location where a keyword appears within the converted character data.
[0302] 5. "Database" refers to a storage means for storing and managing converted text data, extracted keywords, location information, and emotional state.
[0303] 6. "User Interface" means the display means by which a user accesses the system and checks and manipulates data.
[0304] 7. "Emotional state" refers to the results of analyzing voice data to infer the user's emotions.
[0305] 8. "Means of analysis" includes processes and technologies for extracting and identifying text data and emotional states from audio data.
[0306] 9. "Real-time" means that data is processed and displayed with little or no delay after it is generated.
[0307] 10. "Quality control" refers to activities to maintain and improve the quality of products and work processes on the factory floor.
[0308] 11. "Work efficiency" is an indicator of how efficiently work is being carried out on the factory floor.
[0309] 12. "Monitoring" refers to the continuous observation and recording of meetings and work situations.
[0310] The present invention is a system for improving the efficiency of meetings and communication between workers at a factory. Specific examples are shown below.
[0311] First, microphones installed in the factory record audio data of meetings and between workers. Users manage this audio data on their devices and use the system's "audio file upload" function to send it to the server. The server then saves the uploaded audio data in a specified directory.
[0312] The server receives the stored voice data and converts it into text data using the Google Cloud Speech-to-Text API, which has the ability to convert voice to text with high accuracy.
[0313] Next, the server uses a predefined list of important keywords to extract specific keywords from the converted text data. It checks whether any of the keywords in this list are present in the text data and records their location information. For example, if keywords such as "quality," "risk," and "work" are present, it records the words and their locations in the list.
[0314] The server then analyzes the recorded voice data and identifies the user's emotional state using the Dialogflow API. The emotion engine estimates emotions by analyzing the tone, speed, and strength of the voice. As a result of the analysis, the emotional state of the meeting or worker communication is identified as "tension" or "attention," etc.
[0315] The server stores the converted text data, extracted keywords, their location information, and analyzed emotional state in a database. This data also includes a current timestamp to ensure data integrity.
[0316] Users can access the system from their devices and view text data, key keywords, and emotional states in real time through the interface. For example, if a meeting content reads, "We should strengthen quality inspections for the next production batch. Risks are increasing," the system will identify the keywords "quality" and "risk" along with the text data, and display the emotional states of the relevant parts of the conversation as "tension" or "caution."
[0317] This system not only improves the efficiency of in-plant meetings and communication between workers, but also incorporates emotional information to enable early detection of problems and appropriate countermeasures, allowing managers to focus on specific problems and take prompt action.
[0318] Example inputs to a generative AI model:
[0319] Project Management / Meeting Content:
[0320] "Quality inspections should be strengthened for the next production batch. The risk is increasing."
[0321] Keywords: "quality" and "risk"
[0322] Emotional state: "tension" "attention"
[0323] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0324] Step 1:
[0325] Microphones installed in factories record audio data during meetings and between workers. Users manage this audio data on their devices and send it to the server using the system's "audio file upload" function. The input data is an audio file, and the output data is an audio file saved on the server.
[0326] Step 2:
[0327] The server saves the uploaded audio data in the specified directory. The input data is the audio file received from the user, and the output data is the path to the audio file saved on the server. The specific operation is to specify the file save path and save the file in storage.
[0328] Step 3:
[0329] The server converts the saved audio data into text data using the Google Cloud Speech-to-Text API. The input data is the path to the saved audio file, and the output data is the converted text data. Specifically, the API is called to convert the audio file into text format.
[0330] Step 4:
[0331] The server extracts specific keywords from the converted text data. The input data is the text data and a predefined keyword list, and the output data is the extracted keywords and their location information. Specifically, the server searches for the locations where keywords appear in the text data and records the location information.
[0332] Step 5:
[0333] The server analyzes the recorded voice data and identifies the user's emotional state using the Dialogflow API. The input data is the voice data, and the output data is the analyzed emotional state. Specifically, the server passes the voice data to the Dialogflow API, which estimates the user's emotional state.
[0334] Step 6:
[0335] The server stores the converted text data, extracted keywords and their location information, and analyzed emotional states in a database. The input data is the text data, keyword information, and emotional states, and the output data is the information stored in the database. Specifically, the server inserts the data into the database in an appropriate format.
[0336] Step 7:
[0337] Users access the system from their devices and check text data, important keywords, and emotional states in real time through the interface. The input data is information retrieved from the database, and the output data is information displayed on the user interface. Specific operations include updating the UI to display the information retrieved from the database in real time.
[0338] These steps will enable efficient management of factory meetings and communication between workers, including emotional information, and enable early detection and countermeasures for problems.
[0339] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0340] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0341] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0342] [Second embodiment]
[0343] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0344] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0345] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0346] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0347] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0348] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0349] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0350] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0351] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0352] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0353] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0354] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0355] The present invention provides a system that receives voice data, converts it into text data, extracts specific keywords, stores them in a database, and displays them on a user interface in order to improve the efficiency of project management.
[0356] 1. Acceptance of audio data
[0357] First, the user uploads the recording of the meeting or conference from their own device to the server. Specifically, they click the "File Upload" button on the user interface and select the desired audio file from the file system. The server receives the audio file and saves it in the specified directory.
[0358] 2. Converting voice data to text data
[0359] Next, the server reads the received voice data as audio data, initializes a speech recognition library, and converts the voice data into text using a speech recognition service such as the Google Speech Recognition API. During this process, the server works to convert the voice content into text data with high accuracy.
[0360] 3. Keyword extraction
[0361] To extract specific keywords from the converted text data, the server uses a predefined list of important keywords. It checks whether these keywords exist in the text data and records their location information. For example, if keywords such as "important," "risk," and "deadline" are included, it records the words and their locations in a list.
[0362] 4. Saving to the database
[0363] The extracted keywords and their location information are stored in a database by the server. To maintain data integrity, the data is recorded with a current timestamp. This makes it clear what was discussed and when.
[0364] 5. Display in the user interface
[0365] Finally, the server displays the saved data on a user interface. Through the interface, users can check the text data and important keywords extracted from the recording data of meetings and discussions. For example, if a user uploads an audio file called "project_meeting.wav," the server analyzes the audio data and displays the text data "We need to discuss the risks and important deadlines of the next project" along with the keywords "risk" and "deadline" to the user.
[0366] In this way, the user can efficiently grasp the contents of the meeting, quickly confirm important points, and take necessary measures.
[0367] The processing flow will be explained below.
[0368] Step 1:
[0369] A user records an audio file of a meeting or conference and uses the system's "audio file upload" function on their device. They click the "file upload" button on the user interface and select the desired audio file from the file system.
[0370] Step 2:
[0371] The device uploads the selected audio file to the server via an HTTP request, and the server receives the request and saves the audio file in the specified directory.
[0372] Step 3:
[0373] The server reads the saved audio file and captures the audio data as audio data. Specifically, it initializes the speech recognition library and loads the audio file as an audio object.
[0374] Step 4:
[0375] The server calls a speech recognition service such as the Google Speech Recognition API to convert the audio data into text data. If the conversion process is successful, the content of the audio data is obtained in text format.
[0376] Step 5:
[0377] The server extracts specific keywords from the acquired text data, checks them against a predefined list of important keywords, identifies the keywords contained in the text data, and records their location information.
[0378] Step 6:
[0379] The server stores the converted text data, extracted keywords, and their location information in a database. When storing the data, it records the data with the current timestamp to maintain data integrity.
[0380] Step 7:
[0381] The server displays the stored data on a user interface. Users access the system from their devices and can view text data and important keywords in real time through the interface. The display includes specific discussion content, risks, and important points.
[0382] Step 8:
[0383] Users can check the interface and quickly take necessary measures and adjust resource allocation. The system records user interactions, contributing to the efficiency of future meetings.
[0384] Example 1
[0385] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0386] There is a demand for efficient recording of meeting and conference content so that important points can be quickly grasped. However, conventional systems require manual conversion of audio data into text data and manual extraction of important keywords, which is extremely time-consuming. In addition, because data is also saved and displayed manually, there are issues with information leaks and inconsistency.
[0387] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0388] In this invention, the server includes means for receiving voice data, means for converting the received voice data into text data, means for extracting specific keywords from the converted text data, means for saving the extracted keywords and their location information in a database, means for displaying the saved data on a user interface, means for converting voice data into text data using a voice recognition service, means for loading a predefined keyword list to extract specific keywords, means for recording a current timestamp when saving the text data including the location information of the extracted keywords in the database, and means for uploading voice data and checking the results of the analysis process via the user interface. This makes it possible to automatically and efficiently record the contents of meetings and discussions and easily grasp important points.
[0389] "Audio data" refers to data in which the contents of a meeting, conference, etc. are recorded in audio format.
[0390] "Character data" is data in text format that has been converted from voice data using voice recognition technology.
[0391] "Keywords" refer to specific words or phrases that are considered important in the content of a meeting or discussion.
[0392] A "database" is a digital system for efficiently storing and managing information, and in this case, it plays a role in storing text data based on voice data and extracted keywords.
[0393] A "user interface" is the part of a system that contains the screens and controls that allow a user to interact with the system.
[0394] A "means for receiving" is a technique or mechanism for capturing audio data from a user.
[0395] "Means for converting" refers to the technology or method for converting voice data into text data.
[0396] "Extraction means" refers to techniques or methods for extracting specific keywords from the converted character data.
[0397] "Storage means" refers to the technology or method for storing the extracted keywords and their location information in a database.
[0398] "Display means" refers to the techniques and methods for presenting the stored data in a user interface so that the user can check it.
[0399] "Speech Recognition Service" means a cloud-based or local service for analyzing and converting voice data into text data.
[0400] A "keyword list" is a collection of predefined important keywords that are used in text analysis.
[0401] A "timestamp" is time information that indicates that data was saved or processed at a specific time.
[0402] "Upload" refers to the operation or process by which a user sends audio data from their device to a server.
[0403] "Analysis" refers to a series of processes that converts voice data into text data and then extracts keywords.
[0404] The present invention aims to improve the efficiency of project management by providing a system that receives voice data, converts it into text data, extracts specific keywords, stores them in a database, and displays them on a user interface. Specifically, the invention can be implemented in the following forms.
[0405] First, the user uploads the recording of a meeting or conference from their own device to the server. The user clicks the "File Upload" button on the user interface and selects the desired audio file from the file system. The uploaded audio file is sent from the device to the server, and the server receives it and saves it in the specified directory. At this time, the file's metadata (e.g., file name, upload date and time, user ID, etc.) is also saved.
[0406] Next, the server reads the received voice data. It treats the voice data as audio data and initializes a voice recognition library. The system uses a voice recognition service such as the Google Speech Recognition API to convert the voice data into text. Through the conversion process, the content of the voice data is obtained as text data. For example, the resulting text data is, "We need to discuss the risks and important deadlines for the next project."
[0407] The server then extracts specific keywords from the converted text data using a predefined list of important keywords (e.g., "important," "risk," and "deadline"). The server uses a text analysis algorithm to identify the presence of keywords in the converted text data and record their location (e.g., which sentence or word). For example, if the words "risk" and "deadline" appear on the third line of the text, the server will add their location to the list.
[0408] The server then saves the extracted keywords and their location information, along with the text data, to a database. When saving, the server also records the current timestamp. This ensures data integrity. The saved data is stored in the "meeting_transcripts" table.
[0409] Finally, the server displays the stored data on a user interface. The server retrieves the stored text data and keywords from the database, formats them, and displays them to the user. Through the user interface, the user can view the text data and important keywords extracted from the recording of a meeting or discussion. For example, if a user uploads an audio file called "project_meeting.wav," the server analyzes the audio data and displays the text data, "We need to discuss the risks and important deadlines for the next project," along with the keywords "risk" and "deadline" to the user.
[0410] Specific examples
[0411] A user uploads an audio file called "project_meeting.wav." The server then converts it into text using the Google Speech Recognition API and extracts the keywords "risk" and "deadline." The extracted information is stored in a database and the user interface displays the message, "We need to discuss the risks and important deadlines for the next project."
[0412] Example of input prompt for generative AI model
[0413] An example of a prompt is, "Please upload the recording of the project meeting, 'project_meeting.wav'. The contents of this audio file will be converted into text data, and important keywords will be extracted and displayed. The service used is the Google Speech Recognition API."
[0414] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0415] Step 1:
[0416] The user uploads the audio data.
[0417] A user accesses the user interface of the project management system using a terminal. He clicks the "File Upload" button on the user interface and selects the desired audio file from the file system. The input is the audio file selected by the user (e.g., project_meeting.wav). The terminal sends the audio file to the server, and the audio data is sent to the server as the output.
[0418] Step 2:
[0419] The server receives and stores the audio data.
[0420] The server receives audio data sent from the device. It saves the received audio data in a specified directory (e.g., / uploaded_audio). At this time, it also saves metadata such as the file name, upload date and time, and user ID. The input is the audio data sent by the user. The output is an audio file saved in a specified directory on the server.
[0421] Step 3:
[0422] The server initializes the speech recognition library and reads the audio data.
[0423] The server initializes a speech recognition library (e.g., Google Speech Recognition API). It then reads the saved audio data and loads it into memory as audio data. The input is the saved audio data file. The output is the audio data loaded into memory.
[0424] Step 4:
[0425] The server converts the audio data into text data.
[0426] The server converts the audio data into text using an initialized speech recognition library. The speech recognition library analyzes the audio data and outputs the speech in text format. The input is the audio data loaded in memory. The output is text (e.g., "We need to discuss the risks and important deadlines for the next project").
[0427] Step 5:
[0428] The server loads the keyword list and extracts the keywords.
[0429] The server loads a predefined list of important keywords (e.g., "important," "risk," "deadline"), then checks for the presence of keywords in the converted text data and records location information. The input is the converted text data and the keyword list. The output is the extracted keywords (e.g., "risk," "deadline") and their location information.
[0430] Step 6:
[0431] The server stores the extracted data in a database.
[0432] The server stores the extracted keywords and text data, including their location information, in a database. The stored data also includes the current timestamp. The input is text data with keywords and location information added. The output is the information stored in the database.
[0433] Step 7:
[0434] The server displays the stored data in a user interface.
[0435] The server retrieves the stored text data and keywords from the database, formats them, and displays them on the user interface. Through the user interface, users can check the text data and important keywords extracted from the recording data of meetings and conferences. The input is the text data and keyword information retrieved from the database. The output is the text data and keywords displayed on the user interface.
[0436] (Application example 1)
[0437] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0438] Currently, in factory meetings and discussions, creating minutes and extracting important keywords is often done manually, which is extremely time-consuming. There is also the risk that important points may be overlooked, creating challenges for efficient project management and quality control. Large factories, in particular, where many meetings are held, require efficient and accurate management of meeting content. Therefore, a system is needed to streamline in-factory meeting records and quickly grasp important information.
[0439] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0440] In this invention, the server includes means for receiving voice data, means for converting the received voice data into text data, means for extracting specific keywords from the converted text data, means for storing the extracted keywords and their location information in a database, means for displaying the stored data on a user interface, means for collecting voice data from multifunction devices installed in an industrial environment, and means for displaying the text data and its keywords on an industrial user interface, thereby making it possible to efficiently record meetings in a factory and quickly grasp important information.
[0441] "Voice data" means data in digital form that records what a user says.
[0442] "Character data" refers to audio data converted into text format.
[0443] A "keyword" is a word or phrase that has an important meaning within the text data.
[0444] "Position information" is information indicating the position where a keyword appears within character data.
[0445] A "database" is a digital storage system for storing text data, keywords, location information, etc.
[0446] A "user interface" is a collection of screens and input devices that allow a user to interact with a system.
[0447] A "multifunction device" is a device for voice collection and data processing installed in an industrial environment.
[0448] An "industrial user interface" is an interface that displays data in a format suitable for business processes within a factory.
[0449] The system according to the present invention is for efficiently managing voice data from meetings and conferences in industrial environments and quickly grasping important information. Specifically, it includes the following components:
[0450] Hardware and software usage examples
[0451] Hardware:
[0452] Smart speaker: A device for collecting recordings of meetings and discussions within the factory.
[0453] Conference recording device: A device that records speech during a conference and provides the data to a server.
[0454] Robot: A machine installed in a factory that supports voice data collection and data management (e.g., Pepper).
[0455] software:
[0456] Python: A programming language for implementing server-side programs.
[0457] Google Speech Recognition API: A speech recognition service for converting voice data into text data.
[0458] SQLite: A database system for storing converted character data and keywords.
[0459] Flask: A web application framework for building user interfaces and displaying data.
[0460] System operation explanation
[0461] 1. Audio data collection:
[0462] When users hold meetings or conferences, smart speakers, conference recording devices, and robots collect audio data, which is then stored digitally and uploaded to a server.
[0463] 2. Audio to text conversion:
[0464] The server uses a Python program to read the received audio data and converts it into text using the Google Speech Recognition API. Through this process, the meeting contents are recorded in text format.
[0465] 3. Keyword extraction:
[0466] Predefined important keywords (e.g., "production line," "defective product," "maintenance") are automatically extracted from the converted text data, and the positions in the text data where the keywords appear are also recorded.
[0467] 4. Save to database:
[0468] The extracted keywords, their location information, and the converted text data are stored in an SQLite database. A timestamp is added to the data when it is saved to ensure data integrity.
[0469] 5. Display in the user interface:
[0470] Through a web interface built with Flask, users can view saved meeting data and keywords. The displayed content includes text data and important keywords, allowing users to efficiently understand the content of the meeting.
[0471] Specific examples
[0472] Recorded data from quality control meetings held within the factory is collected using a smart speaker, and important keywords such as "defective products" and "maintenance" are automatically extracted and displayed on the server's web interface.
[0473] Prompt Sentence Examples
[0474] "Convert the meeting recording data into text data, extract specific keywords (e.g., "production line," "defective product," "maintenance"), and store them in a database. Create a web application that displays this in a user interface."
[0475] This will make it possible to efficiently record meetings within the factory and quickly grasp important information.
[0476] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0477] Step 1:
[0478] When a user starts a meeting, audio data is collected by a smart speaker, a meeting recording device, or a robot.
[0479] Input: Audio data during the meeting
[0480] Output: Digitally saved audio data file
[0481] What it does: A user starts a meeting, and the device collects audio data in real time and stores it in its internal storage.
[0482] Step 2:
[0483] The collected voice data is uploaded to a server, and the user uses a terminal to send the voice data to the server through a dedicated user interface.
[0484] Input: Audio data file saved on the device
[0485] Output: Audio data file uploaded to the server
[0486] Specific operation: The user clicks the "File Upload" button on the user interface to send the audio data file to the server.
[0487] Step 3:
[0488] The server reads the received audio data using a Python program.
[0489] Input: Uploaded audio data
[0490] Output: Audio data object
[0491] Specific operation: The server saves the audio file in a specific directory and reads the file using a Python program.
[0492] Step 4:
[0493] The server uses the Google Speech Recognition API to convert the audio data into text data.
[0494] Input: Audio data object
[0495] Output: Converted character data (text)
[0496] Specific operation: The server calls the Google Speech Recognition API and converts the voice data into text data.
[0497] Step 5:
[0498] The server extracts specific keywords from the converted character data.
[0499] Input: Character data
[0500] Output: Extracted keywords and their locations
[0501] Specific operation: The server matches the character data with a predefined keyword list and records the keywords and their location information.
[0502] Step 6:
[0503] The server stores the extracted keywords, their location information, and the converted text data in an SQLite database.
[0504] Input: Keywords, location information, text data
[0505] Output: Data stored in the database
[0506] Specific operation: The server connects to an SQLite database and stores keywords, location information, and text data along with timestamps.
[0507] Step 7:
[0508] The server displays the stored data in a web interface built with Flask.
[0509] Input: Data stored in a database
[0510] Output: Data displayed in the user interface
[0511] Specific operation: The server generates a user interface through a Flask application, allowing users to view the meeting content and keywords.
[0512] This allows users to efficiently manage meeting records within the factory and quickly grasp important information.
[0513] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0514] This invention combines a system that receives voice data, converts it into text data, extracts specific keywords, stores them in a database, and displays them on a user interface in order to improve the efficiency of project management, with an emotion engine that recognizes the user's emotions. This makes it possible to analyze not only the content of meetings and discussions, but also the emotional states of team members, to provide an optimal working environment.
[0515] Audio data reception
[0516] First, the user records the audio data of a meeting or conference and uses the system's "audio file upload" function on their device. They click the "File Upload" button on the user interface and select the desired audio file from the file system. The server receives the audio file and saves it in the specified directory.
[0517] Converting voice data to text
[0518] Next, the server reads the received voice data as audio data, initializes the voice recognition library, and converts it into text data using the Google Speech Recognition API, etc. This allows the voice content to be obtained in text format.
[0519] Keyword extraction
[0520] To extract specific keywords from the converted text data, the server uses a predefined list of important keywords. It checks whether these keywords exist in the text data and records their location information. For example, if keywords such as "important," "risk," and "deadline" are included, it records the words and their locations in a list.
[0521] Emotional state analysis using emotion engine
[0522] The server then analyzes the recorded voice data and identifies the user's emotional state using an emotion engine, which uses voice analysis technology to infer emotions from the tone, speed, and strength of the voice.
[0523] Saving to a database
[0524] The server stores the converted text data, extracted keywords and their location information, as well as the analyzed emotional state in a database. When storing the data, it records the current timestamp to maintain data integrity.
[0525] Display in the user interface
[0526] Finally, the server displays the saved data on a user interface. Users access the system from their devices and view text data, important keywords, and emotional states in real time through the interface. For example, the text data for a meeting might read, "We need to discuss the risks and important deadlines for the next project," along with the keywords "risk," "importance," and "deadline." Furthermore, the emotional states of the participants at the relevant points in the conversation might be displayed as "tense" or "excited."
[0527] In this way, users and teams can not only understand the content of the meeting, but also take into account emotional information to make appropriate responses and allocate resources.The system also records user interactions, contributing to the efficiency of future meetings.
[0528] The processing flow will be explained below.
[0529] Step 1:
[0530] A user records an audio file of a meeting or conference and uses the system's "audio file upload" function on their device. They click the "file upload" button on the user interface and select the desired audio file from the file system.
[0531] Step 2:
[0532] The device uploads the selected audio file to the server via an HTTP request, and the server receives the request and saves the audio file in the specified directory.
[0533] Step 3:
[0534] The server reads the saved audio file and captures the audio data as audio data. Specifically, it initializes the speech recognition library and loads the audio file as an audio object.
[0535] Step 4:
[0536] The server calls a speech recognition service such as the Google Speech Recognition API to convert the audio data into text data. If the conversion process is successful, the content of the audio data is obtained in text format.
[0537] Step 5:
[0538] The server extracts specific keywords from the acquired text data, checks them against a predefined list of important keywords, identifies the keywords contained in the text data, and records their location information.
[0539] Step 6:
[0540] The server inputs the voice data into the emotion engine to analyze the emotional state of the user or speaker. The emotion engine analyzes the tone, rate, intensity, etc. of the voice to determine emotions (e.g., tension, excitement, satisfaction, etc.).
[0541] Step 7:
[0542] The server stores the converted text data, extracted keywords, and analyzed emotional states in a database, adding a current timestamp to each piece of data to maintain data integrity.
[0543] Step 8:
[0544] The server displays the stored data on a user interface, and the user accesses the system from a terminal and checks the text data, important keywords, and emotional state in real time through the interface.
[0545] Step 9:
[0546] Users can review the displayed data to understand the content of the meeting and the emotional state of the speakers, which allows them to take appropriate action and adjust resource allocation.
[0547] Step 10:
[0548] The user can take necessary measures and allocate resources to improve the efficiency of future meetings. The system records these interactions and uses them for future analysis and improvement.
[0549] Example 2
[0550] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0551] Conventional project management systems make it difficult to effectively grasp the content of meetings and discussions. In particular, they do not support extracting information from voice data or analyzing users' emotional states, making it difficult to grasp the psychological state and important points of team members in real time.
[0552] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0553] In this invention, the server includes means for receiving voice data, means for converting the received voice data into character data, means for extracting specific keywords from the converted character data, means for saving the extracted keywords and their location information in a database, means for analyzing the emotional state of a user from the voice data using an emotion engine, means for saving the analyzed emotional state in a database, and means for displaying the saved data on a user interface. This makes it possible to grasp in detail not only the content of meetings and discussions but also the emotional state of the speakers, thereby making project management more efficient.
[0554] "Audio data" refers to digital data of audio recorded during a meeting or conference.
[0555] The "receiving means" is a device or software for capturing voice data sent from a user into the system.
[0556] "Means for converting into text data" refers to a device or program that includes an algorithm or voice recognition technology for analyzing voice data and converting it into text format data.
[0557] A "keyword extraction means" is a device or software that detects and extracts specified important words or phrases from character data.
[0558] An "emotion engine" is an algorithm or software that estimates emotional states through the analysis of voice data.
[0559] The "means for storing in a database" is a storage system for organizing and recording the converted character data, extracted keywords, and analyzed emotional states.
[0560] The "means for displaying on the user interface" refers to a screen or software for displaying the saved data in a form that can be viewed by the user.
[0561] The present invention aims to improve the efficiency of project management by providing a system that receives voice data, converts it into text data, extracts keywords, analyzes emotional states, stores the data in a database, and displays it on a user interface.
[0562] Hardware and software used
[0563] The main components of the system are the server, the terminal, and the user. The specific hardware and software used are specified below:
[0564] 1. Hardware
[0565] Server: A powerful computer for data processing and storage
[0566] Terminal: The device used by the user to access the site (PC, tablet, smartphone, etc.)
[0567] Microphone: A device for recording audio data.
[0568] 2. Software
[0569] Speech recognition library: Google Speech Recognition API
[0570] Database management system: MySQL or PostgreSQL
[0571] Emotion engine: a program with a voice analysis algorithm
[0572] User Interface: Web applications implemented using HTML, CSS, and JavaScript
[0573] Program processing
[0574] Audio data reception
[0575] Users record audio from meetings and conferences and upload it to the server via their devices. Specifically, they click the "File Upload" button on the user interface and select an audio file from the file system. The server receives the audio file and saves it in the specified directory.
[0576] Converting voice data to text
[0577] The server reads the received voice data as audio data and initializes a voice recognition library such as the Google Speech Recognition API. The voice data is converted into text data using the voice recognition library, and the voice content is obtained in text format.
[0578] Keyword extraction
[0579] The server extracts specific keywords from the text data. Using a predefined list of important keywords (e.g., "important," "risk," "deadline"), it searches for the presence of these keywords in the text data. It records the keywords found and their location information in a list.
[0580] Emotional state analysis using emotion engine
[0581] The server feeds the recorded voice data to the emotion engine, which analyzes factors such as tone, speed, and strength of the voice. Based on the analysis results, the user's emotional state is identified and the results are stored in a database.
[0582] Saving to a database
[0583] The server stores the converted text data, extracted keywords and their location information, and analyzed emotional state in a database, where the data is recorded with the current timestamp to maintain data integrity.
[0584] Display in the user interface
[0585] The server displays the stored data on a user interface, and users can access the system from their devices and check the text data, important keywords, and emotional state in real time through the interface.
[0586] Specific examples
[0587] For example, the text data for a meeting might read, "We need to discuss risks and important deadlines for the next project," along with keywords like "risk," "important," and "deadline." Emotional states are also displayed as "tension" or "excitement." Based on this information, users and teams can take appropriate action and allocate resources.
[0588] Prompt Sentence Examples
[0589] Prompt: "Transcribe what is said in the next meeting, extract key keywords, and analyze the emotional state of the conversation. For example, display keywords such as risk, importance, and deadline, along with the associated emotions."
[0590] In this way, the system analyzes the content of meetings and discussions in detail, aiming to improve the efficiency of project management.
[0591] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0592] Step 1: Record and upload audio data
[0593] The user records the audio of a meeting or conference and selects the audio file from the device's file system. Next, the user uses the system's "file upload" function from the device to send the audio file to the server. The input is the audio file, and the output is an audio file saved in a specified directory on the server. Specifically, the user presses the record button, and then clicks the upload button after recording is complete.
[0594] Step 2: Receive and save the audio file
[0595] The server receives audio files sent through the user interface and saves them in the specified directory. The received audio files are recorded in a log, along with metadata to maintain consistency. The input is the uploaded audio file, and the output is a file saved on the server. Specifically, the server receives a send request and processes the saving of the audio file.
[0596] Step 3: Loading the audio data
[0597] The server reads the saved audio file as audio data. During this process, the file format and encoding are checked. The input is the saved audio file, and the output is the read audio data. Specifically, the server uses Python or JavaScript libraries to open the audio file.
[0598] Step 4: Initializing the speech recognition library
[0599] The server initializes a speech recognition library such as the Google Speech Recognition API, sets an API key, and runs the authentication process. The input is the API key, and the output is the initialized speech recognition library. Specifically, the server sets the API key and initializes the library.
[0600] Step 5: Converting audio data to text
[0601] The server uses a speech recognition library to convert voice data into text data. The input is the read audio data, and the output is the converted text data. Specifically, the server sends the voice data to the speech recognition API and stores the returned text data in memory.
[0602] Step 6: Keyword extraction
[0603] The server loads a predefined list of important keywords (e.g., "important," "risk," "deadline") and extracts these keywords from the text data. The input is the converted text data and a keyword list, and the output is the extracted keywords and their location information. Specifically, it performs regular expression and pattern matching on the text data.
[0604] Step 7: Emotional state analysis by the emotion engine
[0605] The server supplies the recorded voice data to the emotion engine, which analyzes the tone, speed, and strength of the voice. The input is the voice data, and the output is information about the user's emotional state. Specifically, the engine uses a voice analysis algorithm to estimate the emotional state.
[0606] Step 8: Saving to the Database
[0607] The server stores the converted text data, extracted keywords and their location information, and analyzed emotional state in a database. During this process, the current timestamp is also recorded. The input is the text data, keywords, location information, and emotional state, and the output is the complete dataset stored in the database. Specifically, it generates an SQL query and inserts it into the database.
[0608] Step 9: Display in the User Interface
[0609] The server retrieves the necessary data from the database and displays it on the user interface. The input is the information stored in the database, and the output is the displayed text data, keywords, and emotional state. Specifically, the interface is constructed using HTML and JavaScript, and the data is displayed dynamically.
[0610] In this way, the system can sequentially perform specific tasks at each processing step, ultimately providing the information the user needs.
[0611] (Application example 2)
[0612] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0613] In today's factories, meetings and communication between workers are not managed efficiently, and improvements are needed, particularly in quality control and work efficiency. There is a lack of means to grasp the situation in real time using voice data, and there is no system that analyzes the emotional state of workers and provides appropriate feedback. This makes it difficult to detect problems early and take appropriate measures.
[0614] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0615] In this invention, the server includes means for receiving voice data, means for converting the received voice data into character data, means for extracting specific keywords from the converted character data, means for saving the extracted keywords and their location information in a database, means for displaying the saved data on a user interface, means for analyzing an emotional state from the voice data, means for saving the analyzed emotional state in a database, and means for displaying the saved data and the emotional state on a user interface. This enables real-time monitoring and efficient management of meetings and work status in a factory, and realizes quality control and improved work efficiency based on the emotional state of workers.
[0616] 1. "Voice data" refers to data that includes voice information from meetings or communications between workers.
[0617] 2. "Text data" means text-format data converted from voice data using voice recognition technology.
[0618] 3. "Keywords" refer to words or phrases that are particularly important in a meeting or work content.
[0619] 4. "Location information" is information that indicates the location where a keyword appears within the converted character data.
[0620] 5. "Database" refers to a storage means for storing and managing converted text data, extracted keywords, location information, and emotional state.
[0621] 6. "User Interface" means the display means by which a user accesses the system and checks and manipulates data.
[0622] 7. "Emotional state" refers to the results of analyzing voice data to infer the user's emotions.
[0623] 8. "Means of analysis" includes processes and technologies for extracting and identifying text data and emotional states from audio data.
[0624] 9. "Real-time" means that data is processed and displayed with little or no delay after it is generated.
[0625] 10. "Quality control" refers to activities to maintain and improve the quality of products and work processes on the factory floor.
[0626] 11. "Work efficiency" is an indicator of how efficiently work is being carried out on the factory floor.
[0627] 12. "Monitoring" refers to the continuous observation and recording of meetings and work situations.
[0628] The present invention is a system for improving the efficiency of meetings and communication between workers at a factory. Specific examples are shown below.
[0629] First, microphones installed in the factory record audio data of meetings and between workers. Users manage this audio data on their devices and use the system's "audio file upload" function to send it to the server. The server then saves the uploaded audio data in a specified directory.
[0630] The server receives the stored voice data and converts it into text data using the Google Cloud Speech-to-Text API, which has the ability to convert voice to text with high accuracy.
[0631] Next, the server uses a predefined list of important keywords to extract specific keywords from the converted text data. It checks whether any of the keywords in this list are present in the text data and records their location information. For example, if keywords such as "quality," "risk," and "work" are present, it records the words and their locations in the list.
[0632] The server then analyzes the recorded voice data and identifies the user's emotional state using the Dialogflow API. The emotion engine estimates emotions by analyzing the tone, speed, and strength of the voice. As a result of the analysis, the emotional state of the meeting or worker communication is identified as "tension" or "attention," etc.
[0633] The server stores the converted text data, extracted keywords, their location information, and analyzed emotional state in a database. This data also includes a current timestamp to ensure data integrity.
[0634] Users can access the system from their devices and view text data, key keywords, and emotional states in real time through the interface. For example, if a meeting content reads, "We should strengthen quality inspections for the next production batch. Risks are increasing," the system will identify the keywords "quality" and "risk" along with the text data, and display the emotional states of the relevant parts of the conversation as "tension" or "caution."
[0635] This system not only improves the efficiency of in-plant meetings and communication between workers, but also incorporates emotional information to enable early detection of problems and appropriate countermeasures, allowing managers to focus on specific problems and take prompt action.
[0636] Example inputs to a generative AI model:
[0637] Project Management / Meeting Content:
[0638] "Quality inspections should be strengthened for the next production batch. The risk is increasing."
[0639] Keywords: "quality" and "risk"
[0640] Emotional state: "tension" "attention"
[0641] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0642] Step 1:
[0643] Microphones installed in factories record audio data during meetings and between workers. Users manage this audio data on their devices and send it to the server using the system's "audio file upload" function. The input data is an audio file, and the output data is an audio file saved on the server.
[0644] Step 2:
[0645] The server saves the uploaded audio data in the specified directory. The input data is the audio file received from the user, and the output data is the path to the audio file saved on the server. The specific operation is to specify the file save path and save the file in storage.
[0646] Step 3:
[0647] The server converts the saved audio data into text data using the Google Cloud Speech-to-Text API. The input data is the path to the saved audio file, and the output data is the converted text data. Specifically, the API is called to convert the audio file into text format.
[0648] Step 4:
[0649] The server extracts specific keywords from the converted text data. The input data is the text data and a predefined keyword list, and the output data is the extracted keywords and their location information. Specifically, the server searches for the locations where keywords appear in the text data and records the location information.
[0650] Step 5:
[0651] The server analyzes the recorded voice data and identifies the user's emotional state using the Dialogflow API. The input data is the voice data, and the output data is the analyzed emotional state. Specifically, the server passes the voice data to the Dialogflow API, which estimates the user's emotional state.
[0652] Step 6:
[0653] The server stores the converted text data, extracted keywords and their location information, and analyzed emotional states in a database. The input data is the text data, keyword information, and emotional states, and the output data is the information stored in the database. Specifically, the server inserts the data into the database in an appropriate format.
[0654] Step 7:
[0655] Users access the system from their devices and check text data, important keywords, and emotional states in real time through the interface. The input data is information retrieved from the database, and the output data is information displayed on the user interface. Specific operations include updating the UI to display the information retrieved from the database in real time.
[0656] These steps will enable efficient management of factory meetings and communication between workers, including emotional information, and enable early detection and countermeasures for problems.
[0657] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0658] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0659] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0660] [Third embodiment]
[0661] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0662] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[0663] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0664] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0665] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0666] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0667] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0668] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0669] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0670] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0671] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0672] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0673] The present invention provides a system that receives voice data, converts it into text data, extracts specific keywords, stores them in a database, and displays them on a user interface in order to improve the efficiency of project management.
[0674] 1. Acceptance of audio data
[0675] First, the user uploads the recording of the meeting or conference from their own device to the server. Specifically, they click the "File Upload" button on the user interface and select the desired audio file from the file system. The server receives the audio file and saves it in the specified directory.
[0676] 2. Converting voice data to text data
[0677] Next, the server reads the received voice data as audio data, initializes a speech recognition library, and converts the voice data into text using a speech recognition service such as the Google Speech Recognition API. During this process, the server works to convert the voice content into text data with high accuracy.
[0678] 3. Keyword extraction
[0679] To extract specific keywords from the converted text data, the server uses a predefined list of important keywords. It checks whether these keywords exist in the text data and records their location information. For example, if keywords such as "important," "risk," and "deadline" are included, it records the words and their locations in a list.
[0680] 4. Saving to the database
[0681] The extracted keywords and their location information are stored in a database by the server. To maintain data integrity, the data is recorded with a current timestamp. This makes it clear what was discussed and when.
[0682] 5. Display in the user interface
[0683] Finally, the server displays the saved data on a user interface. Through the interface, users can check the text data and important keywords extracted from the recording data of meetings and discussions. For example, if a user uploads an audio file called "project_meeting.wav," the server analyzes the audio data and displays the text data "We need to discuss the risks and important deadlines of the next project" along with the keywords "risk" and "deadline" to the user.
[0684] In this way, the user can efficiently grasp the contents of the meeting, quickly confirm important points, and take necessary measures.
[0685] The processing flow will be explained below.
[0686] Step 1:
[0687] A user records an audio file of a meeting or conference and uses the system's "audio file upload" function on their device. They click the "file upload" button on the user interface and select the desired audio file from the file system.
[0688] Step 2:
[0689] The device uploads the selected audio file to the server via an HTTP request, and the server receives the request and saves the audio file in the specified directory.
[0690] Step 3:
[0691] The server reads the saved audio file and captures the audio data as audio data. Specifically, it initializes the speech recognition library and loads the audio file as an audio object.
[0692] Step 4:
[0693] The server calls a speech recognition service such as the Google Speech Recognition API to convert the audio data into text data. If the conversion process is successful, the content of the audio data is obtained in text format.
[0694] Step 5:
[0695] The server extracts specific keywords from the acquired text data, checks them against a predefined list of important keywords, identifies the keywords contained in the text data, and records their location information.
[0696] Step 6:
[0697] The server stores the converted text data, extracted keywords, and their location information in a database. When storing the data, it records the data with the current timestamp to maintain data integrity.
[0698] Step 7:
[0699] The server displays the stored data on a user interface. Users access the system from their devices and can view text data and important keywords in real time through the interface. The display includes specific discussion content, risks, and important points.
[0700] Step 8:
[0701] Users can check the interface and quickly take necessary measures and adjust resource allocation. The system records user interactions, contributing to the efficiency of future meetings.
[0702] Example 1
[0703] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0704] There is a demand for efficient recording of meeting and conference content so that important points can be quickly grasped. However, conventional systems require manual conversion of audio data into text data and manual extraction of important keywords, which is extremely time-consuming. In addition, because data is also saved and displayed manually, there are issues with information leaks and inconsistency.
[0705] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0706] In this invention, the server includes means for receiving voice data, means for converting the received voice data into text data, means for extracting specific keywords from the converted text data, means for saving the extracted keywords and their location information in a database, means for displaying the saved data on a user interface, means for converting voice data into text data using a voice recognition service, means for loading a predefined keyword list to extract specific keywords, means for recording a current timestamp when saving the text data including the location information of the extracted keywords in the database, and means for uploading voice data and checking the results of the analysis process via the user interface. This makes it possible to automatically and efficiently record the contents of meetings and discussions and easily grasp important points.
[0707] "Audio data" refers to data in which the contents of a meeting, conference, etc. are recorded in audio format.
[0708] "Character data" is data in text format that has been converted from voice data using voice recognition technology.
[0709] "Keywords" refer to specific words or phrases that are considered important in the content of a meeting or discussion.
[0710] A "database" is a digital system for efficiently storing and managing information, and in this case, it plays a role in storing text data based on voice data and extracted keywords.
[0711] A "user interface" is the part of a system that contains the screens and controls that allow a user to interact with the system.
[0712] A "means for receiving" is a technique or mechanism for capturing audio data from a user.
[0713] "Means for converting" refers to the technology or method for converting voice data into text data.
[0714] "Extraction means" refers to techniques or methods for extracting specific keywords from the converted character data.
[0715] "Storage means" refers to the technology or method for storing the extracted keywords and their location information in a database.
[0716] "Display means" refers to the techniques and methods for presenting the stored data in a user interface so that the user can check it.
[0717] "Speech Recognition Service" means a cloud-based or local service for analyzing and converting voice data into text data.
[0718] A "keyword list" is a collection of predefined important keywords that are used in text analysis.
[0719] A "timestamp" is time information that indicates that data was saved or processed at a specific time.
[0720] "Upload" refers to the operation or process by which a user sends audio data from their device to a server.
[0721] "Analysis" refers to a series of processes that converts voice data into text data and then extracts keywords.
[0722] The present invention aims to improve the efficiency of project management by providing a system that receives voice data, converts it into text data, extracts specific keywords, stores them in a database, and displays them on a user interface. Specifically, the invention can be implemented in the following forms.
[0723] First, the user uploads the recording of a meeting or conference from their own device to the server. The user clicks the "File Upload" button on the user interface and selects the desired audio file from the file system. The uploaded audio file is sent from the device to the server, and the server receives it and saves it in the specified directory. At this time, the file's metadata (e.g., file name, upload date and time, user ID, etc.) is also saved.
[0724] Next, the server reads the received voice data. It treats the voice data as audio data and initializes a voice recognition library. The system uses a voice recognition service such as the Google Speech Recognition API to convert the voice data into text. Through the conversion process, the content of the voice data is obtained as text data. For example, the resulting text data is, "We need to discuss the risks and important deadlines for the next project."
[0725] The server then extracts specific keywords from the converted text data using a predefined list of important keywords (e.g., "important," "risk," and "deadline"). The server uses a text analysis algorithm to identify the presence of keywords in the converted text data and record their location (e.g., which sentence or word). For example, if the words "risk" and "deadline" appear on the third line of the text, the server will add their location to the list.
[0726] The server then saves the extracted keywords and their location information, along with the text data, to a database. When saving, the server also records the current timestamp. This ensures data integrity. The saved data is stored in the "meeting_transcripts" table.
[0727] Finally, the server displays the stored data on a user interface. The server retrieves the stored text data and keywords from the database, formats them, and displays them to the user. Through the user interface, the user can view the text data and important keywords extracted from the recording of a meeting or discussion. For example, if a user uploads an audio file called "project_meeting.wav," the server analyzes the audio data and displays the text data, "We need to discuss the risks and important deadlines for the next project," along with the keywords "risk" and "deadline" to the user.
[0728] Specific examples
[0729] A user uploads an audio file called "project_meeting.wav." The server then converts it into text using the Google Speech Recognition API and extracts the keywords "risk" and "deadline." The extracted information is stored in a database and the user interface displays the message, "We need to discuss the risks and important deadlines for the next project."
[0730] Example of input prompt for generative AI model
[0731] An example of a prompt is, "Please upload the recording of the project meeting, 'project_meeting.wav'. The contents of this audio file will be converted into text data, and important keywords will be extracted and displayed. The service used is the Google Speech Recognition API."
[0732] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0733] Step 1:
[0734] The user uploads the audio data.
[0735] A user accesses the user interface of the project management system using a terminal. He clicks the "File Upload" button on the user interface and selects the desired audio file from the file system. The input is the audio file selected by the user (e.g., project_meeting.wav). The terminal sends the audio file to the server, and the audio data is sent to the server as the output.
[0736] Step 2:
[0737] The server receives and stores the audio data.
[0738] The server receives audio data sent from the device. It saves the received audio data in a specified directory (e.g., / uploaded_audio). At this time, it also saves metadata such as the file name, upload date and time, and user ID. The input is the audio data sent by the user. The output is an audio file saved in a specified directory on the server.
[0739] Step 3:
[0740] The server initializes the speech recognition library and reads the audio data.
[0741] The server initializes a speech recognition library (e.g., Google Speech Recognition API). It then reads the saved audio data and loads it into memory as audio data. The input is the saved audio data file. The output is the audio data loaded into memory.
[0742] Step 4:
[0743] The server converts the audio data into text data.
[0744] The server converts the audio data into text using an initialized speech recognition library. The speech recognition library analyzes the audio data and outputs the speech in text format. The input is the audio data loaded in memory. The output is text (e.g., "We need to discuss the risks and important deadlines for the next project").
[0745] Step 5:
[0746] The server loads the keyword list and extracts the keywords.
[0747] The server loads a predefined list of important keywords (e.g., "important," "risk," "deadline"), then checks for the presence of keywords in the converted text data and records location information. The input is the converted text data and the keyword list. The output is the extracted keywords (e.g., "risk," "deadline") and their location information.
[0748] Step 6:
[0749] The server stores the extracted data in a database.
[0750] The server stores the extracted keywords and text data, including their location information, in a database. The stored data also includes the current timestamp. The input is text data with keywords and location information added. The output is the information stored in the database.
[0751] Step 7:
[0752] The server displays the stored data in a user interface.
[0753] The server retrieves the stored text data and keywords from the database, formats them, and displays them on the user interface. Through the user interface, users can check the text data and important keywords extracted from the recording data of meetings and conferences. The input is the text data and keyword information retrieved from the database. The output is the text data and keywords displayed on the user interface.
[0754] (Application example 1)
[0755] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0756] Currently, in factory meetings and discussions, creating minutes and extracting important keywords is often done manually, which is extremely time-consuming. There is also the risk that important points may be overlooked, creating challenges for efficient project management and quality control. Large factories, in particular, where many meetings are held, require efficient and accurate management of meeting content. Therefore, a system is needed to streamline in-factory meeting records and quickly grasp important information.
[0757] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0758] In this invention, the server includes means for receiving voice data, means for converting the received voice data into text data, means for extracting specific keywords from the converted text data, means for storing the extracted keywords and their location information in a database, means for displaying the stored data on a user interface, means for collecting voice data from multifunction devices installed in an industrial environment, and means for displaying the text data and its keywords on an industrial user interface, thereby making it possible to efficiently record meetings in a factory and quickly grasp important information.
[0759] "Voice data" means data in digital form that records what a user says.
[0760] "Character data" refers to audio data converted into text format.
[0761] A "keyword" is a word or phrase that has an important meaning within the text data.
[0762] "Position information" is information indicating the position where a keyword appears within character data.
[0763] A "database" is a digital storage system for storing text data, keywords, location information, etc.
[0764] A "user interface" is a collection of screens and input devices that allow a user to interact with a system.
[0765] A "multifunction device" is a device for voice collection and data processing installed in an industrial environment.
[0766] An "industrial user interface" is an interface that displays data in a format suitable for business processes within a factory.
[0767] The system according to the present invention is for efficiently managing voice data from meetings and conferences in industrial environments and quickly grasping important information. Specifically, it includes the following components:
[0768] Hardware and software usage examples
[0769] Hardware:
[0770] Smart speaker: A device for collecting recordings of meetings and discussions within the factory.
[0771] Conference recording device: A device that records speech during a conference and provides the data to a server.
[0772] Robot: A machine installed in a factory that supports voice data collection and data management (e.g., Pepper).
[0773] software:
[0774] Python: A programming language for implementing server-side programs.
[0775] Google Speech Recognition API: A speech recognition service for converting voice data into text data.
[0776] SQLite: A database system for storing converted character data and keywords.
[0777] Flask: A web application framework for building user interfaces and displaying data.
[0778] System operation explanation
[0779] 1. Audio data collection:
[0780] When users hold meetings or conferences, smart speakers, conference recording devices, and robots collect audio data, which is then stored digitally and uploaded to a server.
[0781] 2. Audio to text conversion:
[0782] The server uses a Python program to read the received audio data and converts it into text using the Google Speech Recognition API. Through this process, the meeting contents are recorded in text format.
[0783] 3. Keyword extraction:
[0784] Predefined important keywords (e.g., "production line," "defective product," "maintenance") are automatically extracted from the converted text data, and the position in the text data where the keywords appear is also recorded.
[0785] 4. Save to database:
[0786] The extracted keywords, their location information, and the converted text data are stored in an SQLite database. A timestamp is added to the data when it is saved to ensure data integrity.
[0787] 5. Display in the user interface:
[0788] Through a web interface built with Flask, users can view saved meeting data and keywords. The displayed content includes text data and important keywords, allowing users to efficiently understand the content of the meeting.
[0789] Specific examples
[0790] Recorded data from quality control meetings held within the factory is collected using a smart speaker, and important keywords such as "defective products" and "maintenance" are automatically extracted and displayed on the server's web interface.
[0791] Prompt Sentence Examples
[0792] "Convert the meeting recording data into text data, extract specific keywords (e.g., "production line," "defective product," "maintenance"), and store them in a database. Create a web application that displays this in a user interface."
[0793] This will make it possible to efficiently record meetings within the factory and quickly grasp important information.
[0794] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0795] Step 1:
[0796] When a user starts a meeting, audio data is collected by a smart speaker, a meeting recording device, or a robot.
[0797] Input: Audio data during the meeting
[0798] Output: Digitally saved audio data file
[0799] What it does: A user starts a meeting, and the device collects audio data in real time and stores it in its internal storage.
[0800] Step 2:
[0801] The collected voice data is uploaded to a server, and the user uses a terminal to send the voice data to the server through a dedicated user interface.
[0802] Input: Audio data file saved on the device
[0803] Output: Audio data file uploaded to the server
[0804] Specific operation: The user clicks the "File Upload" button on the user interface to send the audio data file to the server.
[0805] Step 3:
[0806] The server reads the received audio data using a Python program.
[0807] Input: Uploaded audio data
[0808] Output: Audio data object
[0809] Specific operation: The server saves the audio file in a specific directory and reads the file using a Python program.
[0810] Step 4:
[0811] The server uses the Google Speech Recognition API to convert the audio data into text data.
[0812] Input: Audio data object
[0813] Output: Converted character data (text)
[0814] Specific operation: The server calls the Google Speech Recognition API and converts the voice data into text data.
[0815] Step 5:
[0816] The server extracts specific keywords from the converted character data.
[0817] Input: Character data
[0818] Output: Extracted keywords and their locations
[0819] Specific operation: The server matches the character data with a predefined keyword list and records the keywords and their location information.
[0820] Step 6:
[0821] The server stores the extracted keywords, their location information, and the converted text data in an SQLite database.
[0822] Input: Keywords, location information, text data
[0823] Output: Data stored in the database
[0824] Specific operation: The server connects to an SQLite database and stores keywords, location information, and text data along with timestamps.
[0825] Step 7:
[0826] The server displays the stored data in a web interface built with Flask.
[0827] Input: Data stored in a database
[0828] Output: Data displayed in the user interface
[0829] Specific operation: The server generates a user interface through a Flask application, allowing users to view the meeting content and keywords.
[0830] This allows users to efficiently manage meeting records within the factory and quickly grasp important information.
[0831] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0832] This invention combines a system that receives voice data, converts it into text data, extracts specific keywords, stores them in a database, and displays them on a user interface in order to improve the efficiency of project management, with an emotion engine that recognizes the user's emotions. This makes it possible to analyze not only the content of meetings and discussions, but also the emotional states of team members, to provide an optimal working environment.
[0833] Audio data reception
[0834] First, the user records the audio data of a meeting or conference and uses the system's "audio file upload" function on their device. They click the "File Upload" button on the user interface and select the desired audio file from the file system. The server receives the audio file and saves it in the specified directory.
[0835] Converting voice data to text
[0836] Next, the server reads the received voice data as audio data, initializes the voice recognition library, and converts it into text data using the Google Speech Recognition API, etc. This allows the voice content to be obtained in text format.
[0837] Keyword extraction
[0838] To extract specific keywords from the converted text data, the server uses a predefined list of important keywords. It checks whether these keywords exist in the text data and records their location information. For example, if keywords such as "important," "risk," and "deadline" are included, it records the words and their locations in a list.
[0839] Emotional state analysis using emotion engine
[0840] The server then analyzes the recorded voice data and identifies the user's emotional state using an emotion engine, which uses voice analysis technology to infer emotions from the tone, speed, and strength of the voice.
[0841] Saving to a database
[0842] The server stores the converted text data, extracted keywords and their location information, as well as the analyzed emotional state in a database. When storing the data, it records the current timestamp to maintain data integrity.
[0843] Display in the user interface
[0844] Finally, the server displays the saved data on a user interface. Users access the system from their devices and view text data, important keywords, and emotional states in real time through the interface. For example, the text data for a meeting might read, "We need to discuss the risks and important deadlines for the next project," along with the keywords "risk," "importance," and "deadline." Furthermore, the emotional states of the participants at the relevant points in the conversation might be displayed as "tense" or "excited."
[0845] In this way, users and teams can not only understand the content of the meeting, but also take into account emotional information to make appropriate responses and allocate resources.The system also records user interactions, contributing to the efficiency of future meetings.
[0846] The processing flow will be explained below.
[0847] Step 1:
[0848] A user records an audio file of a meeting or conference and uses the system's "audio file upload" function on their device. They click the "file upload" button on the user interface and select the desired audio file from the file system.
[0849] Step 2:
[0850] The device uploads the selected audio file to the server via an HTTP request, and the server receives the request and saves the audio file in the specified directory.
[0851] Step 3:
[0852] The server reads the saved audio file and captures the audio data as audio data. Specifically, it initializes the speech recognition library and loads the audio file as an audio object.
[0853] Step 4:
[0854] The server calls a speech recognition service such as the Google Speech Recognition API to convert the audio data into text data. If the conversion process is successful, the content of the audio data is obtained in text format.
[0855] Step 5:
[0856] The server extracts specific keywords from the acquired text data, checks them against a predefined list of important keywords, identifies the keywords contained in the text data, and records their location information.
[0857] Step 6:
[0858] The server inputs the voice data into the emotion engine to analyze the emotional state of the user or speaker. The emotion engine analyzes the tone, rate, intensity, etc. of the voice to determine emotions (e.g., tension, excitement, satisfaction, etc.).
[0859] Step 7:
[0860] The server stores the converted text data, extracted keywords, and analyzed emotional states in a database, adding a current timestamp to each piece of data to maintain data integrity.
[0861] Step 8:
[0862] The server displays the stored data on a user interface, and the user accesses the system from a terminal and checks the text data, important keywords, and emotional state in real time through the interface.
[0863] Step 9:
[0864] Users can review the displayed data to understand the content of the meeting and the emotional state of the speakers, which allows them to take appropriate action and adjust resource allocation.
[0865] Step 10:
[0866] The user can take necessary measures and allocate resources to improve the efficiency of future meetings. The system records these interactions and uses them for future analysis and improvement.
[0867] Example 2
[0868] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0869] Conventional project management systems make it difficult to effectively grasp the content of meetings and discussions. In particular, they do not support extracting information from voice data or analyzing users' emotional states, making it difficult to grasp the psychological state and important points of team members in real time.
[0870] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0871] In this invention, the server includes means for receiving voice data, means for converting the received voice data into character data, means for extracting specific keywords from the converted character data, means for saving the extracted keywords and their location information in a database, means for analyzing the emotional state of a user from the voice data using an emotion engine, means for saving the analyzed emotional state in a database, and means for displaying the saved data on a user interface. This makes it possible to grasp in detail not only the content of meetings and discussions but also the emotional state of the speakers, thereby making project management more efficient.
[0872] "Audio data" refers to digital data of audio recorded during a meeting or conference.
[0873] The "receiving means" is a device or software for capturing voice data sent from a user into the system.
[0874] "Means for converting into text data" refers to a device or program that includes an algorithm or voice recognition technology for analyzing voice data and converting it into text format data.
[0875] A "keyword extraction means" is a device or software that detects and extracts specified important words or phrases from character data.
[0876] An "emotion engine" is an algorithm or software that estimates emotional states through the analysis of voice data.
[0877] The "means for storing in a database" is a storage system for organizing and recording the converted character data, extracted keywords, and analyzed emotional states.
[0878] The "means for displaying on the user interface" refers to a screen or software for displaying the saved data in a form that can be viewed by the user.
[0879] The present invention aims to improve the efficiency of project management by providing a system that receives voice data, converts it into text data, extracts keywords, analyzes emotional states, stores the data in a database, and displays it on a user interface.
[0880] Hardware and software used
[0881] The main components of the system are the server, the terminal, and the user. The specific hardware and software used are specified below:
[0882] 1. Hardware
[0883] Server: A powerful computer for data processing and storage
[0884] Terminal: The device used by the user to access the site (PC, tablet, smartphone, etc.)
[0885] Microphone: A device for recording audio data.
[0886] 2. Software
[0887] Speech recognition library: Google Speech Recognition API
[0888] Database management system: MySQL or PostgreSQL
[0889] Emotion engine: a program with a voice analysis algorithm
[0890] User Interface: Web applications implemented using HTML, CSS, and JavaScript
[0891] Program processing
[0892] Audio data reception
[0893] Users record audio from meetings and conferences and upload it to the server via their devices. Specifically, they click the "File Upload" button on the user interface and select an audio file from the file system. The server receives the audio file and saves it in the specified directory.
[0894] Converting voice data to text
[0895] The server reads the received voice data as audio data and initializes a voice recognition library such as the Google Speech Recognition API. The voice data is converted into text data using the voice recognition library, and the voice content is obtained in text format.
[0896] Keyword extraction
[0897] The server extracts specific keywords from the text data. Using a predefined list of important keywords (e.g., "important," "risk," "deadline"), it searches for the presence of these keywords in the text data. It records the keywords found and their location information in a list.
[0898] Emotional state analysis using emotion engine
[0899] The server feeds the recorded voice data to the emotion engine, which analyzes factors such as tone, speed, and strength of the voice. Based on the analysis results, the user's emotional state is identified and the results are stored in a database.
[0900] Saving to a database
[0901] The server stores the converted text data, extracted keywords and their location information, and analyzed emotional state in a database, where the data is recorded with the current timestamp to maintain data integrity.
[0902] Display in the user interface
[0903] The server displays the stored data on a user interface, and users can access the system from their devices and check the text data, important keywords, and emotional state in real time through the interface.
[0904] Specific examples
[0905] For example, the text data for a meeting might read, "We need to discuss risks and important deadlines for the next project," along with keywords like "risk," "important," and "deadline." Emotional states are also displayed as "tension" or "excitement." Based on this information, users and teams can take appropriate action and allocate resources.
[0906] Prompt Sentence Examples
[0907] Prompt: "Transcribe what is said in the next meeting, extract key keywords, and analyze the emotional state of the conversation. For example, display keywords such as risk, importance, and deadline, along with the associated emotions."
[0908] In this way, the system analyzes the content of meetings and discussions in detail, aiming to improve the efficiency of project management.
[0909] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0910] Step 1: Record and upload audio data
[0911] The user records the audio of a meeting or conference and selects the audio file from the device's file system. Next, the user uses the system's "file upload" function from the device to send the audio file to the server. The input is the audio file, and the output is an audio file saved in a specified directory on the server. Specifically, the user presses the record button, and then clicks the upload button after recording is complete.
[0912] Step 2: Receive and save the audio file
[0913] The server receives audio files sent through the user interface and saves them in the specified directory. The received audio files are recorded in a log, along with metadata to maintain consistency. The input is the uploaded audio file, and the output is a file saved on the server. Specifically, the server receives a send request and processes the saving of the audio file.
[0914] Step 3: Loading the audio data
[0915] The server reads the saved audio file as audio data. During this process, the file format and encoding are checked. The input is the saved audio file, and the output is the read audio data. Specifically, the server uses Python or JavaScript libraries to open the audio file.
[0916] Step 4: Initializing the speech recognition library
[0917] The server initializes a speech recognition library such as the Google Speech Recognition API, sets an API key, and runs the authentication process. The input is the API key, and the output is the initialized speech recognition library. Specifically, the server sets the API key and initializes the library.
[0918] Step 5: Converting audio data to text
[0919] The server uses a speech recognition library to convert voice data into text data. The input is the read audio data, and the output is the converted text data. Specifically, the server sends the voice data to the speech recognition API and stores the returned text data in memory.
[0920] Step 6: Keyword extraction
[0921] The server loads a predefined list of important keywords (e.g., "important," "risk," "deadline") and extracts these keywords from the text data. The input is the converted text data and a keyword list, and the output is the extracted keywords and their location information. Specifically, it performs regular expression and pattern matching on the text data.
[0922] Step 7: Emotional state analysis by the emotion engine
[0923] The server supplies the recorded voice data to the emotion engine, which analyzes the tone, speed, and strength of the voice. The input is the voice data, and the output is information about the user's emotional state. Specifically, the engine uses a voice analysis algorithm to estimate the emotional state.
[0924] Step 8: Saving to the Database
[0925] The server stores the converted text data, extracted keywords and their location information, and analyzed emotional state in a database. During this process, the current timestamp is also recorded. The input is the text data, keywords, location information, and emotional state, and the output is the complete dataset stored in the database. Specifically, it generates an SQL query and inserts it into the database.
[0926] Step 9: Display in the User Interface
[0927] The server retrieves the necessary data from the database and displays it on the user interface. The input is the information stored in the database, and the output is the displayed text data, keywords, and emotional state. Specifically, the interface is constructed using HTML and JavaScript, and the data is displayed dynamically.
[0928] In this way, the system can sequentially perform specific tasks at each processing step, ultimately providing the information the user needs.
[0929] (Application example 2)
[0930] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0931] In today's factories, meetings and communication between workers are not managed efficiently, and improvements are needed, particularly in quality control and work efficiency. There is a lack of means to grasp the situation in real time using voice data, and there is no system that analyzes the emotional state of workers and provides appropriate feedback. This makes it difficult to detect problems early and take appropriate measures.
[0932] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0933] In this invention, the server includes means for receiving voice data, means for converting the received voice data into character data, means for extracting specific keywords from the converted character data, means for saving the extracted keywords and their location information in a database, means for displaying the saved data on a user interface, means for analyzing an emotional state from the voice data, means for saving the analyzed emotional state in a database, and means for displaying the saved data and the emotional state on a user interface. This enables real-time monitoring and efficient management of meetings and work status in a factory, and realizes quality control and improved work efficiency based on the emotional state of workers.
[0934] 1. "Voice data" refers to data that includes voice information from meetings or communications between workers.
[0935] 2. "Text data" means text-format data converted from voice data using voice recognition technology.
[0936] 3. "Keywords" refer to words or phrases that are particularly important in a meeting or work content.
[0937] 4. "Location information" is information that indicates the location where a keyword appears within the converted character data.
[0938] 5. "Database" refers to a storage means for storing and managing converted text data, extracted keywords, location information, and emotional state.
[0939] 6. "User Interface" means the display means by which a user accesses the system and checks and manipulates data.
[0940] 7. "Emotional state" refers to the results of analyzing voice data to infer the user's emotions.
[0941] 8. "Means of analysis" includes processes and technologies for extracting and identifying text data and emotional states from audio data.
[0942] 9. "Real-time" means that data is processed and displayed with little or no delay after it is generated.
[0943] 10. "Quality control" refers to activities to maintain and improve the quality of products and work processes on the factory floor.
[0944] 11. "Work efficiency" is an indicator of how efficiently work is being carried out on the factory floor.
[0945] 12. "Monitoring" refers to the continuous observation and recording of meetings and work situations.
[0946] The present invention is a system for improving the efficiency of meetings and communication between workers at a factory. Specific examples are shown below.
[0947] First, microphones installed in the factory record audio data of meetings and between workers. Users manage this audio data on their devices and use the system's "audio file upload" function to send it to the server. The server then saves the uploaded audio data in a specified directory.
[0948] The server receives the stored voice data and converts it into text data using the Google Cloud Speech-to-Text API, which has the ability to convert voice to text with high accuracy.
[0949] Next, the server uses a predefined list of important keywords to extract specific keywords from the converted text data. It checks whether any of the keywords in this list are present in the text data and records their location information. For example, if keywords such as "quality," "risk," and "work" are present, it records the words and their locations in the list.
[0950] The server then analyzes the recorded voice data and identifies the user's emotional state using the Dialogflow API. The emotion engine estimates emotions by analyzing the tone, speed, and strength of the voice. As a result of the analysis, the emotional state of the meeting or worker communication is identified as "tension" or "attention," etc.
[0951] The server stores the converted text data, extracted keywords, their location information, and analyzed emotional state in a database. This data also includes a current timestamp to ensure data integrity.
[0952] Users can access the system from their devices and view text data, key keywords, and emotional states in real time through the interface. For example, if a meeting content reads, "We should strengthen quality inspections for the next production batch. Risks are increasing," the system will identify the keywords "quality" and "risk" along with the text data, and display the emotional states of the relevant parts of the conversation as "tension" or "caution."
[0953] This system not only improves the efficiency of in-plant meetings and communication between workers, but also incorporates emotional information to enable early detection of problems and appropriate countermeasures, allowing managers to focus on specific problems and take prompt action.
[0954] Example inputs to a generative AI model:
[0955] Project Management / Meeting Content:
[0956] "Quality inspections should be strengthened for the next production batch. The risk is increasing."
[0957] Keywords: "quality" and "risk"
[0958] Emotional state: "tension" "attention"
[0959] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0960] Step 1:
[0961] Microphones installed in factories record audio data during meetings and between workers. Users manage this audio data on their devices and send it to the server using the system's "audio file upload" function. The input data is an audio file, and the output data is an audio file saved on the server.
[0962] Step 2:
[0963] The server saves the uploaded audio data in the specified directory. The input data is the audio file received from the user, and the output data is the path to the audio file saved on the server. The specific operation is to specify the file save path and save the file in storage.
[0964] Step 3:
[0965] The server converts the saved audio data into text data using the Google Cloud Speech-to-Text API. The input data is the path to the saved audio file, and the output data is the converted text data. Specifically, the API is called to convert the audio file into text format.
[0966] Step 4:
[0967] The server extracts specific keywords from the converted text data. The input data is the text data and a predefined keyword list, and the output data is the extracted keywords and their location information. Specifically, the server searches for the locations where keywords appear in the text data and records the location information.
[0968] Step 5:
[0969] The server analyzes the recorded voice data and identifies the user's emotional state using the Dialogflow API. The input data is the voice data, and the output data is the analyzed emotional state. Specifically, the server passes the voice data to the Dialogflow API, which estimates the user's emotional state.
[0970] Step 6:
[0971] The server stores the converted text data, extracted keywords and their location information, and analyzed emotional states in a database. The input data is the text data, keyword information, and emotional states, and the output data is the information stored in the database. Specifically, the server inserts the data into the database in an appropriate format.
[0972] Step 7:
[0973] Users access the system from their devices and check text data, important keywords, and emotional states in real time through the interface. The input data is information retrieved from the database, and the output data is information displayed on the user interface. Specific operations include updating the UI to display the information retrieved from the database in real time.
[0974] These steps will enable efficient management of factory meetings and communication between workers, including emotional information, and enable early detection and countermeasures for problems.
[0975] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0976] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0977] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[0978] [Fourth embodiment]
[0979] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[0980] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0981] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0982] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[0983] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0984] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0985] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0986] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[0987] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0988] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0989] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0990] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0991] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[0992] The present invention provides a system that receives voice data, converts it into text data, extracts specific keywords, stores them in a database, and displays them on a user interface in order to improve the efficiency of project management.
[0993] 1. Acceptance of audio data
[0994] First, the user uploads the recording of the meeting or conference from their own device to the server. Specifically, they click the "File Upload" button on the user interface and select the desired audio file from the file system. The server receives the audio file and saves it in the specified directory.
[0995] 2. Converting voice data to text data
[0996] Next, the server reads the received voice data as audio data, initializes a speech recognition library, and converts the voice data into text using a speech recognition service such as the Google Speech Recognition API. During this process, the server works to convert the voice content into text data with high accuracy.
[0997] 3. Keyword extraction
[0998] To extract specific keywords from the converted text data, the server uses a predefined list of important keywords. It checks whether these keywords exist in the text data and records their location information. For example, if keywords such as "important," "risk," and "deadline" are included, it records the words and their locations in a list.
[0999] 4. Saving to the database
[1000] The extracted keywords and their location information are stored in a database by the server. To maintain data integrity, the data is recorded with a current timestamp. This makes it clear what was discussed and when.
[1001] 5. Display in the user interface
[1002] Finally, the server displays the saved data on a user interface. Through the interface, users can check the text data and important keywords extracted from the recording data of meetings and discussions. For example, if a user uploads an audio file called "project_meeting.wav," the server analyzes the audio data and displays the text data "We need to discuss the risks and important deadlines of the next project" along with the keywords "risk" and "deadline" to the user.
[1003] In this way, the user can efficiently grasp the contents of the meeting, quickly confirm important points, and take necessary measures.
[1004] The processing flow will be explained below.
[1005] Step 1:
[1006] A user records an audio file of a meeting or conference and uses the system's "audio file upload" function on their device. They click the "file upload" button on the user interface and select the desired audio file from the file system.
[1007] Step 2:
[1008] The device uploads the selected audio file to the server via an HTTP request, and the server receives the request and saves the audio file in the specified directory.
[1009] Step 3:
[1010] The server reads the saved audio file and captures the audio data as audio data. Specifically, it initializes the speech recognition library and loads the audio file as an audio object.
[1011] Step 4:
[1012] The server calls a speech recognition service such as the Google Speech Recognition API to convert the audio data into text data. If the conversion process is successful, the content of the audio data is obtained in text format.
[1013] Step 5:
[1014] The server extracts specific keywords from the acquired text data, checks them against a predefined list of important keywords, identifies the keywords contained in the text data, and records their location information.
[1015] Step 6:
[1016] The server stores the converted text data, extracted keywords, and their location information in a database. When storing the data, it records the data with the current timestamp to maintain data integrity.
[1017] Step 7:
[1018] The server displays the stored data on a user interface. Users access the system from their devices and can view text data and important keywords in real time through the interface. The display includes specific discussion content, risks, and important points.
[1019] Step 8:
[1020] Users can check the interface and quickly take necessary measures and adjust resource allocation. The system records user interactions, contributing to the efficiency of future meetings.
[1021] Example 1
[1022] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1023] There is a demand for efficient recording of meeting and conference content so that important points can be quickly grasped. However, conventional systems require manual conversion of audio data into text data and manual extraction of important keywords, which is extremely time-consuming. In addition, because data is also saved and displayed manually, there are issues with information leaks and inconsistency.
[1024] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1025] In this invention, the server includes means for receiving voice data, means for converting the received voice data into text data, means for extracting specific keywords from the converted text data, means for saving the extracted keywords and their location information in a database, means for displaying the saved data on a user interface, means for converting voice data into text data using a voice recognition service, means for loading a predefined keyword list to extract specific keywords, means for recording a current timestamp when saving the text data including the location information of the extracted keywords in the database, and means for uploading voice data and checking the results of the analysis process via the user interface. This makes it possible to automatically and efficiently record the contents of meetings and discussions and easily grasp important points.
[1026] "Audio data" refers to data in which the contents of a meeting, conference, etc. are recorded in audio format.
[1027] "Character data" is data in text format that has been converted from voice data using voice recognition technology.
[1028] "Keywords" refer to specific words or phrases that are considered important in the content of a meeting or discussion.
[1029] A "database" is a digital system for efficiently storing and managing information, and in this case, it plays a role in storing text data based on voice data and extracted keywords.
[1030] A "user interface" is the part of a system that contains the screens and controls that allow a user to interact with the system.
[1031] A "means for receiving" is a technique or mechanism for capturing audio data from a user.
[1032] "Means for converting" refers to the technology or method for converting voice data into text data.
[1033] "Extraction means" refers to techniques or methods for extracting specific keywords from the converted character data.
[1034] "Storage means" refers to the technology or method for storing the extracted keywords and their location information in a database.
[1035] "Display means" refers to the techniques and methods for presenting the stored data in a user interface so that the user can check it.
[1036] "Speech Recognition Service" means a cloud-based or local service for analyzing and converting voice data into text data.
[1037] A "keyword list" is a collection of predefined important keywords that are used in text analysis.
[1038] A "timestamp" is time information that indicates that data was saved or processed at a specific time.
[1039] "Upload" refers to the operation or process by which a user sends audio data from their device to a server.
[1040] "Analysis" refers to a series of processes that converts voice data into text data and then extracts keywords.
[1041] The present invention aims to improve the efficiency of project management by providing a system that receives voice data, converts it into text data, extracts specific keywords, stores them in a database, and displays them on a user interface. Specifically, the invention can be implemented in the following forms.
[1042] First, the user uploads the recording of a meeting or conference from their own device to the server. The user clicks the "File Upload" button on the user interface and selects the desired audio file from the file system. The uploaded audio file is sent from the device to the server, and the server receives it and saves it in the specified directory. At this time, the file's metadata (e.g., file name, upload date and time, user ID, etc.) is also saved.
[1043] Next, the server reads the received voice data. It treats the voice data as audio data and initializes a voice recognition library. The system uses a voice recognition service such as the Google Speech Recognition API to convert the voice data into text. Through the conversion process, the content of the voice data is obtained as text data. For example, the resulting text data is, "We need to discuss the risks and important deadlines for the next project."
[1044] The server then extracts specific keywords from the converted text data using a predefined list of important keywords (e.g., "important," "risk," and "deadline"). The server uses a text analysis algorithm to identify the presence of keywords in the converted text data and record their location (e.g., which sentence or word). For example, if the words "risk" and "deadline" appear on the third line of the text, the server will add their location to the list.
[1045] The server then saves the extracted keywords and their location information, along with the text data, to a database. When saving, the server also records the current timestamp. This ensures data integrity. The saved data is stored in the "meeting_transcripts" table.
[1046] Finally, the server displays the stored data on a user interface. The server retrieves the stored text data and keywords from the database, formats them, and displays them to the user. Through the user interface, the user can view the text data and important keywords extracted from the recording of a meeting or discussion. For example, if a user uploads an audio file called "project_meeting.wav," the server analyzes the audio data and displays the text data, "We need to discuss the risks and important deadlines for the next project," along with the keywords "risk" and "deadline" to the user.
[1047] Specific examples
[1048] A user uploads an audio file called "project_meeting.wav." The server then converts it into text using the Google Speech Recognition API and extracts the keywords "risk" and "deadline." The extracted information is stored in a database and the user interface displays the message, "We need to discuss the risks and important deadlines for the next project."
[1049] Example of input prompt for generative AI model
[1050] An example of a prompt is, "Please upload the recording of the project meeting, 'project_meeting.wav'. The contents of this audio file will be converted into text data, and important keywords will be extracted and displayed. The service used is the Google Speech Recognition API."
[1051] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1052] Step 1:
[1053] The user uploads the audio data.
[1054] A user accesses the user interface of the project management system using a terminal. He clicks the "File Upload" button on the user interface and selects the desired audio file from the file system. The input is the audio file selected by the user (e.g., project_meeting.wav). The terminal sends the audio file to the server, and the audio data is sent to the server as the output.
[1055] Step 2:
[1056] The server receives and stores the audio data.
[1057] The server receives audio data sent from the device. It saves the received audio data in a specified directory (e.g., / uploaded_audio). At this time, it also saves metadata such as the file name, upload date and time, and user ID. The input is the audio data sent by the user. The output is an audio file saved in a specified directory on the server.
[1058] Step 3:
[1059] The server initializes the speech recognition library and reads the audio data.
[1060] The server initializes a speech recognition library (e.g., Google Speech Recognition API). It then reads the saved audio data and loads it into memory as audio data. The input is the saved audio data file. The output is the audio data loaded into memory.
[1061] Step 4:
[1062] The server converts the audio data into text data.
[1063] The server converts the audio data into text using an initialized speech recognition library. The speech recognition library analyzes the audio data and outputs the speech in text format. The input is the audio data loaded in memory. The output is text (e.g., "We need to discuss the risks and important deadlines for the next project").
[1064] Step 5:
[1065] The server loads the keyword list and extracts the keywords.
[1066] The server loads a predefined list of important keywords (e.g., "important," "risk," "deadline"), then checks for the presence of keywords in the converted text data and records location information. The input is the converted text data and the keyword list. The output is the extracted keywords (e.g., "risk," "deadline") and their location information.
[1067] Step 6:
[1068] The server stores the extracted data in a database.
[1069] The server stores the extracted keywords and text data, including their location information, in a database. The stored data also includes the current timestamp. The input is text data with keywords and location information added. The output is the information stored in the database.
[1070] Step 7:
[1071] The server displays the stored data in a user interface.
[1072] The server retrieves the stored text data and keywords from the database, formats them, and displays them on the user interface. Through the user interface, users can check the text data and important keywords extracted from the recording data of meetings and conferences. The input is the text data and keyword information retrieved from the database. The output is the text data and keywords displayed on the user interface.
[1073] (Application example 1)
[1074] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1075] Currently, in factory meetings and discussions, creating minutes and extracting important keywords is often done manually, which is extremely time-consuming. There is also the risk that important points may be overlooked, creating challenges for efficient project management and quality control. Large factories, in particular, where many meetings are held, require efficient and accurate management of meeting content. Therefore, a system is needed to streamline in-factory meeting records and quickly grasp important information.
[1076] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1077] In this invention, the server includes means for receiving voice data, means for converting the received voice data into text data, means for extracting specific keywords from the converted text data, means for storing the extracted keywords and their location information in a database, means for displaying the stored data on a user interface, means for collecting voice data from multifunction devices installed in an industrial environment, and means for displaying the text data and its keywords on an industrial user interface, thereby making it possible to efficiently record meetings in a factory and quickly grasp important information.
[1078] "Voice data" means data in digital form that records what a user says.
[1079] "Character data" refers to audio data converted into text format.
[1080] A "keyword" is a word or phrase that has an important meaning within the text data.
[1081] "Position information" is information indicating the position where a keyword appears within character data.
[1082] A "database" is a digital storage system for storing text data, keywords, location information, etc.
[1083] A "user interface" is a collection of screens and input devices that allow a user to interact with a system.
[1084] A "multifunction device" is a device for voice collection and data processing installed in an industrial environment.
[1085] An "industrial user interface" is an interface that displays data in a format suitable for business processes within a factory.
[1086] The system according to the present invention is for efficiently managing voice data from meetings and conferences in industrial environments and quickly grasping important information. Specifically, it includes the following components:
[1087] Hardware and software usage examples
[1088] Hardware:
[1089] Smart speaker: A device for collecting recordings of meetings and discussions within the factory.
[1090] Conference recording device: A device that records speech during a conference and provides the data to a server.
[1091] Robot: A machine installed in a factory that supports voice data collection and data management (e.g., Pepper).
[1092] software:
[1093] Python: A programming language for implementing server-side programs.
[1094] Google Speech Recognition API: A speech recognition service for converting voice data into text data.
[1095] SQLite: A database system for storing converted character data and keywords.
[1096] Flask: A web application framework for building user interfaces and displaying data.
[1097] System operation explanation
[1098] 1. Audio data collection:
[1099] When users hold meetings or conferences, smart speakers, conference recording devices, and robots collect audio data, which is then stored digitally and uploaded to a server.
[1100] 2. Audio to text conversion:
[1101] The server uses a Python program to read the received audio data and converts it into text using the Google Speech Recognition API. Through this process, the meeting contents are recorded in text format.
[1102] 3. Keyword extraction:
[1103] Predefined important keywords (e.g., "production line," "defective product," "maintenance") are automatically extracted from the converted text data, and the position in the text data where the keywords appear is also recorded.
[1104] 4. Save to database:
[1105] The extracted keywords, their location information, and the converted text data are stored in an SQLite database. A timestamp is added to the data when it is saved to ensure data integrity.
[1106] 5. Display in the user interface:
[1107] Through a web interface built with Flask, users can view saved meeting data and keywords. The displayed content includes text data and important keywords, allowing users to efficiently understand the content of the meeting.
[1108] Specific examples
[1109] Recorded data from quality control meetings held within the factory is collected using a smart speaker, and important keywords such as "defective products" and "maintenance" are automatically extracted and displayed on the server's web interface.
[1110] Prompt Sentence Examples
[1111] "Convert the meeting recording data into text data, extract specific keywords (e.g., "production line," "defective product," "maintenance"), and store them in a database. Create a web application that displays this in a user interface."
[1112] This will make it possible to efficiently record meetings within the factory and quickly grasp important information.
[1113] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1114] Step 1:
[1115] When a user starts a meeting, audio data is collected by a smart speaker, a meeting recording device, or a robot.
[1116] Input: Audio data during the meeting
[1117] Output: Digitally saved audio data file
[1118] What it does: A user starts a meeting, and the device collects audio data in real time and stores it in its internal storage.
[1119] Step 2:
[1120] The collected voice data is uploaded to a server, and the user uses a terminal to send the voice data to the server through a dedicated user interface.
[1121] Input: Audio data file saved on the device
[1122] Output: Audio data file uploaded to the server
[1123] Specific operation: The user clicks the "File Upload" button on the user interface to send the audio data file to the server.
[1124] Step 3:
[1125] The server reads the received audio data using a Python program.
[1126] Input: Uploaded audio data
[1127] Output: Audio data object
[1128] Specific operation: The server saves the audio file in a specific directory and reads the file using a Python program.
[1129] Step 4:
[1130] The server uses the Google Speech Recognition API to convert the audio data into text data.
[1131] Input: Audio data object
[1132] Output: Converted character data (text)
[1133] Specific operation: The server calls the Google Speech Recognition API and converts the voice data into text data.
[1134] Step 5:
[1135] The server extracts specific keywords from the converted character data.
[1136] Input: Character data
[1137] Output: Extracted keywords and their locations
[1138] Specific operation: The server matches the character data with a predefined keyword list and records the keywords and their location information.
[1139] Step 6:
[1140] The server stores the extracted keywords, their location information, and the converted text data in an SQLite database.
[1141] Input: Keywords, location information, text data
[1142] Output: Data stored in the database
[1143] Specific operation: The server connects to an SQLite database and stores keywords, location information, and text data along with timestamps.
[1144] Step 7:
[1145] The server displays the stored data in a web interface built with Flask.
[1146] Input: Data stored in a database
[1147] Output: Data displayed in the user interface
[1148] Specific operation: The server generates a user interface through a Flask application, allowing users to view the meeting content and keywords.
[1149] This allows users to efficiently manage meeting records within the factory and quickly grasp important information.
[1150] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1151] This invention combines a system that receives voice data, converts it into text data, extracts specific keywords, stores them in a database, and displays them on a user interface in order to improve the efficiency of project management, with an emotion engine that recognizes the user's emotions. This makes it possible to analyze not only the content of meetings and discussions, but also the emotional states of team members, to provide an optimal working environment.
[1152] Audio data reception
[1153] First, the user records the audio data of a meeting or conference and uses the system's "audio file upload" function on their device. They click the "File Upload" button on the user interface and select the desired audio file from the file system. The server receives the audio file and saves it in the specified directory.
[1154] Converting voice data to text
[1155] Next, the server reads the received voice data as audio data, initializes the voice recognition library, and converts it into text data using the Google Speech Recognition API, etc. This allows the voice content to be obtained in text format.
[1156] Keyword extraction
[1157] To extract specific keywords from the converted text data, the server uses a predefined list of important keywords. It checks whether these keywords exist in the text data and records their location information. For example, if keywords such as "important," "risk," and "deadline" are included, it records the words and their locations in a list.
[1158] Emotional state analysis using emotion engine
[1159] The server then analyzes the recorded voice data and identifies the user's emotional state using an emotion engine, which uses voice analysis technology to infer emotions from the tone, speed, and strength of the voice.
[1160] Saving to a database
[1161] The server stores the converted text data, extracted keywords and their location information, as well as the analyzed emotional state in a database. When storing the data, it records the current timestamp to maintain data integrity.
[1162] Display in the user interface
[1163] Finally, the server displays the saved data on a user interface. Users access the system from their devices and view text data, important keywords, and emotional states in real time through the interface. For example, the text data for a meeting might read, "We need to discuss the risks and important deadlines for the next project," along with the keywords "risk," "importance," and "deadline." Furthermore, the emotional states of the participants at the relevant points in the conversation might be displayed as "tense" or "excited."
[1164] In this way, users and teams can not only understand the content of the meeting, but also take into account emotional information to make appropriate responses and allocate resources.The system also records user interactions, contributing to the efficiency of future meetings.
[1165] The processing flow will be explained below.
[1166] Step 1:
[1167] A user records an audio file of a meeting or conference and uses the system's "audio file upload" function on their device. They click the "file upload" button on the user interface and select the desired audio file from the file system.
[1168] Step 2:
[1169] The device uploads the selected audio file to the server via an HTTP request, and the server receives the request and saves the audio file in the specified directory.
[1170] Step 3:
[1171] The server reads the saved audio file and captures the audio data as audio data. Specifically, it initializes the speech recognition library and loads the audio file as an audio object.
[1172] Step 4:
[1173] The server calls a speech recognition service such as the Google Speech Recognition API to convert the audio data into text data. If the conversion process is successful, the content of the audio data is obtained in text format.
[1174] Step 5:
[1175] The server extracts specific keywords from the acquired text data, checks them against a predefined list of important keywords, identifies the keywords contained in the text data, and records their location information.
[1176] Step 6:
[1177] The server inputs the voice data into the emotion engine to analyze the emotional state of the user or speaker. The emotion engine analyzes the tone, rate, intensity, etc. of the voice to determine emotions (e.g., tension, excitement, satisfaction, etc.).
[1178] Step 7:
[1179] The server stores the converted text data, extracted keywords, and analyzed emotional states in a database, adding a current timestamp to each piece of data to maintain data integrity.
[1180] Step 8:
[1181] The server displays the stored data on a user interface, and the user accesses the system from a terminal and checks the text data, important keywords, and emotional state in real time through the interface.
[1182] Step 9:
[1183] Users can review the displayed data to understand the content of the meeting and the emotional state of the speakers, which allows them to take appropriate action and adjust resource allocation.
[1184] Step 10:
[1185] The user can take necessary measures and allocate resources to improve the efficiency of future meetings. The system records these interactions and uses them for future analysis and improvement.
[1186] Example 2
[1187] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1188] Conventional project management systems make it difficult to effectively grasp the content of meetings and discussions. In particular, they do not support extracting information from voice data or analyzing users' emotional states, making it difficult to grasp the psychological state and important points of team members in real time.
[1189] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1190] In this invention, the server includes means for receiving voice data, means for converting the received voice data into character data, means for extracting specific keywords from the converted character data, means for saving the extracted keywords and their location information in a database, means for analyzing the emotional state of a user from the voice data using an emotion engine, means for saving the analyzed emotional state in a database, and means for displaying the saved data on a user interface. This makes it possible to grasp in detail not only the content of meetings and discussions but also the emotional state of the speakers, thereby making project management more efficient.
[1191] "Audio data" refers to digital data of audio recorded during a meeting or conference.
[1192] The "receiving means" is a device or software for capturing voice data sent from a user into the system.
[1193] "Means for converting into text data" refers to a device or program that includes an algorithm or voice recognition technology for analyzing voice data and converting it into text format data.
[1194] A "keyword extraction means" is a device or software that detects and extracts specified important words or phrases from character data.
[1195] An "emotion engine" is an algorithm or software that estimates emotional states through the analysis of voice data.
[1196] The "means for storing in a database" is a storage system for organizing and recording the converted character data, extracted keywords, and analyzed emotional states.
[1197] The "means for displaying on the user interface" refers to a screen or software for displaying the saved data in a form that can be viewed by the user.
[1198] The present invention aims to improve the efficiency of project management by providing a system that receives voice data, converts it into text data, extracts keywords, analyzes emotional states, stores the data in a database, and displays it on a user interface.
[1199] Hardware and software used
[1200] The main components of the system are the server, the terminal, and the user. The specific hardware and software used are specified below:
[1201] 1. Hardware
[1202] Server: A powerful computer for data processing and storage
[1203] Terminal: The device used by the user to access the site (PC, tablet, smartphone, etc.)
[1204] Microphone: A device for recording audio data.
[1205] 2. Software
[1206] Speech recognition library: Google Speech Recognition API
[1207] Database management system: MySQL or PostgreSQL
[1208] Emotion engine: a program with a voice analysis algorithm
[1209] User Interface: Web applications implemented using HTML, CSS, and JavaScript
[1210] Program processing
[1211] Audio data reception
[1212] Users record audio from meetings and conferences and upload it to the server via their devices. Specifically, they click the "File Upload" button on the user interface and select an audio file from the file system. The server receives the audio file and saves it in the specified directory.
[1213] Converting voice data to text
[1214] The server reads the received voice data as audio data and initializes a voice recognition library such as the Google Speech Recognition API. The voice data is converted into text data using the voice recognition library, and the voice content is obtained in text format.
[1215] Keyword extraction
[1216] The server extracts specific keywords from the text data. Using a predefined list of important keywords (e.g., "important," "risk," "deadline"), it searches for the presence of these keywords in the text data. It records the keywords found and their location information in a list.
[1217] Emotional state analysis using emotion engine
[1218] The server feeds the recorded voice data to the emotion engine, which analyzes factors such as tone, speed, and strength of the voice. Based on the analysis results, the user's emotional state is identified and the results are stored in a database.
[1219] Saving to a database
[1220] The server stores the converted text data, extracted keywords and their location information, and analyzed emotional state in a database, where the data is recorded with the current timestamp to maintain data integrity.
[1221] Display in the user interface
[1222] The server displays the stored data on a user interface, and users can access the system from their devices and check the text data, important keywords, and emotional state in real time through the interface.
[1223] Specific examples
[1224] For example, the text data for a meeting might read, "We need to discuss risks and important deadlines for the next project," along with keywords like "risk," "important," and "deadline." Emotional states are also displayed as "tension" or "excitement." Based on this information, users and teams can take appropriate action and allocate resources.
[1225] Prompt Sentence Examples
[1226] Prompt: "Transcribe what is said in the next meeting, extract key keywords, and analyze the emotional state of the conversation. For example, display keywords such as risk, importance, and deadline, along with the associated emotions."
[1227] In this way, the system analyzes the content of meetings and discussions in detail, aiming to improve the efficiency of project management.
[1228] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1229] Step 1: Record and upload audio data
[1230] The user records the audio of a meeting or conference and selects the audio file from the device's file system. Next, the user uses the system's "file upload" function from the device to send the audio file to the server. The input is the audio file, and the output is an audio file saved in a specified directory on the server. Specifically, the user presses the record button, and then clicks the upload button after recording is complete.
[1231] Step 2: Receive and save the audio file
[1232] The server receives audio files sent through the user interface and saves them in the specified directory. The received audio files are recorded in a log, along with metadata to maintain consistency. The input is the uploaded audio file, and the output is a file saved on the server. Specifically, the server receives a send request and processes the saving of the audio file.
[1233] Step 3: Loading the audio data
[1234] The server reads the saved audio file as audio data. During this process, the file format and encoding are checked. The input is the saved audio file, and the output is the read audio data. Specifically, the server uses Python or JavaScript libraries to open the audio file.
[1235] Step 4: Initializing the speech recognition library
[1236] The server initializes a speech recognition library such as the Google Speech Recognition API, sets an API key, and runs the authentication process. The input is the API key, and the output is the initialized speech recognition library. Specifically, the server sets the API key and initializes the library.
[1237] Step 5: Converting audio data to text
[1238] The server uses a speech recognition library to convert voice data into text data. The input is the read audio data, and the output is the converted text data. Specifically, the server sends the voice data to the speech recognition API and stores the returned text data in memory.
[1239] Step 6: Keyword extraction
[1240] The server loads a predefined list of important keywords (e.g., "important," "risk," "deadline") and extracts these keywords from the text data. The input is the converted text data and a keyword list, and the output is the extracted keywords and their location information. Specifically, it performs regular expression and pattern matching on the text data.
[1241] Step 7: Emotional state analysis by the emotion engine
[1242] The server supplies the recorded voice data to the emotion engine, which analyzes the tone, speed, and strength of the voice. The input is the voice data, and the output is information about the user's emotional state. Specifically, the engine uses a voice analysis algorithm to estimate the emotional state.
[1243] Step 8: Saving to the Database
[1244] The server stores the converted text data, extracted keywords and their location information, and analyzed emotional state in a database. During this process, the current timestamp is also recorded. The input is the text data, keywords, location information, and emotional state, and the output is the complete dataset stored in the database. Specifically, it generates an SQL query and inserts it into the database.
[1245] Step 9: Display in the User Interface
[1246] The server retrieves the necessary data from the database and displays it on the user interface. The input is the information stored in the database, and the output is the displayed text data, keywords, and emotional state. Specifically, the interface is constructed using HTML and JavaScript, and the data is displayed dynamically.
[1247] In this way, the system can sequentially perform specific tasks at each processing step, ultimately providing the information the user needs.
[1248] (Application example 2)
[1249] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1250] In today's factories, meetings and communication between workers are not managed efficiently, and improvements are needed, particularly in quality control and work efficiency. There is a lack of means to grasp the situation in real time using voice data, and there is no system that analyzes the emotional state of workers and provides appropriate feedback. This makes it difficult to detect problems early and take appropriate measures.
[1251] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1252] In this invention, the server includes means for receiving voice data, means for converting the received voice data into character data, means for extracting specific keywords from the converted character data, means for saving the extracted keywords and their location information in a database, means for displaying the saved data on a user interface, means for analyzing an emotional state from the voice data, means for saving the analyzed emotional state in a database, and means for displaying the saved data and the emotional state on a user interface. This enables real-time monitoring and efficient management of meetings and work status in a factory, and realizes quality control and improved work efficiency based on the emotional state of workers.
[1253] 1. "Voice data" refers to data that includes voice information from meetings or communications between workers.
[1254] 2. "Text data" means text-format data converted from voice data using voice recognition technology.
[1255] 3. "Keywords" refer to words or phrases that are particularly important in a meeting or work content.
[1256] 4. "Location information" is information that indicates the location where a keyword appears within the converted character data.
[1257] 5. "Database" refers to a storage means for storing and managing converted text data, extracted keywords, location information, and emotional state.
[1258] 6. "User Interface" means the display means by which a user accesses the system and checks and manipulates data.
[1259] 7. "Emotional state" refers to the results of analyzing voice data to infer the user's emotions.
[1260] 8. "Means of analysis" includes processes and technologies for extracting and identifying text data and emotional states from audio data.
[1261] 9. "Real-time" means that data is processed and displayed with little or no delay after it is generated.
[1262] 10. "Quality control" refers to activities to maintain and improve the quality of products and work processes on the factory floor.
[1263] 11. "Work efficiency" is an indicator of how efficiently work is being carried out on the factory floor.
[1264] 12. "Monitoring" refers to the continuous observation and recording of meetings and work situations.
[1265] The present invention is a system for improving the efficiency of meetings and communication between workers at a factory. Specific examples are shown below.
[1266] First, microphones installed in the factory record audio data of meetings and between workers. Users manage this audio data on their devices and use the system's "audio file upload" function to send it to the server. The server then saves the uploaded audio data in a specified directory.
[1267] The server receives the stored voice data and converts it into text data using the Google Cloud Speech-to-Text API, which has the ability to convert voice to text with high accuracy.
[1268] Next, the server uses a predefined list of important keywords to extract specific keywords from the converted text data. It checks whether any of the keywords in this list are present in the text data and records their location information. For example, if keywords such as "quality," "risk," and "work" are present, it records the words and their locations in the list.
[1269] The server then analyzes the recorded voice data and identifies the user's emotional state using the Dialogflow API. The emotion engine estimates emotions by analyzing the tone, speed, and strength of the voice. As a result of the analysis, the emotional state of the meeting or worker communication is identified as "tension" or "attention," etc.
[1270] The server stores the converted text data, extracted keywords, their location information, and the analyzed emotional state in a database. This data also includes a current timestamp to ensure data integrity.
[1271] Users can access the system from their devices and view text data, key keywords, and emotional states in real time through the interface. For example, if a meeting content reads, "We should strengthen quality inspections for the next production batch. Risks are increasing," the system will identify the keywords "quality" and "risk" along with the text data, and display the emotional states of the relevant parts of the conversation as "tension" or "caution."
[1272] This system not only improves the efficiency of in-plant meetings and communication between workers, but also enables early detection of problems and appropriate countermeasures by incorporating emotional information, allowing managers to focus on specific problems and take prompt action.
[1273] Example inputs to a generative AI model:
[1274] Project Management / Meeting Content:
[1275] "Quality inspections should be strengthened for the next production batch. The risk is increasing."
[1276] Keywords: "quality" and "risk"
[1277] Emotional state: "tension" "attention"
[1278] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1279] Step 1:
[1280] Microphones installed in factories record audio data during meetings and between workers. Users manage this audio data on their devices and send it to the server using the system's "audio file upload" function. The input data is an audio file, and the output data is an audio file saved on the server.
[1281] Step 2:
[1282] The server saves the uploaded audio data in the specified directory. The input data is the audio file received from the user, and the output data is the path to the audio file saved on the server. The specific operation is to specify the file save path and save the file in storage.
[1283] Step 3:
[1284] The server converts the saved audio data into text data using the Google Cloud Speech-to-Text API. The input data is the path to the saved audio file, and the output data is the converted text data. Specifically, the API is called to convert the audio file into text format.
[1285] Step 4:
[1286] The server extracts specific keywords from the converted text data. The input data is the text data and a predefined keyword list, and the output data is the extracted keywords and their location information. Specifically, the server searches for the locations where keywords appear in the text data and records the location information.
[1287] Step 5:
[1288] The server analyzes the recorded voice data and identifies the user's emotional state using the Dialogflow API. The input data is the voice data, and the output data is the analyzed emotional state. Specifically, the server passes the voice data to the Dialogflow API, which estimates the user's emotional state.
[1289] Step 6:
[1290] The server stores the converted text data, extracted keywords and their location information, and analyzed emotional states in a database. The input data is the text data, keyword information, and emotional states, and the output data is the information stored in the database. Specifically, the server inserts the data into the database in an appropriate format.
[1291] Step 7:
[1292] Users access the system from their devices and check text data, important keywords, and emotional states in real time through the interface. The input data is information retrieved from the database, and the output data is information displayed on the user interface. Specific operations include updating the UI to display the information retrieved from the database in real time.
[1293] These steps will enable efficient management of factory meetings and communication between workers, including emotional information, and enable early detection and countermeasures for problems.
[1294] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1295] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1296] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1297] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1298] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1299] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1300] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1301] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1302] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1303] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1304] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1305] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1306] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1307] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1308] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1309] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1310] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1311] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1312] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1313] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1314] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1315] The following is further disclosed regarding the above embodiment.
[1316] (Claim 1)
[1317] means for receiving audio data;
[1318] means for converting received voice data into text data;
[1319] A means for extracting specific keywords from the converted character data;
[1320] a means for storing the extracted keywords and their location information in a database;
[1321] means for displaying the stored data in a user interface;
[1322] A system including:
[1323] (Claim 2)
[1324] 10. The system according to claim 1, further comprising means for analyzing the voice data in real time and displaying text data containing a specific keyword on the user interface in real time.
[1325] (Claim 3)
[1326] 2. The system according to claim 1, further comprising means for analyzing the audio data of a meeting or conference, automatically extracting important points from the converted text data, and storing and displaying the extracted important points.
[1327] "Example 1"
[1328] (Claim 1)
[1329] means for receiving audio data;
[1330] means for converting received voice data into text data;
[1331] A means for extracting specific keywords from the converted character data;
[1332] a means for storing the extracted keywords and their location information in a database;
[1333] means for displaying the stored data in a user interface;
[1334] A system including:
[1335] (Claim 2)
[1336] 10. The system according to claim 1, further comprising means for analyzing the voice data in real time and displaying text data containing a specific keyword on the user interface in real time.
[1337] (Claim 3)
[1338] 2. The system according to claim 1, further comprising means for analyzing the audio data of a meeting or conference, automatically extracting important points from the converted text data, and storing and displaying the extracted important points.
[1339] (Claim 4)
[1340] a means for converting voice data into text data using a voice recognition service;
[1341] means for loading a predefined keyword list for extracting specific keywords;
[1342] 2. The system according to claim 1, further comprising a means for recording a current timestamp when character data including position information of the extracted keyword is stored in the database.
[1343] (Claim 5)
[1344] 2. The system according to claim 1, further comprising means for uploading voice data and checking analysis processing results via a user interface.
[1345] "Application Example 1"
[1346] (Claim 1)
[1347] means for receiving audio data;
[1348] means for converting received voice data into text data;
[1349] A means for extracting specific keywords from the converted character data;
[1350] a means for storing the extracted keywords and their location information in a database;
[1351] means for displaying the stored data in a user interface;
[1352] means for collecting audio data from a multi-function device installed in an industrial environment;
[1353] means for displaying the text data and its keywords on an industrial user interface;
[1354] A system including:
[1355] (Claim 2)
[1356] 10. The system according to claim 1, further comprising means for analyzing the voice data in real time and displaying text data containing a specific keyword on the user interface in real time.
[1357] (Claim 3)
[1358] 2. The system according to claim 1, further comprising means for analyzing the audio data of a meeting or conference, automatically extracting important points from the converted text data, and storing and displaying the extracted important points.
[1359] "Example 2: Combining Emotion Engines"
[1360] (Claim 1)
[1361] means for receiving audio data;
[1362] means for converting received voice data into text data;
[1363] A means for extracting specific keywords from the converted character data;
[1364] a means for storing the extracted keywords and their location information in a database;
[1365] means for analyzing a user's emotional state from the voice data using an emotion engine;
[1366] means for storing the analyzed emotional state in a database;
[1367] means for displaying the stored data in a user interface;
[1368] A system including:
[1369] (Claim 2)
[1370] 10. The system of claim 1, further comprising means for analyzing the voice data in real time and displaying text data containing specific keywords and emotional states on the user interface in real time.
[1371] (Claim 3)
[1372] 2. The system according to claim 1, further comprising means for analyzing the audio data of a meeting or conference, automatically extracting important points and emotional states from the converted text data, and storing and displaying them.
[1373] "Application example 2 when combining emotion engines"
[1374] (Claim 1)
[1375] means for receiving audio data;
[1376] means for converting received voice data into text data;
[1377] A means for extracting specific keywords from the converted character data;
[1378] a means for storing the extracted keywords and their location information in a database;
[1379] means for displaying the stored data in a user interface;
[1380] means for analyzing emotional state from audio data;
[1381] means for storing the analyzed emotional state in a database;
[1382] The system includes a means for displaying the stored data and emotional state in a user interface.
[1383] (Claim 2)
[1384] 10. The system according to claim 1, further comprising means for analyzing the voice data in real time and displaying character data containing specific keywords and emotional states on a user interface in real time.
[1385] (Claim 3)
[1386] 2. The system according to claim 1, further comprising means for analyzing the audio data of a meeting or conference, automatically extracting important points from the converted text data, and storing and displaying the extracted important points. [Explanation of symbols]
[1387] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. means for receiving audio data; means for converting received voice data into text data; A means for extracting specific keywords from the converted character data; a means for storing the extracted keywords and their location information in a database; means for displaying the stored data in a user interface; A system including:
2. 2. The system according to claim 1, further comprising means for analyzing the voice data in real time and displaying text data containing a specific keyword on the user interface in real time.
3. 2. The system according to claim 1, further comprising means for analyzing the voice data of a conference or meeting, automatically extracting important points from the converted text data, and storing and displaying the extracted important points.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A