System
The meeting efficiency system uses a generative AI model with a web conferencing tool for real-time transcription and action generation, addressing the challenges of complex automation tools and enhancing meeting efficiency and participation.
Patent Information
- Application Number
- JP2024118998
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-24
- Publication Date
- 2026-02-05
AI Technical Summary
Meetings often require significant time for preparation and struggle with poor progress and lack of active participation due to the complexity of existing automation tools, which are difficult for average users to use without specialized knowledge.
A meeting efficiency improvement system utilizing a generative AI model integrated with a web conferencing tool for real-time transcription, analysis, and action generation, enhanced by a cloud-based system for continuous improvement through user feedback.
The system enhances meeting efficiency by providing real-time transcription, analysis, and action generation, enabling all participants to contribute proactively and improve the overall meeting progress.
Smart Images

Figure 2026017937000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] Meetings often require a lot of time to prepare, leading to issues such as poor progress and a lack of active exchange of opinions. While the importance of online meetings has increased with the recent spread of remote work, it has also become more difficult to ensure efficient progress and active participation by all participants. Conventional automation tools require extensive data sets and specialized knowledge, making them difficult for average users to use. To address these issues, there is a need to provide generative AI models that can be easily used by users without specialized knowledge, thereby improving the efficiency of meetings. [Means for solving the problem]
[0005] The present invention provides a meeting efficiency improvement system that uses a generative AI model that works in conjunction with a web conferencing tool. Specifically, meeting efficiency is improved by building a system that includes a means for transcribing participants' voices in real time, a means for analyzing the text transcribed by the generative AI model and generating the next action to proceed with the meeting, and a means for presenting the generated action to participants in real time. Furthermore, the accuracy and reliability of the system can be improved by utilizing a cloud-based system to add a means for real-time processing and a means for continuously improving the accuracy of the generative AI model by reflecting user feedback. As a result, meeting preparation and progress are made more efficient, enabling all participants to contribute proactively.
[0006] "Web conferencing tool" refers to software or a platform for real-time audio, video, and text-based communication over the Internet.
[0007] A "generative AI model" refers to an algorithm or system that uses artificial intelligence techniques to analyze input data and generate output based on a predetermined purpose.
[0008] "Means of transcribing speech in real time" refers to speech recognition technology and its implementation means for instantly converting speech data into text data.
[0009] "Transcribed text" refers to text data that includes text information extracted from audio data.
[0010] "Meeting action" refers to specific actions or suggestions that are necessary to smoothly advance the flow of discussion and dialogue within the meeting.
[0011] "Means for presenting to participants in real time" refers to a display device or interface that allows participants to instantly see the progress of the generated conference.
[0012] A "cloud-based system" refers to a system architecture that processes and stores data on remote servers over the Internet.
[0013] "User feedback" refers to user opinions, evaluations, and suggestions for improvement obtained during or after use of the system.
[0014] "Measures to improve accuracy" refers to methods and techniques for continuously improving the performance of a system or model. [Brief explanation of the drawings]
[0015] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram illustrating a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0016] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0017] First, the terms used in the following description will be explained.
[0018] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0019] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0020] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0021] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0022] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0023] [First embodiment]
[0024] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0025] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0026] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0027] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0028] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0029] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0030] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0031] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0032] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0033] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0034] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0035] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0036] The present invention is a system that improves the efficiency of meetings by using a generative AI model that works in conjunction with a web conferencing tool. This system has the function of transcribing participants' voices in real time, analyzing the text data, and generating actions to progress the meeting. In addition, by presenting the generated actions to participants in real time, the progress of the meeting can be made smoother. A specific embodiment of this system is described below.
[0037] Key Components Used
[0038] 1. Server: Transcribing audio data and analyzing text data using a generative AI model.
[0039] 2. Terminal: Receives input from the user, sends it to the server, and presents the generated progress actions to the user.
[0040] 3. User: Set the purpose and goals of the meeting and proceed with the meeting based on progress actions.
[0041] Program processing explanation
[0042] 1. User Input
[0043] Users set the purpose and goals of the meeting through the interface of the web conferencing tool, for example, entering "planning the next project" and "defining the main tasks and deciding who will be responsible."
[0044] The user also enters the path to the conference audio file (eg "audio.wav").
[0045] 2. Transmission of information
[0046] The terminal transmits the purpose and goal of the meeting and the path of the audio file input by the user to the server.
[0047] 3. Analysis of audio data
[0048] The server transcribes the received audio file in real time, using a speech recognition library (e.g., Google Speech Recognition API).
[0049] Transcribed text data is generated and fed into a generative AI model.
[0050] 4. Proceeding Action Generation
[0051] The server uses a generative AI model to analyze the transcribed text data and generate next steps based on the purpose and goals of the meeting, such as "encourage discussion on which specific markets to focus on."
[0052] 5. Presenting Progress Actions
[0053] The server transmits the generated progress action to the terminal.
[0054] The terminal presents this ongoing action to the user in real time.
[0055] Specific examples
[0056] For example, suppose a user wants to hold a meeting titled "New Product Launch Plan." In this case, the user enters the following information into the web conferencing tool's interface:
[0057] Meeting purpose: "New product launch plan"
[0058] Goal: "Formulate market strategies and assess risks"
[0059] Audio file: "audio_new_product_meeting.wav"
[0060] This information is sent from the device to the server. The server transcribes the audio file and sends it to the generative AI model for analysis. As a result, the generative AI model generates a progress action such as "Please propose specific measures for market strategy." This action is sent to the device and displayed to the user through the device.
[0061] This system allows users to efficiently organize what is being said during a meeting and receive appropriate proceeding actions in real time, thereby improving the efficiency and effectiveness of meetings.
[0062] The processing flow will be explained below.
[0063] Step 1:
[0064] The user inputs the purpose of the meeting (e.g., "Planning the next project") and goal (e.g., "Define main tasks and determine who is responsible") into the interface of the web conferencing tool. The user also inputs the path to the audio file for the meeting (e.g., "audio.wav").
[0065] Step 2:
[0066] The terminal transmits the purpose and goal of the meeting and the path of the audio file input by the user to the server.
[0067] Step 3:
[0068] The server loads the audio file based on the path of the audio file received, and transcribes the audio data in real time using a speech recognition library (e.g., Google Speech Recognition API).
[0069] Step 4:
[0070] The server sends the transcribed text data to the generative AI model, which analyzes the transcribed data and the purpose and goals of the meeting to generate the next action to take.
[0071] Step 5:
[0072] The server transmits the generated progress action to the terminal.
[0073] Step 6:
[0074] The terminal displays the progress actions received from the server to the user in real time, and the user proceeds with the conference based on the presented progress actions.
[0075] Step 7:
[0076] After the conference ends, the server stores the transcript log of the entire conference and the history of the ongoing actions generated.
[0077] Step 8:
[0078] After the meeting ends, users can download the transcription log and progress action history stored on the server via their device and use them for summarizing and providing feedback.
[0079] Example 1
[0080] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0081] Modern meetings are increasingly being held online, especially with the spread of remote work. However, it remains difficult to efficiently understand what participants are saying during a meeting and decide on the appropriate next steps. As a result, the meeting often stalls and delays the achievement of its objectives and goals. Maintaining meeting efficiency and effectiveness becomes even more difficult when there are many participants or a wide range of information.
[0082] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0083] In this invention, the server includes a means for a user to set a purpose and goal of a meeting, a means for a user to upload audio data of the meeting, a means for transcribing the audio data received from the target in real time, a means for analyzing the transcribed text data by a generative AI model and generating the next action to proceed with the meeting, and a means for presenting the generated action to proceed with the meeting to the user in real time, thereby enabling real-time analysis of utterances during a meeting and the presentation of action to proceed with the meeting.
[0084] "User" refers to a person who uses this system to set the purpose and goal of a meeting and upload audio data.
[0085] The "server" is a device that transcribes voice data received from users in real time, analyzes the text using a generative AI model, and generates actions to progress the meeting.
[0086] A "terminal" is a device that allows a user to input the purpose and goal of a meeting and transmit audio files to a server, and is also a device that presents the generated progress actions to the user.
[0087] "Audio data" refers to a digital data file that records what participants say during a meeting.
[0088] "Real-time" refers to processing and displaying data immediately, with minimal time delay.
[0089] "Transcription" refers to the process of converting audio data into text form.
[0090] "Generative AI model" refers to the artificial intelligence technology used to analyze transcribed text data and generate actions that help guide meetings.
[0091] "Progression actions" refer to the next actions or discussion content suggested by the generative AI model based on the purpose and goals of the meeting.
[0092] "Upload" refers to the act of transferring data from a terminal to a server.
[0093] "Analysis" refers to the process by which generative AI models process transcribed text data, extract information, and generate appropriate actions.
[0094] This invention is a system that improves the efficiency of meetings by using a generative AI model that works in conjunction with a web conferencing tool. This system operates using multiple means, primarily servers, terminals, and users. The detailed configuration and operation of each means are explained below.
[0095] First, the user sets the purpose and goals of the meeting through the interface of the web conferencing tool. A dedicated form is provided for entering this information. For example, the purpose of the meeting can be "Planning the next project" and the goal can be "Define the main tasks and determine who will be responsible." In addition, a field is provided for uploading audio data from the meeting (e.g., "audio.wav"). The user selects the required audio file and clicks the "Send" button.
[0096] Next, the terminal has a means for sending the meeting purpose, goal, and audio file path entered by the user together to the server. This information is packetized and sent to the server as an HTTP request. This generally uses the HTTPS protocol.
[0097] The server has a means to transcribe the received voice data in real time. Specifically, it uses a speech recognition library such as the Google Speech Recognition API to convert the voice data into text data. To do this, it processes the received audio file (e.g., "audio.wav") and passes it to the speech recognition library. The speech recognition library converts the voice into text data and returns the result to the server, which temporarily stores it.
[0098] The server then sends the transcribed text data to a generative AI model for analysis. The generative AI model uses an interactive generative model (e.g., GPT-3.5). Based on the transcribed text data, this model generates the next progress action appropriate to the purpose and goal of the meeting. For example, a progress action such as "Please propose specific measures for market strategy" may be generated. This generated progress action is temporarily stored on the server.
[0099] Finally, the generated progress actions are sent from the server to the terminal and presented to the user in real time. The terminal displays the progress actions received from the server on the screen interface. For example, the action display area of the interface may display "Please propose specific measures for the market strategy."
[0100] Specific examples
[0101] For example, if a user wants to hold a meeting titled "New Product Launch Plan," they would enter the following information into the web conferencing tool's interface:
[0102] Meeting purpose: "New product launch plan"
[0103] Goal: "Formulate market strategies and assess risks"
[0104] Audio file: "audio_new_product_meeting.wav"
[0105] This information is sent from the device to the server. The server transcribes the audio file using the Google Speech Recognition API and passes it to the generative AI model for analysis. As a result, the generative AI model generates a progress action such as "Please propose specific measures for our market strategy." This action is sent to the device and displayed to the user in real time.
[0106] Example prompt sentence:
[0107] "The purpose of this meeting is to plan the market launch of a new product. We want to develop a market strategy and assess risks. Based on this objective, what are the next steps we should take?"
[0108] This system allows users to efficiently organize what is being said during a meeting and receive appropriate proceeding actions in real time, improving the efficiency and effectiveness of meetings.
[0109] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0110] Step 1: User Input
[0111] The user accesses the interface of the web conferencing tool and inputs the purpose and goals of the meeting. For example, they might set "Planning the next project" and "Define main tasks and determine who will be responsible." They also specify the path to the audio file for the meeting (e.g., "audio.wav"). After inputting this information, the user clicks the "Send" button.
[0112] Input: Meeting purpose, goal, audio file path
[0113] Output: Configuration information passed to the terminal
[0114] Step 2: Submit your information
[0115] The terminal packetizes the user-entered meeting purpose, goal, and audio file path, and sends the packet as an HTTP request to the server using the HTTPS protocol.
[0116] Input: Information entered by the user (meeting purpose, goal, audio file path)
[0117] Output: HTTP request sent to the server
[0118] Step 3: Receiving and transcribing audio data
[0119] The server receives the HTTP request and gets the path to the audio file. The server loads this audio file and transcribes it in real time using a speech recognition library such as the Google Speech Recognition API. This process converts the audio data into text data.
[0120] Input: The path to the audio file sent to the server
[0121] Output: Transcribed text data
[0122] Step 4: Analyzing the text data
[0123] The server temporarily stores the transcribed text data and then passes it to a generative AI model. The generative AI model analyzes the data and generates the next action based on the purpose and goals of the meeting. For example, it might generate an instruction such as, "Please propose specific measures for market strategy."
[0124] Input: Transcribed text data
[0125] Output: Generated progress actions
[0126] Step 5: Present progress actions
[0127] The server packets the generated progress actions and sends them to the terminal. The terminal receives the progress actions and displays them to the user in real time. For example, the action display area of the interface may display "Please propose specific measures for market strategy."
[0128] Input: Generated progress action
[0129] Output: Progression actions presented to the user
[0130] (Application example 1)
[0131] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0132] Modern conference systems and driver assistance systems require technologies that can streamline decision-making for participants and drivers and suggest appropriate actions in real time. However, conventional systems have issues such as not being able to transcribe speech in real time, not being able to fully utilize generative AI models, and not being able to quickly suggest appropriate actions. This can sometimes hinder smooth progress in meetings and driving assistance, resulting in issues of reduced efficiency.
[0133] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0134] In this invention, the server includes a means for transcribing the speech of participants and the driver in real time, a means for analyzing the transcribed text using a generative AI model and generating the next action for proceeding with the meeting or driving assistance action, and a means for presenting the generated action for proceeding with the meeting or driving assistance action to the participants and the driver in real time. This enables quick and appropriate decision-making during meetings and driving, and enables efficient progress and assistance.
[0135] A "web conferencing tool" is software that allows multiple participants in remote locations to hold meetings by sharing audio and video over the Internet.
[0136] A "generative AI model" is an artificial intelligence algorithm that generates output, such as text or actions, based on input data.
[0137] The "meeting efficiency system" is a system that transcribes and analyzes audio to generate and suggest the next action in order to make meetings proceed smoothly and effectively.
[0138] A "driving assistance system" is a system that assists vehicle drivers by transcribing and analyzing voice and providing appropriate instructions and information to the driver in real time.
[0139] "Participant" or "Driver" refers to a person attending a meeting or driving a vehicle.
[0140] "Real time" means that processing is carried out almost immediately after an event occurs.
[0141] "Transcription" is the process of analyzing audio data and converting its content into text format.
[0142] "Means" refer to the methods or tools used to achieve a particular goal.
[0143] "Next Actions" are specific actions or proposals to be taken next, generated based on the purpose and progress of the meeting.
[0144] "Driving assistance actions" are instructions or information provided to the driver in real time to assist driving.
[0145] A "cloud-based system" is a system in which data is processed and stored on a remote server via the Internet.
[0146] "User feedback" refers to opinions and evaluation data provided by users of the system, and is used to improve the system.
[0147] The present invention is an efficiency improvement system using a generative AI model that can be applied to web conferencing tools and driving assistance systems for autonomous vehicles. This system has the function of transcribing the voices of participants and drivers in real time, analyzing the text data, and generating the next progress action or driving assistance action. In addition, by presenting the generated actions to participants and drivers in real time, the progress of the meeting or driving can be made smoother. Specific embodiments for implementing the present invention are described below.
[0148] Key Components Used
[0149] 1. Server:
[0150] The audio data is transcribed and the text data is analyzed using a generative AI model.
[0151] Use a speech recognition library (e.g., Google Speech Recognition API) for transcription.
[0152] 2. Terminal:
[0153] It receives input from the user and sends it to the server.
[0154] The generated progress actions and driving assistance actions are presented to the user.
[0155] 3. User:
[0156] Set the purpose or goal of the meeting or provide driving context.
[0157] Proceed with meetings or driving based on the generated actions.
[0158] Program processing explanation
[0159] The server transcribes speech from participants and drivers in real time. To do this, it uses a speech recognition library to convert the speech data into text format. It then uses a generative AI model to analyze the transcribed text data and generate next actions for the meeting or driver assistance actions. This generative AI model processes the data in real time on a cloud-based system.
[0160] The generated actions are sent from the server to the terminal, and the terminal presents these actions to the user in real time. The user then takes the next action based on the presented actions.
[0161] Specific examples
[0162] For example, if a user wants to hold a meeting titled "New Product Launch Plan," the user enters the following information into the web conferencing tool's interface:
[0163] Meeting purpose: "New product launch plan"
[0164] Goal: "Formulate market strategies and assess risks"
[0165] Audio file path: "audio_new_product_meeting.wav"
[0166] This information is sent from the device to the server. The server transcribes the audio file and sends it to the generative AI model for analysis. As a result, the generative AI model generates a progress action such as "Please propose specific measures for market strategy." This action is sent to the device and displayed to the user through the device.
[0167] In a driver assistance system, if a driver wants to know the location of the next exit while driving, they can utter the voice command "Tell me where the next exit is." This speech is transcribed in real time and analyzed by a generative AI model. The server generates a driving assistance action such as "Please use the exit on the right 2 km ahead," and sends it to the device. This action is displayed on the driver's smart glasses or head-mounted display.
[0168] Example prompt sentence:
[0169] You are a driving assistance bot. Please interpret the following voice commands and suggest appropriate actions.
[0170] Voice command: What's the next entrance / exit?
[0171] In this way, the system allows users to efficiently organize what is said during meetings or while driving and receive appropriate proceeding actions and driving assistance actions in real time, thereby improving meeting efficiency and driving safety.
[0172] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0173] Step 1:
[0174] User input
[0175] The user uses the interface of the web conferencing tool or the voice input device of the driving assistance system to input the purpose and goal of the meeting, as well as the path of the audio file and voice commands while driving, including input in text format or voice input, which results in specific data.
[0176] Input: purpose of the meeting, goal, audio file path, voice command
[0177] Output: User input data (text data and audio data)
[0178] Step 2:
[0179] Sending information
[0180] The device sends the entered meeting purpose, goal, audio file path, and voice command to the server, which then sends the data over the network to a cloud-based server for processing.
[0181] Input: User-entered data
[0182] Output: Data sent to the server
[0183] Step 3:
[0184] Analysis of audio data
[0185] The server transcribes the received audio file using a speech recognition library (such as the Google Speech Recognition API). The audio data is analyzed and converted into text. This transcribed text data is then sent to the generative AI model.
[0186] Input: Audio data
[0187] Output: Transcribed text data
[0188] Step 4:
[0189] Analysis using generative AI models
[0190] The server analyzes the transcribed text data using a generative AI model and generates the next progress action or driving assistance action based on the prompt sentence. The generative AI model analyzes the text data using a specific algorithm and suggests the optimal action.
[0191] Input: Transcribed text data
[0192] Output: Generated progress actions and driver assistance actions
[0193] Step 5:
[0194] Suggesting Actions
[0195] The server sends the generated progress actions and driving assistance actions to the device, which then presents these actions to the user in real time: if in a meeting, they are displayed in the meeting interface, and if driving, they are displayed to the driver through smart glasses or a head-mounted display.
[0196] Input: Generated progress actions and driver assistance actions
[0197] Output: Real-time suggested actions
[0198] Step 6:
[0199] User Feedback
[0200] The user continues to conduct a meeting or drive based on the generated actions, and user feedback provides the system with data for improvement, allowing the accuracy of the generative AI model to be continuously improved.
[0201] Input: User feedback
[0202] Output: Improve the accuracy of the generative AI model.
[0203] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0204] The present invention is a system that uses a generative AI model that works with a web conferencing tool to improve the efficiency of meetings and, by combining it with an emotion engine, makes the meeting proceed more effectively. This system has the function of transcribing participants' voices in real time and analyzing the text data to generate meeting proceeding actions. Furthermore, by presenting the generated proceeding actions to participants in real time and combining it with an emotion engine that recognizes the user's emotions, the meeting can proceed smoothly. A specific embodiment of this system is described below.
[0205] Key Components Used
[0206] 1. Server: Transcribing voice data, analyzing text data using a generative AI model, and recognizing user emotions using an emotion engine.
[0207] 2. Terminal: Receives input from the user, sends it to the server, and provides feedback to the user based on the generated progress actions and emotions.
[0208] 3. User: Sets the purpose and goals of the meeting, provides emotional data, and advances the meeting based on progress actions.
[0209] 4. Emotion engine: Analyzes the user's voice and facial expressions and sends the emotion data to the generative AI model.
[0210] Program processing explanation
[0211] 1. User Input
[0212] The user sets the purpose and goals of the meeting through the interface of the web conferencing tool. For example, they input "Plan the next project" and "Define the main tasks and determine who will be responsible." The user also inputs the path to the meeting audio file (e.g., "audio.wav").
[0213] 2. Transmission of information
[0214] The terminal transmits the purpose and goal of the meeting and the path of the audio file input by the user to the server.
[0215] 3. Analysis of audio data
[0216] The server loads the audio file based on the path of the audio file received, and transcribes the audio data in real time using a speech recognition library (e.g., Google Speech Recognition API).
[0217] 4. Acquiring Emotion Data
[0218] The server uses an emotion engine to analyze the user's emotional data from the audio and video data, based on the tone of voice, facial expressions, and content of speech.
[0219] 5. Proceeding Action Generation
[0220] The server sends the transcribed text data and emotion data to the generative AI model. The generative AI model analyzes the transcribed data and emotion data and generates the next action to take. For example, an action such as "If the user is nervous, suggest questions to help them relax" may be generated.
[0221] 6. Presenting Progressive Actions and Emotional Feedback
[0222] The server transmits feedback based on the generated progress actions and emotions to the terminal.
[0223] The device displays this information to the user in real time, and the user can proceed with the meeting based on the presented actions and emotional feedback.
[0224] Specific examples
[0225] For example, suppose a user wants to hold a meeting titled "New Product Launch Plan." In this case, the user enters the following information into the web conferencing tool's interface:
[0226] Meeting purpose: "New product launch plan"
[0227] Goal: "Formulate market strategies and assess risks"
[0228] Audio file: "audio_new_product_meeting.wav"
[0229] This information is sent from the device to the server. The server transcribes the audio file and analyzes the user's emotions using an emotion engine. The generative AI model then analyzes this data and generates a progress action, such as "Please propose specific measures for market strategy." If the user is nervous, it will also suggest questions to relax them (e.g., "Tell me about a recent success story"). This action and emotional feedback are sent to the device and displayed to the user.
[0230] This system allows users to efficiently organize statements and emotional data during meetings and receive appropriate proceeding actions in real time, further improving the efficiency and effectiveness of meetings.
[0231] The processing flow will be explained below.
[0232] Step 1:
[0233] The user inputs the purpose of the meeting (e.g., "Plan the next project") and goal (e.g., "Define main tasks and determine who is responsible") into the interface of the web conferencing tool. The user also inputs the path to the audio file for the meeting (e.g., "audio.wav").
[0234] Step 2:
[0235] The terminal transmits the purpose and goal of the meeting and the path of the audio file input by the user to the server.
[0236] Step 3:
[0237] The server loads the audio file based on the path of the audio file received, and transcribes the audio data in real time using a speech recognition library (e.g., Google Speech Recognition API).
[0238] Step 4:
[0239] The server sends the transcribed text data to the generative AI model, which analyzes the transcribed data and the purpose and goals of the meeting to generate the next action to take.
[0240] Step 5:
[0241] The server captures the user's voice and facial expressions, and uses an emotion engine to analyze the user's emotional data. Emotions are recognized based on the tone of voice, changes in facial expressions, and the content of speech.
[0242] Step 6:
[0243] The server inputs the emotion data into the generative AI model and takes it into consideration when generating the next action to proceed, such as "suggesting questions to help a nervous user relax."
[0244] Step 7:
[0245] The server sends the generated progress actions and emotional feedback to the device.
[0246] Step 8:
[0247] The device receives the progress actions and emotional feedback from the server and presents them to the user in real time. The user then proceeds with the conference based on the progress actions and emotional feedback presented to them.
[0248] Example 2
[0249] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0250] In today's business environment, there is a demand for improving the efficiency and effectiveness of meetings. However, it is difficult to instantly analyze participants' emotions and comments and provide appropriate proceeding actions in real time. Furthermore, conventional conferencing systems are unable to quickly respond to changes in emotions, which can hinder the smooth progress of the meeting. Therefore, there is a need for a method that can analyze participants' emotions in real time and optimize the progress of the meeting based on that data.
[0251] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0252] In this invention, the server includes a means for a user to set the purpose and goal of the meeting, a means for transcribing the participants' voices in real time, a means for analyzing the transcribed text and the users' emotional data using a generative AI model to generate the next action to proceed with the meeting, and a means for presenting feedback to the participants in real time based on the generated action to proceed with the meeting and their emotions. This increases the efficiency and effectiveness of the meeting, and enables the meeting to proceed smoothly by quickly responding to changes in the participants' emotions.
[0253] "Means for users to set the purpose and goals of a meeting" means allowing users to input the main purpose and goals of a meeting using the interface of the web conferencing tool.
[0254] "Means for transcribing participants' voices in real time" refers to means for converting the speech of conference participants into text data in real time using voice recognition technology.
[0255] A "generative AI model" refers to an algorithm or program that uses natural language processing and machine learning to generate meaningful information from input data.
[0256] The "means for analyzing the transcribed text and user emotion data and generating the next action to proceed with the meeting" refers to a means for generating instructions and suggestions for determining the progress of the meeting based on the transcribed text data and emotion analysis data.
[0257] An "emotion engine" refers to technology or algorithms that analyze voice tone, speech rate, facial expression changes, etc. to identify the emotional state of users or participants.
[0258] The "means for presenting participants with feedback based on generated meeting progress actions and emotions in real time" is a means for instantly providing analyzed and generated information to meeting participants to guide the progress of the meeting.
[0259] "Cloud-based processing" refers to the technology of processing data on remote servers via the Internet, and is used to efficiently perform large-scale data analysis and real-time processing.
[0260] "Means for incorporating user feedback and continuously improving the accuracy of the generative AI model" refers to means for improving the performance and accuracy of the generative AI model over time by incorporating user opinions and evaluations into the model.
[0261] This invention relates to a system that improves the efficiency and effectiveness of meetings by utilizing a generative AI model that works in conjunction with a web conferencing tool. Its key feature is that it transcribes participants' speech in real time, analyzes the text data to generate meeting progress actions, and provides feedback to participants, thereby helping to ensure smooth progress of the meeting. Furthermore, by combining it with an emotion engine, it is possible to analyze participants' emotions and reflect them in the progress actions.
[0262] Key Components Used
[0263] The following describes the major components for implementing the present invention.
[0264] server
[0265] The server is the core of this system and has the following main functions:
[0266] 1. Ability to transcribe voice data in real time: Uses a speech recognition library (e.g., Google Speech Recognition API) to convert participants' voice data into text data.
[0267] 2. The ability to analyze text data and emotional data using a generative AI model and generate next steps: The generative AI model uses natural language processing and machine learning based on the input data to derive the appropriate next action.
[0268] 3. Ability to analyze emotional data using an emotion engine: Identify the user's emotional state from voice tone, speaking rate, facial expression changes, etc.
[0269] Terminal
[0270] The terminal is a device that provides a user interface and has the following functions:
[0271] 1. Function to receive input from the user and send it to the server: Provides an interface for inputting the purpose and goals of the meeting, the path of the audio file, etc.
[0272] 2. A function to present feedback to the user based on the generated progress actions and emotions: Feedback sent from the server is displayed in real time.
[0273] User
[0274] The user sets the purpose and goal of the meeting and is responsible for running the meeting. The user also provides data on comments and emotions made during the meeting, and the system guides the meeting according to the actions taken by the system.
[0275] Specific examples
[0276] For example, consider a user who wants to hold a meeting titled "New Product Launch Plan." The user enters the following information into the web conferencing tool's interface:
[0277] Meeting purpose: "New product launch plan"
[0278] Goal: "Formulate market strategies and assess risks"
[0279] Audio file: "audio_new_product_meeting.wav"
[0280] Prompt Sentence Examples
[0281] "Planning the next project"
[0282] "Define key tasks and determine who is responsible"
[0283] "Marketing Strategy Meeting"
[0284] "Target Market Definition and Competitive Analysis"
[0285] The device sends this information to a server. The server reads the audio file and transcribes it using a speech recognition library. It then uses an emotion engine to analyze the user's emotions from the acquired transcription and video data. The analysis results are sent to a generative AI model, which generates the next action to proceed.
[0286] For example, if a user is nervous, the generative AI model will generate a progression action such as "suggesting questions to relax the user." It may also suggest questions to dig deeper into marketing strategies. The generated progression actions and feedback are sent to the device and displayed to the user in real time.
[0287] This allows users to smoothly progress with the conference, and improves the efficiency and effectiveness of the conference.
[0288] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0289] Step 1:
[0290] The user sets the purpose and goal of the meeting through the interface of the web conferencing tool and inputs the path of the audio file. As input, the user provides the purpose of the meeting (e.g., "Plan for the market launch of a new product"), the goal (e.g., "Formulate market strategy and assess risks"), and the path of the audio file (e.g., "audio_new_product_meeting.wav"). This information is saved on the device for the next step.
[0291] Step 2:
[0292] The terminal sends the information input by the user to the server. The input information includes the purpose and goal of the meeting, and the path to the audio file. The terminal appropriately packets this information and sends it to the server. The output is the meeting information sent to the server.
[0293] Step 3:
[0294] The server receives the audio file path and loads it. Then, it uses a speech recognition library (e.g., Google Speech Recognition API) to transcribe the audio data in real time. The input is the audio file path "audio_new_product_meeting.wav" and the output is the transcribed text data (e.g., "The discussion about the new product has begun...").
[0295] Step 4:
[0296] The server inputs the acquired transcription data and video data into the emotion engine. The emotion engine analyzes the tone of voice, speech rate, and facial expression changes to generate the user's emotion data. The inputs are transcription data and video data, and the output is the analyzed emotion data (e.g., "The user is nervous").
[0297] Step 5:
[0298] The server sends the transcribed text data and emotion data to the generative AI model, which generates the next progress action based on this data. The input is the transcribed text and emotion data, and the output is the generated progress action (e.g., "Suggest questions to help the user relax").
[0299] Step 6:
[0300] The server transmits the generated progress action and emotion-based feedback to the terminal. The input is the progress action and emotion feedback, and the output is the data transmitted to the terminal.
[0301] Step 7:
[0302] The device displays the progress actions and emotional feedback received from the server to the user. Specifically, they are displayed as notifications or pop-ups on the interface and provided to the user in real time. The input is the progress actions and emotional feedback from the server, and the output is the information displayed to the user.
[0303] (Application example 2)
[0304] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0305] With conventional web conferencing tools, meeting efficiency and progress management are largely manual, which means that it takes time to organize what is being said and decide on actions to take. There are also concerns that the quality of meetings may decline due to the difficulty of properly understanding participants' emotions and smoothly progressing the meeting. Furthermore, when dealing with customers in virtual stores, it is difficult to suggest appropriate actions to take in response to customer questions, making it difficult to increase customer satisfaction.
[0306] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0307] In this invention, the server includes means for transcribing the participants' voices in real time, means for analyzing the transcribed text using a generative AI model and generating the next action to proceed with the conference, means for presenting the generated action to proceed with the conference to the participants in real time, means for analyzing the emotions of the participants using an emotion engine and providing appropriate feedback based on the analysis results, means for monitoring the customer service process in the virtual store and generating appropriate actions in response to customer questions, and means for presenting the generated response actions and emotion feedback to the customers in real time. This makes it possible to efficiently manage the progress of the conference and provide optimal actions in response to the emotions of the participants, as well as improve the quality of customer service in the virtual store and increase customer satisfaction.
[0308] "Web conferencing tool" refers to software or a platform that allows multiple participants to hold real-time conferences using audio and video over the Internet.
[0309] A "generative AI model" refers to an algorithm or system that uses artificial intelligence to analyze input data and automatically generate appropriate responses or actions.
[0310] A "meeting efficiency system" refers to a system that combines technologies and means to support the progress and management of meetings and improve their efficiency.
[0311] "Real-time transcription means" refers to technology or equipment for instantly converting audio data into text data.
[0312] "Progressive actions" refer to specific actions or instructions that should be taken next in the progress of a meeting.
[0313] An "emotion engine" refers to technology and algorithms that analyze data such as participants' voices and facial expressions to estimate their emotional state.
[0314] "Feedback" refers to information or instructions provided to users or participants based on analysis results, progress actions, etc.
[0315] "Virtual store" refers to a virtual store that offers products and services via the Internet.
[0316] A "customer service process" refers to a set of procedures or methods for appropriately responding to customer inquiries and questions.
[0317] This invention is a system that improves the efficiency of meetings by using a generative AI model that works in conjunction with a web conferencing tool, and further combines it with an emotion engine to more effectively progress the meeting. This system has the function of transcribing participants' voices in real time, analyzing the text data, and generating the next progress action. Furthermore, by presenting the generated progress action to participants in real time and combining it with an emotion engine that recognizes the user's emotions, the meeting can be progressed smoothly.
[0318] Key Components Used
[0319] 1. Server: Transcribing voice data, analyzing text data using a generative AI model, and recognizing user emotions using an emotion engine. Specifically, it uses a speech recognition library (e.g., Google Speech Recognition API) and an emotion analysis engine (e.g., EmotionEngine). It also generates progress actions using a generative AI model (e.g., ActionsGenerator).
[0320] 2. Device: Receives input from the user and sends it to the server. It also provides feedback to the user based on the generated progress actions and emotions. Devices can be smartphones, head-mounted displays, smart glasses, etc.
[0321] 3. User: Sets the purpose and goals of the meeting, provides emotional data and advances the meeting based on progress actions.
[0322] 4. Emotion engine: Analyzes the user's voice and facial expressions and sends the emotional data to a generative AI model.
[0323] Program processing explanation
[0324] 1. The server transcribes the user's voice data using the Google Speech Recognition API. This process converts the voice data into text data.
[0325] 2. The emotion engine then analyzes the voice tone and facial expressions to generate the user's emotion data. This is done using the EmotionEngine.
[0326] 3. The server uses a generative AI model to analyze the transcribed text data and emotion data, and generates the next action to proceed. Using ActionsGenerator, the generated action is provided in real time.
[0327] 4. The device presents the progress actions and emotional feedback received from the server to the user in real time. This information is displayed via a smartphone, head-mounted display, or smart glasses.
[0328] Specific examples
[0329] For example, suppose a user wants to hold a meeting titled "New Product Launch Plan." In this case, the user enters the following information into the web conferencing tool's interface:
[0330] Meeting purpose: "New product launch plan"
[0331] Goal: "Formulate market strategies and assess risks"
[0332] Audio file: "audio_new_product_meeting.wav"
[0333] This information is sent from the device to the server. The server transcribes the audio file and analyzes the user's emotions using an emotion engine. The generative AI model then analyzes this data and generates a progress action, such as "Please propose specific measures for market strategy." If the user is nervous, it will also suggest questions to relax them (e.g., "Tell me about a recent success story"). This action and emotional feedback are sent to the device and displayed to the user.
[0334] Prompt Sentence Examples
[0335] Below are some example prompts to use with generative AI models:
[0336] text
[0337] Voice input: "I want to know about my new TV."
[0338] Emotional data: interesting but a bit confusing
[0339] Generate Action: Describe the features of their new TV. Ask about their specific needs.
[0340] This system efficiently organizes statements and emotional data during meetings and allows appropriate actions to be taken in real time, further improving the efficiency and effectiveness of meetings. It also makes it possible to generate and provide appropriate actions in real time when dealing with customers in virtual stores, thereby increasing customer satisfaction.
[0341] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0342] Step 1:
[0343] The server receives the user's voice data. The user inputs the purpose, goal, and audio file of the meeting into the terminal. The terminal sends this information to the server. Specific inputs include the audio file (e.g., "audio_new_product_meeting.wav"), the purpose of the meeting, "a plan to launch a new product," and the goal, "formulation of a market strategy and risk assessment." The server receives this data and begins processing.
[0344] Step 2:
[0345] The server transcribes the received voice data in real time using the Google Speech Recognition API. Text data (output) is generated from the voice data (input). This transcribed text data serves as the basis for generating actions to progress the meeting. Converting the input voice data into text makes subsequent data analysis easier.
[0346] Step 3:
[0347] The server uses an emotion engine to analyze the user's emotional data along with the transcribed voice data. Emotion analysis is performed using the Emotion Engine. Specifically, it analyzes changes in voice tone and facial expressions. Through this process, the server obtains emotional data (output) from the voice data (input). For example, it determines whether the user is nervous, relaxed, or interested. At this stage, data enrichment is performed.
[0348] Step 4:
[0349] The server uses a generative AI model (e.g., ActionsGenerator) to analyze the transcribed text data and emotional data. Specifically, the generative AI model generates the next action to proceed (output) based on the text data (input) and emotional data (input). For example, an action such as "Please propose specific measures for market strategy" may be generated. This action provides specific instructions to ensure the smooth flow of the meeting.
[0350] Step 5:
[0351] The server sends the generated progress action and emotional feedback to the device. Specifically, the generated action (output) and emotional feedback (output) are sent. For example, the action may be "Please propose specific measures for the market strategy" and the emotional feedback may be "The user is nervous." This allows the device to process the next step.
[0352] Step 6:
[0353] The device presents the progress actions and emotional feedback received from the server to the user in real time. Specifically, information is displayed to the user (output) via a smartphone, head-mounted display, or smart glasses. For example, the device may provide the user with instructions such as "Please propose specific measures for market strategy" or feedback such as "The user seems nervous, so we suggest questions to help them relax." This allows the user to effectively proceed with the next action in the meeting.
[0354] Step 7:
[0355] The user controls the meeting based on the progress actions and emotional feedback presented on the device. By taking specific actions, the user can effectively move forward toward the purpose and goal of the meeting. The user uses the presented information to execute the next step instructions. This is the stage at which the actual meeting progresses.
[0356] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0357] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0358] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0359] [Second embodiment]
[0360] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0361] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0362] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0363] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0364] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0365] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0366] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0367] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0368] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0369] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0370] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0371] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0372] The present invention is a system that improves the efficiency of meetings by using a generative AI model that works in conjunction with a web conferencing tool. This system has the function of transcribing participants' voices in real time, analyzing the text data, and generating actions to progress the meeting. In addition, by presenting the generated actions to participants in real time, the progress of the meeting can be made smoother. A specific embodiment of this system is described below.
[0373] Key Components Used
[0374] 1. Server: Transcribing audio data and analyzing text data using a generative AI model.
[0375] 2. Terminal: Receives input from the user, sends it to the server, and presents the generated progress actions to the user.
[0376] 3. User: Set the purpose and goals of the meeting and proceed with the meeting based on progress actions.
[0377] Program processing explanation
[0378] 1. User Input
[0379] Users set the purpose and goals of the meeting through the interface of the web conferencing tool, for example, entering "planning the next project" and "defining the main tasks and deciding who will be responsible."
[0380] The user also enters the path to the conference audio file (eg "audio.wav").
[0381] 2. Transmission of information
[0382] The terminal transmits the purpose and goal of the meeting and the path of the audio file input by the user to the server.
[0383] 3. Analysis of audio data
[0384] The server transcribes the received audio file in real time, using a speech recognition library (e.g., Google Speech Recognition API).
[0385] Transcribed text data is generated and fed into a generative AI model.
[0386] 4. Proceeding Action Generation
[0387] The server uses a generative AI model to analyze the transcribed text data and generate next steps based on the purpose and goals of the meeting, such as "encourage discussion on which specific markets to focus on."
[0388] 5. Presenting Progress Actions
[0389] The server transmits the generated progress action to the terminal.
[0390] The terminal presents this ongoing action to the user in real time.
[0391] Specific examples
[0392] For example, suppose a user wants to hold a meeting titled "New Product Launch Plan." In this case, the user enters the following information into the web conferencing tool's interface:
[0393] Meeting purpose: "New product launch plan"
[0394] Goal: "Formulate market strategies and assess risks"
[0395] Audio file: "audio_new_product_meeting.wav"
[0396] This information is sent from the device to the server. The server transcribes the audio file and sends it to the generative AI model for analysis. As a result, the generative AI model generates a progress action such as "Please propose specific measures for market strategy." This action is sent to the device and displayed to the user through the device.
[0397] This system allows users to efficiently organize what is being said during a meeting and receive appropriate proceeding actions in real time, thereby improving the efficiency and effectiveness of meetings.
[0398] The processing flow will be explained below.
[0399] Step 1:
[0400] The user inputs the purpose of the meeting (e.g., "Planning the next project") and goal (e.g., "Define main tasks and determine who is responsible") into the interface of the web conferencing tool. The user also inputs the path to the audio file for the meeting (e.g., "audio.wav").
[0401] Step 2:
[0402] The terminal transmits the purpose and goal of the meeting and the path of the audio file input by the user to the server.
[0403] Step 3:
[0404] The server loads the audio file based on the path of the audio file received, and transcribes the audio data in real time using a speech recognition library (e.g., Google Speech Recognition API).
[0405] Step 4:
[0406] The server sends the transcribed text data to the generative AI model, which analyzes the transcribed data and the purpose and goals of the meeting to generate the next action to take.
[0407] Step 5:
[0408] The server transmits the generated progress action to the terminal.
[0409] Step 6:
[0410] The terminal displays the progress actions received from the server to the user in real time, and the user proceeds with the conference based on the presented progress actions.
[0411] Step 7:
[0412] After the conference ends, the server stores the transcript log of the entire conference and the history of the ongoing actions generated.
[0413] Step 8:
[0414] After the meeting ends, users can download the transcription log and progress action history stored on the server via their device and use them for summarizing and providing feedback.
[0415] Example 1
[0416] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0417] Modern meetings are increasingly being held online, especially with the spread of remote work. However, it remains difficult to efficiently understand what participants are saying during a meeting and decide on the appropriate next steps. As a result, the meeting often stalls and delays the achievement of its objectives and goals. Maintaining meeting efficiency and effectiveness becomes even more difficult when there are many participants or a wide range of information.
[0418] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0419] In this invention, the server includes a means for a user to set a purpose and goal of a meeting, a means for a user to upload audio data of the meeting, a means for transcribing the audio data received from the target in real time, a means for analyzing the transcribed text data by a generative AI model and generating the next action to proceed with the meeting, and a means for presenting the generated action to proceed with the meeting to the user in real time, thereby enabling real-time analysis of utterances during a meeting and the presentation of action to proceed with the meeting.
[0420] "User" refers to a person who uses this system to set the purpose and goal of a meeting and upload audio data.
[0421] The "server" is a device that transcribes voice data received from users in real time, analyzes the text using a generative AI model, and generates actions to progress the meeting.
[0422] A "terminal" is a device that allows a user to input the purpose and goal of a meeting and transmit audio files to a server, and is also a device that presents the generated progress actions to the user.
[0423] "Audio data" refers to a digital data file that records what participants say during a meeting.
[0424] "Real-time" refers to processing and displaying data immediately, with minimal time delay.
[0425] "Transcription" refers to the process of converting audio data into text form.
[0426] "Generative AI model" refers to the artificial intelligence technology used to analyze transcribed text data and generate actions that help guide meetings.
[0427] "Progression actions" refer to the next actions or discussion content suggested by the generative AI model based on the purpose and goals of the meeting.
[0428] "Upload" refers to the act of transferring data from a terminal to a server.
[0429] "Analysis" refers to the process by which generative AI models process transcribed text data, extract information, and generate appropriate actions.
[0430] This invention is a system that improves the efficiency of meetings by using a generative AI model that works in conjunction with a web conferencing tool. This system operates using multiple means, primarily servers, terminals, and users. The detailed configuration and operation of each means are explained below.
[0431] First, the user sets the purpose and goals of the meeting through the interface of the web conferencing tool. A dedicated form is provided for entering this information. For example, the purpose of the meeting can be "Planning the next project" and the goal can be "Define the main tasks and determine who will be responsible." In addition, a field is provided for uploading audio data from the meeting (e.g., "audio.wav"). The user selects the required audio file and clicks the "Send" button.
[0432] Next, the terminal has a means for sending the meeting purpose, goal, and audio file path entered by the user together to the server. This information is packetized and sent to the server as an HTTP request. This generally uses the HTTPS protocol.
[0433] The server has a means to transcribe the received voice data in real time. Specifically, it uses a speech recognition library such as the Google Speech Recognition API to convert the voice data into text data. To do this, it processes the received audio file (e.g., "audio.wav") and passes it to the speech recognition library. The speech recognition library converts the voice into text data and returns the result to the server, which temporarily stores it.
[0434] The server then sends the transcribed text data to a generative AI model for analysis. The generative AI model uses an interactive generative model (e.g., GPT-3.5). Based on the transcribed text data, this model generates the next progress action appropriate to the purpose and goal of the meeting. For example, a progress action such as "Please propose specific measures for market strategy" may be generated. This generated progress action is temporarily stored on the server.
[0435] Finally, the generated progress actions are sent from the server to the terminal and presented to the user in real time. The terminal displays the progress actions received from the server on the screen interface. For example, the action display area of the interface may display "Please propose specific measures for the market strategy."
[0436] Specific examples
[0437] For example, if a user wants to hold a meeting titled "New Product Launch Plan," they would enter the following information into the web conferencing tool's interface:
[0438] Meeting purpose: "New product launch plan"
[0439] Goal: "Formulate market strategies and assess risks"
[0440] Audio file: "audio_new_product_meeting.wav"
[0441] This information is sent from the device to the server. The server transcribes the audio file using the Google Speech Recognition API and passes it to the generative AI model for analysis. As a result, the generative AI model generates a progress action such as "Please propose specific measures for our market strategy." This action is sent to the device and displayed to the user in real time.
[0442] Example prompt sentence:
[0443] "The purpose of this meeting is to plan the market launch of a new product. We want to develop a market strategy and assess risks. Based on this objective, what are the next steps we should take?"
[0444] This system allows users to efficiently organize what is being said during a meeting and receive appropriate proceeding actions in real time, improving the efficiency and effectiveness of meetings.
[0445] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0446] Step 1: User Input
[0447] The user accesses the interface of the web conferencing tool and inputs the purpose and goals of the meeting. For example, they might set "Planning the next project" and "Define main tasks and determine who will be responsible." They also specify the path to the audio file for the meeting (e.g., "audio.wav"). After inputting this information, the user clicks the "Send" button.
[0448] Input: Meeting purpose, goal, audio file path
[0449] Output: Configuration information passed to the terminal
[0450] Step 2: Submit your information
[0451] The terminal packetizes the user-entered meeting purpose, goal, and audio file path, and sends the packet as an HTTP request to the server using the HTTPS protocol.
[0452] Input: Information entered by the user (meeting purpose, goal, audio file path)
[0453] Output: HTTP request sent to the server
[0454] Step 3: Receiving and transcribing audio data
[0455] The server receives the HTTP request and gets the path to the audio file. The server loads this audio file and transcribes it in real time using a speech recognition library such as the Google Speech Recognition API. This process converts the audio data into text data.
[0456] Input: The path to the audio file sent to the server
[0457] Output: Transcribed text data
[0458] Step 4: Analyzing the text data
[0459] The server temporarily stores the transcribed text data and then passes it to a generative AI model. The generative AI model analyzes the data and generates the next action based on the purpose and goals of the meeting. For example, it might generate an instruction such as, "Please propose specific measures for market strategy."
[0460] Input: Transcribed text data
[0461] Output: Generated progress actions
[0462] Step 5: Present progress actions
[0463] The server packets the generated progress actions and sends them to the terminal. The terminal receives the progress actions and displays them to the user in real time. For example, the action display area of the interface may display "Please propose specific measures for market strategy."
[0464] Input: Generated progress action
[0465] Output: Progression actions presented to the user
[0466] (Application example 1)
[0467] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0468] Modern conference systems and driver assistance systems require technologies that can streamline decision-making for participants and drivers and suggest appropriate actions in real time. However, conventional systems have issues such as not being able to transcribe speech in real time, not being able to fully utilize generative AI models, and not being able to quickly suggest appropriate actions. This can sometimes hinder smooth progress in meetings and driving assistance, resulting in issues of reduced efficiency.
[0469] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0470] In this invention, the server includes a means for transcribing the speech of participants and the driver in real time, a means for analyzing the transcribed text using a generative AI model and generating the next action for proceeding with the meeting or driving assistance action, and a means for presenting the generated action for proceeding with the meeting or driving assistance action to the participants and the driver in real time. This enables quick and appropriate decision-making during meetings and driving, and enables efficient progress and assistance.
[0471] A "web conferencing tool" is software that allows multiple participants in remote locations to hold meetings by sharing audio and video over the Internet.
[0472] A "generative AI model" is an artificial intelligence algorithm that generates output, such as text or actions, based on input data.
[0473] The "meeting efficiency system" is a system that transcribes and analyzes audio to generate and suggest the next action in order to make meetings proceed smoothly and effectively.
[0474] A "driving assistance system" is a system that assists vehicle drivers by transcribing and analyzing voice and providing appropriate instructions and information to the driver in real time.
[0475] "Participant" or "Driver" refers to a person attending a meeting or driving a vehicle.
[0476] "Real time" means that processing is carried out almost immediately after an event occurs.
[0477] "Transcription" is the process of analyzing audio data and converting its content into text format.
[0478] "Means" refer to the methods or tools used to achieve a particular goal.
[0479] "Next Actions" are specific actions or proposals to be taken next, generated based on the purpose and progress of the meeting.
[0480] "Driving assistance actions" are instructions or information provided to the driver in real time to assist driving.
[0481] A "cloud-based system" is a system in which data is processed and stored on a remote server via the Internet.
[0482] "User feedback" refers to opinions and evaluation data provided by users of the system, and is used to improve the system.
[0483] The present invention is an efficiency improvement system using a generative AI model that can be applied to web conferencing tools and driving assistance systems for autonomous vehicles. This system has the function of transcribing the voices of participants and drivers in real time, analyzing the text data, and generating the next progress action or driving assistance action. In addition, by presenting the generated actions to participants and drivers in real time, the progress of the meeting or driving can be made smoother. Specific embodiments for implementing the present invention are described below.
[0484] Key Components Used
[0485] 1. Server:
[0486] The audio data is transcribed and the text data is analyzed using a generative AI model.
[0487] Use a speech recognition library (e.g., Google Speech Recognition API) for transcription.
[0488] 2. Terminal:
[0489] It receives input from the user and sends it to the server.
[0490] The generated progress actions and driving assistance actions are presented to the user.
[0491] 3. User:
[0492] Set the purpose or goal of the meeting or provide driving context.
[0493] Proceed with meetings or driving based on the generated actions.
[0494] Program processing explanation
[0495] The server transcribes speech from participants and drivers in real time. To do this, it uses a speech recognition library to convert the speech data into text format. It then uses a generative AI model to analyze the transcribed text data and generate next actions for the meeting or driver assistance actions. This generative AI model processes the data in real time on a cloud-based system.
[0496] The generated actions are sent from the server to the terminal, and the terminal presents these actions to the user in real time. The user then takes the next action based on the presented actions.
[0497] Specific examples
[0498] For example, if a user wants to hold a meeting titled "New Product Launch Plan," the user enters the following information into the web conferencing tool's interface:
[0499] Meeting purpose: "New product launch plan"
[0500] Goal: "Formulate market strategies and assess risks"
[0501] Audio file path: "audio_new_product_meeting.wav"
[0502] This information is sent from the device to the server. The server transcribes the audio file and sends it to the generative AI model for analysis. As a result, the generative AI model generates a progress action such as "Please propose specific measures for market strategy." This action is sent to the device and displayed to the user through the device.
[0503] In a driver assistance system, if a driver wants to know the location of the next exit while driving, they can utter the voice command "Tell me where the next exit is." This speech is transcribed in real time and analyzed by a generative AI model. The server generates a driving assistance action such as "Please use the exit on the right 2 km ahead," and sends it to the device. This action is displayed on the driver's smart glasses or head-mounted display.
[0504] Example prompt sentence:
[0505] You are a driving assistance bot. Please interpret the following voice commands and suggest appropriate actions.
[0506] Voice command: What's the next entrance / exit?
[0507] In this way, the system allows users to efficiently organize what is said during meetings or while driving and receive appropriate proceeding actions and driving assistance actions in real time, thereby improving meeting efficiency and driving safety.
[0508] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0509] Step 1:
[0510] User input
[0511] The user uses the interface of the web conferencing tool or the voice input device of the driving assistance system to input the purpose and goal of the meeting, as well as the path of the audio file and voice commands while driving, including input in text format or voice input, which results in specific data.
[0512] Input: purpose of the meeting, goal, audio file path, voice command
[0513] Output: User input data (text data and audio data)
[0514] Step 2:
[0515] Sending information
[0516] The device sends the entered meeting purpose, goal, audio file path, and voice command to the server, which then sends the data over the network to a cloud-based server for processing.
[0517] Input: User-entered data
[0518] Output: Data sent to the server
[0519] Step 3:
[0520] Analysis of audio data
[0521] The server transcribes the received audio file using a speech recognition library (such as the Google Speech Recognition API). The audio data is analyzed and converted into text. This transcribed text data is then sent to the generative AI model.
[0522] Input: Audio data
[0523] Output: Transcribed text data
[0524] Step 4:
[0525] Analysis using generative AI models
[0526] The server analyzes the transcribed text data using a generative AI model and generates the next progress action or driving assistance action based on the prompt sentence. The generative AI model analyzes the text data using a specific algorithm and suggests the optimal action.
[0527] Input: Transcribed text data
[0528] Output: Generated progress actions and driver assistance actions
[0529] Step 5:
[0530] Suggesting Actions
[0531] The server sends the generated progress actions and driving assistance actions to the device, which then presents these actions to the user in real time: if in a meeting, they are displayed in the meeting interface, and if driving, they are displayed to the driver through smart glasses or a head-mounted display.
[0532] Input: Generated progress actions and driver assistance actions
[0533] Output: Real-time suggested actions
[0534] Step 6:
[0535] User Feedback
[0536] The user continues to conduct a meeting or drive based on the generated actions, and user feedback provides the system with data for improvement, allowing the accuracy of the generative AI model to be continuously improved.
[0537] Input: User feedback
[0538] Output: Improve the accuracy of the generative AI model.
[0539] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0540] The present invention is a system that uses a generative AI model that works with a web conferencing tool to improve the efficiency of meetings and, by combining it with an emotion engine, makes the meeting proceed more effectively. This system has the function of transcribing participants' voices in real time and analyzing the text data to generate meeting proceeding actions. Furthermore, by presenting the generated proceeding actions to participants in real time and combining it with an emotion engine that recognizes the user's emotions, the meeting can proceed smoothly. A specific embodiment of this system is described below.
[0541] Key Components Used
[0542] 1. Server: Transcribing voice data, analyzing text data using a generative AI model, and recognizing user emotions using an emotion engine.
[0543] 2. Terminal: Receives input from the user, sends it to the server, and provides feedback to the user based on the generated progress actions and emotions.
[0544] 3. User: Sets the purpose and goals of the meeting, provides emotional data, and advances the meeting based on progress actions.
[0545] 4. Emotion engine: Analyzes the user's voice and facial expressions and sends the emotion data to the generative AI model.
[0546] Program processing explanation
[0547] 1. User Input
[0548] The user sets the purpose and goals of the meeting through the interface of the web conferencing tool. For example, they input "Plan the next project" and "Define the main tasks and determine who will be responsible." The user also inputs the path to the meeting audio file (e.g., "audio.wav").
[0549] 2. Transmission of information
[0550] The terminal transmits the purpose and goal of the meeting and the path of the audio file input by the user to the server.
[0551] 3. Analysis of audio data
[0552] The server loads the audio file based on the path of the audio file received, and transcribes the audio data in real time using a speech recognition library (e.g., Google Speech Recognition API).
[0553] 4. Acquiring Emotion Data
[0554] The server uses an emotion engine to analyze the user's emotional data from the audio and video data, based on the tone of voice, facial expressions, and content of speech.
[0555] 5. Proceeding Action Generation
[0556] The server sends the transcribed text data and emotion data to the generative AI model. The generative AI model analyzes the transcribed data and emotion data and generates the next action to take. For example, an action such as "If the user is nervous, suggest questions to help them relax" may be generated.
[0557] 6. Presenting Progressive Actions and Emotional Feedback
[0558] The server transmits feedback based on the generated progress actions and emotions to the terminal.
[0559] The device displays this information to the user in real time, and the user can proceed with the meeting based on the presented actions and emotional feedback.
[0560] Specific examples
[0561] For example, suppose a user wants to hold a meeting titled "New Product Launch Plan." In this case, the user enters the following information into the web conferencing tool's interface:
[0562] Meeting purpose: "New product launch plan"
[0563] Goal: "Formulate market strategies and assess risks"
[0564] Audio file: "audio_new_product_meeting.wav"
[0565] This information is sent from the device to the server. The server transcribes the audio file and analyzes the user's emotions using an emotion engine. The generative AI model then analyzes this data and generates a progress action, such as "Please propose specific measures for market strategy." If the user is nervous, it will also suggest questions to relax them (e.g., "Tell me about a recent success story"). This action and emotional feedback are sent to the device and displayed to the user.
[0566] This system allows users to efficiently organize statements and emotional data during meetings and receive appropriate proceeding actions in real time, further improving the efficiency and effectiveness of meetings.
[0567] The processing flow will be explained below.
[0568] Step 1:
[0569] The user inputs the purpose of the meeting (e.g., "Plan the next project") and goal (e.g., "Define main tasks and determine who is responsible") into the interface of the web conferencing tool. The user also inputs the path to the audio file for the meeting (e.g., "audio.wav").
[0570] Step 2:
[0571] The terminal transmits the purpose and goal of the meeting and the path of the audio file input by the user to the server.
[0572] Step 3:
[0573] The server loads the audio file based on the path of the audio file received, and transcribes the audio data in real time using a speech recognition library (e.g., Google Speech Recognition API).
[0574] Step 4:
[0575] The server sends the transcribed text data to the generative AI model, which analyzes the transcribed data and the purpose and goals of the meeting to generate the next action to take.
[0576] Step 5:
[0577] The server captures the user's voice and facial expressions, and uses an emotion engine to analyze the user's emotional data. Emotions are recognized based on the tone of voice, changes in facial expressions, and the content of speech.
[0578] Step 6:
[0579] The server inputs the emotion data into the generative AI model and takes it into consideration when generating the next action to proceed, such as "suggesting questions to help a nervous user relax."
[0580] Step 7:
[0581] The server sends the generated progress actions and emotional feedback to the device.
[0582] Step 8:
[0583] The device receives the progress actions and emotional feedback from the server and presents them to the user in real time. The user then proceeds with the conference based on the progress actions and emotional feedback presented to them.
[0584] Example 2
[0585] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0586] In today's business environment, there is a demand for improving the efficiency and effectiveness of meetings. However, it is difficult to instantly analyze participants' emotions and comments and provide appropriate proceeding actions in real time. Furthermore, conventional conferencing systems are unable to quickly respond to changes in emotions, which can hinder the smooth progress of the meeting. Therefore, there is a need for a method that can analyze participants' emotions in real time and optimize the progress of the meeting based on that data.
[0587] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0588] In this invention, the server includes a means for a user to set the purpose and goal of the meeting, a means for transcribing the participants' voices in real time, a means for analyzing the transcribed text and the users' emotional data using a generative AI model to generate the next action to proceed with the meeting, and a means for presenting feedback to the participants in real time based on the generated action to proceed with the meeting and their emotions. This increases the efficiency and effectiveness of the meeting, and enables the meeting to proceed smoothly by quickly responding to changes in the participants' emotions.
[0589] "Means for users to set the purpose and goals of a meeting" means allowing users to input the main purpose and goals of a meeting using the interface of the web conferencing tool.
[0590] "Means for transcribing participants' voices in real time" refers to means for converting the speech of conference participants into text data in real time using voice recognition technology.
[0591] A "generative AI model" refers to an algorithm or program that uses natural language processing and machine learning to generate meaningful information from input data.
[0592] The "means for analyzing the transcribed text and user emotion data and generating the next action to proceed with the meeting" refers to a means for generating instructions and suggestions for determining the progress of the meeting based on the transcribed text data and emotion analysis data.
[0593] An "emotion engine" refers to technology or algorithms that analyze voice tone, speech rate, facial expression changes, etc. to identify the emotional state of users or participants.
[0594] The "means for presenting participants with feedback based on generated meeting progress actions and emotions in real time" is a means for instantly providing analyzed and generated information to meeting participants to guide the progress of the meeting.
[0595] "Cloud-based processing" refers to the technology of processing data on remote servers via the Internet, and is used to efficiently perform large-scale data analysis and real-time processing.
[0596] "Means for incorporating user feedback and continuously improving the accuracy of the generative AI model" refers to means for improving the performance and accuracy of the generative AI model over time by incorporating user opinions and evaluations into the model.
[0597] This invention relates to a system that improves the efficiency and effectiveness of meetings by utilizing a generative AI model that works in conjunction with a web conferencing tool. Its key feature is that it transcribes participants' speech in real time, analyzes the text data to generate meeting progress actions, and provides feedback to participants, thereby helping to ensure smooth progress of the meeting. Furthermore, by combining it with an emotion engine, it is possible to analyze participants' emotions and reflect them in the progress actions.
[0598] Key Components Used
[0599] The following describes the major components for implementing the present invention.
[0600] server
[0601] The server is the core of this system and has the following main functions:
[0602] 1. Ability to transcribe voice data in real time: Uses a speech recognition library (e.g., Google Speech Recognition API) to convert participants' voice data into text data.
[0603] 2. The ability to analyze text data and emotional data using a generative AI model and generate next steps: The generative AI model uses natural language processing and machine learning based on the input data to derive the appropriate next action.
[0604] 3. Ability to analyze emotional data using an emotion engine: Identify the user's emotional state from voice tone, speaking rate, facial expression changes, etc.
[0605] Terminal
[0606] The terminal is a device that provides a user interface and has the following functions:
[0607] 1. Function to receive input from the user and send it to the server: Provides an interface for inputting the purpose and goals of the meeting, the path of the audio file, etc.
[0608] 2. A function to present feedback to the user based on the generated progress actions and emotions: Feedback sent from the server is displayed in real time.
[0609] User
[0610] The user sets the purpose and goal of the meeting and is responsible for running the meeting. The user also provides data on comments and emotions made during the meeting, and the system guides the meeting according to the actions taken by the system.
[0611] Specific examples
[0612] For example, consider a user who wants to hold a meeting titled "New Product Launch Plan." The user enters the following information into the web conferencing tool's interface:
[0613] Meeting purpose: "New product launch plan"
[0614] Goal: "Formulate market strategies and assess risks"
[0615] Audio file: "audio_new_product_meeting.wav"
[0616] Prompt Sentence Examples
[0617] "Planning the next project"
[0618] "Define key tasks and determine who is responsible"
[0619] "Marketing Strategy Meeting"
[0620] "Target Market Definition and Competitive Analysis"
[0621] The device sends this information to a server. The server reads the audio file and transcribes it using a speech recognition library. It then uses an emotion engine to analyze the user's emotions from the acquired transcription and video data. The analysis results are sent to a generative AI model, which generates the next action to proceed.
[0622] For example, if a user is nervous, the generative AI model will generate a progression action such as "suggesting questions to relax the user." It may also suggest questions to dig deeper into marketing strategies. The generated progression actions and feedback are sent to the device and displayed to the user in real time.
[0623] This allows users to smoothly progress with the conference, and improves the efficiency and effectiveness of the conference.
[0624] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0625] Step 1:
[0626] The user sets the purpose and goal of the meeting through the interface of the web conferencing tool and inputs the path of the audio file. As input, the user provides the purpose of the meeting (e.g., "Plan for the market launch of a new product"), the goal (e.g., "Formulate market strategy and assess risks"), and the path of the audio file (e.g., "audio_new_product_meeting.wav"). This information is saved on the device for the next step.
[0627] Step 2:
[0628] The terminal sends the information input by the user to the server. The input information includes the purpose and goal of the meeting, and the path to the audio file. The terminal appropriately packets this information and sends it to the server. The output is the meeting information sent to the server.
[0629] Step 3:
[0630] The server receives the audio file path and loads it. Then, it uses a speech recognition library (e.g., Google Speech Recognition API) to transcribe the audio data in real time. The input is the audio file path "audio_new_product_meeting.wav" and the output is the transcribed text data (e.g., "The discussion about the new product has begun...").
[0631] Step 4:
[0632] The server inputs the acquired transcription data and video data into the emotion engine. The emotion engine analyzes the tone of voice, speech rate, and facial expression changes to generate the user's emotion data. The inputs are transcription data and video data, and the output is the analyzed emotion data (e.g., "The user is nervous").
[0633] Step 5:
[0634] The server sends the transcribed text data and emotion data to the generative AI model, which generates the next progress action based on this data. The input is the transcribed text and emotion data, and the output is the generated progress action (e.g., "Suggest questions to help the user relax").
[0635] Step 6:
[0636] The server transmits the generated progress action and emotion-based feedback to the terminal. The input is the progress action and emotion feedback, and the output is the data transmitted to the terminal.
[0637] Step 7:
[0638] The device displays the progress actions and emotional feedback received from the server to the user. Specifically, they are displayed as notifications or pop-ups on the interface and provided to the user in real time. The input is the progress actions and emotional feedback from the server, and the output is the information displayed to the user.
[0639] (Application example 2)
[0640] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0641] With conventional web conferencing tools, meeting efficiency and progress management are largely manual, which means that it takes time to organize what is being said and decide on actions to take. There are also concerns that the quality of meetings may decline due to the difficulty of properly understanding participants' emotions and smoothly progressing the meeting. Furthermore, when dealing with customers in virtual stores, it is difficult to suggest appropriate actions to take in response to customer questions, making it difficult to increase customer satisfaction.
[0642] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0643] In this invention, the server includes means for transcribing the participants' voices in real time, means for analyzing the transcribed text using a generative AI model and generating the next action to proceed with the conference, means for presenting the generated action to proceed with the conference to the participants in real time, means for analyzing the emotions of the participants using an emotion engine and providing appropriate feedback based on the analysis results, means for monitoring the customer service process in the virtual store and generating appropriate actions in response to customer questions, and means for presenting the generated response actions and emotion feedback to the customers in real time. This makes it possible to efficiently manage the progress of the conference and provide optimal actions in response to the emotions of the participants, as well as improve the quality of customer service in the virtual store and increase customer satisfaction.
[0644] "Web conferencing tool" refers to software or a platform that allows multiple participants to hold real-time conferences using audio and video over the Internet.
[0645] A "generative AI model" refers to an algorithm or system that uses artificial intelligence to analyze input data and automatically generate appropriate responses or actions.
[0646] A "meeting efficiency system" refers to a system that combines technologies and means to support the progress and management of meetings and improve their efficiency.
[0647] "Real-time transcription means" refers to technology or equipment for instantly converting audio data into text data.
[0648] "Progressive actions" refer to specific actions or instructions that should be taken next in the progress of a meeting.
[0649] An "emotion engine" refers to technology and algorithms that analyze data such as participants' voices and facial expressions to estimate their emotional state.
[0650] "Feedback" refers to information or instructions provided to users or participants based on analysis results, progress actions, etc.
[0651] "Virtual store" refers to a virtual store that offers products and services via the Internet.
[0652] A "customer service process" refers to a set of procedures or methods for appropriately responding to customer inquiries and questions.
[0653] This invention is a system that improves the efficiency of meetings by using a generative AI model that works in conjunction with a web conferencing tool, and further combines it with an emotion engine to more effectively progress the meeting. This system has the function of transcribing participants' voices in real time, analyzing the text data, and generating the next progress action. Furthermore, by presenting the generated progress action to participants in real time and combining it with an emotion engine that recognizes the user's emotions, the meeting can be progressed smoothly.
[0654] Key Components Used
[0655] 1. Server: Transcribing voice data, analyzing text data using a generative AI model, and recognizing user emotions using an emotion engine. Specifically, it uses a speech recognition library (e.g., Google Speech Recognition API) and an emotion analysis engine (e.g., EmotionEngine). It also generates progress actions using a generative AI model (e.g., ActionsGenerator).
[0656] 2. Device: Receives input from the user and sends it to the server. It also provides feedback to the user based on the generated progress actions and emotions. Devices can be smartphones, head-mounted displays, smart glasses, etc.
[0657] 3. User: Sets the purpose and goals of the meeting, provides emotional data and advances the meeting based on progress actions.
[0658] 4. Emotion engine: Analyzes the user's voice and facial expressions and sends the emotional data to a generative AI model.
[0659] Program processing explanation
[0660] 1. The server transcribes the user's voice data using the Google Speech Recognition API. This process converts the voice data into text data.
[0661] 2. The emotion engine then analyzes the voice tone and facial expressions to generate the user's emotion data. This is done using the EmotionEngine.
[0662] 3. The server uses a generative AI model to analyze the transcribed text data and emotion data, and generates the next action to proceed. Using ActionsGenerator, the generated action is provided in real time.
[0663] 4. The device presents the progress actions and emotional feedback received from the server to the user in real time. This information is displayed via a smartphone, head-mounted display, or smart glasses.
[0664] Specific examples
[0665] For example, suppose a user wants to hold a meeting titled "New Product Launch Plan." In this case, the user enters the following information into the web conferencing tool's interface:
[0666] Meeting purpose: "New product launch plan"
[0667] Goal: "Formulate market strategies and assess risks"
[0668] Audio file: "audio_new_product_meeting.wav"
[0669] This information is sent from the device to the server. The server transcribes the audio file and analyzes the user's emotions using an emotion engine. The generative AI model then analyzes this data and generates a progress action, such as "Please propose specific measures for market strategy." If the user is nervous, it will also suggest questions to relax them (e.g., "Tell me about a recent success story"). This action and emotional feedback are sent to the device and displayed to the user.
[0670] Prompt Sentence Examples
[0671] Below are some example prompts to use with generative AI models:
[0672] text
[0673] Voice input: "I want to know about my new TV."
[0674] Emotional data: interesting but a bit confusing
[0675] Generate Action: Describe the features of their new TV. Ask about their specific needs.
[0676] This system efficiently organizes statements and emotional data during meetings and allows appropriate actions to be taken in real time, further improving the efficiency and effectiveness of meetings. It also makes it possible to generate and provide appropriate actions in real time when dealing with customers in virtual stores, thereby increasing customer satisfaction.
[0677] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0678] Step 1:
[0679] The server receives the user's voice data. The user inputs the purpose, goal, and audio file of the meeting into the terminal. The terminal sends this information to the server. Specific inputs include the audio file (e.g., "audio_new_product_meeting.wav"), the purpose of the meeting, "a plan to launch a new product," and the goal, "formulation of a market strategy and risk assessment." The server receives this data and begins processing.
[0680] Step 2:
[0681] The server transcribes the received voice data in real time using the Google Speech Recognition API. Text data (output) is generated from the voice data (input). This transcribed text data serves as the basis for generating actions to progress the meeting. Converting the input voice data into text makes subsequent data analysis easier.
[0682] Step 3:
[0683] The server uses an emotion engine to analyze the user's emotional data along with the transcribed voice data. Emotion analysis is performed using the Emotion Engine. Specifically, it analyzes changes in voice tone and facial expressions. Through this process, the server obtains emotional data (output) from the voice data (input). For example, it determines whether the user is nervous, relaxed, or interested. At this stage, data enrichment is performed.
[0684] Step 4:
[0685] The server uses a generative AI model (e.g., ActionsGenerator) to analyze the transcribed text data and emotional data. Specifically, the generative AI model generates the next action to proceed (output) based on the text data (input) and emotional data (input). For example, an action such as "Please propose specific measures for market strategy" may be generated. This action provides specific instructions to ensure the smooth flow of the meeting.
[0686] Step 5:
[0687] The server sends the generated progress action and emotional feedback to the device. Specifically, the generated action (output) and emotional feedback (output) are sent. For example, the action may be "Please propose specific measures for the market strategy" and the emotional feedback may be "The user is nervous." This allows the device to process the next step.
[0688] Step 6:
[0689] The device presents the progress actions and emotional feedback received from the server to the user in real time. Specifically, information is displayed to the user (output) via a smartphone, head-mounted display, or smart glasses. For example, the device may provide the user with instructions such as "Please propose specific measures for market strategy" or feedback such as "The user seems nervous, so we suggest questions to help them relax." This allows the user to effectively proceed with the next action in the meeting.
[0690] Step 7:
[0691] The user controls the meeting based on the progress actions and emotional feedback presented on the device. By taking specific actions, the user can effectively move forward toward the purpose and goal of the meeting. The user uses the presented information to execute the next step instructions. This is the stage at which the actual meeting progresses.
[0692] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0693] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0694] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0695] [Third embodiment]
[0696] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0697] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[0698] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0699] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0700] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0701] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0702] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0703] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0704] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0705] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0706] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0707] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0708] The present invention is a system that improves the efficiency of meetings by using a generative AI model that works in conjunction with a web conferencing tool. This system has the function of transcribing participants' voices in real time, analyzing the text data, and generating actions to progress the meeting. In addition, by presenting the generated actions to participants in real time, the progress of the meeting can be made smoother. A specific embodiment of this system is described below.
[0709] Key Components Used
[0710] 1. Server: Transcribing audio data and analyzing text data using a generative AI model.
[0711] 2. Terminal: Receives input from the user, sends it to the server, and presents the generated progress actions to the user.
[0712] 3. User: Set the purpose and goals of the meeting and proceed with the meeting based on progress actions.
[0713] Program processing explanation
[0714] 1. User Input
[0715] Users set the purpose and goals of the meeting through the interface of the web conferencing tool, for example, entering "planning the next project" and "defining the main tasks and deciding who will be responsible."
[0716] The user also enters the path to the conference audio file (eg "audio.wav").
[0717] 2. Transmission of information
[0718] The terminal transmits the purpose and goal of the meeting and the path of the audio file input by the user to the server.
[0719] 3. Analysis of audio data
[0720] The server transcribes the received audio file in real time, using a speech recognition library (e.g., Google Speech Recognition API).
[0721] Transcribed text data is generated and fed into a generative AI model.
[0722] 4. Proceeding Action Generation
[0723] The server uses a generative AI model to analyze the transcribed text data and generate next steps based on the purpose and goals of the meeting, such as "encourage discussion on which specific markets to focus on."
[0724] 5. Presenting Progress Actions
[0725] The server transmits the generated progress action to the terminal.
[0726] The terminal presents this ongoing action to the user in real time.
[0727] Specific examples
[0728] For example, suppose a user wants to hold a meeting titled "New Product Launch Plan." In this case, the user enters the following information into the web conferencing tool's interface:
[0729] Meeting purpose: "New product launch plan"
[0730] Goal: "Formulate market strategies and assess risks"
[0731] Audio file: "audio_new_product_meeting.wav"
[0732] This information is sent from the device to the server. The server transcribes the audio file and sends it to the generative AI model for analysis. As a result, the generative AI model generates a progress action such as "Please propose specific measures for market strategy." This action is sent to the device and displayed to the user through the device.
[0733] This system allows users to efficiently organize what is being said during a meeting and receive appropriate proceeding actions in real time, thereby improving the efficiency and effectiveness of meetings.
[0734] The processing flow will be explained below.
[0735] Step 1:
[0736] The user inputs the purpose of the meeting (e.g., "Planning the next project") and goal (e.g., "Define main tasks and determine who is responsible") into the interface of the web conferencing tool. The user also inputs the path to the audio file for the meeting (e.g., "audio.wav").
[0737] Step 2:
[0738] The terminal transmits the purpose and goal of the meeting and the path of the audio file input by the user to the server.
[0739] Step 3:
[0740] The server loads the audio file based on the path of the audio file received, and transcribes the audio data in real time using a speech recognition library (e.g., Google Speech Recognition API).
[0741] Step 4:
[0742] The server sends the transcribed text data to the generative AI model, which analyzes the transcribed data and the purpose and goals of the meeting to generate the next action to take.
[0743] Step 5:
[0744] The server transmits the generated progress action to the terminal.
[0745] Step 6:
[0746] The terminal displays the progress actions received from the server to the user in real time, and the user proceeds with the conference based on the presented progress actions.
[0747] Step 7:
[0748] After the conference ends, the server stores the transcript log of the entire conference and the history of the ongoing actions generated.
[0749] Step 8:
[0750] After the meeting ends, users can download the transcription log and progress action history stored on the server via their device and use them for summarizing and providing feedback.
[0751] Example 1
[0752] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0753] Modern meetings are increasingly being held online, especially with the spread of remote work. However, it remains difficult to efficiently understand what participants are saying during a meeting and decide on the appropriate next steps. As a result, the meeting often stalls and delays the achievement of its objectives and goals. Maintaining meeting efficiency and effectiveness becomes even more difficult when there are many participants or a wide range of information.
[0754] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0755] In this invention, the server includes a means for a user to set a purpose and goal of a meeting, a means for a user to upload audio data of the meeting, a means for transcribing the audio data received from the target in real time, a means for analyzing the transcribed text data by a generative AI model and generating the next action to proceed with the meeting, and a means for presenting the generated action to proceed with the meeting to the user in real time, thereby enabling real-time analysis of utterances during a meeting and the presentation of action to proceed with the meeting.
[0756] "User" refers to a person who uses this system to set the purpose and goal of a meeting and upload audio data.
[0757] The "server" is a device that transcribes voice data received from users in real time, analyzes the text using a generative AI model, and generates actions to progress the meeting.
[0758] A "terminal" is a device that allows a user to input the purpose and goal of a meeting and transmit audio files to a server, and is also a device that presents the generated progress actions to the user.
[0759] "Audio data" refers to a digital data file that records what participants say during a meeting.
[0760] "Real-time" refers to processing and displaying data immediately, with minimal time delay.
[0761] "Transcription" refers to the process of converting audio data into text form.
[0762] "Generative AI model" refers to the artificial intelligence technology used to analyze transcribed text data and generate actions that help guide meetings.
[0763] "Progression actions" refer to the next actions or discussion content suggested by the generative AI model based on the purpose and goals of the meeting.
[0764] "Upload" refers to the act of transferring data from a terminal to a server.
[0765] "Analysis" refers to the process by which generative AI models process transcribed text data, extract information, and generate appropriate actions.
[0766] This invention is a system that improves the efficiency of meetings by using a generative AI model that works in conjunction with a web conferencing tool. This system operates using multiple means, primarily servers, terminals, and users. The detailed configuration and operation of each means are explained below.
[0767] First, the user sets the purpose and goals of the meeting through the interface of the web conferencing tool. A dedicated form is provided for entering this information. For example, the purpose of the meeting can be "Planning the next project" and the goal can be "Define the main tasks and determine who will be responsible." In addition, a field is provided for uploading audio data from the meeting (e.g., "audio.wav"). The user selects the required audio file and clicks the "Send" button.
[0768] Next, the terminal has a means for sending the meeting purpose, goal, and audio file path entered by the user together to the server. This information is packetized and sent to the server as an HTTP request. This generally uses the HTTPS protocol.
[0769] The server has a means to transcribe the received voice data in real time. Specifically, it uses a speech recognition library such as the Google Speech Recognition API to convert the voice data into text data. To do this, it processes the received audio file (e.g., "audio.wav") and passes it to the speech recognition library. The speech recognition library converts the voice into text data and returns the result to the server, which temporarily stores it.
[0770] The server then sends the transcribed text data to a generative AI model for analysis. The generative AI model uses an interactive generative model (e.g., GPT-3.5). Based on the transcribed text data, this model generates the next progress action appropriate to the purpose and goal of the meeting. For example, a progress action such as "Please propose specific measures for market strategy" may be generated. This generated progress action is temporarily stored on the server.
[0771] Finally, the generated progress actions are sent from the server to the terminal and presented to the user in real time. The terminal displays the progress actions received from the server on the screen interface. For example, the action display area of the interface may display "Please propose specific measures for the market strategy."
[0772] Specific examples
[0773] For example, if a user wants to hold a meeting titled "New Product Launch Plan," they would enter the following information into the web conferencing tool's interface:
[0774] Meeting purpose: "New product launch plan"
[0775] Goal: "Formulate market strategies and assess risks"
[0776] Audio file: "audio_new_product_meeting.wav"
[0777] This information is sent from the device to the server. The server transcribes the audio file using the Google Speech Recognition API and passes it to the generative AI model for analysis. As a result, the generative AI model generates a progress action such as "Please propose specific measures for our market strategy." This action is sent to the device and displayed to the user in real time.
[0778] Example prompt sentence:
[0779] "The purpose of this meeting is to plan the market launch of a new product. We want to develop a market strategy and assess risks. Based on this objective, what are the next steps we should take?"
[0780] This system allows users to efficiently organize what is being said during a meeting and receive appropriate proceeding actions in real time, improving the efficiency and effectiveness of meetings.
[0781] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0782] Step 1: User Input
[0783] The user accesses the interface of the web conferencing tool and inputs the purpose and goals of the meeting. For example, they might set "Planning the next project" and "Define main tasks and determine who will be responsible." They also specify the path to the audio file for the meeting (e.g., "audio.wav"). After inputting this information, the user clicks the "Send" button.
[0784] Input: Meeting purpose, goal, audio file path
[0785] Output: Configuration information passed to the terminal
[0786] Step 2: Submit your information
[0787] The terminal packetizes the user-entered meeting purpose, goal, and audio file path, and sends the packet as an HTTP request to the server using the HTTPS protocol.
[0788] Input: Information entered by the user (meeting purpose, goal, audio file path)
[0789] Output: HTTP request sent to the server
[0790] Step 3: Receiving and transcribing audio data
[0791] The server receives the HTTP request and gets the path to the audio file. The server loads this audio file and transcribes it in real time using a speech recognition library such as the Google Speech Recognition API. This process converts the audio data into text data.
[0792] Input: The path to the audio file sent to the server
[0793] Output: Transcribed text data
[0794] Step 4: Analyzing the text data
[0795] The server temporarily stores the transcribed text data and then passes it to a generative AI model. The generative AI model analyzes the data and generates the next action based on the purpose and goals of the meeting. For example, it might generate an instruction such as, "Please propose specific measures for market strategy."
[0796] Input: Transcribed text data
[0797] Output: Generated progress actions
[0798] Step 5: Present progress actions
[0799] The server packets the generated progress actions and sends them to the terminal. The terminal receives the progress actions and displays them to the user in real time. For example, the action display area of the interface may display "Please propose specific measures for market strategy."
[0800] Input: Generated progress action
[0801] Output: Progression actions presented to the user
[0802] (Application example 1)
[0803] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0804] Modern conference systems and driver assistance systems require technologies that can streamline decision-making for participants and drivers and suggest appropriate actions in real time. However, conventional systems have issues such as not being able to transcribe speech in real time, not being able to fully utilize generative AI models, and not being able to quickly suggest appropriate actions. This can sometimes hinder smooth progress in meetings and driving assistance, resulting in issues of reduced efficiency.
[0805] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0806] In this invention, the server includes a means for transcribing the speech of participants and the driver in real time, a means for analyzing the transcribed text using a generative AI model and generating the next action for proceeding with the meeting or driving assistance action, and a means for presenting the generated action for proceeding with the meeting or driving assistance action to the participants and the driver in real time. This enables quick and appropriate decision-making during meetings and driving, and enables efficient progress and assistance.
[0807] A "web conferencing tool" is software that allows multiple participants in remote locations to hold meetings by sharing audio and video over the Internet.
[0808] A "generative AI model" is an artificial intelligence algorithm that generates output, such as text or actions, based on input data.
[0809] The "meeting efficiency system" is a system that transcribes and analyzes audio to generate and suggest the next action in order to make meetings proceed smoothly and effectively.
[0810] A "driving assistance system" is a system that assists vehicle drivers by transcribing and analyzing voice and providing appropriate instructions and information to the driver in real time.
[0811] "Participant" or "Driver" refers to a person attending a meeting or driving a vehicle.
[0812] "Real time" means that processing is carried out almost immediately after an event occurs.
[0813] "Transcription" is the process of analyzing audio data and converting its content into text format.
[0814] "Means" refer to the methods or tools used to achieve a particular goal.
[0815] "Next Actions" are specific actions or proposals to be taken next, generated based on the purpose and progress of the meeting.
[0816] "Driving assistance actions" are instructions or information provided to the driver in real time to assist driving.
[0817] A "cloud-based system" is a system in which data is processed and stored on a remote server via the Internet.
[0818] "User feedback" refers to opinions and evaluation data provided by users of the system, and is used to improve the system.
[0819] The present invention is an efficiency improvement system using a generative AI model that can be applied to web conferencing tools and driving assistance systems for autonomous vehicles. This system has the function of transcribing the voices of participants and drivers in real time, analyzing the text data, and generating the next progress action or driving assistance action. In addition, by presenting the generated actions to participants and drivers in real time, the progress of the meeting or driving can be made smoother. Specific embodiments for implementing the present invention are described below.
[0820] Key Components Used
[0821] 1. Server:
[0822] The audio data is transcribed and the text data is analyzed using a generative AI model.
[0823] Use a speech recognition library (e.g., Google Speech Recognition API) for transcription.
[0824] 2. Terminal:
[0825] It receives input from the user and sends it to the server.
[0826] The generated progress actions and driving assistance actions are presented to the user.
[0827] 3. User:
[0828] Set the purpose or goal of the meeting or provide driving context.
[0829] Proceed with meetings or driving based on the generated actions.
[0830] Program processing explanation
[0831] The server transcribes speech from participants and drivers in real time. To do this, it uses a speech recognition library to convert the speech data into text format. It then uses a generative AI model to analyze the transcribed text data and generate next actions for the meeting or driver assistance actions. This generative AI model processes the data in real time on a cloud-based system.
[0832] The generated actions are sent from the server to the terminal, and the terminal presents these actions to the user in real time. The user then takes the next action based on the presented actions.
[0833] Specific examples
[0834] For example, if a user wants to hold a meeting titled "New Product Launch Plan," the user enters the following information into the web conferencing tool's interface:
[0835] Meeting purpose: "New product launch plan"
[0836] Goal: "Formulate market strategies and assess risks"
[0837] Audio file path: "audio_new_product_meeting.wav"
[0838] This information is sent from the device to the server. The server transcribes the audio file and sends it to the generative AI model for analysis. As a result, the generative AI model generates a progress action such as "Please propose specific measures for market strategy." This action is sent to the device and displayed to the user through the device.
[0839] In a driver assistance system, if a driver wants to know the location of the next exit while driving, they can utter the voice command "Tell me where the next exit is." This speech is transcribed in real time and analyzed by a generative AI model. The server generates a driving assistance action such as "Please use the exit on the right 2 km ahead," and sends it to the device. This action is displayed on the driver's smart glasses or head-mounted display.
[0840] Example prompt sentence:
[0841] You are a driving assistance bot. Please interpret the following voice commands and suggest appropriate actions.
[0842] Voice command: What's the next entrance / exit?
[0843] In this way, the system allows users to efficiently organize what is said during meetings or while driving and receive appropriate proceeding actions and driving assistance actions in real time, thereby improving meeting efficiency and driving safety.
[0844] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0845] Step 1:
[0846] User input
[0847] The user uses the interface of the web conferencing tool or the voice input device of the driving assistance system to input the purpose and goal of the meeting, as well as the path of the audio file and voice commands while driving, including input in text format or voice input, which results in specific data.
[0848] Input: purpose of the meeting, goal, audio file path, voice command
[0849] Output: User input data (text data and audio data)
[0850] Step 2:
[0851] Sending information
[0852] The device sends the entered meeting purpose, goal, audio file path, and voice command to the server, which then sends the data over the network to a cloud-based server for processing.
[0853] Input: User-entered data
[0854] Output: Data sent to the server
[0855] Step 3:
[0856] Analysis of audio data
[0857] The server transcribes the received audio file using a speech recognition library (such as the Google Speech Recognition API). The audio data is analyzed and converted into text. This transcribed text data is then sent to the generative AI model.
[0858] Input: Audio data
[0859] Output: Transcribed text data
[0860] Step 4:
[0861] Analysis using generative AI models
[0862] The server analyzes the transcribed text data using a generative AI model and generates the next progress action or driving assistance action based on the prompt sentence. The generative AI model analyzes the text data using a specific algorithm and suggests the optimal action.
[0863] Input: Transcribed text data
[0864] Output: Generated progress actions and driver assistance actions
[0865] Step 5:
[0866] Suggesting Actions
[0867] The server sends the generated progress actions and driving assistance actions to the device, which then presents these actions to the user in real time: if in a meeting, they are displayed in the meeting interface, and if driving, they are displayed to the driver through smart glasses or a head-mounted display.
[0868] Input: Generated progress actions and driver assistance actions
[0869] Output: Real-time suggested actions
[0870] Step 6:
[0871] User Feedback
[0872] The user continues to conduct a meeting or drive based on the generated actions, and user feedback provides the system with data for improvement, allowing the accuracy of the generative AI model to be continuously improved.
[0873] Input: User feedback
[0874] Output: Improve the accuracy of the generative AI model.
[0875] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0876] The present invention is a system that uses a generative AI model that works with a web conferencing tool to improve the efficiency of meetings and, by combining it with an emotion engine, makes the meeting proceed more effectively. This system has the function of transcribing participants' voices in real time and analyzing the text data to generate meeting proceeding actions. Furthermore, by presenting the generated proceeding actions to participants in real time and combining it with an emotion engine that recognizes the user's emotions, the meeting can proceed smoothly. A specific embodiment of this system is described below.
[0877] Key Components Used
[0878] 1. Server: Transcribing voice data, analyzing text data using a generative AI model, and recognizing user emotions using an emotion engine.
[0879] 2. Terminal: Receives input from the user, sends it to the server, and provides feedback to the user based on the generated progress actions and emotions.
[0880] 3. User: Sets the purpose and goals of the meeting, provides emotional data, and advances the meeting based on progress actions.
[0881] 4. Emotion engine: Analyzes the user's voice and facial expressions and sends the emotion data to the generative AI model.
[0882] Program processing explanation
[0883] 1. User Input
[0884] The user sets the purpose and goals of the meeting through the interface of the web conferencing tool. For example, they input "Plan the next project" and "Define the main tasks and determine who will be responsible." The user also inputs the path to the meeting audio file (e.g., "audio.wav").
[0885] 2. Transmission of information
[0886] The terminal transmits the purpose and goal of the meeting and the path of the audio file input by the user to the server.
[0887] 3. Analysis of audio data
[0888] The server loads the audio file based on the path of the audio file received, and transcribes the audio data in real time using a speech recognition library (e.g., Google Speech Recognition API).
[0889] 4. Acquiring Emotion Data
[0890] The server uses an emotion engine to analyze the user's emotional data from the audio and video data, based on the tone of voice, facial expressions, and content of speech.
[0891] 5. Proceeding Action Generation
[0892] The server sends the transcribed text data and emotion data to the generative AI model. The generative AI model analyzes the transcribed data and emotion data and generates the next action to take. For example, an action such as "If the user is nervous, suggest questions to help them relax" may be generated.
[0893] 6. Presenting Progressive Actions and Emotional Feedback
[0894] The server transmits feedback based on the generated progress actions and emotions to the terminal.
[0895] The device displays this information to the user in real time, and the user can proceed with the meeting based on the presented actions and emotional feedback.
[0896] Specific examples
[0897] For example, suppose a user wants to hold a meeting titled "New Product Launch Plan." In this case, the user enters the following information into the web conferencing tool's interface:
[0898] Meeting purpose: "New product launch plan"
[0899] Goal: "Formulate market strategies and assess risks"
[0900] Audio file: "audio_new_product_meeting.wav"
[0901] This information is sent from the device to the server. The server transcribes the audio file and analyzes the user's emotions using an emotion engine. The generative AI model then analyzes this data and generates a progress action, such as "Please propose specific measures for market strategy." If the user is nervous, it will also suggest questions to relax them (e.g., "Tell me about a recent success story"). This action and emotional feedback are sent to the device and displayed to the user.
[0902] This system allows users to efficiently organize statements and emotional data during meetings and receive appropriate proceeding actions in real time, further improving the efficiency and effectiveness of meetings.
[0903] The processing flow will be explained below.
[0904] Step 1:
[0905] The user inputs the purpose of the meeting (e.g., "Plan the next project") and goal (e.g., "Define main tasks and determine who is responsible") into the interface of the web conferencing tool. The user also inputs the path to the audio file for the meeting (e.g., "audio.wav").
[0906] Step 2:
[0907] The terminal transmits the purpose and goal of the meeting and the path of the audio file input by the user to the server.
[0908] Step 3:
[0909] The server loads the audio file based on the path of the audio file received, and transcribes the audio data in real time using a speech recognition library (e.g., Google Speech Recognition API).
[0910] Step 4:
[0911] The server sends the transcribed text data to the generative AI model, which analyzes the transcribed data and the purpose and goals of the meeting to generate the next action to take.
[0912] Step 5:
[0913] The server captures the user's voice and facial expressions, and uses an emotion engine to analyze the user's emotional data. Emotions are recognized based on the tone of voice, changes in facial expressions, and the content of speech.
[0914] Step 6:
[0915] The server inputs the emotion data into the generative AI model and takes it into consideration when generating the next action to proceed, such as "suggesting questions to help a nervous user relax."
[0916] Step 7:
[0917] The server sends the generated progress actions and emotional feedback to the device.
[0918] Step 8:
[0919] The device receives the progress actions and emotional feedback from the server and presents them to the user in real time. The user then proceeds with the conference based on the progress actions and emotional feedback presented to them.
[0920] Example 2
[0921] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0922] In today's business environment, there is a demand for improving the efficiency and effectiveness of meetings. However, it is difficult to instantly analyze participants' emotions and comments and provide appropriate proceeding actions in real time. Furthermore, conventional conferencing systems are unable to quickly respond to changes in emotions, which can hinder the smooth progress of the meeting. Therefore, there is a need for a method that can analyze participants' emotions in real time and optimize the progress of the meeting based on that data.
[0923] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0924] In this invention, the server includes a means for a user to set the purpose and goal of the meeting, a means for transcribing the participants' voices in real time, a means for analyzing the transcribed text and the users' emotional data using a generative AI model to generate the next action to proceed with the meeting, and a means for presenting feedback to the participants in real time based on the generated action to proceed with the meeting and their emotions. This increases the efficiency and effectiveness of the meeting, and enables the meeting to proceed smoothly by quickly responding to changes in the participants' emotions.
[0925] "Means for users to set the purpose and goals of a meeting" means allowing users to input the main purpose and goals of a meeting using the interface of the web conferencing tool.
[0926] "Means for transcribing participants' voices in real time" refers to means for converting the speech of conference participants into text data in real time using voice recognition technology.
[0927] A "generative AI model" refers to an algorithm or program that uses natural language processing and machine learning to generate meaningful information from input data.
[0928] The "means for analyzing the transcribed text and user emotion data and generating the next action to proceed with the meeting" refers to a means for generating instructions and suggestions for determining the progress of the meeting based on the transcribed text data and emotion analysis data.
[0929] An "emotion engine" refers to technology or algorithms that analyze voice tone, speech rate, facial expression changes, etc. to identify the emotional state of users or participants.
[0930] The "means for presenting participants with feedback based on generated meeting progress actions and emotions in real time" is a means for instantly providing analyzed and generated information to meeting participants to guide the progress of the meeting.
[0931] "Cloud-based processing" refers to the technology of processing data on remote servers via the Internet, and is used to efficiently perform large-scale data analysis and real-time processing.
[0932] "Means for incorporating user feedback and continuously improving the accuracy of the generative AI model" refers to means for improving the performance and accuracy of the generative AI model over time by incorporating user opinions and evaluations into the model.
[0933] This invention relates to a system that improves the efficiency and effectiveness of meetings by utilizing a generative AI model that works in conjunction with a web conferencing tool. Its key feature is that it transcribes participants' speech in real time, analyzes the text data to generate meeting progress actions, and provides feedback to participants, thereby helping to ensure smooth progress of the meeting. Furthermore, by combining it with an emotion engine, it is possible to analyze participants' emotions and reflect them in the progress actions.
[0934] Key Components Used
[0935] The following describes the major components for implementing the present invention.
[0936] server
[0937] The server is the core of this system and has the following main functions:
[0938] 1. Ability to transcribe voice data in real time: Uses a speech recognition library (e.g., Google Speech Recognition API) to convert participants' voice data into text data.
[0939] 2. The ability to analyze text data and emotional data using a generative AI model and generate next steps: The generative AI model uses natural language processing and machine learning based on the input data to derive the appropriate next action.
[0940] 3. Ability to analyze emotional data using an emotion engine: Identify the user's emotional state from voice tone, speaking rate, facial expression changes, etc.
[0941] Terminal
[0942] The terminal is a device that provides a user interface and has the following functions:
[0943] 1. Function to receive input from the user and send it to the server: Provides an interface for inputting the purpose and goals of the meeting, the path of the audio file, etc.
[0944] 2. A function to present feedback to the user based on the generated progress actions and emotions: Feedback sent from the server is displayed in real time.
[0945] User
[0946] The user sets the purpose and goal of the meeting and is responsible for running the meeting. The user also provides data on comments and emotions made during the meeting, and the system guides the meeting according to the actions taken by the system.
[0947] Specific examples
[0948] For example, consider a user who wants to hold a meeting titled "New Product Launch Plan." The user enters the following information into the web conferencing tool's interface:
[0949] Meeting purpose: "New product launch plan"
[0950] Goal: "Formulate market strategies and assess risks"
[0951] Audio file: "audio_new_product_meeting.wav"
[0952] Prompt Sentence Examples
[0953] "Planning the next project"
[0954] "Define key tasks and determine who is responsible"
[0955] "Marketing Strategy Meeting"
[0956] "Target Market Definition and Competitive Analysis"
[0957] The device sends this information to a server. The server reads the audio file and transcribes it using a speech recognition library. It then uses an emotion engine to analyze the user's emotions from the acquired transcription and video data. The analysis results are sent to a generative AI model, which generates the next action to proceed.
[0958] For example, if a user is nervous, the generative AI model will generate a progression action such as "suggesting questions to relax the user." It may also suggest questions to dig deeper into marketing strategies. The generated progression actions and feedback are sent to the device and displayed to the user in real time.
[0959] This allows users to smoothly progress with the conference, and improves the efficiency and effectiveness of the conference.
[0960] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0961] Step 1:
[0962] The user sets the purpose and goal of the meeting through the interface of the web conferencing tool and inputs the path of the audio file. As input, the user provides the purpose of the meeting (e.g., "Plan for the market launch of a new product"), the goal (e.g., "Formulate market strategy and assess risks"), and the path of the audio file (e.g., "audio_new_product_meeting.wav"). This information is saved on the device for the next step.
[0963] Step 2:
[0964] The terminal sends the information input by the user to the server. The input information includes the purpose and goal of the meeting, and the path to the audio file. The terminal appropriately packets this information and sends it to the server. The output is the meeting information sent to the server.
[0965] Step 3:
[0966] The server receives the audio file path and loads it. Then, it uses a speech recognition library (e.g., Google Speech Recognition API) to transcribe the audio data in real time. The input is the audio file path "audio_new_product_meeting.wav" and the output is the transcribed text data (e.g., "The discussion about the new product has begun...").
[0967] Step 4:
[0968] The server inputs the acquired transcription data and video data into the emotion engine. The emotion engine analyzes the tone of voice, speech rate, and facial expression changes to generate the user's emotion data. The inputs are transcription data and video data, and the output is the analyzed emotion data (e.g., "The user is nervous").
[0969] Step 5:
[0970] The server sends the transcribed text data and emotion data to the generative AI model, which generates the next progress action based on this data. The input is the transcribed text and emotion data, and the output is the generated progress action (e.g., "Suggest questions to help the user relax").
[0971] Step 6:
[0972] The server transmits the generated progress action and emotion-based feedback to the terminal. The input is the progress action and emotion feedback, and the output is the data transmitted to the terminal.
[0973] Step 7:
[0974] The device displays the progress actions and emotional feedback received from the server to the user. Specifically, they are displayed as notifications or pop-ups on the interface and provided to the user in real time. The input is the progress actions and emotional feedback from the server, and the output is the information displayed to the user.
[0975] (Application example 2)
[0976] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0977] With conventional web conferencing tools, meeting efficiency and progress management are largely manual, which means that it takes time to organize what is being said and decide on actions to take. There are also concerns that the quality of meetings may decline due to the difficulty of properly understanding participants' emotions and smoothly progressing the meeting. Furthermore, when dealing with customers in virtual stores, it is difficult to suggest appropriate actions to take in response to customer questions, making it difficult to increase customer satisfaction.
[0978] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0979] In this invention, the server includes means for transcribing the participants' voices in real time, means for analyzing the transcribed text using a generative AI model and generating the next action to proceed with the conference, means for presenting the generated action to proceed with the conference to the participants in real time, means for analyzing the emotions of the participants using an emotion engine and providing appropriate feedback based on the analysis results, means for monitoring the customer service process in the virtual store and generating appropriate actions in response to customer questions, and means for presenting the generated response actions and emotion feedback to the customers in real time. This makes it possible to efficiently manage the progress of the conference and provide optimal actions in response to the emotions of the participants, as well as improve the quality of customer service in the virtual store and increase customer satisfaction.
[0980] "Web conferencing tool" refers to software or a platform that allows multiple participants to hold real-time conferences using audio and video over the Internet.
[0981] A "generative AI model" refers to an algorithm or system that uses artificial intelligence to analyze input data and automatically generate appropriate responses or actions.
[0982] A "meeting efficiency system" refers to a system that combines technologies and means to support the progress and management of meetings and improve their efficiency.
[0983] "Real-time transcription means" refers to technology or equipment for instantly converting audio data into text data.
[0984] "Progressive actions" refer to specific actions or instructions that should be taken next in the progress of a meeting.
[0985] An "emotion engine" refers to technology and algorithms that analyze data such as participants' voices and facial expressions to estimate their emotional state.
[0986] "Feedback" refers to information or instructions provided to users or participants based on analysis results, progress actions, etc.
[0987] "Virtual store" refers to a virtual store that offers products and services via the Internet.
[0988] A "customer service process" refers to a set of procedures or methods for appropriately responding to customer inquiries and questions.
[0989] This invention is a system that improves the efficiency of meetings by using a generative AI model that works in conjunction with a web conferencing tool, and further combines it with an emotion engine to more effectively progress the meeting. This system has the function of transcribing participants' voices in real time, analyzing the text data, and generating the next progress action. Furthermore, by presenting the generated progress action to participants in real time and combining it with an emotion engine that recognizes the user's emotions, the meeting can be progressed smoothly.
[0990] Key Components Used
[0991] 1. Server: Transcribing voice data, analyzing text data using a generative AI model, and recognizing user emotions using an emotion engine. Specifically, it uses a speech recognition library (e.g., Google Speech Recognition API) and an emotion analysis engine (e.g., EmotionEngine). It also generates progress actions using a generative AI model (e.g., ActionsGenerator).
[0992] 2. Device: Receives input from the user and sends it to the server. It also provides feedback to the user based on the generated progress actions and emotions. Devices can be smartphones, head-mounted displays, smart glasses, etc.
[0993] 3. User: Sets the purpose and goals of the meeting, provides emotional data and advances the meeting based on progress actions.
[0994] 4. Emotion engine: Analyzes the user's voice and facial expressions and sends the emotional data to a generative AI model.
[0995] Program processing explanation
[0996] 1. The server transcribes the user's voice data using the Google Speech Recognition API. This process converts the voice data into text data.
[0997] 2. The emotion engine then analyzes the voice tone and facial expressions to generate the user's emotion data. This is done using the EmotionEngine.
[0998] 3. The server uses a generative AI model to analyze the transcribed text data and emotion data, and generates the next action to proceed. Using ActionsGenerator, the generated action is provided in real time.
[0999] 4. The device presents the progress actions and emotional feedback received from the server to the user in real time. This information is displayed via a smartphone, head-mounted display, or smart glasses.
[1000] Specific examples
[1001] For example, suppose a user wants to hold a meeting titled "New Product Launch Plan." In this case, the user enters the following information into the web conferencing tool's interface:
[1002] Meeting purpose: "New product launch plan"
[1003] Goal: "Formulate market strategies and assess risks"
[1004] Audio file: "audio_new_product_meeting.wav"
[1005] This information is sent from the device to the server. The server transcribes the audio file and analyzes the user's emotions using an emotion engine. The generative AI model then analyzes this data and generates a progress action, such as "Please propose specific measures for market strategy." If the user is nervous, it will also suggest questions to relax them (e.g., "Tell me about a recent success story"). This action and emotional feedback are sent to the device and displayed to the user.
[1006] Prompt Sentence Examples
[1007] Below are some example prompts to use with generative AI models:
[1008] text
[1009] Voice input: "I want to know about my new TV."
[1010] Emotional data: interesting but a bit confusing
[1011] Generate Action: Describe the features of their new TV. Ask about their specific needs.
[1012] This system efficiently organizes statements and emotional data during meetings and allows appropriate actions to be taken in real time, further improving the efficiency and effectiveness of meetings. It also makes it possible to generate and provide appropriate actions in real time when dealing with customers in virtual stores, thereby increasing customer satisfaction.
[1013] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1014] Step 1:
[1015] The server receives the user's voice data. The user inputs the purpose, goal, and audio file of the meeting into the terminal. The terminal sends this information to the server. Specific inputs include the audio file (e.g., "audio_new_product_meeting.wav"), the purpose of the meeting, "a plan to launch a new product," and the goal, "formulation of a market strategy and risk assessment." The server receives this data and begins processing.
[1016] Step 2:
[1017] The server transcribes the received voice data in real time using the Google Speech Recognition API. Text data (output) is generated from the voice data (input). This transcribed text data serves as the basis for generating actions to progress the meeting. Converting the input voice data into text makes subsequent data analysis easier.
[1018] Step 3:
[1019] The server uses an emotion engine to analyze the user's emotional data along with the transcribed voice data. Emotion analysis is performed using the Emotion Engine. Specifically, it analyzes changes in voice tone and facial expressions. Through this process, the server obtains emotional data (output) from the voice data (input). For example, it determines whether the user is nervous, relaxed, or interested. At this stage, data enrichment is performed.
[1020] Step 4:
[1021] The server uses a generative AI model (e.g., ActionsGenerator) to analyze the transcribed text data and emotional data. Specifically, the generative AI model generates the next action to proceed (output) based on the text data (input) and emotional data (input). For example, an action such as "Please propose specific measures for market strategy" may be generated. This action provides specific instructions to ensure the smooth flow of the meeting.
[1022] Step 5:
[1023] The server sends the generated progress action and emotional feedback to the device. Specifically, the generated action (output) and emotional feedback (output) are sent. For example, the action may be "Please propose specific measures for the market strategy" and the emotional feedback may be "The user is nervous." This allows the device to process the next step.
[1024] Step 6:
[1025] The device presents the progress actions and emotional feedback received from the server to the user in real time. Specifically, information is displayed to the user (output) via a smartphone, head-mounted display, or smart glasses. For example, the device may provide the user with instructions such as "Please propose specific measures for market strategy" or feedback such as "The user seems nervous, so we suggest questions to help them relax." This allows the user to effectively proceed with the next action in the meeting.
[1026] Step 7:
[1027] The user controls the meeting based on the progress actions and emotional feedback presented on the device. By taking specific actions, the user can effectively move forward toward the purpose and goal of the meeting. The user uses the presented information to execute the next step instructions. This is the stage at which the actual meeting progresses.
[1028] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1029] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1030] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1031] [Fourth embodiment]
[1032] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1033] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1034] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1035] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1036] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1037] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1038] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1039] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1040] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1041] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1042] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1043] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1044] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1045] The present invention is a system that improves the efficiency of meetings by using a generative AI model that works in conjunction with a web conferencing tool. This system has the function of transcribing participants' voices in real time, analyzing the text data, and generating actions to progress the meeting. In addition, by presenting the generated actions to participants in real time, the progress of the meeting can be made smoother. A specific embodiment of this system is described below.
[1046] Key Components Used
[1047] 1. Server: Transcribing audio data and analyzing text data using a generative AI model.
[1048] 2. Terminal: Receives input from the user, sends it to the server, and presents the generated progress actions to the user.
[1049] 3. User: Set the purpose and goals of the meeting and proceed with the meeting based on progress actions.
[1050] Program processing explanation
[1051] 1. User Input
[1052] Users set the purpose and goals of the meeting through the interface of the web conferencing tool, for example, entering "planning the next project" and "defining the main tasks and deciding who will be responsible."
[1053] The user also enters the path to the conference audio file (eg "audio.wav").
[1054] 2. Transmission of information
[1055] The terminal transmits the purpose and goal of the meeting and the path of the audio file input by the user to the server.
[1056] 3. Analysis of audio data
[1057] The server transcribes the received audio file in real time, using a speech recognition library (e.g., Google Speech Recognition API).
[1058] Transcribed text data is generated and fed into a generative AI model.
[1059] 4. Proceeding Action Generation
[1060] The server uses a generative AI model to analyze the transcribed text data and generate next steps based on the purpose and goals of the meeting, such as "encourage discussion on which specific markets to focus on."
[1061] 5. Presenting Progress Actions
[1062] The server transmits the generated progress action to the terminal.
[1063] The terminal presents this ongoing action to the user in real time.
[1064] Specific examples
[1065] For example, suppose a user wants to hold a meeting titled "New Product Launch Plan." In this case, the user enters the following information into the web conferencing tool's interface:
[1066] Meeting purpose: "New product launch plan"
[1067] Goal: "Formulate market strategies and assess risks"
[1068] Audio file: "audio_new_product_meeting.wav"
[1069] This information is sent from the device to the server. The server transcribes the audio file and sends it to the generative AI model for analysis. As a result, the generative AI model generates a progress action such as "Please propose specific measures for market strategy." This action is sent to the device and displayed to the user through the device.
[1070] This system allows users to efficiently organize what is being said during a meeting and receive appropriate proceeding actions in real time, thereby improving the efficiency and effectiveness of meetings.
[1071] The processing flow will be explained below.
[1072] Step 1:
[1073] The user inputs the purpose of the meeting (e.g., "Planning the next project") and goal (e.g., "Define main tasks and determine who is responsible") into the interface of the web conferencing tool. The user also inputs the path to the audio file for the meeting (e.g., "audio.wav").
[1074] Step 2:
[1075] The terminal transmits the purpose and goal of the meeting and the path of the audio file input by the user to the server.
[1076] Step 3:
[1077] The server loads the audio file based on the path of the audio file received, and transcribes the audio data in real time using a speech recognition library (e.g., Google Speech Recognition API).
[1078] Step 4:
[1079] The server sends the transcribed text data to the generative AI model, which analyzes the transcribed data and the purpose and goals of the meeting to generate the next action to take.
[1080] Step 5:
[1081] The server transmits the generated progress action to the terminal.
[1082] Step 6:
[1083] The terminal displays the progress actions received from the server to the user in real time, and the user proceeds with the conference based on the presented progress actions.
[1084] Step 7:
[1085] After the conference ends, the server stores the transcript log of the entire conference and the history of the ongoing actions generated.
[1086] Step 8:
[1087] After the meeting ends, users can download the transcription log and progress action history stored on the server via their device and use them for summarizing and providing feedback.
[1088] Example 1
[1089] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1090] Modern meetings are increasingly being held online, especially with the spread of remote work. However, it remains difficult to efficiently understand what participants are saying during a meeting and decide on the appropriate next steps. As a result, the meeting often stalls and delays the achievement of its objectives and goals. Maintaining meeting efficiency and effectiveness becomes even more difficult when there are many participants or a wide range of information.
[1091] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1092] In this invention, the server includes a means for a user to set a purpose and goal of a meeting, a means for a user to upload audio data of the meeting, a means for transcribing the audio data received from the target in real time, a means for analyzing the transcribed text data by a generative AI model and generating the next action to proceed with the meeting, and a means for presenting the generated action to proceed with the meeting to the user in real time, thereby enabling real-time analysis of utterances during a meeting and the presentation of action to proceed with the meeting.
[1093] "User" refers to a person who uses this system to set the purpose and goal of a meeting and upload audio data.
[1094] The "server" is a device that transcribes voice data received from users in real time, analyzes the text using a generative AI model, and generates actions to progress the meeting.
[1095] A "terminal" is a device that allows a user to input the purpose and goal of a meeting and transmit audio files to a server, and is also a device that presents the generated progress actions to the user.
[1096] "Audio data" refers to a digital data file that records what participants say during a meeting.
[1097] "Real-time" refers to processing and displaying data immediately, with minimal time delay.
[1098] "Transcription" refers to the process of converting audio data into text form.
[1099] "Generative AI model" refers to the artificial intelligence technology used to analyze transcribed text data and generate actions that help guide meetings.
[1100] "Progression actions" refer to the next actions or discussion content suggested by the generative AI model based on the purpose and goals of the meeting.
[1101] "Upload" refers to the act of transferring data from a terminal to a server.
[1102] "Analysis" refers to the process by which generative AI models process transcribed text data, extract information, and generate appropriate actions.
[1103] This invention is a system that improves the efficiency of meetings by using a generative AI model that works in conjunction with a web conferencing tool. This system operates using multiple means, primarily servers, terminals, and users. The detailed configuration and operation of each means are explained below.
[1104] First, the user sets the purpose and goals of the meeting through the interface of the web conferencing tool. A dedicated form is provided for entering this information. For example, the purpose of the meeting can be "Planning the next project" and the goal can be "Define the main tasks and determine who will be responsible." In addition, a field is provided for uploading audio data from the meeting (e.g., "audio.wav"). The user selects the required audio file and clicks the "Send" button.
[1105] Next, the terminal has a means for sending the meeting purpose, goal, and audio file path entered by the user together to the server. This information is packetized and sent to the server as an HTTP request. This generally uses the HTTPS protocol.
[1106] The server has a means to transcribe the received voice data in real time. Specifically, it uses a speech recognition library such as the Google Speech Recognition API to convert the voice data into text data. To do this, it processes the received audio file (e.g., "audio.wav") and passes it to the speech recognition library. The speech recognition library converts the voice into text data and returns the result to the server, which temporarily stores it.
[1107] The server then sends the transcribed text data to a generative AI model for analysis. The generative AI model uses an interactive generative model (e.g., GPT-3.5). Based on the transcribed text data, this model generates the next progress action appropriate to the purpose and goal of the meeting. For example, a progress action such as "Please propose specific measures for market strategy" may be generated. This generated progress action is temporarily stored on the server.
[1108] Finally, the generated progress actions are sent from the server to the terminal and presented to the user in real time. The terminal displays the progress actions received from the server on the screen interface. For example, the action display area of the interface may display "Please propose specific measures for the market strategy."
[1109] Specific examples
[1110] For example, if a user wants to hold a meeting titled "New Product Launch Plan," they would enter the following information into the web conferencing tool's interface:
[1111] Meeting purpose: "New product launch plan"
[1112] Goal: "Formulate market strategies and assess risks"
[1113] Audio file: "audio_new_product_meeting.wav"
[1114] This information is sent from the device to the server. The server transcribes the audio file using the Google Speech Recognition API and passes it to the generative AI model for analysis. As a result, the generative AI model generates a progress action such as "Please propose specific measures for our market strategy." This action is sent to the device and displayed to the user in real time.
[1115] Example prompt sentence:
[1116] "The purpose of this meeting is to plan the market launch of a new product. We want to develop a market strategy and assess risks. Based on this objective, what are the next steps we should take?"
[1117] This system allows users to efficiently organize what is being said during a meeting and receive appropriate proceeding actions in real time, improving the efficiency and effectiveness of meetings.
[1118] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1119] Step 1: User Input
[1120] The user accesses the interface of the web conferencing tool and inputs the purpose and goals of the meeting. For example, they might set "Planning the next project" and "Define main tasks and determine who will be responsible." They also specify the path to the audio file for the meeting (e.g., "audio.wav"). After inputting this information, the user clicks the "Send" button.
[1121] Input: Meeting purpose, goal, audio file path
[1122] Output: Configuration information passed to the terminal
[1123] Step 2: Submit your information
[1124] The terminal packetizes the user-entered meeting purpose, goal, and audio file path, and sends the packet as an HTTP request to the server using the HTTPS protocol.
[1125] Input: Information entered by the user (meeting purpose, goal, audio file path)
[1126] Output: HTTP request sent to the server
[1127] Step 3: Receiving and transcribing audio data
[1128] The server receives the HTTP request and gets the path to the audio file. The server loads this audio file and transcribes it in real time using a speech recognition library such as the Google Speech Recognition API. This process converts the audio data into text data.
[1129] Input: The path to the audio file sent to the server
[1130] Output: Transcribed text data
[1131] Step 4: Analyzing the text data
[1132] The server temporarily stores the transcribed text data and then passes it to a generative AI model. The generative AI model analyzes the data and generates the next action based on the purpose and goals of the meeting. For example, it might generate an instruction such as, "Please propose specific measures for market strategy."
[1133] Input: Transcribed text data
[1134] Output: Generated progress actions
[1135] Step 5: Present progress actions
[1136] The server packets the generated progress actions and sends them to the terminal. The terminal receives the progress actions and displays them to the user in real time. For example, the action display area of the interface may display "Please propose specific measures for market strategy."
[1137] Input: Generated progress action
[1138] Output: Progression actions presented to the user
[1139] (Application example 1)
[1140] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1141] Modern conference systems and driver assistance systems require technologies that can streamline decision-making for participants and drivers and suggest appropriate actions in real time. However, conventional systems have issues such as not being able to transcribe speech in real time, not being able to fully utilize generative AI models, and not being able to quickly suggest appropriate actions. This can sometimes hinder smooth progress in meetings and driving assistance, resulting in issues of reduced efficiency.
[1142] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1143] In this invention, the server includes a means for transcribing the speech of participants and the driver in real time, a means for analyzing the transcribed text using a generative AI model and generating the next action for proceeding with the meeting or driving assistance action, and a means for presenting the generated action for proceeding with the meeting or driving assistance action to the participants and the driver in real time. This enables quick and appropriate decision-making during meetings and driving, and enables efficient progress and assistance.
[1144] A "web conferencing tool" is software that allows multiple participants in remote locations to hold meetings by sharing audio and video over the Internet.
[1145] A "generative AI model" is an artificial intelligence algorithm that generates output, such as text or actions, based on input data.
[1146] The "meeting efficiency system" is a system that transcribes and analyzes audio to generate and suggest the next action in order to make meetings proceed smoothly and effectively.
[1147] A "driving assistance system" is a system that assists vehicle drivers by transcribing and analyzing voice and providing appropriate instructions and information to the driver in real time.
[1148] "Participant" or "Driver" refers to a person attending a meeting or driving a vehicle.
[1149] "Real time" means that processing is carried out almost immediately after an event occurs.
[1150] "Transcription" is the process of analyzing audio data and converting its content into text format.
[1151] "Means" refer to the methods or tools used to achieve a particular goal.
[1152] "Next Actions" are specific actions or proposals to be taken next, generated based on the purpose and progress of the meeting.
[1153] "Driving assistance actions" are instructions or information provided to the driver in real time to assist driving.
[1154] A "cloud-based system" is a system in which data is processed and stored on a remote server via the Internet.
[1155] "User feedback" refers to opinions and evaluation data provided by users of the system, and is used to improve the system.
[1156] The present invention is an efficiency improvement system using a generative AI model that can be applied to web conferencing tools and driving assistance systems for autonomous vehicles. This system has the function of transcribing the voices of participants and drivers in real time, analyzing the text data, and generating the next progress action or driving assistance action. In addition, by presenting the generated actions to participants and drivers in real time, the progress of the meeting or driving can be made smoother. Specific embodiments for implementing the present invention are described below.
[1157] Key Components Used
[1158] 1. Server:
[1159] The audio data is transcribed and the text data is analyzed using a generative AI model.
[1160] Use a speech recognition library (e.g., Google Speech Recognition API) for transcription.
[1161] 2. Terminal:
[1162] It receives input from the user and sends it to the server.
[1163] The generated progress actions and driving assistance actions are presented to the user.
[1164] 3. User:
[1165] Set the purpose or goal of the meeting or provide driving context.
[1166] Proceed with meetings or driving based on the generated actions.
[1167] Program processing explanation
[1168] The server transcribes speech from participants and drivers in real time. To do this, it uses a speech recognition library to convert the speech data into text format. It then uses a generative AI model to analyze the transcribed text data and generate next actions for the meeting or driver assistance actions. This generative AI model processes the data in real time on a cloud-based system.
[1169] The generated actions are sent from the server to the terminal, and the terminal presents these actions to the user in real time. The user then takes the next action based on the presented actions.
[1170] Specific examples
[1171] For example, if a user wants to hold a meeting titled "New Product Launch Plan," the user enters the following information into the web conferencing tool's interface:
[1172] Meeting purpose: "New product launch plan"
[1173] Goal: "Formulate market strategies and assess risks"
[1174] Audio file path: "audio_new_product_meeting.wav"
[1175] This information is sent from the device to the server. The server transcribes the audio file and sends it to the generative AI model for analysis. As a result, the generative AI model generates a progress action such as "Please propose specific measures for market strategy." This action is sent to the device and displayed to the user through the device.
[1176] In a driver assistance system, if a driver wants to know the location of the next exit while driving, they can utter the voice command "Tell me where the next exit is." This speech is transcribed in real time and analyzed by a generative AI model. The server generates a driving assistance action such as "Please use the exit on the right 2 km ahead," and sends it to the device. This action is displayed on the driver's smart glasses or head-mounted display.
[1177] Example prompt sentence:
[1178] You are a driving assistance bot. Please interpret the following voice commands and suggest appropriate actions.
[1179] Voice command: What's the next entrance / exit?
[1180] In this way, the system allows users to efficiently organize what is said during meetings or while driving and receive appropriate proceeding actions and driving assistance actions in real time, thereby improving meeting efficiency and driving safety.
[1181] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1182] Step 1:
[1183] User input
[1184] The user uses the interface of the web conferencing tool or the voice input device of the driving assistance system to input the purpose and goal of the meeting, as well as the path of the audio file and voice commands while driving, including input in text format or voice input, which results in specific data.
[1185] Input: purpose of the meeting, goal, audio file path, voice command
[1186] Output: User input data (text data and audio data)
[1187] Step 2:
[1188] Sending information
[1189] The device sends the entered meeting purpose, goal, audio file path, and voice command to the server, which then sends the data over the network to a cloud-based server for processing.
[1190] Input: User-entered data
[1191] Output: Data sent to the server
[1192] Step 3:
[1193] Analysis of audio data
[1194] The server transcribes the received audio file using a speech recognition library (such as the Google Speech Recognition API). The audio data is analyzed and converted into text. This transcribed text data is then sent to the generative AI model.
[1195] Input: Audio data
[1196] Output: Transcribed text data
[1197] Step 4:
[1198] Analysis using generative AI models
[1199] The server analyzes the transcribed text data using a generative AI model and generates the next progress action or driving assistance action based on the prompt sentence. The generative AI model analyzes the text data using a specific algorithm and suggests the optimal action.
[1200] Input: Transcribed text data
[1201] Output: Generated progress actions and driver assistance actions
[1202] Step 5:
[1203] Suggesting Actions
[1204] The server sends the generated progress actions and driving assistance actions to the device, which then presents these actions to the user in real time: if in a meeting, they are displayed in the meeting interface, and if driving, they are displayed to the driver through smart glasses or a head-mounted display.
[1205] Input: Generated progress actions and driver assistance actions
[1206] Output: Real-time suggested actions
[1207] Step 6:
[1208] User Feedback
[1209] The user continues to conduct a meeting or drive based on the generated actions, and user feedback provides the system with data for improvement, allowing the accuracy of the generative AI model to be continuously improved.
[1210] Input: User feedback
[1211] Output: Improve the accuracy of the generative AI model.
[1212] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1213] The present invention is a system that uses a generative AI model that works with a web conferencing tool to improve the efficiency of meetings and, by combining it with an emotion engine, makes the meeting proceed more effectively. This system has the function of transcribing participants' voices in real time and analyzing the text data to generate meeting proceeding actions. Furthermore, by presenting the generated proceeding actions to participants in real time and combining it with an emotion engine that recognizes the user's emotions, the meeting can proceed smoothly. A specific embodiment of this system is described below.
[1214] Key Components Used
[1215] 1. Server: Transcribing voice data, analyzing text data using a generative AI model, and recognizing user emotions using an emotion engine.
[1216] 2. Terminal: Receives input from the user, sends it to the server, and provides feedback to the user based on the generated progress actions and emotions.
[1217] 3. User: Sets the purpose and goals of the meeting, provides emotional data, and advances the meeting based on progress actions.
[1218] 4. Emotion engine: Analyzes the user's voice and facial expressions and sends the emotion data to the generative AI model.
[1219] Program processing explanation
[1220] 1. User Input
[1221] The user sets the purpose and goals of the meeting through the interface of the web conferencing tool. For example, they input "Plan the next project" and "Define the main tasks and determine who will be responsible." The user also inputs the path to the meeting audio file (e.g., "audio.wav").
[1222] 2. Transmission of information
[1223] The terminal transmits the purpose and goal of the meeting and the path of the audio file input by the user to the server.
[1224] 3. Analysis of audio data
[1225] The server loads the audio file based on the path of the audio file received, and transcribes the audio data in real time using a speech recognition library (e.g., Google Speech Recognition API).
[1226] 4. Acquiring Emotion Data
[1227] The server uses an emotion engine to analyze the user's emotional data from the audio and video data, based on the tone of voice, facial expressions, and content of speech.
[1228] 5. Proceeding Action Generation
[1229] The server sends the transcribed text data and emotion data to the generative AI model. The generative AI model analyzes the transcribed data and emotion data and generates the next action to take. For example, an action such as "If the user is nervous, suggest questions to help them relax" may be generated.
[1230] 6. Presenting Progressive Actions and Emotional Feedback
[1231] The server transmits feedback based on the generated progress actions and emotions to the terminal.
[1232] The device displays this information to the user in real time, and the user can proceed with the meeting based on the presented actions and emotional feedback.
[1233] Specific examples
[1234] For example, suppose a user wants to hold a meeting titled "New Product Launch Plan." In this case, the user enters the following information into the web conferencing tool's interface:
[1235] Meeting purpose: "New product launch plan"
[1236] Goal: "Formulate market strategies and assess risks"
[1237] Audio file: "audio_new_product_meeting.wav"
[1238] This information is sent from the device to the server. The server transcribes the audio file and analyzes the user's emotions using an emotion engine. The generative AI model then analyzes this data and generates a progress action, such as "Please propose specific measures for market strategy." If the user is nervous, it will also suggest questions to relax them (e.g., "Tell me about a recent success story"). This action and emotional feedback are sent to the device and displayed to the user.
[1239] This system allows users to efficiently organize statements and emotional data during meetings and receive appropriate proceeding actions in real time, further improving the efficiency and effectiveness of meetings.
[1240] The processing flow will be explained below.
[1241] Step 1:
[1242] The user inputs the purpose of the meeting (e.g., "Plan the next project") and goal (e.g., "Define main tasks and determine who is responsible") into the interface of the web conferencing tool. The user also inputs the path to the audio file for the meeting (e.g., "audio.wav").
[1243] Step 2:
[1244] The terminal transmits the purpose and goal of the meeting and the path of the audio file input by the user to the server.
[1245] Step 3:
[1246] The server loads the audio file based on the path of the audio file received, and transcribes the audio data in real time using a speech recognition library (e.g., Google Speech Recognition API).
[1247] Step 4:
[1248] The server sends the transcribed text data to the generative AI model, which analyzes the transcribed data and the purpose and goals of the meeting to generate the next action to take.
[1249] Step 5:
[1250] The server captures the user's voice and facial expressions, and uses an emotion engine to analyze the user's emotional data. Emotions are recognized based on the tone of voice, changes in facial expressions, and the content of speech.
[1251] Step 6:
[1252] The server inputs the emotion data into the generative AI model and takes it into consideration when generating the next action to proceed, such as "suggesting questions to help a nervous user relax."
[1253] Step 7:
[1254] The server sends the generated progress actions and emotional feedback to the device.
[1255] Step 8:
[1256] The device receives the progress actions and emotional feedback from the server and presents them to the user in real time. The user then proceeds with the conference based on the progress actions and emotional feedback presented to them.
[1257] Example 2
[1258] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1259] In today's business environment, there is a demand for improving the efficiency and effectiveness of meetings. However, it is difficult to instantly analyze participants' emotions and comments and provide appropriate proceeding actions in real time. Furthermore, conventional conferencing systems are unable to quickly respond to changes in emotions, which can hinder the smooth progress of the meeting. Therefore, there is a need for a method that can analyze participants' emotions in real time and optimize the progress of the meeting based on that data.
[1260] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1261] In this invention, the server includes a means for a user to set the purpose and goal of the meeting, a means for transcribing the participants' voices in real time, a means for analyzing the transcribed text and the users' emotional data using a generative AI model to generate the next action to proceed with the meeting, and a means for presenting feedback to the participants in real time based on the generated action to proceed with the meeting and their emotions. This increases the efficiency and effectiveness of the meeting, and enables the meeting to proceed smoothly by quickly responding to changes in the participants' emotions.
[1262] "Means for users to set the purpose and goals of a meeting" means allowing users to input the main purpose and goals of a meeting using the interface of the web conferencing tool.
[1263] "Means for transcribing participants' voices in real time" refers to means for converting the speech of conference participants into text data in real time using voice recognition technology.
[1264] A "generative AI model" refers to an algorithm or program that uses natural language processing and machine learning to generate meaningful information from input data.
[1265] The "means for analyzing the transcribed text and user emotion data and generating the next action to proceed with the meeting" refers to a means for generating instructions and suggestions for determining the progress of the meeting based on the transcribed text data and emotion analysis data.
[1266] An "emotion engine" refers to technology or algorithms that analyze voice tone, speech rate, facial expression changes, etc. to identify the emotional state of users or participants.
[1267] The "means for presenting participants with feedback based on generated meeting progress actions and emotions in real time" is a means for instantly providing analyzed and generated information to meeting participants to guide the progress of the meeting.
[1268] "Cloud-based processing" refers to the technology of processing data on remote servers via the Internet, and is used to efficiently perform large-scale data analysis and real-time processing.
[1269] "Means for incorporating user feedback and continuously improving the accuracy of the generative AI model" refers to means for improving the performance and accuracy of the generative AI model over time by incorporating user opinions and evaluations into the model.
[1270] This invention relates to a system that improves the efficiency and effectiveness of meetings by utilizing a generative AI model that works in conjunction with a web conferencing tool. Its key feature is that it transcribes participants' speech in real time, analyzes the text data to generate meeting progress actions, and provides feedback to participants, thereby helping to ensure smooth progress of the meeting. Furthermore, by combining it with an emotion engine, it is possible to analyze participants' emotions and reflect them in the progress actions.
[1271] Key Components Used
[1272] The following describes the major components for implementing the present invention.
[1273] server
[1274] The server is the core of this system and has the following main functions:
[1275] 1. Ability to transcribe voice data in real time: Uses a speech recognition library (e.g., Google Speech Recognition API) to convert participants' voice data into text data.
[1276] 2. The ability to analyze text data and emotional data using a generative AI model and generate next steps: The generative AI model uses natural language processing and machine learning based on the input data to derive the appropriate next action.
[1277] 3. Ability to analyze emotional data using an emotion engine: Identify the user's emotional state from voice tone, speaking rate, facial expression changes, etc.
[1278] Terminal
[1279] The terminal is a device that provides a user interface and has the following functions:
[1280] 1. Function to receive input from the user and send it to the server: Provides an interface for inputting the purpose and goals of the meeting, the path of the audio file, etc.
[1281] 2. A function to present feedback to the user based on the generated progress actions and emotions: Feedback sent from the server is displayed in real time.
[1282] User
[1283] The user sets the purpose and goal of the meeting and is responsible for running the meeting. The user also provides data on comments and emotions made during the meeting, and the system guides the meeting according to the actions taken by the system.
[1284] Specific examples
[1285] For example, consider a user who wants to hold a meeting titled "New Product Launch Plan." The user enters the following information into the web conferencing tool's interface:
[1286] Meeting purpose: "New product launch plan"
[1287] Goal: "Formulate market strategies and assess risks"
[1288] Audio file: "audio_new_product_meeting.wav"
[1289] Prompt Sentence Examples
[1290] "Planning the next project"
[1291] "Define key tasks and determine who is responsible"
[1292] "Marketing Strategy Meeting"
[1293] "Target Market Definition and Competitive Analysis"
[1294] The device sends this information to a server. The server reads the audio file and transcribes it using a speech recognition library. It then uses an emotion engine to analyze the user's emotions from the acquired transcription and video data. The analysis results are sent to a generative AI model, which generates the next action to proceed.
[1295] For example, if a user is nervous, the generative AI model will generate a progression action such as "suggesting questions to relax the user." It may also suggest questions to dig deeper into marketing strategies. The generated progression actions and feedback are sent to the device and displayed to the user in real time.
[1296] This allows users to smoothly progress with the conference, and improves the efficiency and effectiveness of the conference.
[1297] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1298] Step 1:
[1299] The user sets the purpose and goal of the meeting through the interface of the web conferencing tool and inputs the path of the audio file. As input, the user provides the purpose of the meeting (e.g., "Plan for the market launch of a new product"), the goal (e.g., "Formulate market strategy and assess risks"), and the path of the audio file (e.g., "audio_new_product_meeting.wav"). This information is saved on the device for the next step.
[1300] Step 2:
[1301] The terminal sends the information input by the user to the server. The input information includes the purpose and goal of the meeting, and the path to the audio file. The terminal appropriately packets this information and sends it to the server. The output is the meeting information sent to the server.
[1302] Step 3:
[1303] The server receives the audio file path and loads it. Then, it uses a speech recognition library (e.g., Google Speech Recognition API) to transcribe the audio data in real time. The input is the audio file path "audio_new_product_meeting.wav" and the output is the transcribed text data (e.g., "The discussion about the new product has begun...").
[1304] Step 4:
[1305] The server inputs the acquired transcription data and video data into the emotion engine. The emotion engine analyzes the tone of voice, speech rate, and facial expression changes to generate the user's emotion data. The inputs are transcription data and video data, and the output is the analyzed emotion data (e.g., "The user is nervous").
[1306] Step 5:
[1307] The server sends the transcribed text data and emotion data to the generative AI model, which generates the next progress action based on this data. The input is the transcribed text and emotion data, and the output is the generated progress action (e.g., "Suggest questions to help the user relax").
[1308] Step 6:
[1309] The server transmits the generated progress action and emotion-based feedback to the terminal. The input is the progress action and emotion feedback, and the output is the data transmitted to the terminal.
[1310] Step 7:
[1311] The device displays the progress actions and emotional feedback received from the server to the user. Specifically, they are displayed as notifications or pop-ups on the interface and provided to the user in real time. The input is the progress actions and emotional feedback from the server, and the output is the information displayed to the user.
[1312] (Application example 2)
[1313] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1314] With conventional web conferencing tools, meeting efficiency and progress management are largely manual, which means that it takes time to organize what is being said and decide on actions to take. There are also concerns that the quality of meetings may decline due to the difficulty of properly understanding participants' emotions and smoothly progressing the meeting. Furthermore, when dealing with customers in virtual stores, it is difficult to suggest appropriate actions to take in response to customer questions, making it difficult to increase customer satisfaction.
[1315] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1316] In this invention, the server includes means for transcribing the participants' voices in real time, means for analyzing the transcribed text using a generative AI model and generating the next action to proceed with the conference, means for presenting the generated action to proceed with the conference to the participants in real time, means for analyzing the emotions of the participants using an emotion engine and providing appropriate feedback based on the analysis results, means for monitoring the customer service process in the virtual store and generating appropriate actions in response to customer questions, and means for presenting the generated response actions and emotion feedback to the customers in real time. This makes it possible to efficiently manage the progress of the conference and provide optimal actions in response to the emotions of the participants, as well as improve the quality of customer service in the virtual store and increase customer satisfaction.
[1317] "Web conferencing tool" refers to software or a platform that allows multiple participants to hold real-time conferences using audio and video over the Internet.
[1318] A "generative AI model" refers to an algorithm or system that uses artificial intelligence to analyze input data and automatically generate appropriate responses or actions.
[1319] A "meeting efficiency system" refers to a system that combines technologies and means to support the progress and management of meetings and improve their efficiency.
[1320] "Real-time transcription means" refers to technology or equipment for instantly converting audio data into text data.
[1321] "Progressive actions" refer to specific actions or instructions that should be taken next in the progress of a meeting.
[1322] An "emotion engine" refers to technology and algorithms that analyze data such as participants' voices and facial expressions to estimate their emotional state.
[1323] "Feedback" refers to information or instructions provided to users or participants based on analysis results, progress actions, etc.
[1324] "Virtual store" refers to a virtual store that offers products and services via the Internet.
[1325] A "customer service process" refers to a set of procedures or methods for appropriately responding to customer inquiries and questions.
[1326] This invention is a system that improves the efficiency of meetings by using a generative AI model that works in conjunction with a web conferencing tool, and further combines it with an emotion engine to more effectively progress the meeting. This system has the function of transcribing participants' voices in real time, analyzing the text data, and generating the next progress action. Furthermore, by presenting the generated progress action to participants in real time and combining it with an emotion engine that recognizes the user's emotions, the meeting can be progressed smoothly.
[1327] Key Components Used
[1328] 1. Server: Transcribing voice data, analyzing text data using a generative AI model, and recognizing user emotions using an emotion engine. Specifically, it uses a speech recognition library (e.g., Google Speech Recognition API) and an emotion analysis engine (e.g., EmotionEngine). It also generates progress actions using a generative AI model (e.g., ActionsGenerator).
[1329] 2. Device: Receives input from the user and sends it to the server. It also provides feedback to the user based on the generated progress actions and emotions. Devices can be smartphones, head-mounted displays, smart glasses, etc.
[1330] 3. User: Sets the purpose and goals of the meeting, provides emotional data and advances the meeting based on progress actions.
[1331] 4. Emotion engine: Analyzes the user's voice and facial expressions and sends the emotional data to a generative AI model.
[1332] Program processing explanation
[1333] 1. The server transcribes the user's voice data using the Google Speech Recognition API. This process converts the voice data into text data.
[1334] 2. The emotion engine then analyzes the voice tone and facial expressions to generate the user's emotion data. This is done using the EmotionEngine.
[1335] 3. The server uses a generative AI model to analyze the transcribed text data and emotion data, and generates the next action to proceed. Using ActionsGenerator, the generated action is provided in real time.
[1336] 4. The device presents the progress actions and emotional feedback received from the server to the user in real time. This information is displayed via a smartphone, head-mounted display, or smart glasses.
[1337] Specific examples
[1338] For example, suppose a user wants to hold a meeting titled "New Product Launch Plan." In this case, the user enters the following information into the web conferencing tool's interface:
[1339] Meeting purpose: "New product launch plan"
[1340] Goal: "Formulate market strategies and assess risks"
[1341] Audio file: "audio_new_product_meeting.wav"
[1342] This information is sent from the device to the server. The server transcribes the audio file and analyzes the user's emotions using an emotion engine. The generative AI model then analyzes this data and generates a progress action, such as "Please propose specific measures for market strategy." If the user is nervous, it will also suggest questions to relax them (e.g., "Tell me about a recent success story"). This action and emotional feedback are sent to the device and displayed to the user.
[1343] Prompt Sentence Examples
[1344] Below are some example prompts to use with generative AI models:
[1345] text
[1346] Voice input: "I want to know about my new TV."
[1347] Emotional data: interesting but a bit confusing
[1348] Generate Action: Describe the features of their new TV. Ask about their specific needs.
[1349] This system efficiently organizes statements and emotional data during meetings and allows appropriate actions to be taken in real time, further improving the efficiency and effectiveness of meetings. It also makes it possible to generate and provide appropriate actions in real time when dealing with customers in virtual stores, thereby increasing customer satisfaction.
[1350] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1351] Step 1:
[1352] The server receives the user's voice data. The user inputs the purpose, goal, and audio file of the meeting into the terminal. The terminal sends this information to the server. Specific inputs include the audio file (e.g., "audio_new_product_meeting.wav"), the purpose of the meeting, "a plan to launch a new product," and the goal, "formulation of a market strategy and risk assessment." The server receives this data and begins processing.
[1353] Step 2:
[1354] The server transcribes the received voice data in real time using the Google Speech Recognition API. Text data (output) is generated from the voice data (input). This transcribed text data serves as the basis for generating actions to progress the meeting. Converting the input voice data into text makes subsequent data analysis easier.
[1355] Step 3:
[1356] The server uses an emotion engine to analyze the user's emotional data along with the transcribed voice data. Emotion analysis is performed using the Emotion Engine. Specifically, it analyzes changes in voice tone and facial expressions. Through this process, the server obtains emotional data (output) from the voice data (input). For example, it determines whether the user is nervous, relaxed, or interested. At this stage, data enrichment is performed.
[1357] Step 4:
[1358] The server uses a generative AI model (e.g., ActionsGenerator) to analyze the transcribed text data and emotional data. Specifically, the generative AI model generates the next action to proceed (output) based on the text data (input) and emotional data (input). For example, an action such as "Please propose specific measures for market strategy" may be generated. This action provides specific instructions to ensure the smooth flow of the meeting.
[1359] Step 5:
[1360] The server sends the generated progress action and emotional feedback to the device. Specifically, the generated action (output) and emotional feedback (output) are sent. For example, the action may be "Please propose specific measures for the market strategy" and the emotional feedback may be "The user is nervous." This allows the device to process the next step.
[1361] Step 6:
[1362] The device presents the progress actions and emotional feedback received from the server to the user in real time. Specifically, information is displayed to the user (output) via a smartphone, head-mounted display, or smart glasses. For example, the device may provide the user with instructions such as "Please propose specific measures for market strategy" or feedback such as "The user seems nervous, so we suggest questions to help them relax." This allows the user to effectively proceed with the next action in the meeting.
[1363] Step 7:
[1364] The user controls the meeting based on the progress actions and emotional feedback presented on the device. By taking specific actions, the user can effectively move forward toward the purpose and goal of the meeting. The user uses the presented information to execute the next step instructions. This is the stage at which the actual meeting progresses.
[1365] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1366] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1367] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1368] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1369] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1370] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1371] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1372] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1373] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1374] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1375] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1376] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1377] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1378] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1379] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1380] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1381] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1382] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1383] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1384] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1385] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1386] The following is further disclosed regarding the above embodiment.
[1387] (Claim 1)
[1388] A meeting efficiency improvement system that uses a generative AI model that works in conjunction with a web conferencing tool,
[1389] A means of transcribing participants' voices in real time;
[1390] A means for analyzing the transcribed text by a generative AI model and generating next actions for proceeding with the meeting;
[1391] A means for presenting the generated conference progress actions to participants in real time;
[1392] A system including:
[1393] (Claim 2)
[1394] 10. The system of claim 1, further comprising means for utilizing a cloud-based system to perform processing in real time.
[1395] (Claim 3)
[1396] 10. The system of claim 1, further comprising means for incorporating user feedback and continuously improving the accuracy of the generated AI model.
[1397] "Example 1"
[1398] (Claim 1)
[1399] a means for users to set meeting objectives and goals;
[1400] means for a user to upload audio data of a conference;
[1401] means for transcribing audio data received from a subject in real time;
[1402] A means for analyzing the transcribed text data by the generative AI model and generating the next action to proceed with the meeting;
[1403] means for presenting the generated conference progress actions to a user in real time;
[1404] A system including:
[1405] (Claim 2)
[1406] 10. The system of claim 1, further comprising means for utilizing a cloud-based system to perform processing in real time.
[1407] (Claim 3)
[1408] 10. The system of claim 1, further comprising means for incorporating user feedback and continuously improving the accuracy of the generated AI model.
[1409] "Application Example 1"
[1410] (Claim 1)
[1411] A meeting efficiency improvement system and a driving assistance system that use a generative AI model that works in conjunction with a web conferencing tool,
[1412] A means of transcribing participant and driver voices in real time;
[1413] A means for analyzing the transcribed text by the generative AI model and generating next actions for proceeding with the meeting or driving assistance actions;
[1414] A means for presenting the generated meeting progress actions and driving assistance actions to participants and drivers in real time;
[1415] A system including:
[1416] (Claim 2)
[1417] 10. The system of claim 1, further comprising means for utilizing a cloud-based system to perform processing in real time.
[1418] (Claim 3)
[1419] 10. The system of claim 1, further comprising means for incorporating user feedback and continuously improving the accuracy of the generated AI model.
[1420] "Example 2: Combining Emotion Engines"
[1421] (Claim 1)
[1422] a means for a user to set the purpose and goals of the meeting;
[1423] A means of transcribing participants' voices in real time;
[1424] A means for analyzing the transcribed text and user sentiment data by the generative AI model and generating the next action to proceed with the meeting;
[1425] a means for presenting feedback based on the generated meeting progress actions and emotions to participants in real time;
[1426] A system including:
[1427] (Claim 2)
[1428] 10. The system of claim 1, further comprising means for utilizing cloud-based processing to process data in real time.
[1429] (Claim 3)
[1430] 10. The system of claim 1, further comprising means for incorporating user feedback and continuously improving the accuracy of the generative AI model.
[1431] "Application example 2 when combining emotion engines"
[1432] New Claims
[1433] (Claim 1)
[1434] A meeting efficiency improvement system that uses a generative AI model that works in conjunction with a web conferencing tool,
[1435] A means of transcribing participants' voices in real time;
[1436] A means for analyzing the transcribed text by a generative AI model and generating next actions for proceeding with the meeting;
[1437] A means for presenting the generated conference progress actions to participants in real time;
[1438] A means for analyzing the emotions of participants using an emotion engine and providing appropriate feedback based on the analysis results;
[1439] A means for monitoring the customer service process in the virtual store and generating appropriate actions in response to customer inquiries;
[1440] A means for presenting the generated response actions and emotional feedback to the customer in real time;
[1441] A system including:
[1442] (Claim 2)
[1443] 10. The system of claim 1, further comprising means for utilizing a cloud-based system to perform processing in real time.
[1444] (Claim 3)
[1445] 10. The system of claim 1, further comprising means for incorporating user feedback and continuously improving the accuracy of the generated AI model. [Explanation of symbols]
[1446] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. A meeting efficiency improvement system that uses a generative AI model that works in conjunction with a web conferencing tool, A means of transcribing participants' voices in real time; A means for analyzing the transcribed text by a generative AI model and generating next actions for proceeding with the meeting; A means for presenting the generated conference progress actions to participants in real time; A system including:
2. The system of claim 1 , further comprising means for utilizing a cloud-based system to perform processing in real time.
3. The system of claim 1 further comprising means for incorporating feedback from a user and continuously improving the accuracy of the generative AI model.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A