system

The system addresses presenter challenges by using an information terminal and server device for real-time audience reaction analysis and feedback, enhancing presentation quality and engagement through immediate adjustments and responses.

JP2026071048APending Publication Date: 2026-04-28SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-16
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Presenters experience stress and difficulty in adjusting presentations based on real-time audience reactions and providing immediate responses to questions, leading to suboptimal presentation quality and engagement.

Method used

A system that includes an information terminal to capture audio and image data, a server device for real-time analysis of audience reactions and feedback generation, and earphones for immediate advice delivery, supporting presentation adjustments and question responses.

Benefits of technology

Enhances presentation quality by allowing presenters to adapt in real-time to audience feedback and provide accurate answers, improving engagement and overall presentation effectiveness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026071048000001_ABST
    Figure 2026071048000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] The information terminal acquires audio data and image data at the start of the presentation, and means for converting the audio data into text, The server device analyzes the text data and image data in real time and provides means for estimating and reporting the audience's reactions. A means by which an information terminal provides advice based on the report during a presentation via earphones, A system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In the preparation and implementation of a presentation, there are problems such as the stress felt by many presenters and forgetting information during the presentation. There is also a problem that it is difficult to appropriately evaluate the reaction of the audience and quickly adjust the presentation content and progress method on the spot.

Means for Solving the Problems

[0005] This invention provides a system in which an information terminal acquires audio and image data at the start of a presentation and transmits them to a server device. The server device analyzes the transmitted data in real time and estimates the audience's reactions. Based on these estimations, it supports the progress of the presentation by providing advice to the presenter through earphones. Furthermore, by instantly generating optimal example answers to questions and providing improvement suggestions after the presentation, it can reduce the presenter's stress and improve the quality of the presentation.

[0006] An "information terminal" is a device that acquires audio and image data and transmits that data to a server device.

[0007] A "server device" is a central processing unit that analyzes received data, estimates audience reactions, and generates advice and improvement suggestions.

[0008] "Audio data" refers to information recorded in digital format from the voice spoken by the presenter during a presentation.

[0009] "Image data" refers to digital data containing visual information captured during a presentation.

[0010] "Real-time" refers to the fact that data acquisition, analysis, and feedback processes are performed instantly.

[0011] "Audience reaction" refers to the emotional state and level of understanding that can be inferred from the audience's facial expressions and posture during a presentation.

[0012] "Earphones" are audio devices used to provide audio feedback from an information terminal to a presenter.

[0013] "Advice" refers to audio instructions or supplementary information provided during a presentation to support the presenter's progress.

[0014] A "database" is an aggregation of information that is referenced to generate optimal answer examples for questions.

[0015] A "review report" is a report provided after the presentation, which includes an evaluation of the entire presentation and suggestions for improvement for the next time.

Brief Description of Drawings

[0016] [Figure 1] It is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] It is a conceptual diagram showing an example of the main functions of a data processing device and a smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which multiple emotions are mapped. [Figure 10] It shows an emotion map to which multiple emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13]It is a sequence diagram showing the processing flow of the data processing system in Example 2 when combined with an emotion engine. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when combined with an emotion engine.

Embodiment for Carrying Out the Invention

[0017] Hereinafter, an example of an embodiment of the system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0018] First, the terms used in the following description will be explained.

[0019] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), etc.

[0020] In the following embodiments, the numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0021] In the following embodiments, the numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, etc.

[0022] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0023] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0024] [First Embodiment]

[0025] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0026] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0027] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0028] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0029] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0030] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0031] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0032] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0033] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0034] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0035] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0036] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0037] This invention integrates an information terminal, a server device, and earphones into a single system to support presentations. At the start of a presentation, this system uses the camera and microphone equipped on the information terminal to capture the presenter's audio and video. This information is transmitted to the server device in real time.

[0038] The server device first utilizes speech recognition technology to convert received audio data into text. Simultaneously, it analyzes video data to understand the slides and the presenter's movements, and analyzes the audience's facial expressions and attitudes to estimate their reactions. This allows it to grasp the progress of the presentation and the audience's level of understanding and interest.

[0039] The server device generates appropriate advice based on the analysis results. This advice includes specific instructions for adjusting the presentation flow and information on points that require further explanation. The generated advice is converted into audio data using speech synthesis technology and transmitted to the earphones via the information terminal.

[0040] Users can receive direct feedback through their earphones and adjust their presentations in real time. For example, if it is estimated that the audience's interest is waning, the presenter can reintroduce topics or ask questions based on advice from the server device.

[0041] In the Q&A session, questions from the audience are collected via information terminals and analyzed by a server device. The server device uses a database to quickly generate optimal answer examples and provides supplementary information to the user through earphones. This allows the user to answer questions with confidence.

[0042] After the presentation ends, the server device organizes all the data and generates a review report. This report includes an overall evaluation of the presentation, audience reaction trends, Q&A performance, and specific suggestions for improvement for the next presentation. Information terminals receive this report and present it to the user, allowing them to use it to prepare for their next presentation.

[0043] This system enables users to deliver effective and smooth presentations, increasing audience engagement and satisfaction.

[0044] The following describes the processing flow.

[0045] Step 1:

[0046] The information terminal activates its camera and microphone at the start of the presentation. This initiates the acquisition of the presenter's audio and video data. The acquired data is temporarily stored in a buffer.

[0047] Step 2:

[0048] The information terminal transmits the collected audio data to the server device in real time. Video data is also transmitted to the server device simultaneously, capturing the presenter's movements and slide content.

[0049] Step 3:

[0050] The server device processes the received audio data using a speech recognition engine and converts it into text data. This creates a sequential record of the words spoken by the presenter.

[0051] Step 4:

[0052] The server device analyzes video data to detect the content of the slides and the presenter's movements. It also identifies the faces of the audience and estimates their reactions by analyzing changes in their facial expressions.

[0053] Step 5:

[0054] The server integrates the presenter's speech, slide data, and audience reaction data to evaluate the progress of the presentation. Based on this evaluation, it generates advice for the presenter.

[0055] Step 6:

[0056] The generated advice is converted into audio data using speech synthesis technology. The server device transmits this audio data to the information terminal, where it is provided to the user as real-time feedback through earphones.

[0057] Step 7:

[0058] The user receives advice through their earphones and adjusts the presentation accordingly. They encourage audience engagement by supplementing slide explanations or asking questions as needed.

[0059] Step 8:

[0060] The audio questions collected during the Q&A session are sent again from the information terminal to the server device. The server device then transcribes and analyzes the questions into text.

[0061] Step 9:

[0062] The server device matches the question against the database and searches for relevant answer information. It generates the best possible answer example, converts it into audio data, and provides it to the user.

[0063] Step 10:

[0064] After the presentation ends, the server analyzes all session data and generates a review report. This report includes a performance evaluation and improvement suggestions, and is presented to the user via an information terminal.

[0065] (Example 1)

[0066] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0067] In modern presentations, it is extremely difficult for speakers to gauge audience reactions in real time and immediately adjust the content and flow of their presentation. This makes it challenging to maintain audience engagement and effectively convey information. Furthermore, there is insufficient support for providing quick and accurate answers during Q&A sessions. These challenges hinder efficient and effective communication and contribute to a decline in presentation quality.

[0068] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0069] In this invention, the server includes means for converting audio information into text format, means for analyzing and reporting participant reactions, and means for generating advice using generative AI technology. This allows the presenter to grasp audience reactions in real time during a presentation and adjust the progress accordingly. Furthermore, by receiving immediate feedback, the presenter can maintain participant interest and achieve effective communication. In addition, by providing quick and accurate answers during Q&A sessions, it enables high-quality presentations.

[0070] An "information terminal" is a device that acquires audio and video information at the start of a presentation and transmits it to a server.

[0071] "Audio information" refers to data used to capture speech during a presentation.

[0072] "Video information" refers to data used to record and analyze video footage from a presentation.

[0073] "Text format" refers to character data obtained by converting audio information.

[0074] A "server" is a device that processes received audio and video information and performs necessary analysis and reporting.

[0075] "Generative AI technology" is a technology that uses artificial intelligence to generate useful advice and information from data.

[0076] "Advice" refers to specific instructions or suggestions for optimizing the flow and content of a presentation.

[0077] "Earphones" are devices that output advice from the server as audio and convey it directly to the speaker.

[0078] This invention is based on a system combining an information processing device, a data processing device, and an audio output device, and aims to improve the quality of presentations. Specific embodiments are shown below.

[0079] The device will use its camera and microphone to acquire audio and video information at the start of the presentation. The acquired audio information will be converted into text format using speech recognition technology. Specifically, it is expected that speech recognition software such as "Google® Speech-to-Text API" will be used.

[0080] The server processes text and video information received from the terminal. During this process, it analyzes participants' reactions using a data processing library (for example, "OpenCV" is used for video processing). This allows for the estimation of participants' levels of interest and understanding. Next, it utilizes generative AI technology (such as "OpenAI® GPT") to automatically generate appropriate advice based on the analysis results. This advice includes instructions and suggestions for adjusting the presentation's progress.

[0081] The generated advice is converted into speech using speech synthesis technology (such as "Amazon Polly") and output directly through the user's earphones. This allows the user to make immediate adjustments based on the feedback received during the presentation.

[0082] For example, if the server determines that the audience's interest is waning during a presentation, it can provide the user with advice via their headphones, such as "Please add more interactive segments."

[0083] An example of a prompt message could be something like, "The audience's attention is waning. Should we add some relevant interactive content?" which could then be input into a generative AI model.

[0084] This system allows users to deliver more dynamic and effective presentations, thereby increasing participant satisfaction.

[0085] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0086] Step 1:

[0087] When the device detects the start of a presentation, it activates the camera and microphone, collecting video and audio information, respectively. The input is the audio and images detected by the device, and the output is this data. The audio information captured by the microphone is converted into text in real time using speech recognition technology. This process prepares the initial data necessary for transmission to the server.

[0088] Step 2:

[0089] The terminal sends the acquired text and video information to the server. The server starts analysis based on the information obtained by converting the received audio data into text. The input is text and video information, and the output is processed data for analysis. The server analyzes the data using a generative AI model and estimates the participants' reactions. Specifically, it analyzes the participants' facial expressions and attitudes from the video and determines their level of interest from the audio content.

[0090] Step 3:

[0091] The server uses an AI model based on the analysis results to generate advice that optimizes the presentation flow. The input is the analysis results, and the output is specific advice on how to proceed. The generated advice is converted into audio data using speech synthesis technology. This process prepares the audio feedback for the user.

[0092] Step 4:

[0093] The audio data generated by the server is transmitted to the user's earphones via the terminal. The user adjusts their presentation based on this advice. In this step, the user is required to follow the instructions received in real time and find ways to maintain the participants' interest.

[0094] Step 5:

[0095] During the Q&A session of the presentation, the device collects question information from participants and sends it to the server. The server analyzes the question, consults a database, and generates the most appropriate example answer. The input is the question information, and the output is the example answer. The user receives this answer through their earphones and can answer the participants with confidence.

[0096] Step 6:

[0097] After the presentation ends, the server aggregates all the data and generates a review report. The input consists of audio, video, and analytical data collected during the presentation, and the output is the evaluation report. The information terminal displays this report to the user, who can then use it to prepare for the next presentation. This final step allows for identifying areas for improvement and enabling the user to continue delivering high-quality presentations.

[0098] (Application Example 1)

[0099] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0100] In product introductions and demonstrations at trade shows and physical stores, there is a need for a system that can instantly grasp participant reactions and allow presenters to receive real-time feedback. Traditional methods have the challenge of making it difficult to respond flexibly to participants' interests and levels of understanding, making it difficult to maximize the effectiveness of the exhibition.

[0101] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0102] In this invention, the server includes means for an information processing device to acquire audio and video information in an exhibit and convert the audio information into text, means for a data processing device to analyze the text and video information in real time, estimate and notify participants of their responses, and means for the information processing device to supply advice based on the notification during the exhibit via an audio output device. This allows presenters to immediately grasp participants' reactions and adjust their presentations on the spot.

[0103] An "information processing device" is a device that acquires audio and video information and converts it into digital signals.

[0104] An "audio output device" is a device that outputs audio data and plays a role in providing real-time audio feedback to the presenter.

[0105] A "data processing device" is a device that analyzes acquired digital signal data in real time and generates necessary reports and advice based on the analysis results.

[0106] "Exhibition" refers to activities that showcase products and information in various ways at physical stores or trade shows.

[0107] "Participants" refers to people who receive product introductions and demonstrations at exhibitions or physical stores.

[0108] "Estimating and notifying responses" means analyzing participants' reactions and behaviors and informing the presenter of the results based on that analysis.

[0109] "Providing advice" means providing the presenter with specific instructions and suggestions for improvement via audio, based on the analysis results.

[0110] The system for carrying out this invention includes an information processing device (such as smart glasses), a data processing device (server computer), and an audio output device (earphones). The role of each device and how they work together to function in this system are described below.

[0111] The information processing device is used in the form of smart glasses worn by the presenter and acquires audio and video information in real time. Using the built-in camera and microphone, it captures the participants' reactions and the presenter's movements, and collects the audio. The acquired data is immediately transmitted to the data processing device.

[0112] The data processing unit is implemented on a cloud server and uses the Google Speech-to-Text engine to convert audio information into text. It also analyzes collected video information using the OpenCV library, estimating participant reactions by analyzing facial expressions and eye movements. Based on the analyzed data, Amazon Polly is used to synthesize necessary advice, generating information to help the presenter decide on their next course of action.

[0113] The audio output device is an earphone worn by the presenter, which provides generated audio feedback in real time. This allows the presenter to instantly adjust the content and flow of the presentation in response to participant reactions.

[0114] For example, imagine a presenter explaining a new product to a customer in a cosmetics store. If the data processing device determines that the participant's expression shows no interest, it will send a prompt to the presenter via an audio output device, such as "Emphatize the product's scent" or "Demonstrate its use." This prompt might take the form of: "The customer's interest is waning, please present a new use example."

[0115] This system allows presenters to instantly incorporate participant reactions and conduct more effective demonstrations.

[0116] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0117] Step 1:

[0118] The information processing terminal uses a camera and microphone to collect audio and video information of the presentation in real time. It captures ambient sounds and visual scenes as input and transmits the data to the server in digital format. The data acquired includes the presenter's words and participants' reactions, which are then transmitted to the server using wireless communication technology.

[0119] Step 2:

[0120] The server converts the received audio information into text using the Google Speech-to-Text engine. It receives audio data as input, performs phoneme analysis and natural language processing, and generates text data as output. The specific actions performed in this step are the analysis of the audio waveform and the identification of phonemes.

[0121] Step 3:

[0122] The server analyzes the received video information using the OpenCV library. It takes a video stream as input, performs face recognition and facial expression analysis frame by frame, and generates metrics indicating the participants' level of interest as output. Specifically, it performs facial feature point detection and gaze direction estimation.

[0123] Step 4:

[0124] The server integrates the transcribed speech text with the analyzed video metrics to estimate the participant's response. The input consists of text data and video metrics, and based on this, it performs a response evaluation and generates advice, including prompts, as output. This step involves statistical data analysis and predictions using generative AI models.

[0125] Step 5:

[0126] The server converts advice generated using Amazon Polly into audio data. It receives advice text as input, synthesizes a natural-sounding waveform, and produces audio data as output. This operation involves applying a speech synthesis algorithm.

[0127] Step 6:

[0128] The generated audio data is provided to the presenter (user) through an audio output device. The user receives this real-time feedback through earphones and adjusts the presentation accordingly. The specific output is a direct audio message prompting the presenter to take action.

[0129] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0130] This invention provides a system that effectively supports presentations by combining an information terminal, a server device, and an emotion engine. The system begins at the start of a presentation when the information terminal acquires the presenter's audio and image data in real time and transmits it to the server device.

[0131] The server device converts received audio data into text using speech recognition technology. It also analyzes video data to understand the presenter's gestures and slide content, and further analyzes the audience's facial expressions and movements to estimate their reactions.

[0132] The emotion engine, a key feature of this invention, analyzes the presenter's emotional state based on data transmitted from an information terminal. This emotion analysis is estimated from the user's tone of voice, facial expressions, and body movements. If the emotion engine determines that the user is feeling stressed, it generates content to promote relaxation and provides the user with advice to regain their composure.

[0133] The server device uses the analysis results from the emotion engine to further refine the advice given to the user. This advice includes suggestions for improving the presentation flow and specific ways to capture the audience's attention. The generated advice is converted into audio data using speech synthesis technology and delivered to the user via earphones through an information terminal.

[0134] For example, if a user receives an unexpected question during a presentation, the emotion engine can sense the user's temporary confusion, and the server device can immediately provide support by offering suggested answers and relevant information, helping the user to calmly answer the question.

[0135] Furthermore, after the presentation ends, the server device analyzes all the data and generates a review report. This report includes user emotional changes, performance evaluations, and suggestions for improvement for the next presentation, and is presented to the user via an information terminal.

[0136] This system allows users to receive real-time feedback tailored to their emotional state, enabling them to deliver more effective and smoother presentations.

[0137] The following describes the processing flow.

[0138] Step 1:

[0139] The device activates its camera and microphone simultaneously with the start of the presentation, and begins acquiring audio and image data. The data collected in real time is temporarily stored in a buffer.

[0140] Step 2:

[0141] The device sends the collected audio data to the server. This data includes the presenter's voice during the presentation. Image data is also sent to the server simultaneously, including presentation slides and audience reactions.

[0142] Step 3:

[0143] The server analyzes the received audio data using a speech recognition engine and converts it into text. Furthermore, it analyzes the presenter's gestures and slide content from the image data.

[0144] Step 4:

[0145] Based on the analyzed audio and image data, the server uses an emotion engine to determine the user's emotional state. The emotion engine estimates stress and tension from the tone of voice and facial features.

[0146] Step 5:

[0147] The server uses information from the emotion engine to generate advice for the user, including suggestions on how to conduct a presentation and recommendations for relaxation based on their emotional state.

[0148] Step 6:

[0149] The generated advice is synthesized into speech, sent to the earphones via the device, and provided to the user in real time. This allows the user to receive immediate feedback as the presentation progresses.

[0150] Step 7:

[0151] The user progresses with their presentation based on advice received through the earphones, adjusting the content as needed. This increases audience engagement and improves the quality of the presentation.

[0152] Step 8:

[0153] After the presentation ends, the server analyzes all collected data and generates a review report. This report includes user emotional changes, an overall summary of the presentation performance, and suggestions for improvement for the next presentation.

[0154] Step 9:

[0155] The generated review report is provided to the user via their device, allowing them to use it to prepare for their next presentation. This process promotes continuous improvement and learning.

[0156] (Example 2)

[0157] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0158] During a presentation, presenters often struggle to manage their own emotions and respond quickly to audience reactions. They are also required to provide accurate answers to unexpected questions, and obtaining appropriate information instantly is difficult. Furthermore, insufficient post-presentation review prevents effective improvement for future presentations.

[0159] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0160] This invention includes a server that provides means for estimating and reporting audience reactions in real time, means for suggesting appropriate information candidates when questions are received during the presentation, and means for creating an evaluation report based on aggregated data after the presentation and sending suggestions for optimizing the next presentation. This allows presenters to adjust their responses according to their emotional state and conduct presentations based on audience reactions. Furthermore, it enables the quality of presentations to be improved by obtaining appropriate answers to questions immediately. In addition, it allows for efficient review and improvement for the next presentation after it has ended.

[0161] An "information processing device" is a device that has the function of acquiring audio signals and visual data and converting them into text information as needed.

[0162] A "processing unit" is a device that analyzes acquired textual and visual data in real time and processes data to estimate the participants' reactions.

[0163] An "auditory transmission device" is a device that transmits instructions provided by an information processing device to the user as audio.

[0164] A "emotion analysis device" is a system that monitors the user's emotional state and generates content that promotes relaxation as needed.

[0165] A "user" is a person who gives a presentation or operates its assistive devices, and who uses the system to solve a problem.

[0166] "Attendees" refers to the people who watch the presentation, and their reactions are the ones whose responses are perceived.

[0167] "Character information" refers to string data converted from audio signals, and it has a format that can be used as text.

[0168] "Visual data" refers to video information acquired by video cameras or similar devices, and is the data format that is subject to analysis.

[0169] "Instructions" refers to advice and guidelines provided to the user through auditory communication devices.

[0170] "Appropriate information candidates" are a collection of highly relevant and useful information presented in response to questions asked by the user.

[0171] An "evaluation report" is a report intended to review a presentation, including areas for improvement and suggestions for optimization for the next presentation.

[0172] An "optimization proposal" is a suggestion that outlines specific areas for improvement and ways to streamline the next presentation.

[0173] This system utilizes advanced information processing and computing power to support presentations. The primary devices used are smart devices (e.g., smartphones, tablets) as information processing units, and the computing power unit paired with them consists of a high-performance server computer.

[0174] Data acquisition by information processing equipment:

[0175] The device uses its built-in camera and microphone to acquire the presenter's audio signal and visual data in real time. A dedicated application on the device is used for this purpose, and the audio signal is compressed in AAC encoding format, while the visual data is compressed in JPEG or H.264 format.

[0176] Audio and video data transmission and analysis:

[0177] The device transmits the acquired data to the server via the internet. The server converts the audio into text using the Google Cloud Speech-to-Text API and analyzes the visual data using OpenCV. This allows for real-time tracking of user gestures and participant reactions.

[0178] Estimation of emotional state:

[0179] The server works in conjunction with an emotion analysis device to evaluate the user's emotional state based on their voice tone and facial movements. For example, if the user is feeling tense, the emotion analysis device generates content to promote relaxation and provides appropriate advice. This utilizes AI-powered natural language processing.

[0180] Actual advice and support:

[0181] The server provides the user with guidelines for the generated audio through an auditory transmission device (e.g., earphones) via the terminal. If the user receives an unexpected question during the presentation, the processing unit immediately presents relevant information candidates and suggested answers.

[0182] Examples of execution and prompts for the generated AI model:

[0183] For example, if a user tends to move their head a lot, the server might offer advice such as, "It would be good to tone down your gestures." Another example of a prompt for the generative AI model could be, "Think of ways to help a nervous user relax during a presentation."

[0184] This system allows presenters to respond appropriately to their emotions, enabling them to deliver more effective presentations and build upon their performance for future events.

[0185] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0186] Step 1:

[0187] Data acquisition and transmission

[0188] The device collects the presenter's audio signal and visual data in real time. During this process, the device records audio using its built-in microphone and captures video with its camera. It acquires audio and video data as input, encodes the audio in AAC format, and compresses the video in JPEG or H.264 format. This compressed data is then transmitted to the server via the internet.

[0189] Step 2:

[0190] Speech recognition and text conversion

[0191] The server analyzes the audio data received from the terminal. The input audio data is converted into text using the Google Cloud Speech-to-Text API. The output is the presenter's speech in text format.

[0192] Step 3:

[0193] Video analysis and reaction estimation

[0194] The server analyzes visual data using OpenCV. Using the input video data, it detects the presenter's gestures and the audience's facial expressions, and estimates their reactions. Specifically, it uses face detection and motion recognition algorithms to determine whether the presenter is pointing to a slide and how the audience is reacting. This generates a text report of the reactions as output.

[0195] Step 4:

[0196] Emotion analysis and relaxation support

[0197] The server uses an emotion analysis device to evaluate the user's emotions. It takes voice tone and facial expression data obtained from video as input, and uses an AI algorithm to estimate the level of stress and tension. As output, it generates content and advice to help the user relax. Specifically, it creates suggestions for calming music and audio guides to encourage deep breathing.

[0198] Step 5:

[0199] Question support

[0200] The server handles questions from users during presentations. It retrieves the question content and related database information as input, performs calculations, and generates the most relevant information candidates. The output presents immediately usable answer suggestions for the user, specifically including information from similar past questions and authoritative sources.

[0201] Step 6:

[0202] Evaluation report generation and feedback

[0203] The server generates an evaluation report based on data collected after the presentation. It uses aggregated data such as audio, video, user emotional states, and audience reactions recorded during the presentation as input. This data is analyzed, and the output is a detailed report including suggestions for optimizing future presentations. Specifically, this involves identifying and analyzing successes and areas for improvement.

[0204] Through this series of steps, the system helps presenters deliver more effective presentations.

[0205] (Application Example 2)

[0206] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0207] In the workplace, the quality of work can be affected by fluctuations in workers' emotions and decreased work efficiency. Therefore, there is a need for a system that monitors workers' emotions and efficiency in real time and provides appropriate advice to improve work efficiency.

[0208] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0209] In this invention, the server includes means for an information terminal to acquire audio and image data at the start of work and convert the audio data into text; means for the server device to analyze the text and image data in real time, estimate and report the worker's emotions and efficiency; and means for the information terminal to provide advice based on the report during work via a presentation device. As a result, the worker can receive feedback in real time that corresponds to their emotional state, enabling improved work efficiency.

[0210] An "information terminal" is an electronic device used to acquire and process audio and image data.

[0211] "Audio data" refers to digital data that includes sound information acquired from workers and the environment.

[0212] "Image data" refers to digital data that includes visual information acquired at a work site.

[0213] A "server device" is a computer system that analyzes acquired data in real time and performs the necessary processing.

[0214] A "presentation device" is a device used to communicate analysis results and advice to workers, and can provide visual or audio output.

[0215] "Emotion and efficiency estimation" is a process of evaluating a worker's emotional state and work efficiency based on data.

[0216] A "report" is information generated based on analyzed data and provided to the worker.

[0217] "Advice" refers to instructions that provide suggestions for improvement based on the worker's situation.

[0218] In this embodiment of the invention, a system is constructed that combines an information terminal, a server device, and a display device in order to maximize the efficiency of workers in factories and work sites.

[0219] The server device receives audio and image data transmitted from information terminals and analyzes them in real time. The audio data is converted to text using a speech recognition engine such as Google Speech-to-Text. The image data is analyzed using the OpenCV library to capture the movements and facial expressions of workers. This enables the estimation of emotions and work efficiency in the work environment.

[0220] Next, the server device generates advice for the worker based on the results of the emotion and efficiency analysis. The Hugging Face NLP library is used for emotion analysis to assess the worker's stress level, etc. Appropriate advice is generated by speech synthesis software and sent to the worker via a presentation device. This presentation device includes smart glasses and tablets.

[0221] For example, a worker may face a sudden problem and experience temporary stress. In this case, the system, after analysis, advises relaxation techniques such as deep breathing and encourages improvements to work procedures. In this process, the system uses the prompt "What advice would you give if work efficiency decreased?" as a prompt using a generative AI model, allowing the AI ​​to output the most appropriate advice.

[0222] This allows workers to receive feedback tailored to their emotional state during work, thereby improving work efficiency.

[0223] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0224] Step 1:

[0225] The information terminal acquires voice and image data of the worker at the start of work. A high-sensitivity microphone and high-resolution camera are used as input data, which is converted to digital format in real time and transmitted to the server. This allows for the recording of dynamic information about the work environment.

[0226] Step 2:

[0227] The server converts the received audio data into text data using the Google Speech-to-Text engine. Converting the audio input into text format enables subsequent natural language processing (NLP). The text output is used to document the work performed and the situation.

[0228] Step 3:

[0229] The server uses the OpenCV library to analyze image data and evaluate the worker's facial expressions and movements. It performs face recognition and gesture identification from the input image and estimates the emotional state from the results. The output provides estimated values ​​for the worker's stress level and work attitude.

[0230] Step 4:

[0231] The server uses the Hugging Face NLP library to perform sentiment analysis on text data. By analyzing the nuances of emotion, it assesses the degree of stress and confusion. The output is a quantified sentiment score.

[0232] Step 5:

[0233] Based on the analysis results, the server sends prompts to the generated AI model to produce advice for improving work efficiency. Using the prompt, "What advice would you provide if work efficiency decreases?", the AI ​​outputs the optimal advice. The output advice is used as a suggestion for the next action.

[0234] Step 6:

[0235] The information terminal uses speech synthesis software to convert advice into voice and provides it to the worker through a display device. By letting the user listen to the synthesized voice, immediate feedback is achieved. The output is the voice instructions used at the work site.

[0236] These steps allow the system to monitor workers' emotions and efficiency in real time and provide timely advice, thereby maximizing work efficiency.

[0237] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0238] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0239] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0240] [Second Embodiment]

[0241] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0242] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0243] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0244] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0245] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0246] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0247] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0248] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0249] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0250] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0251] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0252] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0253] This invention integrates an information terminal, a server device, and earphones into a single system to support presentations. At the start of a presentation, this system uses the camera and microphone equipped on the information terminal to capture the presenter's audio and video. This information is transmitted to the server device in real time.

[0254] The server device first utilizes speech recognition technology to convert received audio data into text. Simultaneously, it analyzes video data to understand the slides and the presenter's movements, and analyzes the audience's facial expressions and attitudes to estimate their reactions. This allows it to grasp the progress of the presentation and the audience's level of understanding and interest.

[0255] The server device generates appropriate advice based on the analysis results. This advice includes specific instructions for adjusting the presentation flow and information on points that require further explanation. The generated advice is converted into audio data using speech synthesis technology and transmitted to the earphones via the information terminal.

[0256] Users can receive direct feedback through their earphones and adjust their presentations in real time. For example, if it is estimated that the audience's interest is waning, the presenter can reintroduce topics or ask questions based on advice from the server device.

[0257] In the Q&A session, questions from the audience are collected via information terminals and analyzed by a server device. The server device uses a database to quickly generate optimal answer examples and provides supplementary information to the user through earphones. This allows the user to answer questions with confidence.

[0258] After the presentation ends, the server device organizes all the data and generates a review report. This report includes an overall evaluation of the presentation, audience reaction trends, Q&A performance, and specific suggestions for improvement for the next presentation. Information terminals receive this report and present it to the user, allowing them to use it to prepare for their next presentation.

[0259] This system enables users to deliver effective and smooth presentations, increasing audience engagement and satisfaction.

[0260] The following describes the processing flow.

[0261] Step 1:

[0262] The information terminal activates its camera and microphone at the start of the presentation. This initiates the acquisition of the presenter's audio and video data. The acquired data is temporarily stored in a buffer.

[0263] Step 2:

[0264] The information terminal transmits the collected audio data to the server device in real time. Video data is also transmitted to the server device simultaneously, capturing the presenter's movements and slide content.

[0265] Step 3:

[0266] The server device processes the received audio data using a speech recognition engine and converts it into text data. This creates a sequential record of the words spoken by the presenter.

[0267] Step 4:

[0268] The server device analyzes video data to detect the content of the slides and the presenter's movements. It also identifies the faces of the audience and estimates their reactions by analyzing changes in their facial expressions.

[0269] Step 5:

[0270] The server integrates the presenter's speech, slide data, and audience reaction data to evaluate the progress of the presentation. Based on this evaluation, it generates advice for the presenter.

[0271] Step 6:

[0272] The generated advice is converted into audio data using speech synthesis technology. The server device transmits this audio data to the information terminal, where it is provided to the user as real-time feedback through earphones.

[0273] Step 7:

[0274] The user receives advice through their earphones and adjusts the presentation accordingly. They encourage audience engagement by supplementing slide explanations or asking questions as needed.

[0275] Step 8:

[0276] The audio questions collected during the Q&A session are sent again from the information terminal to the server device. The server device then transcribes and analyzes the questions into text.

[0277] Step 9:

[0278] The server device matches the question against the database and searches for relevant answer information. It generates the best possible answer example, converts it into audio data, and provides it to the user.

[0279] Step 10:

[0280] After the presentation ends, the server analyzes all session data and generates a review report. This report includes a performance evaluation and improvement suggestions, and is presented to the user via an information terminal.

[0281] (Example 1)

[0282] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0283] In modern presentations, it is extremely difficult for speakers to gauge audience reactions in real time and immediately adjust the content and flow of their presentation. This makes it challenging to maintain audience engagement and effectively convey information. Furthermore, there is insufficient support for providing quick and accurate answers during Q&A sessions. These challenges hinder efficient and effective communication and contribute to a decline in presentation quality.

[0284] The specific processing by the specific processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0285] In this invention, the server includes means for converting voice information into text format, means for analyzing and reporting the reactions of participants, and means for generating advice using generative AI technology. Thereby, during the presentation, the speaker can grasp the reactions of the audience in real time and adjust the progress. Also, by receiving immediate feedback, the speaker can sustain the interest of the participants and achieve effective communication. Furthermore, by providing quick and accurate answers in the Q&A session, a high-quality presentation can be realized.

[0286] An "information terminal" is a device that acquires voice information and video information at the start of a presentation and transmits it to the server.

[0287] "Voice information" is data for capturing the speech during a presentation.

[0288] "Video information" is data for recording and analyzing the video during a presentation.

[0289] "Text format" refers to character data obtained by converting voice information.

[0290] A "server" is a device that processes the received voice information and video information and performs necessary analysis and reporting.

[0291] "Generative AI technology" is a technology that utilizes artificial intelligence to generate useful advice and information from data.

[0292] "Advice" refers to specific instructions or suggestions for optimizing the progress and content during a presentation.

[0293] An "earphone" is a device that outputs the advice from the server as voice and conveys it directly to the speaker.

[0294] This invention is based on a system combining an information processing device, a data processing device, and an audio output device, and aims to improve the quality of presentations. Specific embodiments are shown below.

[0295] The device will use its camera and microphone to acquire audio and video information at the start of the presentation. The acquired audio information will be converted into text format using speech recognition technology. Specifically, it is expected that speech recognition software such as "Google Speech-to-Text API" will be used.

[0296] The server processes text and video information received from the terminal. During this process, it analyzes participants' reactions using a data processing library (for example, "OpenCV" is used for video processing). This allows for the estimation of participants' levels of interest and understanding. Next, it utilizes generative AI technology (such as "OpenAI GPT") to automatically generate appropriate advice based on the analysis results. This advice includes instructions and suggestions for adjusting the presentation's progress.

[0297] The generated advice is converted into speech using speech synthesis technology (such as "Amazon Polly") and output directly through the user's earphones. This allows the user to make immediate adjustments based on the feedback received during the presentation.

[0298] For example, if the server determines that the audience's interest is waning during a presentation, it can provide the user with advice via their headphones, such as "Please add more interactive segments."

[0299] An example of a prompt message could be something like, "The audience's attention is waning. Should we add some relevant interactive content?" which could then be input into a generative AI model.

[0300] With this system, the user can conduct a more dynamic and effective presentation, enhancing the satisfaction of the participants.

[0301] The flow of the specific process in Example 1 will be described using FIG. 11.

[0302] Step 1:

[0303] When the terminal detects the start of a presentation, it activates the camera and microphone to collect video information and audio information respectively. The input is the voice and image detected by the terminal, and the output is these data. The audio information captured from the microphone is converted into text format in real time by speech recognition technology. This process prepares the initial data necessary for transmission to the server.

[0304] Step 2:

[0305] The terminal transmits the acquired text information and video information to the server. The server starts analysis based on the information obtained by converting the received audio data into text. The input is the text information and video information, and the output is the processed data for analysis. The server analyzes the data using the generated AI model to estimate the reactions of the participants. Specifically, it analyzes the expressions and attitudes of the participants from the video and determines the degree of interest from the audio content.

[0306] Step 3:

[0307] Based on the analysis results, the server utilizes the generated AI model to generate advice for optimizing the progress of the presentation. The input is the analysis result, and the output is specific progress advice. The generated advice is converted into audio data by speech synthesis technology. This process prepares the audio feedback for the user.

[0308] Step 4:

[0309] The audio data generated by the server is transmitted to the user's earphones via the terminal. The user adjusts their presentation based on this advice. In this step, the user is required to follow the instructions received in real time and find ways to maintain the participants' interest.

[0310] Step 5:

[0311] During the Q&A session of the presentation, the device collects question information from participants and sends it to the server. The server analyzes the question, consults a database, and generates the most appropriate example answer. The input is the question information, and the output is the example answer. The user receives this answer through their earphones and can answer the participants with confidence.

[0312] Step 6:

[0313] After the presentation ends, the server aggregates all the data and generates a review report. The input consists of audio, video, and analytical data collected during the presentation, and the output is the evaluation report. The information terminal displays this report to the user, who can then use it to prepare for the next presentation. This final step allows for identifying areas for improvement and enabling the user to continue delivering high-quality presentations.

[0314] (Application Example 1)

[0315] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0316] In product introductions and demonstrations at trade shows and physical stores, there is a need for a system that can instantly grasp participant reactions and allow presenters to receive real-time feedback. Traditional methods have the challenge of making it difficult to respond flexibly to participants' interests and levels of understanding, making it difficult to maximize the effectiveness of the exhibition.

[0317] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0318] In this invention, the server includes means for an information processing device to acquire audio and video information in an exhibit and convert the audio information into text, means for a data processing device to analyze the text and video information in real time, estimate and notify participants of their responses, and means for the information processing device to supply advice based on the notification during the exhibit via an audio output device. This allows presenters to immediately grasp participants' reactions and adjust their presentations on the spot.

[0319] An "information processing device" is a device that acquires audio and video information and converts it into digital signals.

[0320] An "audio output device" is a device that outputs audio data and plays a role in providing real-time audio feedback to the presenter.

[0321] A "data processing device" is a device that analyzes acquired digital signal data in real time and generates necessary reports and advice based on the analysis results.

[0322] "Exhibition" refers to activities that showcase products and information in various ways at physical stores or trade shows.

[0323] "Participants" refers to people who receive product introductions and demonstrations at exhibitions or physical stores.

[0324] "Estimating and notifying responses" means analyzing participants' reactions and behaviors and informing the presenter of the results based on that analysis.

[0325] "Providing advice" means providing the presenter with specific instructions and suggestions for improvement via audio, based on the analysis results.

[0326] The system for carrying out this invention includes an information processing device (such as smart glasses), a data processing device (server computer), and an audio output device (earphones). The role of each device and how they work together to function in this system are described below.

[0327] The information processing device is used in the form of smart glasses worn by the presenter and acquires audio and video information in real time. Using the built-in camera and microphone, it captures the participants' reactions and the presenter's movements, and collects the audio. The acquired data is immediately transmitted to the data processing device.

[0328] The data processing unit is implemented on a cloud server and uses the Google Speech-to-Text engine to convert audio information into text. It also analyzes collected video information using the OpenCV library, estimating participant reactions by analyzing facial expressions and eye movements. Based on the analyzed data, Amazon Polly is used to synthesize necessary advice, generating information to help the presenter decide on their next course of action.

[0329] The audio output device is an earphone worn by the presenter, which provides generated audio feedback in real time. This allows the presenter to instantly adjust the content and flow of the presentation in response to participant reactions.

[0330] For example, imagine a presenter explaining a new product to a customer in a cosmetics store. If the data processing device determines that the participant's expression shows no interest, it will send a prompt to the presenter via an audio output device, such as "Emphatize the product's scent" or "Demonstrate its use." This prompt might take the form of: "The customer's interest is waning, please present a new use example."

[0331] This system allows presenters to instantly incorporate participant reactions and conduct more effective demonstrations.

[0332] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0333] Step 1:

[0334] The information processing terminal uses a camera and microphone to collect audio and video information of the presentation in real time. It captures ambient sounds and visual scenes as input and transmits the data to the server in digital format. The data acquired includes the presenter's words and participants' reactions, which are then transmitted to the server using wireless communication technology.

[0335] Step 2:

[0336] The server converts the received audio information into text using the Google Speech-to-Text engine. It receives audio data as input, performs phoneme analysis and natural language processing, and generates text data as output. The specific actions performed in this step are the analysis of the audio waveform and the identification of phonemes.

[0337] Step 3:

[0338] The server analyzes the received video information using the OpenCV library. It takes a video stream as input, performs face recognition and facial expression analysis frame by frame, and generates metrics indicating the participants' level of interest as output. Specifically, it performs facial feature point detection and gaze direction estimation.

[0339] Step 4:

[0340] The server integrates the transcribed speech text with the analyzed video metrics to estimate the participant's response. The input consists of text data and video metrics, and based on this, it performs a response evaluation and generates advice, including prompts, as output. This step involves statistical data analysis and predictions using generative AI models.

[0341] Step 5:

[0342] The server converts advice generated using Amazon Polly into audio data. It receives advice text as input, synthesizes a natural-sounding waveform, and produces audio data as output. This operation involves applying a speech synthesis algorithm.

[0343] Step 6:

[0344] The generated audio data is provided to the presenter (user) through an audio output device. The user receives this real-time feedback through earphones and adjusts the presentation accordingly. The specific output is a direct audio message prompting the presenter to take action.

[0345] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0346] This invention provides a system that effectively supports presentations by combining an information terminal, a server device, and an emotion engine. The system begins at the start of a presentation when the information terminal acquires the presenter's audio and image data in real time and transmits it to the server device.

[0347] The server device converts received audio data into text using speech recognition technology. It also analyzes video data to understand the presenter's gestures and slide content, and further analyzes the audience's facial expressions and movements to estimate their reactions.

[0348] The emotion engine, a key feature of this invention, analyzes the presenter's emotional state based on data transmitted from an information terminal. This emotion analysis is estimated from the user's tone of voice, facial expressions, and body movements. If the emotion engine determines that the user is feeling stressed, it generates content to promote relaxation and provides the user with advice to regain their composure.

[0349] The server device uses the analysis results from the emotion engine to further refine the advice given to the user. This advice includes suggestions for improving the presentation flow and specific ways to capture the audience's attention. The generated advice is converted into audio data using speech synthesis technology and delivered to the user via earphones through an information terminal.

[0350] For example, if a user receives an unexpected question during a presentation, the emotion engine can sense the user's temporary confusion, and the server device can immediately provide support by offering suggested answers and relevant information, helping the user to calmly answer the question.

[0351] Furthermore, after the presentation ends, the server device analyzes all the data and generates a review report. This report includes user emotional changes, performance evaluations, and suggestions for improvement for the next presentation, and is presented to the user via an information terminal.

[0352] This system allows users to receive real-time feedback tailored to their emotional state, enabling them to deliver more effective and smoother presentations.

[0353] The following describes the processing flow.

[0354] Step 1:

[0355] The device activates its camera and microphone simultaneously with the start of the presentation, and begins acquiring audio and image data. The data collected in real time is temporarily stored in a buffer.

[0356] Step 2:

[0357] The device sends the collected audio data to the server. This data includes the presenter's voice during the presentation. Image data is also sent to the server simultaneously, including presentation slides and audience reactions.

[0358] Step 3:

[0359] The server analyzes the received audio data using a speech recognition engine and converts it into text. Furthermore, it analyzes the presenter's gestures and slide content from the image data.

[0360] Step 4:

[0361] Based on the analyzed audio and image data, the server uses an emotion engine to determine the user's emotional state. The emotion engine estimates stress and tension from the tone of voice and facial features.

[0362] Step 5:

[0363] The server uses information from the emotion engine to generate advice for the user, including suggestions on how to conduct a presentation and recommendations for relaxation based on their emotional state.

[0364] Step 6:

[0365] The generated advice is synthesized into speech, sent to the earphones via the device, and provided to the user in real time. This allows the user to receive immediate feedback as the presentation progresses.

[0366] Step 7:

[0367] The user progresses with their presentation based on advice received through the earphones, adjusting the content as needed. This increases audience engagement and improves the quality of the presentation.

[0368] Step 8:

[0369] After the presentation ends, the server analyzes all collected data and generates a review report. This report includes user emotional changes, an overall summary of the presentation performance, and suggestions for improvement for the next presentation.

[0370] Step 9:

[0371] The generated review report is provided to the user via their device, allowing them to use it to prepare for their next presentation. This process promotes continuous improvement and learning.

[0372] (Example 2)

[0373] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0374] During a presentation, presenters often struggle to manage their own emotions and respond quickly to audience reactions. They are also required to provide accurate answers to unexpected questions, and obtaining appropriate information instantly is difficult. Furthermore, insufficient post-presentation review prevents effective improvement for future presentations.

[0375] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0376] This invention includes a server that provides means for estimating and reporting audience reactions in real time, means for suggesting appropriate information candidates when questions are received during the presentation, and means for creating an evaluation report based on aggregated data after the presentation and sending suggestions for optimizing the next presentation. This allows presenters to adjust their responses according to their emotional state and conduct presentations based on audience reactions. Furthermore, it enables the quality of presentations to be improved by obtaining appropriate answers to questions immediately. In addition, it allows for efficient review and improvement for the next presentation after it has ended.

[0377] An "information processing device" is a device that has the function of acquiring audio signals and visual data and converting them into text information as needed.

[0378] A "processing unit" is a device that analyzes acquired textual and visual data in real time and processes data to estimate the participants' reactions.

[0379] An "auditory transmission device" is a device that transmits instructions provided by an information processing device to the user as audio.

[0380] A "emotion analysis device" is a system that monitors the user's emotional state and generates content that promotes relaxation as needed.

[0381] A "user" is a person who gives a presentation or operates its assistive devices, and who uses the system to solve a problem.

[0382] "Attendees" refers to the people who watch the presentation, and their reactions are the ones whose responses are perceived.

[0383] "Character information" refers to string data converted from audio signals, and it has a format that can be used as text.

[0384] "Visual data" refers to video information acquired by video cameras or similar devices, and is the data format that is subject to analysis.

[0385] "Instructions" refers to advice and guidelines provided to the user through auditory communication devices.

[0386] "Appropriate information candidates" are a collection of highly relevant and useful information presented in response to questions asked by the user.

[0387] An "evaluation report" is a report intended to review a presentation, including areas for improvement and suggestions for optimization for the next presentation.

[0388] An "optimization proposal" is a suggestion that outlines specific areas for improvement and ways to streamline the next presentation.

[0389] This system utilizes advanced information processing and computing power to support presentations. The primary devices used are smart devices (e.g., smartphones, tablets) as information processing units, and the computing power unit paired with them consists of a high-performance server computer.

[0390] Data acquisition by information processing equipment:

[0391] The device uses its built-in camera and microphone to acquire the presenter's audio signal and visual data in real time. A dedicated application on the device is used for this purpose, and the audio signal is compressed in AAC encoding format, while the visual data is compressed in JPEG or H.264 format.

[0392] Audio and video data transmission and analysis:

[0393] The device transmits the acquired data to the server via the internet. The server converts the audio into text using the Google Cloud Speech-to-Text API and analyzes the visual data using OpenCV. This allows for real-time tracking of user gestures and participant reactions.

[0394] Estimation of emotional state:

[0395] The server works in conjunction with an emotion analysis device to evaluate the user's emotional state based on their voice tone and facial movements. For example, if the user is feeling tense, the emotion analysis device generates content to promote relaxation and provides appropriate advice. This utilizes AI-powered natural language processing.

[0396] Actual advice and support:

[0397] The server provides the user with guidelines for the generated audio through an auditory transmission device (e.g., earphones) via the terminal. If the user receives an unexpected question during the presentation, the processing unit immediately presents relevant information candidates and suggested answers.

[0398] Examples of execution and prompts for the generated AI model:

[0399] For example, if a user tends to move their head a lot, the server might offer advice such as, "It would be good to tone down your gestures." Another example of a prompt for the generative AI model could be, "Think of ways to help a nervous user relax during a presentation."

[0400] This system allows presenters to respond appropriately to their emotions, enabling them to deliver more effective presentations and build upon their performance for future events.

[0401] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0402] Step 1:

[0403] Data acquisition and transmission

[0404] The device collects the presenter's audio signal and visual data in real time. During this process, the device records audio using its built-in microphone and captures video with its camera. It acquires audio and video data as input, encodes the audio in AAC format, and compresses the video in JPEG or H.264 format. This compressed data is then transmitted to the server via the internet.

[0405] Step 2:

[0406] Speech recognition and text conversion

[0407] The server analyzes the audio data received from the terminal. The input audio data is converted into text using the Google Cloud Speech-to-Text API. The output is the presenter's speech in text format.

[0408] Step 3:

[0409] Video analysis and reaction estimation

[0410] The server analyzes visual data using OpenCV. Using the input video data, it detects the presenter's gestures and the audience's facial expressions, and estimates their reactions. Specifically, it uses face detection and motion recognition algorithms to determine whether the presenter is pointing to a slide and how the audience is reacting. This generates a text report of the reactions as output.

[0411] Step 4:

[0412] Emotion analysis and relaxation support

[0413] The server uses an emotion analysis device to evaluate the user's emotions. It takes voice tone and facial expression data obtained from video as input, and uses an AI algorithm to estimate the level of stress and tension. As output, it generates content and advice to help the user relax. Specifically, it creates suggestions for calming music and audio guides to encourage deep breathing.

[0414] Step 5:

[0415] Question support

[0416] The server handles questions from users during presentations. It retrieves the question content and related database information as input, performs calculations, and generates the most relevant information candidates. The output presents immediately usable answer suggestions for the user, specifically including information from similar past questions and authoritative sources.

[0417] Step 6:

[0418] Evaluation report generation and feedback

[0419] The server generates an evaluation report based on data collected after the presentation. It uses aggregated data such as audio, video, user emotional states, and audience reactions recorded during the presentation as input. This data is analyzed, and the output is a detailed report including suggestions for optimizing future presentations. Specifically, this involves identifying and analyzing successes and areas for improvement.

[0420] Through this series of steps, the system helps presenters deliver more effective presentations.

[0421] (Application Example 2)

[0422] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0423] In the workplace, the quality of work can be affected by fluctuations in workers' emotions and decreased work efficiency. Therefore, there is a need for a system that monitors workers' emotions and efficiency in real time and provides appropriate advice to improve work efficiency.

[0424] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0425] In this invention, the server includes means for an information terminal to acquire audio and image data at the start of work and convert the audio data into text; means for the server device to analyze the text and image data in real time, estimate and report the worker's emotions and efficiency; and means for the information terminal to provide advice based on the report during work via a presentation device. As a result, the worker can receive feedback in real time that corresponds to their emotional state, enabling improved work efficiency.

[0426] An "information terminal" is an electronic device used to acquire and process audio and image data.

[0427] "Audio data" refers to digital data that includes sound information acquired from workers and the environment.

[0428] "Image data" refers to digital data that includes visual information acquired at a work site.

[0429] A "server device" is a computer system that analyzes acquired data in real time and performs the necessary processing.

[0430] A "presentation device" is a device used to communicate analysis results and advice to workers, and can provide visual or audio output.

[0431] "Emotion and efficiency estimation" is a process of evaluating a worker's emotional state and work efficiency based on data.

[0432] A "report" is information generated based on analyzed data and provided to the worker.

[0433] "Advice" refers to instructions that provide suggestions for improvement based on the worker's situation.

[0434] In this embodiment of the invention, a system is constructed that combines an information terminal, a server device, and a display device in order to maximize the efficiency of workers in factories and work sites.

[0435] The server device receives audio and image data transmitted from information terminals and analyzes them in real time. The audio data is converted to text using a speech recognition engine such as Google Speech-to-Text. The image data is analyzed using the OpenCV library to capture the movements and facial expressions of workers. This enables the estimation of emotions and work efficiency in the work environment.

[0436] Next, the server device generates advice for the worker based on the results of the emotion and efficiency analysis. The Hugging Face NLP library is used for emotion analysis to assess the worker's stress level, etc. Appropriate advice is generated by speech synthesis software and sent to the worker via a presentation device. This presentation device includes smart glasses and tablets.

[0437] For example, a worker may face a sudden problem and experience temporary stress. In this case, the system, after analysis, advises relaxation techniques such as deep breathing and encourages improvements to work procedures. In this process, the system uses the prompt "What advice would you give if work efficiency decreased?" as a prompt using a generative AI model, allowing the AI ​​to output the most appropriate advice.

[0438] This allows workers to receive feedback tailored to their emotional state during work, thereby improving work efficiency.

[0439] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0440] Step 1:

[0441] The information terminal acquires voice and image data of the worker at the start of work. A high-sensitivity microphone and high-resolution camera are used as input data, which is converted to digital format in real time and transmitted to the server. This allows for the recording of dynamic information about the work environment.

[0442] Step 2:

[0443] The server converts the received audio data into text data using the Google Speech-to-Text engine. Converting the audio input into text format enables subsequent natural language processing (NLP). The text output is used to document the work performed and the situation.

[0444] Step 3:

[0445] The server uses the OpenCV library to analyze image data and evaluate the worker's facial expressions and movements. It performs face recognition and gesture identification from the input image and estimates the emotional state from the results. The output provides estimated values ​​for the worker's stress level and work attitude.

[0446] Step 4:

[0447] The server uses the Hugging Face NLP library to perform sentiment analysis on text data. By analyzing the nuances of emotion, it assesses the degree of stress and confusion. The output is a quantified sentiment score.

[0448] Step 5:

[0449] Based on the analysis results, the server sends prompts to the generated AI model to produce advice for improving work efficiency. Using the prompt, "What advice would you provide if work efficiency decreases?", the AI ​​outputs the optimal advice. The output advice is used as a suggestion for the next action.

[0450] Step 6:

[0451] The information terminal uses speech synthesis software to convert advice into voice and provides it to the worker through a display device. By letting the user listen to the synthesized voice, immediate feedback is achieved. The output is the voice instructions used at the work site.

[0452] These steps allow the system to monitor workers' emotions and efficiency in real time and provide timely advice, thereby maximizing work efficiency.

[0453] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0454] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet Search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0455] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0456] [Third Embodiment]

[0457] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0458] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0459] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0460] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0461] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0462] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0463] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0464] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0465] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0466] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0467] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0468] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0469] This invention integrates an information terminal, a server device, and earphones into a single system to support presentations. At the start of a presentation, this system uses the camera and microphone equipped on the information terminal to capture the presenter's audio and video. This information is transmitted to the server device in real time.

[0470] The server device first utilizes speech recognition technology to convert received audio data into text. Simultaneously, it analyzes video data to understand the slides and the presenter's movements, and analyzes the audience's facial expressions and attitudes to estimate their reactions. This allows it to grasp the progress of the presentation and the audience's level of understanding and interest.

[0471] The server device generates appropriate advice based on the analysis results. This advice includes specific instructions for adjusting the presentation flow and information on points that require further explanation. The generated advice is converted into audio data using speech synthesis technology and transmitted to the earphones via the information terminal.

[0472] Users can receive direct feedback through their earphones and adjust their presentations in real time. For example, if it is estimated that the audience's interest is waning, the presenter can reintroduce topics or ask questions based on advice from the server device.

[0473] In the Q&A session, questions from the audience are collected via information terminals and analyzed by a server device. The server device uses a database to quickly generate optimal answer examples and provides supplementary information to the user through earphones. This allows the user to answer questions with confidence.

[0474] After the presentation ends, the server device organizes all the data and generates a review report. This report includes an overall evaluation of the presentation, audience reaction trends, Q&A performance, and specific suggestions for improvement for the next presentation. Information terminals receive this report and present it to the user, allowing them to use it to prepare for their next presentation.

[0475] This system enables users to deliver effective and smooth presentations, increasing audience engagement and satisfaction.

[0476] The following describes the processing flow.

[0477] Step 1:

[0478] The information terminal activates its camera and microphone at the start of the presentation. This initiates the acquisition of the presenter's audio and video data. The acquired data is temporarily stored in a buffer.

[0479] Step 2:

[0480] The information terminal transmits the collected audio data to the server device in real time. Video data is also transmitted to the server device simultaneously, capturing the presenter's movements and slide content.

[0481] Step 3:

[0482] The server device processes the received audio data using a speech recognition engine and converts it into text data. This creates a sequential record of the words spoken by the presenter.

[0483] Step 4:

[0484] The server device analyzes video data to detect the content of the slides and the presenter's movements. It also identifies the faces of the audience and estimates their reactions by analyzing changes in their facial expressions.

[0485] Step 5:

[0486] The server integrates the presenter's speech, slide data, and audience reaction data to evaluate the progress of the presentation. Based on this evaluation, it generates advice for the presenter.

[0487] Step 6:

[0488] The generated advice is converted into audio data using speech synthesis technology. The server device transmits this audio data to the information terminal, where it is provided to the user as real-time feedback through earphones.

[0489] Step 7:

[0490] The user receives advice through their earphones and adjusts the presentation accordingly. They encourage audience engagement by supplementing slide explanations or asking questions as needed.

[0491] Step 8:

[0492] The audio questions collected during the Q&A session are sent again from the information terminal to the server device. The server device then transcribes and analyzes the questions into text.

[0493] Step 9:

[0494] The server device matches the question against the database and searches for relevant answer information. It generates the best possible answer example, converts it into audio data, and provides it to the user.

[0495] Step 10:

[0496] After the presentation ends, the server analyzes all session data and generates a review report. This report includes a performance evaluation and improvement suggestions, and is presented to the user via an information terminal.

[0497] (Example 1)

[0498] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0499] In modern presentations, it is extremely difficult for speakers to gauge audience reactions in real time and immediately adjust the content and flow of their presentation. This makes it challenging to maintain audience engagement and effectively convey information. Furthermore, there is insufficient support for providing quick and accurate answers during Q&A sessions. These challenges hinder efficient and effective communication and contribute to a decline in presentation quality.

[0500] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0501] In this invention, the server includes means for converting audio information into text format, means for analyzing and reporting participant reactions, and means for generating advice using generative AI technology. This allows the presenter to grasp audience reactions in real time during a presentation and adjust the progress accordingly. Furthermore, by receiving immediate feedback, the presenter can maintain participant interest and achieve effective communication. In addition, by providing quick and accurate answers during Q&A sessions, it enables high-quality presentations.

[0502] An "information terminal" is a device that acquires audio and video information at the start of a presentation and transmits it to a server.

[0503] "Audio information" refers to data used to capture speech during a presentation.

[0504] "Video information" refers to data used to record and analyze video footage from a presentation.

[0505] "Text format" refers to character data obtained by converting audio information.

[0506] A "server" is a device that processes received audio and video information and performs necessary analysis and reporting.

[0507] "Generative AI technology" is a technology that uses artificial intelligence to generate useful advice and information from data.

[0508] "Advice" refers to specific instructions or suggestions for optimizing the flow and content of a presentation.

[0509] "Earphones" are devices that output advice from the server as audio and convey it directly to the speaker.

[0510] This invention is based on a system combining an information processing device, a data processing device, and an audio output device, and aims to improve the quality of presentations. Specific embodiments are shown below.

[0511] The device will use its camera and microphone to acquire audio and video information at the start of the presentation. The acquired audio information will be converted into text format using speech recognition technology. Specifically, it is expected that speech recognition software such as "Google Speech-to-Text API" will be used.

[0512] The server processes text and video information received from the terminal. During this process, it analyzes participants' reactions using a data processing library (for example, "OpenCV" is used for video processing). This allows for the estimation of participants' levels of interest and understanding. Next, it utilizes generative AI technology (such as "OpenAI GPT") to automatically generate appropriate advice based on the analysis results. This advice includes instructions and suggestions for adjusting the presentation's progress.

[0513] The generated advice is converted into speech using speech synthesis technology (such as "Amazon Polly") and output directly through the user's earphones. This allows the user to make immediate adjustments based on the feedback received during the presentation.

[0514] For example, if the server determines that the audience's interest is waning during a presentation, it can provide the user with advice via their headphones, such as "Please add more interactive segments."

[0515] An example of a prompt message could be something like, "The audience's attention is waning. Should we add some relevant interactive content?" which could then be input into a generative AI model.

[0516] This system allows users to deliver more dynamic and effective presentations, thereby increasing participant satisfaction.

[0517] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0518] Step 1:

[0519] When the device detects the start of a presentation, it activates the camera and microphone, collecting video and audio information, respectively. The input is the audio and images detected by the device, and the output is this data. The audio information captured by the microphone is converted into text in real time using speech recognition technology. This process prepares the initial data necessary for transmission to the server.

[0520] Step 2:

[0521] The terminal sends the acquired text and video information to the server. The server starts analysis based on the information obtained by converting the received audio data into text. The input is text and video information, and the output is processed data for analysis. The server analyzes the data using a generative AI model and estimates the participants' reactions. Specifically, it analyzes the participants' facial expressions and attitudes from the video and determines their level of interest from the audio content.

[0522] Step 3:

[0523] The server uses an AI model based on the analysis results to generate advice that optimizes the presentation flow. The input is the analysis results, and the output is specific advice on how to proceed. The generated advice is converted into audio data using speech synthesis technology. This process prepares the audio feedback for the user.

[0524] Step 4:

[0525] The audio data generated by the server is transmitted to the user's earphones via the terminal. The user adjusts their presentation based on this advice. In this step, the user is required to follow the instructions received in real time and find ways to maintain the participants' interest.

[0526] Step 5:

[0527] During the Q&A session of the presentation, the device collects question information from participants and sends it to the server. The server analyzes the question, consults a database, and generates the most appropriate example answer. The input is the question information, and the output is the example answer. The user receives this answer through their earphones and can answer the participants with confidence.

[0528] Step 6:

[0529] After the presentation ends, the server aggregates all the data and generates a review report. The input consists of audio, video, and analytical data collected during the presentation, and the output is the evaluation report. The information terminal displays this report to the user, who can then use it to prepare for the next presentation. This final step allows for identifying areas for improvement and enabling the user to continue delivering high-quality presentations.

[0530] (Application Example 1)

[0531] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0532] In product introductions and demonstrations at trade shows and physical stores, there is a need for a system that can instantly grasp participant reactions and allow presenters to receive real-time feedback. Traditional methods have the challenge of making it difficult to respond flexibly to participants' interests and levels of understanding, making it difficult to maximize the effectiveness of the exhibition.

[0533] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0534] In this invention, the server includes means for an information processing device to acquire audio and video information in an exhibit and convert the audio information into text, means for a data processing device to analyze the text and video information in real time, estimate and notify participants of their responses, and means for the information processing device to supply advice based on the notification during the exhibit via an audio output device. This allows presenters to immediately grasp participants' reactions and adjust their presentations on the spot.

[0535] An "information processing device" is a device that acquires audio and video information and converts it into digital signals.

[0536] An "audio output device" is a device that outputs audio data and plays a role in providing real-time audio feedback to the presenter.

[0537] A "data processing device" is a device that analyzes acquired digital signal data in real time and generates necessary reports and advice based on the analysis results.

[0538] "Exhibition" refers to activities that showcase products and information in various ways at physical stores or trade shows.

[0539] "Participants" refers to people who receive product introductions and demonstrations at exhibitions or physical stores.

[0540] "Estimating and notifying responses" means analyzing participants' reactions and behaviors and informing the presenter of the results based on that analysis.

[0541] "Providing advice" means providing the presenter with specific instructions and suggestions for improvement via audio, based on the analysis results.

[0542] The system for carrying out this invention includes an information processing device (such as smart glasses), a data processing device (server computer), and an audio output device (earphones). The role of each device and how they work together to function in this system are described below.

[0543] The information processing device is used in the form of smart glasses worn by the presenter and acquires audio and video information in real time. Using the built-in camera and microphone, it captures the participants' reactions and the presenter's movements, and collects the audio. The acquired data is immediately transmitted to the data processing device.

[0544] The data processing unit is implemented on a cloud server and uses the Google Speech-to-Text engine to convert audio information into text. It also analyzes collected video information using the OpenCV library, estimating participant reactions by analyzing facial expressions and eye movements. Based on the analyzed data, Amazon Polly is used to synthesize necessary advice, generating information to help the presenter decide on their next course of action.

[0545] The audio output device is an earphone worn by the presenter, which provides generated audio feedback in real time. This allows the presenter to instantly adjust the content and flow of the presentation in response to participant reactions.

[0546] For example, imagine a presenter explaining a new product to a customer in a cosmetics store. If the data processing device determines that the participant's expression shows no interest, it will send a prompt to the presenter via an audio output device, such as "Emphatize the product's scent" or "Demonstrate its use." This prompt might take the form of: "The customer's interest is waning, please present a new use example."

[0547] This system allows presenters to instantly incorporate participant reactions and conduct more effective demonstrations.

[0548] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0549] Step 1:

[0550] The information processing terminal uses a camera and microphone to collect audio and video information of the presentation in real time. It captures ambient sounds and visual scenes as input and transmits the data to the server in digital format. The data acquired includes the presenter's words and participants' reactions, which are then transmitted to the server using wireless communication technology.

[0551] Step 2:

[0552] The server converts the received audio information into text using the Google Speech-to-Text engine. It receives audio data as input, performs phoneme analysis and natural language processing, and generates text data as output. The specific actions performed in this step are the analysis of the audio waveform and the identification of phonemes.

[0553] Step 3:

[0554] The server analyzes the received video information using the OpenCV library. It takes a video stream as input, performs face recognition and facial expression analysis frame by frame, and generates metrics indicating the participants' level of interest as output. Specifically, it performs facial feature point detection and gaze direction estimation.

[0555] Step 4:

[0556] The server integrates the transcribed speech text with the analyzed video metrics to estimate the participant's response. The input consists of text data and video metrics, and based on this, it performs a response evaluation and generates advice, including prompts, as output. This step involves statistical data analysis and predictions using generative AI models.

[0557] Step 5:

[0558] The server converts advice generated using Amazon Polly into audio data. It receives advice text as input, synthesizes a natural-sounding waveform, and produces audio data as output. This operation involves applying a speech synthesis algorithm.

[0559] Step 6:

[0560] The generated audio data is provided to the presenter (user) through an audio output device. The user receives this real-time feedback through earphones and adjusts the presentation accordingly. The specific output is a direct audio message prompting the presenter to take action.

[0561] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0562] This invention provides a system that effectively supports presentations by combining an information terminal, a server device, and an emotion engine. The system begins at the start of a presentation when the information terminal acquires the presenter's audio and image data in real time and transmits it to the server device.

[0563] The server device converts received audio data into text using speech recognition technology. It also analyzes video data to understand the presenter's gestures and slide content, and further analyzes the audience's facial expressions and movements to estimate their reactions.

[0564] The emotion engine, a key feature of this invention, analyzes the presenter's emotional state based on data transmitted from an information terminal. This emotion analysis is estimated from the user's tone of voice, facial expressions, and body movements. If the emotion engine determines that the user is feeling stressed, it generates content to promote relaxation and provides the user with advice to regain their composure.

[0565] The server device uses the analysis results from the emotion engine to further refine the advice given to the user. This advice includes suggestions for improving the presentation flow and specific ways to capture the audience's attention. The generated advice is converted into audio data using speech synthesis technology and delivered to the user via earphones through an information terminal.

[0566] For example, if a user receives an unexpected question during a presentation, the emotion engine can sense the user's temporary confusion, and the server device can immediately provide support by offering suggested answers and relevant information, helping the user to calmly answer the question.

[0567] Furthermore, after the presentation ends, the server device analyzes all the data and generates a review report. This report includes user emotional changes, performance evaluations, and suggestions for improvement for the next presentation, and is presented to the user via an information terminal.

[0568] This system allows users to receive real-time feedback tailored to their emotional state, enabling them to deliver more effective and smoother presentations.

[0569] The following describes the processing flow.

[0570] Step 1:

[0571] The device activates its camera and microphone simultaneously with the start of the presentation, and begins acquiring audio and image data. The data collected in real time is temporarily stored in a buffer.

[0572] Step 2:

[0573] The device sends the collected audio data to the server. This data includes the presenter's voice during the presentation. Image data is also sent to the server simultaneously, including presentation slides and audience reactions.

[0574] Step 3:

[0575] The server analyzes the received audio data using a speech recognition engine and converts it into text. Furthermore, it analyzes the presenter's gestures and slide content from the image data.

[0576] Step 4:

[0577] Based on the analyzed audio and image data, the server uses an emotion engine to determine the user's emotional state. The emotion engine estimates stress and tension from the tone of voice and facial features.

[0578] Step 5:

[0579] The server uses information from the emotion engine to generate advice for the user, including suggestions on how to conduct a presentation and recommendations for relaxation based on their emotional state.

[0580] Step 6:

[0581] The generated advice is synthesized into speech, sent to the earphones via the device, and provided to the user in real time. This allows the user to receive immediate feedback as the presentation progresses.

[0582] Step 7:

[0583] The user progresses with their presentation based on advice received through the earphones, adjusting the content as needed. This increases audience engagement and improves the quality of the presentation.

[0584] Step 8:

[0585] After the presentation ends, the server analyzes all collected data and generates a review report. This report includes user emotional changes, an overall summary of the presentation performance, and suggestions for improvement for the next presentation.

[0586] Step 9:

[0587] The generated review report is provided to the user via their device, allowing them to use it to prepare for their next presentation. This process promotes continuous improvement and learning.

[0588] (Example 2)

[0589] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0590] During a presentation, presenters often struggle to manage their own emotions and respond quickly to audience reactions. They are also required to provide accurate answers to unexpected questions, and obtaining appropriate information instantly is difficult. Furthermore, insufficient post-presentation review prevents effective improvement for future presentations.

[0591] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0592] This invention includes a server that provides means for estimating and reporting audience reactions in real time, means for suggesting appropriate information candidates when questions are received during the presentation, and means for creating an evaluation report based on aggregated data after the presentation and sending suggestions for optimizing the next presentation. This allows presenters to adjust their responses according to their emotional state and conduct presentations based on audience reactions. Furthermore, it enables the quality of presentations to be improved by obtaining appropriate answers to questions immediately. In addition, it allows for efficient review and improvement for the next presentation after it has ended.

[0593] An "information processing device" is a device that has the function of acquiring audio signals and visual data and converting them into text information as needed.

[0594] A "processing unit" is a device that analyzes acquired textual and visual data in real time and processes data to estimate the participants' reactions.

[0595] An "auditory transmission device" is a device that transmits instructions provided by an information processing device to the user as audio.

[0596] A "emotion analysis device" is a system that monitors the user's emotional state and generates content that promotes relaxation as needed.

[0597] A "user" is a person who gives a presentation or operates its assistive devices, and who uses the system to solve a problem.

[0598] "Attendees" refers to the people who watch the presentation, and their reactions are the ones whose responses are perceived.

[0599] "Character information" refers to string data converted from audio signals, and it has a format that can be used as text.

[0600] "Visual data" refers to video information acquired by video cameras or similar devices, and is the data format that is subject to analysis.

[0601] "Instructions" refers to advice and guidelines provided to the user through auditory communication devices.

[0602] "Appropriate information candidates" are a collection of highly relevant and useful information presented in response to questions asked by the user.

[0603] An "evaluation report" is a report intended to review a presentation, including areas for improvement and suggestions for optimization for the next presentation.

[0604] An "optimization proposal" is a suggestion that outlines specific areas for improvement and ways to streamline the next presentation.

[0605] This system utilizes advanced information processing and computing power to support presentations. The primary devices used are smart devices (e.g., smartphones, tablets) as information processing units, and the computing power unit paired with them consists of a high-performance server computer.

[0606] Data acquisition by information processing equipment:

[0607] The device uses its built-in camera and microphone to acquire the presenter's audio signal and visual data in real time. A dedicated application on the device is used for this purpose, and the audio signal is compressed in AAC encoding format, while the visual data is compressed in JPEG or H.264 format.

[0608] Audio and video data transmission and analysis:

[0609] The device transmits the acquired data to the server via the internet. The server converts the audio into text using the Google Cloud Speech-to-Text API and analyzes the visual data using OpenCV. This allows for real-time tracking of user gestures and participant reactions.

[0610] Estimation of emotional state:

[0611] The server works in conjunction with an emotion analysis device to evaluate the user's emotional state based on their voice tone and facial movements. For example, if the user is feeling tense, the emotion analysis device generates content to promote relaxation and provides appropriate advice. This utilizes AI-powered natural language processing.

[0612] Actual advice and support:

[0613] The server provides the user with guidelines for the generated audio through an auditory transmission device (e.g., earphones) via the terminal. If the user receives an unexpected question during the presentation, the processing unit immediately presents relevant information candidates and suggested answers.

[0614] Examples of execution and prompts for the generated AI model:

[0615] For example, if a user tends to move their head a lot, the server might offer advice such as, "It would be good to tone down your gestures." Another example of a prompt for the generative AI model could be, "Think of ways to help a nervous user relax during a presentation."

[0616] This system allows presenters to respond appropriately to their emotions, enabling them to deliver more effective presentations and build upon their performance for future events.

[0617] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0618] Step 1:

[0619] Data acquisition and transmission

[0620] The device collects the presenter's audio signal and visual data in real time. During this process, the device records audio using its built-in microphone and captures video with its camera. It acquires audio and video data as input, encodes the audio in AAC format, and compresses the video in JPEG or H.264 format. This compressed data is then transmitted to the server via the internet.

[0621] Step 2:

[0622] Speech recognition and text conversion

[0623] The server analyzes the audio data received from the terminal. The input audio data is converted into text using the Google Cloud Speech-to-Text API. The output is the presenter's speech in text format.

[0624] Step 3:

[0625] Video analysis and reaction estimation

[0626] The server analyzes visual data using OpenCV. Using the input video data, it detects the presenter's gestures and the audience's facial expressions, and estimates their reactions. Specifically, it uses face detection and motion recognition algorithms to determine whether the presenter is pointing to a slide and how the audience is reacting. This generates a text report of the reactions as output.

[0627] Step 4:

[0628] Emotion analysis and relaxation support

[0629] The server uses an emotion analysis device to evaluate the user's emotions. It takes voice tone and facial expression data obtained from video as input, and uses an AI algorithm to estimate the level of stress and tension. As output, it generates content and advice to help the user relax. Specifically, it creates suggestions for calming music and audio guides to encourage deep breathing.

[0630] Step 5:

[0631] Question support

[0632] The server handles questions from users during presentations. It retrieves the question content and related database information as input, performs calculations, and generates the most relevant information candidates. The output presents immediately usable answer suggestions for the user, specifically including information from similar past questions and authoritative sources.

[0633] Step 6:

[0634] Evaluation report generation and feedback

[0635] The server generates an evaluation report based on data collected after the presentation. It uses aggregated data such as audio, video, user emotional states, and audience reactions recorded during the presentation as input. This data is analyzed, and the output is a detailed report including suggestions for optimizing future presentations. Specifically, this involves identifying and analyzing successes and areas for improvement.

[0636] Through this series of steps, the system helps presenters deliver more effective presentations.

[0637] (Application Example 2)

[0638] Next, we will explain Application Example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0639] In the workplace, the quality of work can be affected by fluctuations in workers' emotions and decreased work efficiency. Therefore, there is a need for a system that monitors workers' emotions and efficiency in real time and provides appropriate advice to improve work efficiency.

[0640] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0641] In this invention, the server includes means for an information terminal to acquire audio and image data at the start of work and convert the audio data into text; means for the server device to analyze the text and image data in real time, estimate and report the worker's emotions and efficiency; and means for the information terminal to provide advice based on the report during work via a presentation device. As a result, the worker can receive feedback in real time that corresponds to their emotional state, enabling improved work efficiency.

[0642] An "information terminal" is an electronic device used to acquire and process audio and image data.

[0643] "Audio data" refers to digital data that includes sound information acquired from workers and the environment.

[0644] "Image data" refers to digital data that includes visual information acquired at a work site.

[0645] A "server device" is a computer system that analyzes acquired data in real time and performs the necessary processing.

[0646] A "presentation device" is a device used to communicate analysis results and advice to workers, and can provide visual or audio output.

[0647] "Emotion and efficiency estimation" is a process of evaluating a worker's emotional state and work efficiency based on data.

[0648] A "report" is information generated based on analyzed data and provided to the worker.

[0649] "Advice" refers to instructions that provide suggestions for improvement based on the worker's situation.

[0650] In this embodiment of the invention, a system is constructed that combines an information terminal, a server device, and a display device in order to maximize the efficiency of workers in factories and work sites.

[0651] The server device receives audio and image data transmitted from information terminals and analyzes them in real time. The audio data is converted to text using a speech recognition engine such as Google Speech-to-Text. The image data is analyzed using the OpenCV library to capture the movements and facial expressions of workers. This enables the estimation of emotions and work efficiency in the work environment.

[0652] Next, the server device generates advice for the worker based on the results of the emotion and efficiency analysis. The Hugging Face NLP library is used for emotion analysis to assess the worker's stress level, etc. Appropriate advice is generated by speech synthesis software and sent to the worker via a presentation device. This presentation device includes smart glasses and tablets.

[0653] For example, a worker may face a sudden problem and experience temporary stress. In this case, the system, after analysis, advises relaxation techniques such as deep breathing and encourages improvements to work procedures. In this process, the system uses the prompt "What advice would you give if work efficiency decreased?" as a prompt using a generative AI model, allowing the AI ​​to output the most appropriate advice.

[0654] This allows workers to receive feedback tailored to their emotional state during work, thereby improving work efficiency.

[0655] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0656] Step 1:

[0657] The information terminal acquires voice and image data of the worker at the start of work. A high-sensitivity microphone and high-resolution camera are used as input data, which is converted to digital format in real time and transmitted to the server. This allows for the recording of dynamic information about the work environment.

[0658] Step 2:

[0659] The server converts the received audio data into text data using the Google Speech-to-Text engine. Converting the audio input into text format enables subsequent natural language processing (NLP). The text output is used to document the work performed and the situation.

[0660] Step 3:

[0661] The server uses the OpenCV library to analyze image data and evaluate the worker's facial expressions and movements. It performs face recognition and gesture identification from the input image and estimates the emotional state from the results. The output provides estimated values ​​for the worker's stress level and work attitude.

[0662] Step 4:

[0663] The server uses the Hugging Face NLP library to perform sentiment analysis on text data. By analyzing the nuances of emotion, it assesses the degree of stress and confusion. The output is a quantified sentiment score.

[0664] Step 5:

[0665] Based on the analysis results, the server sends prompts to the generated AI model to produce advice for improving work efficiency. Using the prompt, "What advice would you provide if work efficiency decreases?", the AI ​​outputs the optimal advice. The output advice is used as a suggestion for the next action.

[0666] Step 6:

[0667] The information terminal uses speech synthesis software to convert advice into voice and provides it to the worker through a display device. By letting the user listen to the synthesized voice, immediate feedback is achieved. The output is the voice instructions used at the work site.

[0668] These steps allow the system to monitor workers' emotions and efficiency in real time and provide timely advice, thereby maximizing work efficiency.

[0669] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0670] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet Search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0671] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0672] [Fourth Embodiment]

[0673] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0674] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0675] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0676] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0677] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0678] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0679] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0680] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0681] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0682] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0683] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0684] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0685] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0686] This invention integrates an information terminal, a server device, and earphones into a single system to support presentations. At the start of a presentation, this system uses the camera and microphone equipped on the information terminal to capture the presenter's audio and video. This information is transmitted to the server device in real time.

[0687] The server device first utilizes speech recognition technology to convert received audio data into text. Simultaneously, it analyzes video data to understand the slides and the presenter's movements, and analyzes the audience's facial expressions and attitudes to estimate their reactions. This allows it to grasp the progress of the presentation and the audience's level of understanding and interest.

[0688] The server device generates appropriate advice based on the analysis results. This advice includes specific instructions for adjusting the presentation flow and information on points that require further explanation. The generated advice is converted into audio data using speech synthesis technology and transmitted to the earphones via the information terminal.

[0689] Users can receive direct feedback through their earphones and adjust their presentations in real time. For example, if it is estimated that the audience's interest is waning, the presenter can reintroduce topics or ask questions based on advice from the server device.

[0690] In the Q&A session, questions from the audience are collected via information terminals and analyzed by a server device. The server device uses a database to quickly generate optimal answer examples and provides supplementary information to the user through earphones. This allows the user to answer questions with confidence.

[0691] After the presentation ends, the server device organizes all the data and generates a review report. This report includes an overall evaluation of the presentation, audience reaction trends, Q&A performance, and specific suggestions for improvement for the next presentation. Information terminals receive this report and present it to the user, allowing them to use it to prepare for their next presentation.

[0692] This system enables users to deliver effective and smooth presentations, increasing audience engagement and satisfaction.

[0693] The following describes the processing flow.

[0694] Step 1:

[0695] The information terminal activates its camera and microphone at the start of the presentation. This initiates the acquisition of the presenter's audio and video data. The acquired data is temporarily stored in a buffer.

[0696] Step 2:

[0697] The information terminal transmits the collected audio data to the server device in real time. Video data is also transmitted to the server device simultaneously, capturing the presenter's movements and slide content.

[0698] Step 3:

[0699] The server device processes the received audio data using a speech recognition engine and converts it into text data. This creates a sequential record of the words spoken by the presenter.

[0700] Step 4:

[0701] The server device analyzes video data to detect the content of the slides and the presenter's movements. It also identifies the faces of the audience and estimates their reactions by analyzing changes in their facial expressions.

[0702] Step 5:

[0703] The server integrates the presenter's speech, slide data, and audience reaction data to evaluate the progress of the presentation. Based on this evaluation, it generates advice for the presenter.

[0704] Step 6:

[0705] The generated advice is converted into audio data using speech synthesis technology. The server device transmits this audio data to the information terminal, where it is provided to the user as real-time feedback through earphones.

[0706] Step 7:

[0707] The user receives advice through their earphones and adjusts the presentation accordingly. They encourage audience engagement by supplementing slide explanations or asking questions as needed.

[0708] Step 8:

[0709] The audio questions collected during the Q&A session are sent again from the information terminal to the server device. The server device then transcribes and analyzes the questions into text.

[0710] Step 9:

[0711] The server device matches the question against the database and searches for relevant answer information. It generates the best possible answer example, converts it into audio data, and provides it to the user.

[0712] Step 10:

[0713] After the presentation ends, the server analyzes all session data and generates a review report. This report includes a performance evaluation and improvement suggestions, and is presented to the user via an information terminal.

[0714] (Example 1)

[0715] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0716] In modern presentations, it is extremely difficult for speakers to gauge audience reactions in real time and immediately adjust the content and flow of their presentation. This makes it challenging to maintain audience engagement and effectively convey information. Furthermore, there is insufficient support for providing quick and accurate answers during Q&A sessions. These challenges hinder efficient and effective communication and contribute to a decline in presentation quality.

[0717] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0718] In this invention, the server includes means for converting audio information into text format, means for analyzing and reporting participant reactions, and means for generating advice using generative AI technology. This allows the presenter to grasp audience reactions in real time during a presentation and adjust the progress accordingly. Furthermore, by receiving immediate feedback, the presenter can maintain participant interest and achieve effective communication. In addition, by providing quick and accurate answers during Q&A sessions, it enables high-quality presentations.

[0719] An "information terminal" is a device that acquires audio and video information at the start of a presentation and transmits it to a server.

[0720] "Audio information" refers to data used to capture speech during a presentation.

[0721] "Video information" refers to data used to record and analyze video footage from a presentation.

[0722] "Text format" refers to character data obtained by converting audio information.

[0723] A "server" is a device that processes received audio and video information and performs necessary analysis and reporting.

[0724] "Generative AI technology" is a technology that uses artificial intelligence to generate useful advice and information from data.

[0725] "Advice" refers to specific instructions or suggestions for optimizing the flow and content of a presentation.

[0726] "Earphones" are devices that output advice from the server as audio and convey it directly to the speaker.

[0727] This invention is based on a system combining an information processing device, a data processing device, and an audio output device, and aims to improve the quality of presentations. Specific embodiments are shown below.

[0728] The device will use its camera and microphone to acquire audio and video information at the start of the presentation. The acquired audio information will be converted into text format using speech recognition technology. Specifically, it is expected that speech recognition software such as "Google Speech-to-Text API" will be used.

[0729] The server processes text and video information received from the terminal. During this process, it analyzes participants' reactions using a data processing library (for example, "OpenCV" is used for video processing). This allows for the estimation of participants' levels of interest and understanding. Next, it utilizes generative AI technology (such as "OpenAI GPT") to automatically generate appropriate advice based on the analysis results. This advice includes instructions and suggestions for adjusting the presentation's progress.

[0730] The generated advice is converted into speech using speech synthesis technology (such as "Amazon Polly") and output directly through the user's earphones. This allows the user to make immediate adjustments based on the feedback received during the presentation.

[0731] For example, if the server determines that the audience's interest is waning during a presentation, it can provide the user with advice via their headphones, such as "Please add more interactive segments."

[0732] An example of a prompt message could be something like, "The audience's attention is waning. Should we add some relevant interactive content?" which could then be input into a generative AI model.

[0733] This system allows users to deliver more dynamic and effective presentations, thereby increasing participant satisfaction.

[0734] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0735] Step 1:

[0736] When the device detects the start of a presentation, it activates the camera and microphone, collecting video and audio information, respectively. The input is the audio and images detected by the device, and the output is this data. The audio information captured by the microphone is converted into text in real time using speech recognition technology. This process prepares the initial data necessary for transmission to the server.

[0737] Step 2:

[0738] The terminal sends the acquired text and video information to the server. The server starts analysis based on the information obtained by converting the received audio data into text. The input is text and video information, and the output is processed data for analysis. The server analyzes the data using a generative AI model and estimates the participants' reactions. Specifically, it analyzes the participants' facial expressions and attitudes from the video and determines their level of interest from the audio content.

[0739] Step 3:

[0740] The server uses an AI model based on the analysis results to generate advice that optimizes the presentation flow. The input is the analysis results, and the output is specific advice on how to proceed. The generated advice is converted into audio data using speech synthesis technology. This process prepares the audio feedback for the user.

[0741] Step 4:

[0742] The audio data generated by the server is transmitted to the user's earphones via the terminal. The user adjusts their presentation based on this advice. In this step, the user is required to follow the instructions received in real time and find ways to maintain the participants' interest.

[0743] Step 5:

[0744] During the Q&A session of the presentation, the device collects question information from participants and sends it to the server. The server analyzes the question, consults a database, and generates the most appropriate example answer. The input is the question information, and the output is the example answer. The user receives this answer through their earphones and can answer the participants with confidence.

[0745] Step 6:

[0746] After the presentation ends, the server aggregates all the data and generates a review report. The input consists of audio, video, and analytical data collected during the presentation, and the output is the evaluation report. The information terminal displays this report to the user, who can then use it to prepare for the next presentation. This final step allows for identifying areas for improvement and enabling the user to continue delivering high-quality presentations.

[0747] (Application Example 1)

[0748] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0749] In product introductions and demonstrations at trade shows and physical stores, there is a need for a system that can instantly grasp participant reactions and allow presenters to receive real-time feedback. Traditional methods have the challenge of making it difficult to respond flexibly to participants' interests and levels of understanding, making it difficult to maximize the effectiveness of the exhibition.

[0750] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0751] In this invention, the server includes means for an information processing device to acquire audio and video information in an exhibit and convert the audio information into text, means for a data processing device to analyze the text and video information in real time, estimate and notify participants of their responses, and means for the information processing device to supply advice based on the notification during the exhibit via an audio output device. This allows presenters to immediately grasp participants' reactions and adjust their presentations on the spot.

[0752] An "information processing device" is a device that acquires audio and video information and converts it into digital signals.

[0753] An "audio output device" is a device that outputs audio data and plays a role in providing real-time audio feedback to the presenter.

[0754] A "data processing device" is a device that analyzes acquired digital signal data in real time and generates necessary reports and advice based on the analysis results.

[0755] "Exhibition" refers to activities that showcase products and information in various ways at physical stores or trade shows.

[0756] "Participants" refers to people who receive product introductions and demonstrations at exhibitions or physical stores.

[0757] "Estimating and notifying responses" means analyzing participants' reactions and behaviors and informing the presenter of the results based on that analysis.

[0758] "Providing advice" means providing the presenter with specific instructions and suggestions for improvement via audio, based on the analysis results.

[0759] The system for carrying out this invention includes an information processing device (such as smart glasses), a data processing device (server computer), and an audio output device (earphones). The role of each device and how they work together to function in this system are described below.

[0760] The information processing device is used in the form of smart glasses worn by the presenter and acquires audio and video information in real time. Using the built-in camera and microphone, it captures the participants' reactions and the presenter's movements, and collects the audio. The acquired data is immediately transmitted to the data processing device.

[0761] The data processing unit is implemented on a cloud server and uses the Google Speech-to-Text engine to convert audio information into text. It also analyzes collected video information using the OpenCV library, estimating participant reactions by analyzing facial expressions and eye movements. Based on the analyzed data, Amazon Polly is used to synthesize necessary advice, generating information to help the presenter decide on their next course of action.

[0762] The audio output device is an earphone worn by the presenter, which provides generated audio feedback in real time. This allows the presenter to instantly adjust the content and flow of the presentation in response to participant reactions.

[0763] For example, imagine a presenter explaining a new product to a customer in a cosmetics store. If the data processing device determines that the participant's expression shows no interest, it will send a prompt to the presenter via an audio output device, such as "Emphatize the product's scent" or "Demonstrate its use." This prompt might take the form of: "The customer's interest is waning, please present a new use example."

[0764] This system allows presenters to instantly incorporate participant reactions and conduct more effective demonstrations.

[0765] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0766] Step 1:

[0767] The information processing terminal uses a camera and microphone to collect audio and video information of the presentation in real time. It captures ambient sounds and visual scenes as input and transmits the data to the server in digital format. The data acquired includes the presenter's words and participants' reactions, which are then transmitted to the server using wireless communication technology.

[0768] Step 2:

[0769] The server converts the received audio information into text using the Google Speech-to-Text engine. It receives audio data as input, performs phoneme analysis and natural language processing, and generates text data as output. The specific actions performed in this step are the analysis of the audio waveform and the identification of phonemes.

[0770] Step 3:

[0771] The server analyzes the received video information using the OpenCV library. It takes a video stream as input, performs face recognition and facial expression analysis frame by frame, and generates metrics indicating the participants' level of interest as output. Specifically, it performs facial feature point detection and gaze direction estimation.

[0772] Step 4:

[0773] The server integrates the transcribed speech text with the analyzed video metrics to estimate the participant's response. The input consists of text data and video metrics, and based on this, it performs a response evaluation and generates advice, including prompts, as output. This step involves statistical data analysis and predictions using generative AI models.

[0774] Step 5:

[0775] The server converts advice generated using Amazon Polly into audio data. It receives advice text as input, synthesizes a natural-sounding waveform, and produces audio data as output. This operation involves applying a speech synthesis algorithm.

[0776] Step 6:

[0777] The generated audio data is provided to the presenter (user) through an audio output device. The user receives this real-time feedback through earphones and adjusts the presentation accordingly. The specific output is a direct audio message prompting the presenter to take action.

[0778] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0779] This invention provides a system that effectively supports presentations by combining an information terminal, a server device, and an emotion engine. The system begins at the start of a presentation when the information terminal acquires the presenter's audio and image data in real time and transmits it to the server device.

[0780] The server device converts received audio data into text using speech recognition technology. It also analyzes video data to understand the presenter's gestures and slide content, and further analyzes the audience's facial expressions and movements to estimate their reactions.

[0781] The emotion engine, a key feature of this invention, analyzes the presenter's emotional state based on data transmitted from an information terminal. This emotion analysis is estimated from the user's tone of voice, facial expressions, and body movements. If the emotion engine determines that the user is feeling stressed, it generates content to promote relaxation and provides the user with advice to regain their composure.

[0782] The server device uses the analysis results from the emotion engine to further refine the advice given to the user. This advice includes suggestions for improving the presentation flow and specific ways to capture the audience's attention. The generated advice is converted into audio data using speech synthesis technology and delivered to the user via earphones through an information terminal.

[0783] For example, if a user receives an unexpected question during a presentation, the emotion engine can sense the user's temporary confusion, and the server device can immediately provide support by offering suggested answers and relevant information, helping the user to calmly answer the question.

[0784] Furthermore, after the presentation ends, the server device analyzes all the data and generates a review report. This report includes user emotional changes, performance evaluations, and suggestions for improvement for the next presentation, and is presented to the user via an information terminal.

[0785] This system allows users to receive real-time feedback tailored to their emotional state, enabling them to deliver more effective and smoother presentations.

[0786] The following describes the processing flow.

[0787] Step 1:

[0788] The device activates its camera and microphone simultaneously with the start of the presentation, and begins acquiring audio and image data. The data collected in real time is temporarily stored in a buffer.

[0789] Step 2:

[0790] The device sends the collected audio data to the server. This data includes the presenter's voice during the presentation. Image data is also sent to the server simultaneously, including presentation slides and audience reactions.

[0791] Step 3:

[0792] The server analyzes the received audio data using a speech recognition engine and converts it into text. Furthermore, it analyzes the presenter's gestures and slide content from the image data.

[0793] Step 4:

[0794] Based on the analyzed audio and image data, the server uses an emotion engine to determine the user's emotional state. The emotion engine estimates stress and tension from the tone of voice and facial features.

[0795] Step 5:

[0796] The server uses information from the emotion engine to generate advice for the user, including suggestions on how to conduct a presentation and recommendations for relaxation based on their emotional state.

[0797] Step 6:

[0798] The generated advice is synthesized into speech, sent to the earphones via the device, and provided to the user in real time. This allows the user to receive immediate feedback as the presentation progresses.

[0799] Step 7:

[0800] The user progresses with their presentation based on advice received through the earphones, adjusting the content as needed. This increases audience engagement and improves the quality of the presentation.

[0801] Step 8:

[0802] After the presentation ends, the server analyzes all collected data and generates a review report. This report includes user emotional changes, an overall summary of the presentation performance, and suggestions for improvement for the next presentation.

[0803] Step 9:

[0804] The generated review report is provided to the user via their device, allowing them to use it to prepare for their next presentation. This process promotes continuous improvement and learning.

[0805] (Example 2)

[0806] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0807] During a presentation, presenters often struggle to manage their own emotions and respond quickly to audience reactions. They are also required to provide accurate answers to unexpected questions, and obtaining appropriate information instantly is difficult. Furthermore, insufficient post-presentation review prevents effective improvement for future presentations.

[0808] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0809] This invention includes a server that provides means for estimating and reporting audience reactions in real time, means for suggesting appropriate information candidates when questions are received during the presentation, and means for creating an evaluation report based on aggregated data after the presentation and sending suggestions for optimizing the next presentation. This allows presenters to adjust their responses according to their emotional state and conduct presentations based on audience reactions. Furthermore, it enables the quality of presentations to be improved by obtaining appropriate answers to questions immediately. In addition, it allows for efficient review and improvement for the next presentation after it has ended.

[0810] An "information processing device" is a device that has the function of acquiring audio signals and visual data and converting them into text information as needed.

[0811] A "processing unit" is a device that analyzes acquired textual and visual data in real time and processes data to estimate the participants' reactions.

[0812] An "auditory transmission device" is a device that transmits instructions provided by an information processing device to the user as audio.

[0813] A "emotion analysis device" is a system that monitors the user's emotional state and generates content that promotes relaxation as needed.

[0814] A "user" is a person who gives a presentation or operates its assistive devices, and who uses the system to solve a problem.

[0815] "Attendees" refers to the people who watch the presentation, and their reactions are the ones whose responses are perceived.

[0816] "Character information" refers to string data converted from audio signals, and it has a format that can be used as text.

[0817] "Visual data" refers to video information acquired by video cameras or similar devices, and is the data format that is subject to analysis.

[0818] "Instructions" refers to advice and guidelines provided to the user through auditory communication devices.

[0819] "Appropriate information candidates" are a collection of highly relevant and useful information presented in response to questions asked by the user.

[0820] An "evaluation report" is a report intended to review a presentation, including areas for improvement and suggestions for optimization for the next presentation.

[0821] An "optimization proposal" is a suggestion that outlines specific areas for improvement and ways to streamline the next presentation.

[0822] This system utilizes advanced information processing and computing power to support presentations. The primary devices used are smart devices (e.g., smartphones, tablets) as information processing units, and the computing power unit paired with them consists of a high-performance server computer.

[0823] Data acquisition by information processing equipment:

[0824] The device uses its built-in camera and microphone to acquire the presenter's audio signal and visual data in real time. A dedicated application on the device is used for this purpose, and the audio signal is compressed in AAC encoding format, while the visual data is compressed in JPEG or H.264 format.

[0825] Audio and video data transmission and analysis:

[0826] The device transmits the acquired data to the server via the internet. The server converts the audio into text using the Google Cloud Speech-to-Text API and analyzes the visual data using OpenCV. This allows for real-time tracking of user gestures and participant reactions.

[0827] Estimation of emotional state:

[0828] The server works in conjunction with an emotion analysis device to evaluate the user's emotional state based on their voice tone and facial movements. For example, if the user is feeling tense, the emotion analysis device generates content to promote relaxation and provides appropriate advice. This utilizes AI-powered natural language processing.

[0829] Actual advice and support:

[0830] The server provides the user with guidelines for the generated audio through an auditory transmission device (e.g., earphones) via the terminal. If the user receives an unexpected question during the presentation, the processing unit immediately presents relevant information candidates and suggested answers.

[0831] Examples of execution and prompts for the generated AI model:

[0832] For example, if a user tends to move their head a lot, the server might offer advice such as, "It would be good to tone down your gestures." Another example of a prompt for the generative AI model could be, "Think of ways to help a nervous user relax during a presentation."

[0833] This system allows presenters to respond appropriately to their emotions, enabling them to deliver more effective presentations and build upon their performance for future events.

[0834] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0835] Step 1:

[0836] Data acquisition and transmission

[0837] The device collects the presenter's audio signal and visual data in real time. During this process, the device records audio using its built-in microphone and captures video with its camera. It acquires audio and video data as input, encodes the audio in AAC format, and compresses the video in JPEG or H.264 format. This compressed data is then transmitted to the server via the internet.

[0838] Step 2:

[0839] Speech recognition and text conversion

[0840] The server analyzes the audio data received from the terminal. The input audio data is converted into text using the Google Cloud Speech-to-Text API. The output is the presenter's speech in text format.

[0841] Step 3:

[0842] Video analysis and reaction estimation

[0843] The server analyzes visual data using OpenCV. Using the input video data, it detects the presenter's gestures and the audience's facial expressions, and estimates their reactions. Specifically, it uses face detection and motion recognition algorithms to determine whether the presenter is pointing to a slide and how the audience is reacting. This generates a text report of the reactions as output.

[0844] Step 4:

[0845] Emotion analysis and relaxation support

[0846] The server uses an emotion analysis device to evaluate the user's emotions. It takes voice tone and facial expression data obtained from video as input, and uses an AI algorithm to estimate the level of stress and tension. As output, it generates content and advice to help the user relax. Specifically, it creates suggestions for calming music and audio guides to encourage deep breathing.

[0847] Step 5:

[0848] Question support

[0849] The server handles questions from users during presentations. It retrieves the question content and related database information as input, performs calculations, and generates the most relevant information candidates. The output presents immediately usable answer suggestions for the user, specifically including information from similar past questions and authoritative sources.

[0850] Step 6:

[0851] Evaluation report generation and feedback

[0852] The server generates an evaluation report based on data collected after the presentation. It uses aggregated data such as audio, video, user emotional states, and audience reactions recorded during the presentation as input. This data is analyzed, and the output is a detailed report including suggestions for optimizing future presentations. Specifically, this involves identifying and analyzing successes and areas for improvement.

[0853] Through this series of steps, the system helps presenters deliver more effective presentations.

[0854] (Application Example 2)

[0855] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0856] In the workplace, the quality of work can be affected by fluctuations in workers' emotions and decreased work efficiency. Therefore, there is a need for a system that monitors workers' emotions and efficiency in real time and provides appropriate advice to improve work efficiency.

[0857] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0858] In this invention, the server includes means for an information terminal to acquire audio and image data at the start of work and convert the audio data into text; means for the server device to analyze the text and image data in real time, estimate and report the worker's emotions and efficiency; and means for the information terminal to provide advice based on the report during work via a presentation device. As a result, the worker can receive feedback in real time that corresponds to their emotional state, enabling improved work efficiency.

[0859] An "information terminal" is an electronic device used to acquire and process audio and image data.

[0860] "Audio data" refers to digital data that includes sound information acquired from workers and the environment.

[0861] "Image data" refers to digital data that includes visual information acquired at a work site.

[0862] A "server device" is a computer system that analyzes acquired data in real time and performs the necessary processing.

[0863] A "presentation device" is a device used to communicate analysis results and advice to workers, and can provide visual or audio output.

[0864] "Emotion and efficiency estimation" is a process of evaluating a worker's emotional state and work efficiency based on data.

[0865] A "report" is information generated based on analyzed data and provided to the worker.

[0866] "Advice" refers to instructions that provide suggestions for improvement based on the worker's situation.

[0867] In this embodiment of the invention, a system is constructed that combines an information terminal, a server device, and a display device in order to maximize the efficiency of workers in factories and work sites.

[0868] The server device receives audio and image data transmitted from information terminals and analyzes them in real time. The audio data is converted to text using a speech recognition engine such as Google Speech-to-Text. The image data is analyzed using the OpenCV library to capture the movements and facial expressions of workers. This enables the estimation of emotions and work efficiency in the work environment.

[0869] Next, the server device generates advice for the worker based on the results of the emotion and efficiency analysis. The Hugging Face NLP library is used for emotion analysis to assess the worker's stress level, etc. Appropriate advice is generated by speech synthesis software and sent to the worker via a presentation device. This presentation device includes smart glasses and tablets.

[0870] For example, a worker may face a sudden problem and experience temporary stress. In this case, the system, after analysis, advises relaxation techniques such as deep breathing and encourages improvements to work procedures. In this process, the system uses the prompt "What advice would you give if work efficiency decreased?" as a prompt using a generative AI model, allowing the AI ​​to output the most appropriate advice.

[0871] This allows workers to receive feedback tailored to their emotional state during work, thereby improving work efficiency.

[0872] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0873] Step 1:

[0874] The information terminal acquires voice and image data of the worker at the start of work. A high-sensitivity microphone and high-resolution camera are used as input data, which is converted to digital format in real time and transmitted to the server. This allows for the recording of dynamic information about the work environment.

[0875] Step 2:

[0876] The server converts the received audio data into text data using the Google Speech-to-Text engine. Converting the audio input into text format enables subsequent natural language processing (NLP). The text output is used to document the work performed and the situation.

[0877] Step 3:

[0878] The server uses the OpenCV library to analyze image data and evaluate the worker's facial expressions and movements. It performs face recognition and gesture identification from the input image and estimates the emotional state from the results. The output provides estimated values ​​for the worker's stress level and work attitude.

[0879] Step 4:

[0880] The server uses the Hugging Face NLP library to perform sentiment analysis on text data. By analyzing the nuances of emotion, it assesses the degree of stress and confusion. The output is a quantified sentiment score.

[0881] Step 5:

[0882] Based on the analysis results, the server sends prompts to the generated AI model to produce advice for improving work efficiency. Using the prompt, "What advice would you provide if work efficiency decreases?", the AI ​​outputs the optimal advice. The output advice is used as a suggestion for the next action.

[0883] Step 6:

[0884] The information terminal uses speech synthesis software to convert advice into voice and provides it to the worker through a display device. By letting the user listen to the synthesized voice, immediate feedback is achieved. The output is the voice instructions used at the work site.

[0885] These steps allow the system to monitor workers' emotions and efficiency in real time and provide timely advice, thereby maximizing work efficiency.

[0886] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0887] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet Search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0888] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0889] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0890] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0891] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0892] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0893] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0894] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0895] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0896] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0897] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0898] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0899] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0900] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0901] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0902] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0903] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0904] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0905] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0906] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted as being incorporated by reference.

[0907] The following is further disclosed regarding the embodiments described above.

[0908] (Claim 1)

[0909] The information terminal acquires audio data and image data at the start of the presentation, and means for converting the audio data into text,

[0910] The server device analyzes the text data and image data in real time and provides means for estimating and reporting the audience's reactions.

[0911] A means by which an information terminal provides advice based on the report during a presentation via earphones,

[0912] ...

[0913] A system that includes this.

[0914] (Claim 2)

[0915] The system according to claim 1, wherein when a server device receives a question, it refers to a database corresponding to that question and generates an optimal example answer.

[0916] (Claim 3)

[0917] The system according to claim 1, wherein the server device creates a review report based on the collected data after the presentation has finished and sends suggestions for improvement to the next presentation to an information terminal.

[0918] "Example 1"

[0919] (Claim 1)

[0920] The information terminal acquires audio and video information at the start of a presentation, and means for converting the audio information into text format,

[0921] A server processes the aforementioned text and video information in real time, analyzes the participants' reactions, and provides a means for reporting them.

[0922] A means of generating advice to adjust the progress of a presentation based on participant reactions using generative AI technology,

[0923] A means of outputting generated advice to earphones via an information terminal,

[0924] ...

[0925] A system that includes this.

[0926] (Claim 2)

[0927] The system according to claim 1, wherein when the server receives a question, it refers to information resources corresponding to the question and generates an optimal example answer.

[0928] (Claim 3)

[0929] The system according to claim 1, wherein the server creates an evaluation report based on the information collected after the presentation has ended and sends suggestions for improving future presentations to an information terminal.

[0930] "Application Example 1"

[0931] (Claim 1)

[0932] The information processing device acquires audio and video information in the exhibit, and means for converting the audio information into text,

[0933] A data processing device analyzes the text information and video information in real time, and provides means for estimating and notifying participants of their responses.

[0934] Means by which an information processing device provides advice based on the said notice during the exhibition via an audio output device,

[0935] ...

[0936] A system that includes this.

[0937] (Claim 2)

[0938] The system according to claim 1, wherein when a data processing device receives a question, it refers to a knowledge base corresponding to the question and generates an optimal response example.

[0939] (Claim 3)

[0940] The system according to claim 1, wherein the data processing device creates an evaluation report based on the accumulated information after the exhibition ends and transmits suggestions for improvements to the next exhibition to the information processing device.

[0941] "Example 2 of combining an emotion engine"

[0942] (Claim 1)

[0943] The information processing device acquires audio signals and visual data at the start of a presentation, and includes means for converting the audio signals into text information.

[0944] A means for a computing device to analyze the textual information and visual data in real time, estimate the participants' reactions,

[0945] A means by which an information processing device provides instructions based on the report during a presentation via an auditory transmission device,

[0946] A means by which an emotion analysis device monitors the user's emotional state and generates content that promotes relaxation,

[0947] A means for suggesting appropriate information candidates when the computing unit receives a question during the presentation time,

[0948] A system that includes this.

[0949] (Claim 2)

[0950] The system according to claim 1, wherein the computing device creates an evaluation report based on aggregated data after the presentation is completed and transmits suggestions for optimizing the next presentation to the information processing device.

[0951] (Claim 3)

[0952] The system according to claim 1, wherein if the emotion analysis device detects the user's tension, it provides the user with advice on how to regain composure.

[0953] "Application example 2 when combining with an emotional engine"

[0954] (Claim 1)

[0955] The information terminal acquires audio data and image data at the start of operation, and means for converting the audio data into text,

[0956] A server device analyzes the text data and image data in real time and provides means for estimating and reporting the worker's emotions and efficiency.

[0957] A means by which an information terminal provides advice based on the report during work via a display device,

[0958] ...

[0959] A system that includes this.

[0960] (Claim 2)

[0961] The system according to claim 1, which, when a server device detects an anomaly, refers to a database related to the anomaly and generates an optimal improvement example.

[0962] (Claim 3)

[0963] The system according to claim 1, wherein the server device creates a review report based on the collected data after the completion of work and sends improvement suggestions for the next work to an information terminal. [Explanation of symbols]

[0964] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. The information terminal acquires audio data and image data at the start of the presentation, and means for converting the audio data into text, The server device analyzes the text data and image data in real time and provides means for estimating and reporting the audience's reactions. A means by which an information terminal provides advice based on the report during a presentation via earphones, A system that includes this.

2. The system according to claim 1, wherein when a server device receives a question, it refers to a database corresponding to that question and generates an optimal example answer.

3. The system according to claim 1, wherein the server device creates a review report based on the collected data after the presentation has finished and sends suggestions for improvement to the next presentation to an information terminal.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A