System
The system provides real-time feedback and virtual audience simulation to enhance presentation skills by allowing users to practice with smart glasses, addressing the limitations of current methods.
Patent Information
- Application Number
- JP2024118090
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-23
- Publication Date
- 2026-02-04
AI Technical Summary
Current presentation practice methods lack real-time feedback and the ability to simulate interactions with a virtual audience, making it difficult to improve presentation skills effectively.
A system that allows users to input profile and audience information, upload presentation materials, and utilize smart glasses to receive real-time reactions and questions from a virtual audience, enabling simulation and feedback.
Enhances presentation quality by allowing users to practice with realistic feedback, improving their confidence and performance.
Smart Images

Figure 2026017308000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] When speaking at an event or conference for the first time, or when presenting in front of people with different demographics, many people feel anxious about whether their presentation will be conveyed to the audience and what reactions and comments they will receive. Because such anxiety can reduce the quality of a presentation, it is important to practice in advance to improve the accuracy and persuasiveness of the presentation. However, current practice methods do not allow for real-time feedback or questions from a virtual audience, and it is difficult to evaluate the results of practice. [Means for solving the problem]
[0005] To solve the above problems, the present invention provides a system including: a means for a user to input profile information and audience information; a means for uploading presentation materials; a means for a server to analyze the uploaded materials and generate reactions; a means for a terminal to transmit and receive data between the user and the server and operate smart glasses; and a means for the smart glasses to display the reactions and perform a presentation simulation. This system allows users to practice in advance while receiving reactions and questions from a virtual audience, evaluate the results, and receive feedback, thereby improving the quality of their presentations and enabling them to take the stage with confidence.
[0006] "User" refers to an individual or group who uses the system to practice a presentation.
[0007] "Profile information" refers to data that indicates personal characteristics such as a user's name, job title, experience, etc.
[0008] "Audience information" refers to data that indicates characteristics such as occupation, age, and areas of interest of the audience to whom the presentation is directed.
[0009] "Presentation materials" refers to document files such as slides and text that a user uses in a presentation.
[0010] "Upload" refers to the act of a user sending data to a server via a terminal.
[0011] A "server" is a computing device for analyzing presentation materials, generating reactions, and processing user data.
[0012] "Analysis" refers to the process by which the server breaks down the content of the uploaded presentation materials, extracts the information, and analyzes it.
[0013] "Reactions" refers to the interactive feedback, such as reactions, comments, and questions, that your virtual audience gives to your presentation.
[0014] A "terminal" is an electronic device such as a smartphone or PC that is operated by a user and exchanges data with a server.
[0015] "Smart glasses" are a type of head-mounted display worn by users to display a virtual audience and reactions.
[0016] "Presentation simulation" is the process of simulating and practicing a real presentation through a virtual audience. [Brief explanation of the drawings]
[0017] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9]1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0018] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0019] First, the terms used in the following description will be explained.
[0020] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0021] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0022] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0023] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0024] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0025] [First embodiment]
[0026] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0027] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0028] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0029] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0030] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0031] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0032] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0033] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0034] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0035] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0036] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0037] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0038] The present invention relates to a system for allowing a user to practice a presentation. Specific embodiments of the present invention will be described below.
[0039] The system begins with the user entering their profile information and audience information and uploading presentation materials. Next, the server analyzes the materials and generates reactions. The device sends and receives data between the user and the server and operates the smart glasses. Finally, the smart glasses display the reactions and simulate the presentation.
[0040] Program processing
[0041] 1. Enter your user profile and audience information
[0042] A user launches the application and enters their profile information and the audience information for the presentation, allowing the system to build an appropriate simulation scenario based on the characteristics of the user and audience.
[0043] 2. Upload your presentation materials
[0044] Users upload presentation materials (e.g., PowerPoint or PDF files) through the application. These materials are sent to the server and analyzed there. Text and image information for each slide is extracted from the analyzed materials.
[0045] 3. Generating Reactions
[0046] The server uses an AI model to generate reactions based on the uploaded presentation materials, user profile information, and audience information, including comments, reactions, and questions from the virtual audience.
[0047] 4. Connecting to smart glasses
[0048] The user wears the smart glasses, and the application establishes a connection with the smart glasses, which are devices that receive data from the server and display reactions.
[0049] 5. Start the presentation simulation
[0050] The user operates the application to enter a presentation simulation mode, and the reactions of the virtual audience are displayed through the smart glasses. The user advances through the slides, presents, and receives feedback from the virtual audience.
[0051] 6. Practice Q&A
[0052] As part of the simulation, users respond to questions posed by a virtual audience, and their responses are sent via smart glasses to a server that analyzes their responses and generates subsequent reactions and feedback.
[0053] Specific examples
[0054] For example, if a user wants to give a presentation on the topic of "agile development," they first enter their profile information, such as their job title and experience, into the application, and set the audience attributes as "engineers, aged 30-40, interested in agile development." They then upload their PowerPoint presentation materials through the application, and the server analyzes them.
[0055] Once a user puts on the smart glasses and enters the simulation mode, a question is displayed from the virtual audience: "What are the main challenges in adopting agile development?" The user responds to this question, and the response is sent to the server for analysis. The server then generates appropriate feedback and additional reactions, which are then displayed to the user.
[0056] In this way, users can repeatedly simulate their presentation with a virtual audience in advance, improving the quality of their presentation and giving them confidence when they actually take the stage.
[0057] The processing flow will be explained below.
[0058] Step 1:
[0059] A user launches the application and logs in. The user enters their profile information and sets up audience information, including name, job title, experience, audience age, occupation, and areas of interest.
[0060] Step 2:
[0061] Users select presentation materials and upload them from their PC or smartphone through the application, which can be in PowerPoint or PDF format.
[0062] Step 3:
[0063] The server receives and stores the uploaded presentation materials. The server analyzes the materials and extracts text and image information for each slide. The analysis uses natural language processing and image recognition technology.
[0064] Step 4:
[0065] The server uses an AI model to generate reactions based on the user's profile information and audience information, specifically determining what comments and questions the virtual audience will make and when they will react.
[0066] Step 5:
[0067] The user wears the smart glasses and starts the simulation mode in the application. The device connects to the smart glasses and establishes data communication with the server.
[0068] Step 6:
[0069] The server confirms the start of the simulation mode and sends the analysis results of the presentation materials and reaction data to the smart glasses. As the user advances through the slides, the smart glasses display the appropriate reactions.
[0070] Step 7:
[0071] The device sends user operation information to the server, including the user's actions as they advance through slides and the progress of the presentation content via voice recognition. The server then uses this information to generate the next reaction or question.
[0072] Step 8:
[0073] The user can see the reactions of the virtual audience (e.g., "Interesting!", questions like, "What are some applications of this technology?") through the smart glasses and respond appropriately. The user's responses are sent to the server via the device.
[0074] Step 9:
[0075] The server analyzes the user's answers and generates follow-up questions and feedback, which are then sent back to the smart glasses and displayed to the user.
[0076] Step 10:
[0077] Once the user has finished the simulation, the server analyzes their overall performance and generates results, including a grade for each slide, the quality and timing of the user's responses, and their overall progress.
[0078] Step 11:
[0079] The user checks the practice results through the application and refers to the feedback provided by the server to improve the presentation materials and delivery method.
[0080] Step 12:
[0081] Users can re-run the simulation and practice as needed, improving the quality of their presentation and building confidence in speaking.
[0082] Example 1
[0083] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0084] Traditional presentation practice methods often require a live audience or offer limited simulations, meaning users miss out on opportunities to effectively improve their presentation skills. Additionally, it's difficult to immediately identify areas for improvement through real-time feedback and Q&A sessions.
[0085] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0086] In this invention, the server includes a means for a user to input profile information and student information, a means for uploading presentation materials, and a means for an information processing device to analyze the uploaded materials and generate responses, thereby enabling the user to receive real-time feedback and questions from a virtual audience and to practice presentations effectively.
[0087] "User" refers to an individual or group who uses the system to practice a presentation.
[0088] "Profile information" refers to detailed information about a user, such as the user's job title, years of experience, and area of expertise.
[0089] "Audience information" refers to information such as the attributes and areas of interest of the audience to whom the presentation is directed.
[0090] "Presentation materials" refers to data such as slides and documents that a user uses to give a presentation.
[0091] An "information processing device" refers to a server or computer used to analyze presentation materials and generate responses.
[0092] "Responses" refers to interactive feedback such as comments, reactions, and questions from the virtual audience.
[0093] A "terminal" is a device (e.g., a smartphone, tablet, or PC) that transmits and receives data between a user and an information processing device and operates a video display device.
[0094] A "video display device" is a display device (e.g., smart glasses, VR headset, etc.) that allows users to visually check reactions.
[0095] "Practice presentation" refers to an activity in which a user prepares for an actual presentation by simulating a presentation in front of a virtual audience.
[0096] The present invention relates to a system for users to practice presentations. The system begins with the user entering profile information and participant information and uploading presentation materials. Next, an information processing device analyzes the materials and generates responses. A terminal transmits and receives data between the user and the information processing device and operates a video display device. Finally, the video display device displays the responses, and the presentation is practiced.
[0097] In a specific embodiment of the system, a program is generated and processing is performed in the following procedure.
[0098] First, the user launches the application, logs in, and enters their profile information (e.g., name, job title, years of experience, etc.). Then, they enter information about the target audience (e.g., industry, age group, areas of interest, etc.). The system then creates an appropriate simulation scenario based on the characteristics of the user and the audience.
[0099] Next, the user uploads presentation materials (e.g., PowerPoint or PDF files) to the application. These materials are sent to an information processing device via the terminal. The information processing device analyzes the uploaded materials and extracts text and image information for each slide. This analysis is performed using text analysis software and image recognition software.
[0100] The information processing device uses a generative AI model to generate responses based on the extracted material content, user profile information, and participant information. The generated responses include comments, reactions, and questions from the virtual audience. For example, a question might be generated: "What are the main challenges in introducing agile development?"
[0101] The user wears a video display device (e.g., smart glasses) and selects the "Video Display Device Connection" option within the application. The terminal establishes a connection with the video display device using a connection method such as Bluetooth or Wi-Fi. The video display device receives the response data from the information processing device and prepares to display it.
[0102] When the user selects "Start Simulation," the presentation begins, and the virtual audience's reactions and questions are displayed on the video display device. The user then verbally responds to the questions, which are then transmitted to the information processing device via the video display device. The information processing device then uses voice analysis software to analyze the user's responses and generate appropriate feedback or additional responses.
[0103] For example, if a user is giving a presentation on the topic of "Agile Development," the steps would be:
[0104] 1. The user enters profile information such as their job title and experience, and sets the audience attributes as "Engineer, 30-40 years old, interested in agile development."
[0105] 2. Upload PowerPoint presentation materials through the application, and the information processing device analyzes the materials.
[0106] 3. Once the user puts on the video display device and starts the simulation mode, a question is displayed from the virtual audience: "What are the main challenges in adopting agile development?"
[0107] 4. The user answers the question, and the answer is sent to the information processing device and analyzed.
[0108] 5. The information processing device generates further appropriate feedback and additional reactions and displays them to the user.
[0109] Examples of prompts include: "The user's profile information is an engineer in his 30s with five years of experience in agile development," "The audience is engineers aged 30-40 who are interested in agile development," and "The theme of the presentation is 'Introduction to agile development and its challenges.' Please generate questions and feedback from the virtual audience."
[0110] In this way, users can repeatedly simulate their presentation with a virtual audience in advance, improving the quality of their presentation and giving them confidence when they actually take the stage.
[0111] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0112] Step 1:
[0113] Entering user profile and audience information
[0114] Input: User profile information (e.g., name, job title, years of experience) and audience information (e.g., industry, age group, interests)
[0115] Processing: The user starts the application and authenticates by entering their user ID and password on the login screen. After logging in, they enter their profile information and audience information in the form that appears. The server receives this information and stores it in a database. It also designs an appropriate simulation scenario based on the user and audience information.
[0116] Output: User and audience information stored in a database, and simulation scenario designs
[0117] Step 2:
[0118] Upload presentation materials
[0119] Input: Presentation materials (e.g. PowerPoint files, PDF files)
[0120] Processing: The user clicks the "File Upload" button in the application and selects presentation materials from a selection screen. The device reads the selected file and sends it to the server. The server analyzes the received file and extracts text and image data for each slide. This analysis is performed using text analysis software and image recognition software.
[0121] Output: Parsed text and image data for each slide
[0122] Step 3:
[0123] Creating a reaction
[0124] Input: Parsed presentation materials, user profile information, audience information
[0125] Processing: The server uses the generative AI model to generate comments, reactions, and questions from the virtual audience based on the analyzed content and the input user profile and audience information. For example, it dynamically generates questions and reactions related to the content of the slides.
[0126] Output: Generated virtual audience reaction data (comments, questions, responses, etc.)
[0127] Step 4:
[0128] Connecting to smart glasses
[0129] Input: Wi-Fi and Bluetooth information for connection
[0130] Processing: The user wears a video display device (e.g., smart glasses) and selects the "Connect to video display device" option within the application. The device establishes a connection with the video display device using the selected connection method. The device confirms that a connection with the video display device has been established.
[0131] Output: Established connection status with video display device
[0132] Step 5:
[0133] Start of presentation simulation
[0134] Input: Command to start the simulation
[0135] Processing: The user clicks the "Start Simulation" button in the application. The terminal sends the presentation materials and the generated reaction data to the video display device and prepares to start the simulation.
[0136] Output: Presentation materials and reaction data sent to a video display device
[0137] Step 6:
[0138] Q&A practice
[0139] Input: Virtual audience question displayed on a video display device
[0140] Processing: The video display device displays the question from the virtual audience. The user responds verbally to the question, and the response is captured using the video display device's audio input function and sent to the server. The server analyzes the audio data and generates feedback and additional reactions based on the user's response, which are then sent back to the video display device.
[0141] Output: Parsed user response data, generated new feedback and reactions
[0142] (Application example 1)
[0143] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0144] Conventional presentation practice systems primarily require users to practice on their own, lacking the ability to provide real-time feedback or reflect the reactions of a virtual audience. Furthermore, systems for effectively training salespeople in brick-and-mortar stores on product explanations are also inadequate, and there is a need for a system that can respond immediately to customer questions and reactions. The present invention aims to solve these problems and improve the quality of presentation practice and sales training.
[0145] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0146] In this invention, the server includes: a means for a user to input profile information and audience information; a means for uploading presentation materials; a means for the server to analyze the uploaded materials and generate reactions; a means for a terminal to transmit and receive data between the user and the server and operate the smart glasses; a means for the smart glasses to display reactions and perform a presentation simulation; a means for a salesperson to input profile information and customer information and upload product explanation materials; a means for the server to analyze the product explanation materials and generate questions and reactions of virtual customers; and a means for the smart glasses to display the responses and questions of the virtual customers in real time and perform a sales training simulation. This allows the user to make a presentation while receiving reactions from the virtual audience, and enables the salesperson to receive practical training through interactions with virtual customers.
[0147] "User profile information" is information that includes personal information, job information, experience, and the like about the user.
[0148] "Audience information" is attribute information about the audience to whom the presentation is given.
[0149] "Presentation materials" are presentation materials provided in the form of slides or documents.
[0150] A "server" is a computer system that receives data from users, analyzes it, and generates appropriate reactions.
[0151] "Means for generating reactions" is a function that uses an AI model to generate virtual comments and questions based on presentation materials and user information.
[0152] A "terminal" is a device that transmits and receives data between a user and a server.
[0153] "Smart glasses" are eyeglass-type devices that utilize augmented reality technology to display information to users in real time.
[0154] The "means for simulating a presentation" is a function that allows a user to use smart glasses to give a presentation while interacting with a virtual audience.
[0155] "Salesperson profile information" is information including job information, products in charge, experience, etc., related to the salesperson.
[0156] "Customer information" is attribute information about customers who are the target of sales.
[0157] "Product description materials" are materials used to explain the characteristics and functions of a product.
[0158] "Virtual Customer" means a fictitious customer generated using AI technology.
[0159] "Means for generating questions and reactions" refers to a function that uses an AI model to generate virtual customer questions and reactions based on salesperson profile information and product explanation materials.
[0160] The "means for conducting sales training simulations" is a function that allows salespeople to use smart glasses to conduct practical training through interactions with virtual customers.
[0161] The present invention relates to a comprehensive system for users to practice presentations and sales training. The system begins when the user inputs profile information and audience information and uploads presentation or product explanation materials. The server analyzes the uploaded materials and generates reactions. The terminal transmits and receives data between the user and the server and operates the smart glasses. The smart glasses display reactions and perform a presentation simulation or sales training simulation.
[0162] Specific steps
[0163] 1. Enter your profile and audience information
[0164] A user launches the application and enters their profile information (e.g., job title, experience, products) and audience or customer information for the presentation or sale (e.g., audience age range, interests, skill level).
[0165] 2. Upload presentation or product description materials
[0166] Users upload presentation materials (e.g., PowerPoint or PDF files) or product descriptions through the application. These materials are then sent to the server, where they are analyzed. During the analysis process, important text and image information is extracted from the materials.
[0167] 3. Generating Reactions
[0168] The server uses a generative AI model to generate reactions based on the uploaded materials and user and audience information, including comments, reactions, and questions from the virtual audience or virtual customers.
[0169] 4. Connecting to smart glasses
[0170] The user wears the smart glasses, and the application establishes a connection with the smart glasses, which receives data from the server and displays reactions.
[0171] 5. Start a presentation simulation or sales training
[0172] The user interacts with the application and enters the simulation mode, where the smart glasses display the reactions of a virtual audience and virtual customers, providing virtual feedback as the user progresses through slides and product presentations.
[0173] 6. Practice Q&A
[0174] As part of the simulation, users respond to questions posed by a virtual audience or customer, and their answers are sent through the smart glasses to a server that analyzes their responses and generates subsequent reactions and feedback.
[0175] Hardware and software used
[0176] Hardware used: Smart glasses (e.g., Google Glass, Vuzix Blade)
[0177] Software used: Flask (server-side framework), generative AI model (reaction generation), data analysis system
[0178] Specific examples
[0179] For example, when training a salesperson to explain a new smartphone, the following prompts are fed into the generative AI model:
[0180] "User: Salesperson, Experience: 3 years, Product: New smartphone, Target Audience: 20-30 years old, Tech-savvy"
[0181] Feedback from the virtual customer is displayed on the smart glasses with questions such as:
[0182] "How does this smartphone's camera performance compare to other products?"
[0183] In this way, users can improve the quality of their presentations and sales training by repeating virtual simulations.
[0184] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0185] Step 1:
[0186] A user launches the application and enters profile and audience information.
[0187] This information includes the user's job title, experience, products they work on, and target audience demographics (age range, interests, skill level, etc.).
[0188] The server receives these input data and stores them in a database.
[0189] Step 2:
[0190] A user uploads a presentation or product description.
[0191] These materials are provided in formats such as PDF and PowerPoint files.
[0192] The server receives the material and uses text analysis software (e.g., Apache Tika) to extract text and image information from the material.
[0193] The extracted information is stored in a database on the server.
[0194] Step 3:
[0195] The server generates reactions using a generative AI model based on uploaded materials, user information, and audience information.
[0196] This process generates questions and comments from the virtual audience or customer based on the content of the material.
[0197] The generated reactions are temporarily stored on the server.
[0198] Step 4:
[0199] The user puts on smart glasses (e.g., Google Glass, Vuzix Blade), and the application establishes a connection with the smart glasses.
[0200] The server transmits the generated reaction data to the smart glasses.
[0201] The smart glasses display the received data in real time in the user's field of vision.
[0202] Step 5:
[0203] The user operates the application and starts a presentation simulation or a sales training simulation.
[0204] Smart glasses display the reactions of virtual audiences and virtual customers.
[0205] Users receive real-time feedback as they progress through their presentations and product explanations.
[0206] Step 6:
[0207] As part of a presentation simulation or sales training, a user answers questions posed by a virtual audience or virtual customers.
[0208] The user's answers are sent to the server via the smart glasses.
[0209] The server analyzes the user's answers and generates the next reaction or feedback.
[0210] The generated feedback is then sent back to the smart glasses and displayed to the user.
[0211] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0212] The present invention relates to a system for providing more advanced feedback by combining a system for users to practice presentations with an emotion engine that recognizes the emotions of the users. Specific embodiments of the present invention will be described below.
[0213] System configuration
[0214] Entering user profile and audience information
[0215] A user launches the application and logs in. The user enters their profile information (e.g., name, job title, experience) and also sets audience information (e.g., age, occupation, areas of interest) for the presentation.
[0216] Upload presentation materials
[0217] Users upload presentation materials (e.g., PowerPoint or PDF files) from their PC or smartphone through the application. These materials are then sent to the server, where they are analyzed. The analysis includes extracting text and image information for each slide.
[0218] Creating a reaction
[0219] The server uses an AI model to generate reactions based on the uploaded presentation materials, user profile information, and audience information. Specifically, it determines what comments and questions the virtual audience will make, as well as the timing of their reactions.
[0220] Linking the Emotion Engine
[0221] When a user puts on the smart glasses and starts the simulation mode in the application, the emotion engine analyzes the user's facial expressions and tone of voice to recognize the user's emotional state in real time. For example, if the user is nervous, it will provide advice on how to relax.
[0222] Start of presentation simulation
[0223] The user interacts with the application to enter a presentation simulation mode, where the smart glasses display the reactions of the virtual audience. The user advances through the slides, presenting, and receives feedback from the virtual audience.
[0224] Q&A practice
[0225] As part of the simulation, users respond to questions posed by a virtual audience, and their responses are sent via smart glasses to a server that analyzes their responses and generates subsequent reactions and feedback.
[0226] Emotional data feedback and adjustment
[0227] The emotion engine analyzes the user's emotional data as they make their presentation and generates feedback based on their emotional state. For example, if the user feels anxious, the engine will advise them to explain the key points more simply.
[0228] Specific examples
[0229] For example, if a user wants to give a presentation on the theme of "innovative technology," the user first enters their profile information (e.g., experience as an engineer, field of expertise) into the application, and sets the audience attributes as "corporate technical personnel, aged 30-40, interested in technological innovation." Then, the user uploads presentation materials (PowerPoint files).
[0230] When a user puts on the smart glasses and starts the simulation mode, the emotion engine analyzes the user's facial expressions and tone of voice, and displays feedback such as, "You're feeling a little nervous on this slide," while the smart glasses also display advice such as, "Take a deep breath to relax."
[0231] In the presentation simulation, a virtual audience asks the user, "Please tell us some specific examples of how your technology can be applied," and the user responds. The response is sent to the server, where it is analyzed and the results are provided as feedback. For example, the user might receive feedback such as, "The explanation of the application example was easy to understand. However, I would like more specific examples."
[0232] In this way, by linking with the emotion engine, users can receive more detailed and useful feedback, significantly improving the quality of their presentations and giving them more confidence when they take the stage.
[0233] The processing flow will be explained below.
[0234] Step 1:
[0235] A user launches the application and logs in. The user enters their profile information (name, job title, experience, etc.) and also sets audience information (age, occupation, areas of interest, etc.) for the presentation.
[0236] Step 2:
[0237] Users select and upload presentation materials (PowerPoint or PDF files) from their PC or smartphone through the application, which are then sent to the server.
[0238] Step 3:
[0239] The server receives and analyzes the uploaded presentation materials, extracting text and image information from each slide, and using natural language processing (NLP) and image recognition technology to identify key points and locations within the materials.
[0240] Step 4:
[0241] The server uses an AI model to generate reactions based on the user's profile information and audience information. Specifically, it determines what comments and questions the virtual audience will make, as well as the timing of their reactions. The emotion engine also prepares to acquire the user's emotional data.
[0242] Step 5:
[0243] The user puts on the smart glasses and starts the simulation mode in the application. The device connects to the smart glasses and performs data communication with the server. The smart glasses receive reaction data from the server and prepare to display it.
[0244] Step 6:
[0245] The server confirms the start of the simulation mode and sends the analysis results of the presentation materials and reaction data to the smart glasses. As the user advances through the slides, the smart glasses display appropriate reactions. The emotion engine also analyzes the user's facial expressions and tone of voice in real time.
[0246] Step 7:
[0247] The device sends the user's operational information (slide progress and comments) to the server. The server uses this information to generate the next reaction or question. At the same time, the emotion engine analyzes the user's mental state based on the acquired emotional data and creates corresponding feedback.
[0248] Step 8:
[0249] The user can see the reactions of the virtual audience (e.g., "Interesting!", questions like, "What are some applications of this technology?") through the smart glasses and respond appropriately. The user's answers are sent to the server via the device.
[0250] Step 9:
[0251] The server analyzes the user's answers and emotional data to generate follow-up questions and feedback. This information is then sent back to the smart glasses and displayed to the user. For example, if the user is nervous, the system will suggest, "Try taking a deep breath to relax."
[0252] Step 10:
[0253] Once the user has finished the simulation, the server analyzes the overall performance data (e.g., the evaluation of each slide, the quality and timing of the user's responses, overall progress, and emotional data) and generates an evaluation result.
[0254] Step 11:
[0255] The user checks the practice results through the application and refers to the feedback provided by the server to improve the presentation materials and delivery method.
[0256] Step 12:
[0257] Users can re-run the simulation and practice as needed. By linking the emotion engine and reaction generation system, users can significantly improve the quality of their presentations, giving them the confidence to take the stage in real life.
[0258] Example 2
[0259] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0260] Previous presentation practice systems lacked the ability to analyze users' emotional states in real time and provide feedback based on that analysis. As a result, users were unable to adequately manage their own emotions and it was difficult to improve the quality of their presentations. It was also difficult to obtain detailed feedback from a virtual audience. This could lead to users losing confidence in their actual presentations and potentially resulting in poor performance.
[0261] The specification processing by the specification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for analyzing uploaded materials and generating reactions of the virtual audience, means for generating evaluation results and feedback based on the progress of the presentation and the user's emotional data, and an emotion analysis engine for analyzing the user's facial expressions and voice in real time and generating feedback. This allows the user to receive feedback according to their own emotional state in real time, improving the quality of their presentation and giving them confidence in their actual presentation.
[0262] A "user" is an individual who uses the system to practice a presentation through the application.
[0263] "Profile information" is information about personal attributes including the user's name, job title, experience, etc.
[0264] "Audience information" refers to attribute information such as age, occupation, and areas of interest of the virtual audience to whom the presentation is directed.
[0265] "Presentation materials" are presentation content such as slides and PDF files uploaded by users.
[0266] The "server" is a computer system that performs processes such as analyzing uploaded presentation materials, generating reactions from a virtual audience, and analyzing user emotional data.
[0267] A "terminal" is a device that allows a user to send and receive data to and from a server and operate a wearable device.
[0268] A "wearable device" is a wearable device (e.g., smart glasses) that can be worn by the user to display reactions and perform emotional analysis.
[0269] "Reactions" are responses such as comments, feedback, and questions from the virtual audience.
[0270] An "emotion analysis engine" is software that analyzes a user's facial expressions, voice, etc. in real time, recognizes the user's emotional state, and generates feedback.
[0271] "Feedback" is response information that includes evaluations and advice regarding the user's behavior, emotional state, etc.
[0272] A "virtual audience" is an imagined audience that makes comments and asks questions about a user's presentation generated by the system.
[0273] "Real-time" refers to the analysis and feedback process occurring instantly while the presentation is in progress.
[0274] The present invention provides a system for providing more advanced feedback by combining a system for users to practice presentations with an emotion analysis engine. Specific embodiments of this system are described below.
[0275] First, the user launches a dedicated application and logs in. After logging in, the user enters profile information (e.g., name, job title, experience) and audience information for the presentation (e.g., age, occupation, areas of interest). Next, the user uploads presentation materials (e.g., PowerPoint or PDF files). These materials are sent to the server, where they are analyzed. This analysis includes extracting text and image information for each slide.
[0276] The server uses a generative AI model to generate reactions from a virtual audience based on the uploaded presentation materials and user input. These reactions include comments and questions, and the timing of their occurrence is also determined. For example, the server inputs a prompt statement such as "Generate questions that the audience would ask during a presentation on technological innovations" into the generative AI model, and uses the responses.
[0277] Next, the user puts on the wearable device (e.g., smart glasses) and starts the application's simulation mode. In this mode, the emotion analysis engine analyzes the user's facial expressions and voice tone in real time. Based on the analysis results, real-time feedback is provided according to the user's emotional state (e.g., nervousness, anxiety, confidence). For example, the feedback "This slide makes you a little nervous" is displayed along with advice such as "Take a deep breath to relax."
[0278] During the simulation, the user gives a presentation and receives reactions from the virtual audience. When the user responds to the comments and questions of the virtual audience, the response is sent to the server via the wearable device. The server analyzes the user's response and generates the next reaction or feedback based on it. For example, the server may provide feedback such as, "The explanation of the application example was easy to understand, but please add more concrete examples."
[0279] This system allows users to receive detailed and useful feedback based on their emotional state in real time, which will improve the quality of their presentations and increase their confidence in their presentations. This confidence is expected to contribute to the success of their presentations.
[0280] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0281] Step 1:
[0282] The user launches the application and logs in.
[0283] Specific behavior:
[0284] The user taps the application icon to launch it, and enters their user ID and password on the login screen that appears. If they are successfully authenticated, the user profile screen is displayed.
[0285] Input and Output:
[0286] Input: User ID, Password
[0287] Data processing / calculation: The authentication server checks the ID and password
[0288] Output: Notification of successful or failed login
[0289] Step 2:
[0290] The user enters profile and audience information.
[0291] Specific behavior:
[0292] The user enters profile information (name, job title, experience) and audience information (age, occupation, areas of interest) into the form and clicks the "Next" button.
[0293] Input and Output:
[0294] Input: Profile information, audience information
[0295] Data processing / calculation: Input data is saved on the server as a user profile
[0296] Output: Confirmation of input completion and moving to the next step
[0297] Step 3:
[0298] A user uploads a presentation.
[0299] Specific behavior:
[0300] The user clicks the "Upload Document" button and selects a presentation file (e.g., PowerPoint, PDF) from the file selection dialog. The selected file is sent to the server.
[0301] Input and Output:
[0302] Input: Selected presentation file
[0303] Data processing / calculation: File upload and file storage on the server
[0304] Output: File upload completion notification
[0305] Step 4:
[0306] The server analyzes the uploaded presentation materials.
[0307] Specific behavior:
[0308] The server analyzes the received presentation file and extracts the text and image information of each slide, which is then stored in a database.
[0309] Input and Output:
[0310] Input: Uploaded presentation file
[0311] Data processing / calculation: Extraction of text and image information, storage in database
[0312] Output: Analysis data saved
[0313] Step 5:
[0314] The server uses a generative AI model to generate reactions from the virtual audience.
[0315] Specific behavior:
[0316] Based on the extracted data and input information, the server sends prompts to the generative AI model to generate reactions from the virtual audience, such as "Generate questions that the audience would ask during a presentation about technological innovations."
[0317] Input and Output:
[0318] Input: Analytics data, profile information, audience information
[0319] Data processing / calculation: Reaction generation using generative AI models
[0320] Output: Virtual audience comments, feedback, and questions
[0321] Step 6:
[0322] The user puts on the wearable device and starts the simulation mode.
[0323] Specific behavior:
[0324] The user puts on the smart glasses and clicks the "Start Simulation Mode" button in the application, which activates the wearable device and displays the reactions of the virtual audience.
[0325] Input and Output:
[0326] Input: Instruction to start the wearable device
[0327] Data processing / calculation: Connection and data transmission between terminal and wearable device
[0328] Output: Start of simulation mode, display of virtual audience
[0329] Step 7:
[0330] The emotion analysis engine analyzes the user's facial expressions and tone of voice in real time.
[0331] Specific behavior:
[0332] The emotion analysis engine uses the smart glasses' built-in camera and microphone to capture and analyze the user's facial expressions and tone of voice, and the analysis data is fed back to the user in real time.
[0333] Input and Output:
[0334] Input: User's facial expression data, voice data
[0335] Data processing / calculation: Data analysis and feedback generation using a sentiment analysis engine
[0336] Output: Real-time emotional feedback
[0337] Step 8:
[0338] Get reactions from a virtual audience as you deliver your presentation.
[0339] Specific behavior:
[0340] As users advance through the slides, the virtual audience displays comments and questions at designated times, to which users can respond.
[0341] Input and Output:
[0342] Input: User's slide progress data
[0343] Data processing / calculation: Reaction display timing control, reaction display
[0344] Output: Display of virtual audience comments and questions
[0345] Step 9:
[0346] Users practice their answers to questions from a virtual audience.
[0347] Specific behavior:
[0348] The user answers questions posed by the virtual audience, and the answers are sent to a server via smart glasses, where they are analyzed and used to generate the next reaction or feedback.
[0349] Input and Output:
[0350] Input: User's response audio data
[0351] Data processing / calculation: Analysis of voice data by server, generation of next reaction
[0352] Output: Analysis results and feedback
[0353] Step 10:
[0354] The emotion analysis engine provides feedback on the user's emotion data during the presentation.
[0355] Specific behavior:
[0356] The emotion analysis engine analyzes the data as it goes through the presentation and provides emotional feedback to the user, such as "I'm a little nervous about the next slide," allowing the user to make appropriate adjustments (e.g., take a deep breath, slow down speaking speed).
[0357] Input and Output:
[0358] Input: Real-time facial expression data, voice data
[0359] Data processing / calculation: Analysis by emotion analysis engine and feedback generation
[0360] Output: Real-time feedback according to emotional state
[0361] (Application example 2)
[0362] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0363] In employee training, it is difficult to improve skills by simulating actual work situations and customer interactions and receiving appropriate feedback in real time. There is also a lack of technology that can analyze employees' emotional states and provide specific advice accordingly. Therefore, there is a need for a system that can effectively improve employees' work performance before they actually begin work.
[0364] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0365] In this invention, the server includes means for a user to input profile information and subject information, means for uploading training materials, means for the server to analyze the uploaded materials and generate feedback, means for a terminal to transmit and receive data between the user and the server and operate the smart device, means for the smart device to display the feedback and perform a training simulation, and means for the server to generate evaluation results and feedback based on the progress of the training. This allows employees to receive feedback from virtual customers in real time while undergoing job training, and can provide specific advice according to the employee's emotional state, thereby improving their skills and providing a sense of security.
[0366] "Profile information" is basic personal data such as the user's name, job title, experience, etc.
[0367] "Target information" is data such as the age, occupation, and areas of interest of a hypothetical customer or audience.
[0368] "Training materials" are presentation materials and simulation documents used by users.
[0369] "Uploading means" refers to a method or function that allows a user to send training materials to a server via a terminal.
[0370] "Means for analyzing and generating feedback" refers to the function that enables the server to analyze uploaded materials and provide appropriate comments and advice.
[0371] A "terminal" is a device used by a user, such as a smartphone or tablet.
[0372] A "smart device" is a device for displaying feedback to a user, such as smart glasses or a head-mounted display.
[0373] The "means for performing training simulation" is a function that allows users to simulate actual work through dialogue with virtual customers and scenarios.
[0374] The "means for generating evaluation results and feedback" refers to a method by which the server evaluates the user's performance based on the progress of the training and generates appropriate advice and comments.
[0375] This invention relates to a system for effectively conducting employee training. The system begins when a user enters their profile information and target information and uploads training materials. The uploaded materials are analyzed by a server, and feedback is generated. A terminal transmits and receives data between the user and the server and operates a smart device. The smart device displays the feedback and functions as a means for conducting training simulations.
[0376] Hardware and software used
[0377] Smart devices (smart glasses, head-mounted displays, etc.)
[0378] These devices are used to give users real-time feedback during a simulation.
[0379] Devices (smartphones, tablets)
[0380] Used by users to upload training materials and enter profile and audience information.
[0381] server
[0382] The server installed in the cloud environment is responsible for analyzing materials, generating feedback, sending and receiving data, etc. Specifically, a server with an AI model implemented (e.g., a machine learning model using TensorFlow or PyTorch) is used.
[0383] Cloud services (e.g. AWS, Azure)
[0384] It is used to assist in centralized data management and calculation processing.
[0385] System Operation
[0386] 1. User configuration
[0387] The user uses a terminal to enter their profile information (e.g., job title, experience, etc.) and target information (e.g., age and occupation of the expected customers) according to the training scenario, and then uploads presentation and training materials.
[0388] 2. Server analysis
[0389] The server analyzes the uploaded materials and uses a generative AI model to generate hypothetical customer reactions and questions, taking into account the user's past training and performance data to generate evaluation results and feedback.
[0390] 3. Real-time feedback
[0391] When a user puts on a smart device and starts training, feedback is displayed on the smart device via the terminal from the server. For example, if the user feels tense, advice on how to relax or specific advice on questions from the customer is provided in real time.
[0392] Specific examples
[0393] For example, if a user is conducting training on the theme of "new smartphone products," they can enter "new staff member with two months' experience" as their profile information and set "male in his 30s, technically knowledgeable customer" as the target scenario. Next, they select a training scenario and questions from the virtual customer are displayed.
[0394] Example prompt sentence:
[0395] A customer asks, "What are the camera features on this phone?" Observe how your employees react and ask the question, and provide appropriate feedback. If your employees seem nervous, include tips to help them relax.
[0396] The system allows employees to receive real-time feedback from virtual customers while training on the job, improving their skills and providing peace of mind.
[0397] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0398] Step 1:
[0399] Operation: A user uses a terminal to enter their profile information and target information.
[0400] Input: User profile information (e.g., name, job title, experience), audience information (e.g., age, occupation, areas of interest).
[0401] Output: The entered information is sent from the terminal to the server.
[0402] Step 2:
[0403] How it works: A user uploads presentation or training materials via a terminal.
[0404] Input: Presentation materials (e.g. PowerPoint files or PDFs).
[0405] Output: Uploaded materials are sent from the device to the server and saved.
[0406] Step 3:
[0407] How it works: The server receives the uploaded material and analyzes it.
[0408] Input: Presentation materials, user profile information, audience information.
[0409] Data processing / calculation: Extract text and image information from documents using generative AI models.
[0410] Output: Generate virtual customer reactions and questions based on the extracted information.
[0411] Step 4:
[0412] Operation: The server generates feedback for the virtual customer based on the user's input data and the analysis results.
[0413] Input: User profile information, privacy data, analytics results.
[0414] Data processing / calculation: Generate prompt statements and generate feedback from virtual customers using AI models.
[0415] Output: The generated feedback data is sent from the server to the terminal.
[0416] Step 5:
[0417] Actions: A user puts on a smart device and starts a training simulation.
[0418] Input: Generated feedback data, user voice, and facial expression data.
[0419] Output: Virtual customer feedback is displayed in real time on a smart device.
[0420] Step 6:
[0421] Operation: The server receives the user's voice and facial expression data during the simulation and analyzes it using the emotion engine.
[0422] Input: User's voice and facial expression data.
[0423] Data processing / calculation: Analyzed by the emotion engine to evaluate the user's stress level and tension.
[0424] Output: As a result of the analysis, feedback based on the user's emotional state is generated.
[0425] Step 7:
[0426] How it works: Based on the analysis results of the emotion engine, the server generates advice for the user in real time and sends it to the smart device.
[0427] Input: Analysis results of the emotion engine, feedback from virtual customers.
[0428] Output: Specific advice to the user (e.g., "Take deep breaths to relax") is displayed on the smart device.
[0429] Step 8:
[0430] Action: The user answers questions posed by the virtual customer.
[0431] Input: User response data.
[0432] Output: The user's answer data is sent to the server and analyzed.
[0433] Step 9:
[0434] How it works: The server analyzes the user's answers and generates the following feedback and rating results:
[0435] Input: User response data, past training data.
[0436] Data processing / calculation: Analyzed by AI model.
[0437] Output: As a result of the analysis, feedback and evaluation results from the next virtual customer are generated and sent to the smart device.
[0438] Step 10:
[0439] Operation: After the training is completed, the server generates an evaluation result of the entire training and provides it to the user.
[0440] Input: All training data, emotion data.
[0441] Data processing / calculation: Conduct an overall evaluation of training performance.
[0442] Output: The final evaluation results are displayed on the device along with feedback on future improvements.
[0443] These steps allow employees to virtually experience real-world work scenarios and receive appropriate feedback in real time, resulting in improved work skills and a sense of psychological security.
[0444] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0445] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0446] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0447] [Second embodiment]
[0448] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0449] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0450] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0451] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0452] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0453] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0454] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0455] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0456] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0457] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0458] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0459] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0460] The present invention relates to a system for allowing a user to practice a presentation. Specific embodiments of the present invention will be described below.
[0461] The system begins with the user entering their profile information and audience information and uploading presentation materials. Next, the server analyzes the materials and generates reactions. The device sends and receives data between the user and the server and operates the smart glasses. Finally, the smart glasses display the reactions and simulate the presentation.
[0462] Program processing
[0463] 1. Enter your user profile and audience information
[0464] A user launches the application and enters their profile information and the audience information for the presentation, allowing the system to build an appropriate simulation scenario based on the characteristics of the user and audience.
[0465] 2. Upload your presentation materials
[0466] Users upload presentation materials (e.g., PowerPoint or PDF files) through the application. These materials are sent to the server and analyzed there. Text and image information for each slide is extracted from the analyzed materials.
[0467] 3. Generating Reactions
[0468] The server uses an AI model to generate reactions based on the uploaded presentation materials, user profile information, and audience information, including comments, reactions, and questions from the virtual audience.
[0469] 4. Connecting to smart glasses
[0470] The user wears the smart glasses, and the application establishes a connection with the smart glasses, which are devices that receive data from the server and display reactions.
[0471] 5. Start the presentation simulation
[0472] The user operates the application to enter a presentation simulation mode, and the reactions of the virtual audience are displayed through the smart glasses. The user advances through the slides, presents, and receives feedback from the virtual audience.
[0473] 6. Practice Q&A
[0474] As part of the simulation, users respond to questions posed by a virtual audience, and their responses are sent via smart glasses to a server that analyzes their responses and generates subsequent reactions and feedback.
[0475] Specific examples
[0476] For example, if a user wants to give a presentation on the topic of "agile development," they first enter their profile information, such as their job title and experience, into the application, and set the audience attributes as "engineers, aged 30-40, interested in agile development." They then upload their PowerPoint presentation materials through the application, and the server analyzes them.
[0477] Once a user puts on the smart glasses and enters the simulation mode, a question is displayed from the virtual audience: "What are the main challenges in adopting agile development?" The user responds to this question, and the response is sent to the server for analysis. The server then generates appropriate feedback and additional reactions, which are then displayed to the user.
[0478] In this way, users can repeatedly simulate their presentation with a virtual audience in advance, improving the quality of their presentation and giving them confidence when they actually take the stage.
[0479] The processing flow will be explained below.
[0480] Step 1:
[0481] A user launches the application and logs in. The user enters their profile information and sets up audience information, including name, job title, experience, audience age, occupation, and areas of interest.
[0482] Step 2:
[0483] Users select presentation materials and upload them from their PC or smartphone through the application, which can be in PowerPoint or PDF format.
[0484] Step 3:
[0485] The server receives and stores the uploaded presentation materials. The server analyzes the materials and extracts text and image information for each slide. The analysis uses natural language processing and image recognition technology.
[0486] Step 4:
[0487] The server uses an AI model to generate reactions based on the user's profile information and audience information, specifically determining what comments and questions the virtual audience will make and when they will react.
[0488] Step 5:
[0489] The user wears the smart glasses and starts the simulation mode in the application. The device connects to the smart glasses and establishes data communication with the server.
[0490] Step 6:
[0491] The server confirms the start of the simulation mode and sends the analysis results of the presentation materials and reaction data to the smart glasses. As the user advances through the slides, the smart glasses display the appropriate reactions.
[0492] Step 7:
[0493] The device sends user operation information to the server, including the user's actions as they advance through slides and the progress of the presentation content via voice recognition. The server then uses this information to generate the next reaction or question.
[0494] Step 8:
[0495] The user can see the reactions of the virtual audience (e.g., "Interesting!", questions like, "What are some applications of this technology?") through the smart glasses and respond appropriately. The user's responses are sent to the server via the device.
[0496] Step 9:
[0497] The server analyzes the user's answers and generates follow-up questions and feedback, which are then sent back to the smart glasses and displayed to the user.
[0498] Step 10:
[0499] Once the user has finished the simulation, the server analyzes their overall performance and generates results, including a grade for each slide, the quality and timing of the user's responses, and their overall progress.
[0500] Step 11:
[0501] The user checks the practice results through the application and refers to the feedback provided by the server to improve the presentation materials and delivery method.
[0502] Step 12:
[0503] Users can re-run the simulation and practice as needed, improving the quality of their presentation and building confidence in speaking.
[0504] Example 1
[0505] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0506] Traditional presentation practice methods often require a live audience or offer limited simulations, meaning users miss out on opportunities to effectively improve their presentation skills. Additionally, it's difficult to immediately identify areas for improvement through real-time feedback and Q&A sessions.
[0507] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0508] In this invention, the server includes a means for a user to input profile information and student information, a means for uploading presentation materials, and a means for an information processing device to analyze the uploaded materials and generate responses, thereby enabling the user to receive real-time feedback and questions from a virtual audience and to practice presentations effectively.
[0509] "User" refers to an individual or group who uses the system to practice a presentation.
[0510] "Profile information" refers to detailed information about a user, such as the user's job title, years of experience, and area of expertise.
[0511] "Audience information" refers to information such as the attributes and areas of interest of the audience to whom the presentation is directed.
[0512] "Presentation materials" refers to data such as slides and documents that a user uses to give a presentation.
[0513] An "information processing device" refers to a server or computer used to analyze presentation materials and generate responses.
[0514] "Responses" refers to interactive feedback such as comments, reactions, and questions from the virtual audience.
[0515] A "terminal" is a device (e.g., a smartphone, tablet, or PC) that transmits and receives data between a user and an information processing device and operates a video display device.
[0516] A "video display device" is a display device (e.g., smart glasses, VR headset, etc.) that allows users to visually check reactions.
[0517] "Practice presentation" refers to an activity in which a user prepares for an actual presentation by simulating a presentation in front of a virtual audience.
[0518] The present invention relates to a system for users to practice presentations. The system begins with the user entering profile information and participant information and uploading presentation materials. Next, an information processing device analyzes the materials and generates responses. A terminal transmits and receives data between the user and the information processing device and operates a video display device. Finally, the video display device displays the responses, and the presentation is practiced.
[0519] In a specific embodiment of the system, a program is generated and processing is performed in the following procedure.
[0520] First, the user launches the application, logs in, and enters their profile information (e.g., name, job title, years of experience, etc.). Then, they enter information about the target audience (e.g., industry, age group, areas of interest, etc.). The system then creates an appropriate simulation scenario based on the characteristics of the user and the audience.
[0521] Next, the user uploads presentation materials (e.g., PowerPoint or PDF files) to the application. These materials are sent to an information processing device via the terminal. The information processing device analyzes the uploaded materials and extracts text and image information for each slide. This analysis is performed using text analysis software and image recognition software.
[0522] The information processing device uses a generative AI model to generate responses based on the extracted material content, user profile information, and participant information. The generated responses include comments, reactions, and questions from the virtual audience. For example, a question might be generated: "What are the main challenges in introducing agile development?"
[0523] The user wears a video display device (e.g., smart glasses) and selects the "Video Display Device Connection" option within the application. The terminal establishes a connection with the video display device using a connection method such as Bluetooth or Wi-Fi. The video display device receives the response data from the information processing device and prepares to display it.
[0524] When the user selects "Start Simulation," the presentation begins, and the virtual audience's reactions and questions are displayed on the video display device. The user then verbally responds to the questions, which are then transmitted to the information processing device via the video display device. The information processing device then uses voice analysis software to analyze the user's responses and generate appropriate feedback or additional responses.
[0525] For example, if a user is giving a presentation on the topic of "Agile Development," the steps would be:
[0526] 1. The user enters profile information such as their job title and experience, and sets the audience attributes as "Engineer, 30-40 years old, interested in agile development."
[0527] 2. Upload PowerPoint presentation materials through the application, and the information processing device analyzes the materials.
[0528] 3. Once the user puts on the video display device and starts the simulation mode, a question is displayed from the virtual audience: "What are the main challenges in adopting agile development?"
[0529] 4. The user answers the question, and the answer is sent to the information processing device and analyzed.
[0530] 5. The information processing device generates further appropriate feedback and additional reactions and displays them to the user.
[0531] Examples of prompts include: "The user's profile information is an engineer in his 30s with five years of experience in agile development," "The audience is engineers aged 30-40 who are interested in agile development," and "The theme of the presentation is 'Introduction to agile development and its challenges.' Please generate questions and feedback from the virtual audience."
[0532] In this way, users can repeatedly simulate their presentation with a virtual audience in advance, improving the quality of their presentation and giving them confidence when they actually take the stage.
[0533] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0534] Step 1:
[0535] Entering user profile and audience information
[0536] Input: User profile information (e.g., name, job title, years of experience) and audience information (e.g., industry, age group, interests)
[0537] Processing: The user starts the application and authenticates by entering their user ID and password on the login screen. After logging in, they enter their profile information and audience information in the form that appears. The server receives this information and stores it in a database. It also designs an appropriate simulation scenario based on the user and audience information.
[0538] Output: User and audience information stored in a database, and simulation scenario designs
[0539] Step 2:
[0540] Upload presentation materials
[0541] Input: Presentation materials (e.g. PowerPoint files, PDF files)
[0542] Processing: The user clicks the "File Upload" button in the application and selects presentation materials from a selection screen. The device reads the selected file and sends it to the server. The server analyzes the received file and extracts text and image data for each slide. This analysis is performed using text analysis software and image recognition software.
[0543] Output: Parsed text and image data for each slide
[0544] Step 3:
[0545] Creating a reaction
[0546] Input: Parsed presentation materials, user profile information, audience information
[0547] Processing: The server uses the generative AI model to generate comments, reactions, and questions from the virtual audience based on the analyzed content and the input user profile and audience information. For example, it dynamically generates questions and reactions related to the content of the slides.
[0548] Output: Generated virtual audience reaction data (comments, questions, responses, etc.)
[0549] Step 4:
[0550] Connecting to smart glasses
[0551] Input: Wi-Fi and Bluetooth information for connection
[0552] Processing: The user wears a video display device (e.g., smart glasses) and selects the "Connect to video display device" option within the application. The device establishes a connection with the video display device using the selected connection method. The device confirms that a connection with the video display device has been established.
[0553] Output: Established connection status with video display device
[0554] Step 5:
[0555] Start of presentation simulation
[0556] Input: Command to start the simulation
[0557] Processing: The user clicks the "Start Simulation" button in the application. The terminal sends the presentation materials and the generated reaction data to the video display device and prepares to start the simulation.
[0558] Output: Presentation materials and reaction data sent to a video display device
[0559] Step 6:
[0560] Q&A practice
[0561] Input: Virtual audience question displayed on a video display device
[0562] Processing: The video display device displays the question from the virtual audience. The user responds verbally to the question, and the response is captured using the video display device's audio input function and sent to the server. The server analyzes the audio data and generates feedback and additional reactions based on the user's response, which are then sent back to the video display device.
[0563] Output: Parsed user response data, generated new feedback and reactions
[0564] (Application example 1)
[0565] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0566] Conventional presentation practice systems primarily require users to practice on their own, lacking the ability to provide real-time feedback or reflect the reactions of a virtual audience. Furthermore, systems for effectively training salespeople in brick-and-mortar stores on product explanations are also inadequate, and there is a need for a system that can respond immediately to customer questions and reactions. The present invention aims to solve these problems and improve the quality of presentation practice and sales training.
[0567] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0568] In this invention, the server includes: a means for a user to input profile information and audience information; a means for uploading presentation materials; a means for the server to analyze the uploaded materials and generate reactions; a means for a terminal to transmit and receive data between the user and the server and operate the smart glasses; a means for the smart glasses to display reactions and perform a presentation simulation; a means for a salesperson to input profile information and customer information and upload product explanation materials; a means for the server to analyze the product explanation materials and generate questions and reactions of virtual customers; and a means for the smart glasses to display the responses and questions of the virtual customers in real time and perform a sales training simulation. This allows the user to make a presentation while receiving reactions from the virtual audience, and enables the salesperson to receive practical training through interactions with virtual customers.
[0569] "User profile information" is information that includes personal information, job information, experience, and the like about the user.
[0570] "Audience information" is attribute information about the audience to whom the presentation is given.
[0571] "Presentation materials" are presentation materials provided in the form of slides or documents.
[0572] A "server" is a computer system that receives data from users, analyzes it, and generates appropriate reactions.
[0573] "Means for generating reactions" is a function that uses an AI model to generate virtual comments and questions based on presentation materials and user information.
[0574] A "terminal" is a device that transmits and receives data between a user and a server.
[0575] "Smart glasses" are eyeglass-type devices that utilize augmented reality technology to display information to users in real time.
[0576] The "means for simulating a presentation" is a function that allows a user to use smart glasses to give a presentation while interacting with a virtual audience.
[0577] "Salesperson profile information" is information including job information, products in charge, experience, etc., related to the salesperson.
[0578] "Customer information" is attribute information about customers who are the target of sales.
[0579] "Product description materials" are materials used to explain the characteristics and functions of a product.
[0580] "Virtual Customer" means a fictitious customer generated using AI technology.
[0581] "Means for generating questions and reactions" refers to a function that uses an AI model to generate virtual customer questions and reactions based on salesperson profile information and product explanation materials.
[0582] The "means for conducting sales training simulations" is a function that allows salespeople to use smart glasses to conduct practical training through interactions with virtual customers.
[0583] The present invention relates to a comprehensive system for users to practice presentations and sales training. The system begins when the user inputs profile information and audience information and uploads presentation or product explanation materials. The server analyzes the uploaded materials and generates reactions. The terminal transmits and receives data between the user and the server and operates the smart glasses. The smart glasses display reactions and perform a presentation simulation or sales training simulation.
[0584] Specific steps
[0585] 1. Enter your profile and audience information
[0586] A user launches the application and enters their profile information (e.g., job title, experience, products) and audience or customer information for the presentation or sale (e.g., audience age range, interests, skill level).
[0587] 2. Upload presentation or product description materials
[0588] Users upload presentation materials (e.g., PowerPoint or PDF files) or product descriptions through the application. These materials are then sent to the server, where they are analyzed. During the analysis process, important text and image information is extracted from the materials.
[0589] 3. Generating Reactions
[0590] The server uses a generative AI model to generate reactions based on the uploaded materials and user and audience information, including comments, reactions, and questions from the virtual audience or virtual customers.
[0591] 4. Connecting to smart glasses
[0592] The user wears the smart glasses, and the application establishes a connection with the smart glasses, which receives data from the server and displays reactions.
[0593] 5. Start a presentation simulation or sales training
[0594] The user interacts with the application and enters the simulation mode, where the smart glasses display the reactions of a virtual audience and virtual customers, providing virtual feedback as the user progresses through slides and product presentations.
[0595] 6. Practice Q&A
[0596] As part of the simulation, users respond to questions posed by a virtual audience or customer, and their answers are sent through the smart glasses to a server that analyzes their responses and generates subsequent reactions and feedback.
[0597] Hardware and software used
[0598] Hardware used: Smart glasses (e.g., Google Glass, Vuzix Blade)
[0599] Software used: Flask (server-side framework), generative AI model (reaction generation), data analysis system
[0600] Specific examples
[0601] For example, when training a salesperson to explain a new smartphone, the following prompts are fed into the generative AI model:
[0602] "User: Salesperson, Experience: 3 years, Product: New smartphone, Target Audience: 20-30 years old, Tech-savvy"
[0603] Feedback from the virtual customer is displayed on the smart glasses with questions such as:
[0604] "How does this smartphone's camera performance compare to other products?"
[0605] In this way, users can improve the quality of their presentations and sales training by repeating virtual simulations.
[0606] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0607] Step 1:
[0608] A user launches the application and enters profile and audience information.
[0609] This information includes the user's job title, experience, products they work on, and target audience demographics (age range, interests, skill level, etc.).
[0610] The server receives these input data and stores them in a database.
[0611] Step 2:
[0612] A user uploads a presentation or product description.
[0613] These materials are provided in formats such as PDF and PowerPoint files.
[0614] The server receives the material and uses text analysis software (e.g., Apache Tika) to extract text and image information from the material.
[0615] The extracted information is stored in a database on the server.
[0616] Step 3:
[0617] The server generates reactions using a generative AI model based on uploaded materials, user information, and audience information.
[0618] This process generates questions and comments from the virtual audience or customer based on the content of the material.
[0619] The generated reactions are temporarily stored on the server.
[0620] Step 4:
[0621] The user puts on smart glasses (e.g., Google Glass, Vuzix Blade), and the application establishes a connection with the smart glasses.
[0622] The server transmits the generated reaction data to the smart glasses.
[0623] The smart glasses display the received data in real time in the user's field of vision.
[0624] Step 5:
[0625] The user operates the application and starts a presentation simulation or a sales training simulation.
[0626] Smart glasses display the reactions of virtual audiences and virtual customers.
[0627] Users receive real-time feedback as they progress through their presentations and product explanations.
[0628] Step 6:
[0629] As part of a presentation simulation or sales training, a user answers questions posed by a virtual audience or virtual customers.
[0630] The user's answers are sent to the server via the smart glasses.
[0631] The server analyzes the user's answers and generates the next reaction or feedback.
[0632] The generated feedback is then sent back to the smart glasses and displayed to the user.
[0633] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0634] The present invention relates to a system for providing more advanced feedback by combining a system for users to practice presentations with an emotion engine that recognizes the emotions of the users. Specific embodiments of the present invention will be described below.
[0635] System configuration
[0636] Entering user profile and audience information
[0637] A user launches the application and logs in. The user enters their profile information (e.g., name, job title, experience) and also sets audience information (e.g., age, occupation, areas of interest) for the presentation.
[0638] Upload presentation materials
[0639] Users upload presentation materials (e.g., PowerPoint or PDF files) from their PC or smartphone through the application. These materials are then sent to the server, where they are analyzed. The analysis includes extracting text and image information for each slide.
[0640] Creating a reaction
[0641] The server uses an AI model to generate reactions based on the uploaded presentation materials, user profile information, and audience information. Specifically, it determines what comments and questions the virtual audience will make, as well as the timing of their reactions.
[0642] Linking the Emotion Engine
[0643] When a user puts on the smart glasses and starts the simulation mode in the application, the emotion engine analyzes the user's facial expressions and tone of voice to recognize the user's emotional state in real time. For example, if the user is nervous, it will provide advice on how to relax.
[0644] Start of presentation simulation
[0645] The user interacts with the application to enter a presentation simulation mode, where the smart glasses display the reactions of the virtual audience. The user advances through the slides, presenting, and receives feedback from the virtual audience.
[0646] Q&A practice
[0647] As part of the simulation, users respond to questions posed by a virtual audience, and their responses are sent via smart glasses to a server that analyzes their responses and generates subsequent reactions and feedback.
[0648] Emotional data feedback and adjustment
[0649] The emotion engine analyzes the user's emotional data as they make their presentation and generates feedback based on their emotional state. For example, if the user feels anxious, the engine will advise them to explain the key points more simply.
[0650] Specific examples
[0651] For example, if a user wants to give a presentation on the theme of "innovative technology," the user first enters their profile information (e.g., experience as an engineer, field of expertise) into the application, and sets the audience attributes as "corporate technical personnel, aged 30-40, interested in technological innovation." Then, the user uploads presentation materials (PowerPoint files).
[0652] When a user puts on the smart glasses and starts the simulation mode, the emotion engine analyzes the user's facial expressions and tone of voice, and displays feedback such as, "You're feeling a little nervous on this slide," while the smart glasses also display advice such as, "Take a deep breath to relax."
[0653] In the presentation simulation, a virtual audience asks the user, "Please tell us some specific examples of how your technology can be applied," and the user responds. The response is sent to the server, where it is analyzed and the results are provided as feedback. For example, the user might receive feedback such as, "The explanation of the application example was easy to understand. However, I would like more specific examples."
[0654] In this way, by linking with the emotion engine, users can receive more detailed and useful feedback, significantly improving the quality of their presentations and giving them more confidence when they take the stage.
[0655] The processing flow will be explained below.
[0656] Step 1:
[0657] A user launches the application and logs in. The user enters their profile information (name, job title, experience, etc.) and also sets audience information (age, occupation, areas of interest, etc.) for the presentation.
[0658] Step 2:
[0659] Users select and upload presentation materials (PowerPoint or PDF files) from their PC or smartphone through the application, which are then sent to the server.
[0660] Step 3:
[0661] The server receives and analyzes the uploaded presentation materials, extracting text and image information from each slide, and using natural language processing (NLP) and image recognition technology to identify key points and locations within the materials.
[0662] Step 4:
[0663] The server uses an AI model to generate reactions based on the user's profile information and audience information. Specifically, it determines what comments and questions the virtual audience will make, as well as the timing of their reactions. The emotion engine also prepares to acquire the user's emotional data.
[0664] Step 5:
[0665] The user puts on the smart glasses and starts the simulation mode in the application. The device connects to the smart glasses and performs data communication with the server. The smart glasses receive reaction data from the server and prepare to display it.
[0666] Step 6:
[0667] The server confirms the start of the simulation mode and sends the analysis results of the presentation materials and reaction data to the smart glasses. As the user advances through the slides, the smart glasses display appropriate reactions. The emotion engine also analyzes the user's facial expressions and tone of voice in real time.
[0668] Step 7:
[0669] The device sends the user's operational information (slide progress and comments) to the server. The server uses this information to generate the next reaction or question. At the same time, the emotion engine analyzes the user's mental state based on the acquired emotional data and creates corresponding feedback.
[0670] Step 8:
[0671] The user can see the reactions of the virtual audience (e.g., "Interesting!", questions like, "What are some applications of this technology?") through the smart glasses and respond appropriately. The user's answers are sent to the server via the device.
[0672] Step 9:
[0673] The server analyzes the user's answers and emotional data to generate follow-up questions and feedback. This information is then sent back to the smart glasses and displayed to the user. For example, if the user is nervous, the system will suggest, "Try taking a deep breath to relax."
[0674] Step 10:
[0675] Once the user has finished the simulation, the server analyzes the overall performance data (e.g., the evaluation of each slide, the quality and timing of the user's responses, overall progress, and emotional data) and generates an evaluation result.
[0676] Step 11:
[0677] The user checks the practice results through the application and refers to the feedback provided by the server to improve the presentation materials and delivery method.
[0678] Step 12:
[0679] Users can re-run the simulation and practice as needed. By linking the emotion engine and reaction generation system, users can significantly improve the quality of their presentations, giving them the confidence to take the stage in real life.
[0680] Example 2
[0681] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0682] Previous presentation practice systems lacked the ability to analyze users' emotional states in real time and provide feedback based on that analysis. As a result, users were unable to adequately manage their own emotions and it was difficult to improve the quality of their presentations. It was also difficult to obtain detailed feedback from a virtual audience. This could lead to users losing confidence in their actual presentations and potentially resulting in poor performance.
[0683] The specification processing by the specification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for analyzing uploaded materials and generating reactions of the virtual audience, means for generating evaluation results and feedback based on the progress of the presentation and the user's emotional data, and an emotion analysis engine for analyzing the user's facial expressions and voice in real time and generating feedback. This allows the user to receive feedback according to their own emotional state in real time, improving the quality of their presentation and giving them confidence in their actual presentation.
[0684] A "user" is an individual who uses the system to practice a presentation through the application.
[0685] "Profile information" is information about personal attributes including the user's name, job title, experience, etc.
[0686] "Audience information" refers to attribute information such as age, occupation, and areas of interest of the virtual audience to whom the presentation is directed.
[0687] "Presentation materials" are presentation content such as slides and PDF files uploaded by users.
[0688] The "server" is a computer system that performs processes such as analyzing uploaded presentation materials, generating reactions from a virtual audience, and analyzing user emotional data.
[0689] A "terminal" is a device that allows a user to send and receive data to and from a server and operate a wearable device.
[0690] A "wearable device" is a wearable device (e.g., smart glasses) that can be worn by the user to display reactions and perform emotional analysis.
[0691] "Reactions" are responses such as comments, feedback, and questions from the virtual audience.
[0692] An "emotion analysis engine" is software that analyzes a user's facial expressions, voice, etc. in real time, recognizes the user's emotional state, and generates feedback.
[0693] "Feedback" is response information that includes evaluations and advice regarding the user's behavior, emotional state, etc.
[0694] A "virtual audience" is an imagined audience that makes comments and asks questions about a user's presentation generated by the system.
[0695] "Real-time" refers to the analysis and feedback process occurring instantly while the presentation is in progress.
[0696] The present invention provides a system for providing more advanced feedback by combining a system for users to practice presentations with an emotion analysis engine. Specific embodiments of this system are described below.
[0697] First, the user launches a dedicated application and logs in. After logging in, the user enters profile information (e.g., name, job title, experience) and audience information for the presentation (e.g., age, occupation, areas of interest). Next, the user uploads presentation materials (e.g., PowerPoint or PDF files). These materials are sent to the server, where they are analyzed. This analysis includes extracting text and image information for each slide.
[0698] The server uses a generative AI model to generate reactions from a virtual audience based on the uploaded presentation materials and user input. These reactions include comments and questions, and the timing of their occurrence is also determined. For example, the server inputs a prompt statement such as "Generate questions that the audience would ask during a presentation on technological innovations" into the generative AI model, and uses the responses.
[0699] Next, the user puts on the wearable device (e.g., smart glasses) and starts the application's simulation mode. In this mode, the emotion analysis engine analyzes the user's facial expressions and voice tone in real time. Based on the analysis results, real-time feedback is provided according to the user's emotional state (e.g., nervousness, anxiety, confidence). For example, the feedback "This slide makes you a little nervous" is displayed along with advice such as "Take a deep breath to relax."
[0700] During the simulation, the user gives a presentation and receives reactions from the virtual audience. When the user responds to the comments and questions of the virtual audience, the response is sent to the server via the wearable device. The server analyzes the user's response and generates the next reaction or feedback based on it. For example, the server may provide feedback such as, "The explanation of the application example was easy to understand, but please add more concrete examples."
[0701] This system allows users to receive detailed and useful feedback based on their emotional state in real time, which will improve the quality of their presentations and increase their confidence in their presentations. This confidence is expected to contribute to the success of their presentations.
[0702] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0703] Step 1:
[0704] The user launches the application and logs in.
[0705] Specific behavior:
[0706] The user taps the application icon to launch it, and enters their user ID and password on the login screen that appears. If they are successfully authenticated, the user profile screen is displayed.
[0707] Input and Output:
[0708] Input: User ID, Password
[0709] Data processing / calculation: The authentication server checks the ID and password
[0710] Output: Notification of successful or failed login
[0711] Step 2:
[0712] The user enters profile and audience information.
[0713] Specific behavior:
[0714] The user enters profile information (name, job title, experience) and audience information (age, occupation, areas of interest) into the form and clicks the "Next" button.
[0715] Input and Output:
[0716] Input: Profile information, audience information
[0717] Data processing / calculation: Input data is saved on the server as a user profile
[0718] Output: Confirmation of input completion and moving to the next step
[0719] Step 3:
[0720] A user uploads a presentation.
[0721] Specific behavior:
[0722] The user clicks the "Upload Document" button and selects a presentation file (e.g., PowerPoint, PDF) from the file selection dialog. The selected file is sent to the server.
[0723] Input and Output:
[0724] Input: Selected presentation file
[0725] Data processing / calculation: File upload and file storage on the server
[0726] Output: File upload completion notification
[0727] Step 4:
[0728] The server analyzes the uploaded presentation materials.
[0729] Specific behavior:
[0730] The server analyzes the received presentation file and extracts the text and image information of each slide, which is then stored in a database.
[0731] Input and Output:
[0732] Input: Uploaded presentation file
[0733] Data processing / calculation: Extraction of text and image information, storage in database
[0734] Output: Analysis data saved
[0735] Step 5:
[0736] The server uses a generative AI model to generate reactions from the virtual audience.
[0737] Specific behavior:
[0738] Based on the extracted data and input information, the server sends prompts to the generative AI model to generate reactions from the virtual audience, such as "Generate questions that the audience would ask during a presentation about technological innovations."
[0739] Input and Output:
[0740] Input: Analytics data, profile information, audience information
[0741] Data processing / calculation: Reaction generation using generative AI models
[0742] Output: Virtual audience comments, feedback, and questions
[0743] Step 6:
[0744] The user puts on the wearable device and starts the simulation mode.
[0745] Specific behavior:
[0746] The user puts on the smart glasses and clicks the "Start Simulation Mode" button in the application, which activates the wearable device and displays the reactions of the virtual audience.
[0747] Input and Output:
[0748] Input: Instruction to start the wearable device
[0749] Data processing / calculation: Connection and data transmission between terminal and wearable device
[0750] Output: Start of simulation mode, display of virtual audience
[0751] Step 7:
[0752] The emotion analysis engine analyzes the user's facial expressions and tone of voice in real time.
[0753] Specific behavior:
[0754] The emotion analysis engine uses the smart glasses' built-in camera and microphone to capture and analyze the user's facial expressions and tone of voice, and the analysis data is fed back to the user in real time.
[0755] Input and Output:
[0756] Input: User's facial expression data, voice data
[0757] Data processing / calculation: Data analysis and feedback generation using a sentiment analysis engine
[0758] Output: Real-time emotional feedback
[0759] Step 8:
[0760] Get reactions from a virtual audience as you deliver your presentation.
[0761] Specific behavior:
[0762] As users advance through the slides, the virtual audience displays comments and questions at designated times, to which users can respond.
[0763] Input and Output:
[0764] Input: User's slide progress data
[0765] Data processing / calculation: Reaction display timing control, reaction display
[0766] Output: Display of virtual audience comments and questions
[0767] Step 9:
[0768] Users practice their answers to questions from a virtual audience.
[0769] Specific behavior:
[0770] The user answers questions posed by the virtual audience, and the answers are sent to a server via smart glasses, where they are analyzed and used to generate the next reaction or feedback.
[0771] Input and Output:
[0772] Input: User's response audio data
[0773] Data processing / calculation: Analysis of voice data by server, generation of next reaction
[0774] Output: Analysis results and feedback
[0775] Step 10:
[0776] The emotion analysis engine provides feedback on the user's emotion data during the presentation.
[0777] Specific behavior:
[0778] The emotion analysis engine analyzes the data as it goes through the presentation and provides emotional feedback to the user, such as "I'm a little nervous about the next slide," allowing the user to make appropriate adjustments (e.g., take a deep breath, slow down speaking speed).
[0779] Input and Output:
[0780] Input: Real-time facial expression data, voice data
[0781] Data processing / calculation: Analysis by emotion analysis engine and feedback generation
[0782] Output: Real-time feedback according to emotional state
[0783] (Application example 2)
[0784] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0785] In employee training, it is difficult to improve skills by simulating actual work situations and customer interactions and receiving appropriate feedback in real time. There is also a lack of technology that can analyze employees' emotional states and provide specific advice accordingly. Therefore, there is a need for a system that can effectively improve employees' work performance before they actually begin work.
[0786] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0787] In this invention, the server includes means for a user to input profile information and subject information, means for uploading training materials, means for the server to analyze the uploaded materials and generate feedback, means for a terminal to transmit and receive data between the user and the server and operate the smart device, means for the smart device to display the feedback and perform a training simulation, and means for the server to generate evaluation results and feedback based on the progress of the training. This allows employees to receive feedback from virtual customers in real time while undergoing job training, and can provide specific advice according to the employee's emotional state, thereby improving their skills and providing a sense of security.
[0788] "Profile information" is basic personal data such as the user's name, job title, experience, etc.
[0789] "Target information" is data such as the age, occupation, and areas of interest of a hypothetical customer or audience.
[0790] "Training materials" are presentation materials and simulation documents used by users.
[0791] "Uploading means" refers to a method or function that allows a user to send training materials to a server via a terminal.
[0792] "Means for analyzing and generating feedback" refers to the function that enables the server to analyze uploaded materials and provide appropriate comments and advice.
[0793] A "terminal" is a device used by a user, such as a smartphone or tablet.
[0794] A "smart device" is a device for displaying feedback to a user, such as smart glasses or a head-mounted display.
[0795] The "means for performing training simulation" is a function that allows users to simulate actual work through dialogue with virtual customers and scenarios.
[0796] The "means for generating evaluation results and feedback" refers to a method by which the server evaluates the user's performance based on the progress of the training and generates appropriate advice and comments.
[0797] This invention relates to a system for effectively conducting employee training. The system begins when a user enters their profile information and target information and uploads training materials. The uploaded materials are analyzed by a server, and feedback is generated. A terminal transmits and receives data between the user and the server and operates a smart device. The smart device displays the feedback and functions as a means for conducting training simulations.
[0798] Hardware and software used
[0799] Smart devices (smart glasses, head-mounted displays, etc.)
[0800] These devices are used to give users real-time feedback during a simulation.
[0801] Devices (smartphones, tablets)
[0802] Used by users to upload training materials and enter profile and audience information.
[0803] server
[0804] The server installed in the cloud environment is responsible for analyzing materials, generating feedback, sending and receiving data, etc. Specifically, a server with an AI model implemented (e.g., a machine learning model using TensorFlow or PyTorch) is used.
[0805] Cloud services (e.g. AWS, Azure)
[0806] It is used to assist in centralized data management and calculation processing.
[0807] System Operation
[0808] 1. User configuration
[0809] The user uses a terminal to enter their profile information (e.g., job title, experience, etc.) and target information (e.g., age and occupation of the expected customers) according to the training scenario, and then uploads presentation and training materials.
[0810] 2. Server analysis
[0811] The server analyzes the uploaded materials and uses a generative AI model to generate hypothetical customer reactions and questions, taking into account the user's past training and performance data to generate evaluation results and feedback.
[0812] 3. Real-time feedback
[0813] When a user puts on a smart device and starts training, feedback is displayed on the smart device via the terminal from the server. For example, if the user feels tense, advice on how to relax or specific advice on questions from the customer is provided in real time.
[0814] Specific examples
[0815] For example, if a user is conducting training on the theme of "new smartphone products," they can enter "new staff member with two months' experience" as their profile information and set "male in his 30s, technically knowledgeable customer" as the target scenario. Next, they select a training scenario and questions from the virtual customer are displayed.
[0816] Example prompt sentence:
[0817] A customer asks, "What are the camera features on this phone?" Observe how your employees react and ask the question, and provide appropriate feedback. If your employees seem nervous, include tips to help them relax.
[0818] The system allows employees to receive real-time feedback from virtual customers while training on the job, improving their skills and providing peace of mind.
[0819] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0820] Step 1:
[0821] Operation: A user uses a terminal to enter their profile information and target information.
[0822] Input: User profile information (e.g., name, job title, experience), audience information (e.g., age, occupation, areas of interest).
[0823] Output: The entered information is sent from the terminal to the server.
[0824] Step 2:
[0825] How it works: A user uploads presentation or training materials via a terminal.
[0826] Input: Presentation materials (e.g. PowerPoint files or PDFs).
[0827] Output: Uploaded materials are sent from the device to the server and saved.
[0828] Step 3:
[0829] How it works: The server receives the uploaded material and analyzes it.
[0830] Input: Presentation materials, user profile information, audience information.
[0831] Data processing / calculation: Extract text and image information from documents using generative AI models.
[0832] Output: Generate virtual customer reactions and questions based on the extracted information.
[0833] Step 4:
[0834] Operation: The server generates feedback for the virtual customer based on the user's input data and the analysis results.
[0835] Input: User profile information, privacy data, analytics results.
[0836] Data processing / calculation: Generate prompt statements and generate feedback from virtual customers using AI models.
[0837] Output: The generated feedback data is sent from the server to the terminal.
[0838] Step 5:
[0839] Actions: A user puts on a smart device and starts a training simulation.
[0840] Input: Generated feedback data, user voice, and facial expression data.
[0841] Output: Virtual customer feedback is displayed in real time on a smart device.
[0842] Step 6:
[0843] Operation: The server receives the user's voice and facial expression data during the simulation and analyzes it using the emotion engine.
[0844] Input: User's voice and facial expression data.
[0845] Data processing / calculation: Analyzed by the emotion engine to evaluate the user's stress level and tension.
[0846] Output: As a result of the analysis, feedback based on the user's emotional state is generated.
[0847] Step 7:
[0848] How it works: Based on the analysis results of the emotion engine, the server generates advice for the user in real time and sends it to the smart device.
[0849] Input: Analysis results of the emotion engine, feedback from virtual customers.
[0850] Output: Specific advice to the user (e.g., "Take deep breaths to relax") is displayed on the smart device.
[0851] Step 8:
[0852] Action: The user answers questions posed by the virtual customer.
[0853] Input: User response data.
[0854] Output: The user's answer data is sent to the server and analyzed.
[0855] Step 9:
[0856] How it works: The server analyzes the user's answers and generates the following feedback and rating results:
[0857] Input: User response data, past training data.
[0858] Data processing / calculation: Analyzed by AI model.
[0859] Output: As a result of the analysis, feedback and evaluation results from the next virtual customer are generated and sent to the smart device.
[0860] Step 10:
[0861] Operation: After the training is completed, the server generates an evaluation result of the entire training and provides it to the user.
[0862] Input: All training data, emotion data.
[0863] Data processing / calculation: Conduct an overall evaluation of training performance.
[0864] Output: The final evaluation results are displayed on the device along with feedback on future improvements.
[0865] These steps allow employees to virtually experience real-world work scenarios and receive appropriate feedback in real time, resulting in improved work skills and a sense of psychological security.
[0866] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0867] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0868] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0869] [Third embodiment]
[0870] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0871] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0872] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0873] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0874] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0875] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0876] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0877] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0878] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0879] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0880] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0881] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0882] The present invention relates to a system for allowing a user to practice a presentation. Specific embodiments of the present invention will be described below.
[0883] The system begins with the user entering their profile information and audience information and uploading presentation materials. Next, the server analyzes the materials and generates reactions. The device sends and receives data between the user and the server and operates the smart glasses. Finally, the smart glasses display the reactions and simulate the presentation.
[0884] Program processing
[0885] 1. Enter your user profile and audience information
[0886] A user launches the application and enters their profile information and the audience information for the presentation, allowing the system to build an appropriate simulation scenario based on the characteristics of the user and audience.
[0887] 2. Upload your presentation materials
[0888] Users upload presentation materials (e.g., PowerPoint or PDF files) through the application. These materials are sent to the server and analyzed there. Text and image information for each slide is extracted from the analyzed materials.
[0889] 3. Generating Reactions
[0890] The server uses an AI model to generate reactions based on the uploaded presentation materials, user profile information, and audience information, including comments, reactions, and questions from the virtual audience.
[0891] 4. Connecting to smart glasses
[0892] The user wears the smart glasses, and the application establishes a connection with the smart glasses, which are devices that receive data from the server and display reactions.
[0893] 5. Start the presentation simulation
[0894] The user operates the application to enter a presentation simulation mode, and the reactions of the virtual audience are displayed through the smart glasses. The user advances through the slides, presents, and receives feedback from the virtual audience.
[0895] 6. Practice Q&A
[0896] As part of the simulation, users respond to questions posed by a virtual audience, and their responses are sent via smart glasses to a server that analyzes their responses and generates subsequent reactions and feedback.
[0897] Specific examples
[0898] For example, if a user wants to give a presentation on the topic of "agile development," they first enter their profile information, such as their job title and experience, into the application, and set the audience attributes as "engineers, aged 30-40, interested in agile development." They then upload their PowerPoint presentation materials through the application, and the server analyzes them.
[0899] Once a user puts on the smart glasses and enters the simulation mode, a question is displayed from the virtual audience: "What are the main challenges in adopting agile development?" The user responds to this question, and the response is sent to the server for analysis. The server then generates appropriate feedback and additional reactions, which are then displayed to the user.
[0900] In this way, users can repeatedly simulate their presentation with a virtual audience in advance, improving the quality of their presentation and giving them confidence when they actually take the stage.
[0901] The processing flow will be explained below.
[0902] Step 1:
[0903] A user launches the application and logs in. The user enters their profile information and sets up audience information, including name, job title, experience, audience age, occupation, and areas of interest.
[0904] Step 2:
[0905] Users select presentation materials and upload them from their PC or smartphone through the application, which can be in PowerPoint or PDF format.
[0906] Step 3:
[0907] The server receives and stores the uploaded presentation materials. The server analyzes the materials and extracts text and image information for each slide. The analysis uses natural language processing and image recognition technology.
[0908] Step 4:
[0909] The server uses an AI model to generate reactions based on the user's profile information and audience information, specifically determining what comments and questions the virtual audience will make and when they will react.
[0910] Step 5:
[0911] The user wears the smart glasses and starts the simulation mode in the application. The device connects to the smart glasses and establishes data communication with the server.
[0912] Step 6:
[0913] The server confirms the start of the simulation mode and sends the analysis results of the presentation materials and reaction data to the smart glasses. As the user advances through the slides, the smart glasses display the appropriate reactions.
[0914] Step 7:
[0915] The device sends user operation information to the server, including the user's actions as they advance through slides and the progress of the presentation content via voice recognition. The server then uses this information to generate the next reaction or question.
[0916] Step 8:
[0917] The user can see the reactions of the virtual audience (e.g., "Interesting!", questions like, "What are some applications of this technology?") through the smart glasses and respond appropriately. The user's responses are sent to the server via the device.
[0918] Step 9:
[0919] The server analyzes the user's answers and generates follow-up questions and feedback, which are then sent back to the smart glasses and displayed to the user.
[0920] Step 10:
[0921] Once the user has finished the simulation, the server analyzes their overall performance and generates results, including a grade for each slide, the quality and timing of the user's responses, and their overall progress.
[0922] Step 11:
[0923] The user checks the practice results through the application and refers to the feedback provided by the server to improve the presentation materials and delivery method.
[0924] Step 12:
[0925] Users can re-run the simulation and practice as needed, improving the quality of their presentation and building confidence in speaking.
[0926] Example 1
[0927] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0928] Traditional presentation practice methods often require a live audience or offer limited simulations, meaning users miss out on opportunities to effectively improve their presentation skills. Additionally, it's difficult to immediately identify areas for improvement through real-time feedback and Q&A sessions.
[0929] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0930] In this invention, the server includes a means for a user to input profile information and student information, a means for uploading presentation materials, and a means for an information processing device to analyze the uploaded materials and generate responses, thereby enabling the user to receive real-time feedback and questions from a virtual audience and to practice presentations effectively.
[0931] "User" refers to an individual or group who uses the system to practice a presentation.
[0932] "Profile information" refers to detailed information about a user, such as the user's job title, years of experience, and area of expertise.
[0933] "Audience information" refers to information such as the attributes and areas of interest of the audience to whom the presentation is directed.
[0934] "Presentation materials" refers to data such as slides and documents that a user uses to give a presentation.
[0935] An "information processing device" refers to a server or computer used to analyze presentation materials and generate responses.
[0936] "Responses" refers to interactive feedback such as comments, reactions, and questions from the virtual audience.
[0937] A "terminal" is a device (e.g., a smartphone, tablet, or PC) that transmits and receives data between a user and an information processing device and operates a video display device.
[0938] A "video display device" is a display device (e.g., smart glasses, VR headset, etc.) that allows users to visually check reactions.
[0939] "Practice presentation" refers to an activity in which a user prepares for an actual presentation by simulating a presentation in front of a virtual audience.
[0940] The present invention relates to a system for users to practice presentations. The system begins with the user entering profile information and participant information and uploading presentation materials. Next, an information processing device analyzes the materials and generates responses. A terminal transmits and receives data between the user and the information processing device and operates a video display device. Finally, the video display device displays the responses, and the presentation is practiced.
[0941] In a specific embodiment of the system, a program is generated and processing is performed in the following procedure.
[0942] First, the user launches the application, logs in, and enters their profile information (e.g., name, job title, years of experience, etc.). Then, they enter information about the target audience (e.g., industry, age group, areas of interest, etc.). The system then creates an appropriate simulation scenario based on the characteristics of the user and the audience.
[0943] Next, the user uploads presentation materials (e.g., PowerPoint or PDF files) to the application. These materials are sent to an information processing device via the terminal. The information processing device analyzes the uploaded materials and extracts text and image information for each slide. This analysis is performed using text analysis software and image recognition software.
[0944] The information processing device uses a generative AI model to generate responses based on the extracted material content, user profile information, and participant information. The generated responses include comments, reactions, and questions from the virtual audience. For example, a question might be generated: "What are the main challenges in introducing agile development?"
[0945] The user wears a video display device (e.g., smart glasses) and selects the "Video Display Device Connection" option within the application. The terminal establishes a connection with the video display device using a connection method such as Bluetooth or Wi-Fi. The video display device receives the response data from the information processing device and prepares to display it.
[0946] When the user selects "Start Simulation," the presentation begins, and the virtual audience's reactions and questions are displayed on the video display device. The user then verbally responds to the questions, which are then transmitted to the information processing device via the video display device. The information processing device then uses voice analysis software to analyze the user's responses and generate appropriate feedback or additional responses.
[0947] For example, if a user is giving a presentation on the topic of "Agile Development," the steps would be:
[0948] 1. The user enters profile information such as their job title and experience, and sets the audience attributes as "Engineer, 30-40 years old, interested in agile development."
[0949] 2. Upload PowerPoint presentation materials through the application, and the information processing device analyzes the materials.
[0950] 3. Once the user puts on the video display device and starts the simulation mode, a question is displayed from the virtual audience: "What are the main challenges in adopting agile development?"
[0951] 4. The user answers the question, and the answer is sent to the information processing device and analyzed.
[0952] 5. The information processing device generates further appropriate feedback and additional reactions and displays them to the user.
[0953] Examples of prompts include: "The user's profile information is an engineer in his 30s with five years of experience in agile development," "The audience is engineers aged 30-40 who are interested in agile development," and "The theme of the presentation is 'Introduction to agile development and its challenges.' Please generate questions and feedback from the virtual audience."
[0954] In this way, users can repeatedly simulate their presentation with a virtual audience in advance, improving the quality of their presentation and giving them confidence when they actually take the stage.
[0955] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0956] Step 1:
[0957] Entering user profile and audience information
[0958] Input: User profile information (e.g., name, job title, years of experience) and audience information (e.g., industry, age group, interests)
[0959] Processing: The user starts the application and authenticates by entering their user ID and password on the login screen. After logging in, they enter their profile information and audience information in the form that appears. The server receives this information and stores it in a database. It also designs an appropriate simulation scenario based on the user and audience information.
[0960] Output: User and audience information stored in a database, and simulation scenario designs
[0961] Step 2:
[0962] Upload presentation materials
[0963] Input: Presentation materials (e.g. PowerPoint files, PDF files)
[0964] Processing: The user clicks the "File Upload" button in the application and selects presentation materials from a selection screen. The device reads the selected file and sends it to the server. The server analyzes the received file and extracts text and image data for each slide. This analysis is performed using text analysis software and image recognition software.
[0965] Output: Parsed text and image data for each slide
[0966] Step 3:
[0967] Creating a reaction
[0968] Input: Parsed presentation materials, user profile information, audience information
[0969] Processing: The server uses the generative AI model to generate comments, reactions, and questions from the virtual audience based on the analyzed content and the input user profile and audience information. For example, it dynamically generates questions and reactions related to the content of the slides.
[0970] Output: Generated virtual audience reaction data (comments, questions, responses, etc.)
[0971] Step 4:
[0972] Connecting to smart glasses
[0973] Input: Wi-Fi and Bluetooth information for connection
[0974] Processing: The user wears a video display device (e.g., smart glasses) and selects the "Connect to video display device" option within the application. The device establishes a connection with the video display device using the selected connection method. The device confirms that a connection with the video display device has been established.
[0975] Output: Established connection status with video display device
[0976] Step 5:
[0977] Start of presentation simulation
[0978] Input: Command to start the simulation
[0979] Processing: The user clicks the "Start Simulation" button in the application. The terminal sends the presentation materials and the generated reaction data to the video display device and prepares to start the simulation.
[0980] Output: Presentation materials and reaction data sent to a video display device
[0981] Step 6:
[0982] Q&A practice
[0983] Input: Virtual audience question displayed on a video display device
[0984] Processing: The video display device displays the question from the virtual audience. The user responds verbally to the question, and the response is captured using the video display device's audio input function and sent to the server. The server analyzes the audio data and generates feedback and additional reactions based on the user's response, which are then sent back to the video display device.
[0985] Output: Parsed user response data, generated new feedback and reactions
[0986] (Application example 1)
[0987] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0988] Conventional presentation practice systems primarily require users to practice on their own, lacking the ability to provide real-time feedback or reflect the reactions of a virtual audience. Furthermore, systems for effectively training salespeople in brick-and-mortar stores on product explanations are also inadequate, and there is a need for a system that can respond immediately to customer questions and reactions. The present invention aims to solve these problems and improve the quality of presentation practice and sales training.
[0989] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0990] In this invention, the server includes: a means for a user to input profile information and audience information; a means for uploading presentation materials; a means for the server to analyze the uploaded materials and generate reactions; a means for a terminal to transmit and receive data between the user and the server and operate the smart glasses; a means for the smart glasses to display reactions and perform a presentation simulation; a means for a salesperson to input profile information and customer information and upload product explanation materials; a means for the server to analyze the product explanation materials and generate questions and reactions of virtual customers; and a means for the smart glasses to display the responses and questions of the virtual customers in real time and perform a sales training simulation. This allows the user to make a presentation while receiving reactions from the virtual audience, and enables the salesperson to receive practical training through interactions with virtual customers.
[0991] "User profile information" is information that includes personal information, job information, experience, and the like about the user.
[0992] "Audience information" is attribute information about the audience to whom the presentation is given.
[0993] "Presentation materials" are presentation materials provided in the form of slides or documents.
[0994] A "server" is a computer system that receives data from users, analyzes it, and generates appropriate reactions.
[0995] "Means for generating reactions" is a function that uses an AI model to generate virtual comments and questions based on presentation materials and user information.
[0996] A "terminal" is a device that transmits and receives data between a user and a server.
[0997] "Smart glasses" are eyeglass-type devices that utilize augmented reality technology to display information to users in real time.
[0998] The "means for simulating a presentation" is a function that allows a user to use smart glasses to give a presentation while interacting with a virtual audience.
[0999] "Salesperson profile information" is information including job information, products in charge, experience, etc., related to the salesperson.
[1000] "Customer information" is attribute information about customers who are the target of sales.
[1001] "Product description materials" are materials used to explain the characteristics and functions of a product.
[1002] "Virtual Customer" means a fictitious customer generated using AI technology.
[1003] "Means for generating questions and reactions" refers to a function that uses an AI model to generate virtual customer questions and reactions based on salesperson profile information and product explanation materials.
[1004] The "means for conducting sales training simulations" is a function that allows salespeople to use smart glasses to conduct practical training through interactions with virtual customers.
[1005] The present invention relates to a comprehensive system for users to practice presentations and sales training. The system begins when the user inputs profile information and audience information and uploads presentation or product explanation materials. The server analyzes the uploaded materials and generates reactions. The terminal transmits and receives data between the user and the server and operates the smart glasses. The smart glasses display reactions and perform a presentation simulation or sales training simulation.
[1006] Specific steps
[1007] 1. Enter your profile and audience information
[1008] A user launches the application and enters their profile information (e.g., job title, experience, products) and audience or customer information for the presentation or sale (e.g., audience age range, interests, skill level).
[1009] 2. Upload presentation or product description materials
[1010] Users upload presentation materials (e.g., PowerPoint or PDF files) or product descriptions through the application. These materials are then sent to the server, where they are analyzed. During the analysis process, important text and image information is extracted from the materials.
[1011] 3. Generating Reactions
[1012] The server uses a generative AI model to generate reactions based on the uploaded materials and user and audience information, including comments, reactions, and questions from the virtual audience or virtual customers.
[1013] 4. Connecting to smart glasses
[1014] The user wears the smart glasses, and the application establishes a connection with the smart glasses, which receives data from the server and displays reactions.
[1015] 5. Start a presentation simulation or sales training
[1016] The user interacts with the application and enters the simulation mode, where the smart glasses display the reactions of a virtual audience and virtual customers, providing virtual feedback as the user progresses through slides and product presentations.
[1017] 6. Practice Q&A
[1018] As part of the simulation, users respond to questions posed by a virtual audience or customer, and their answers are sent through the smart glasses to a server that analyzes their responses and generates subsequent reactions and feedback.
[1019] Hardware and software used
[1020] Hardware used: Smart glasses (e.g., Google Glass, Vuzix Blade)
[1021] Software used: Flask (server-side framework), generative AI model (reaction generation), data analysis system
[1022] Specific examples
[1023] For example, when training a salesperson to explain a new smartphone, the following prompts are fed into the generative AI model:
[1024] "User: Salesperson, Experience: 3 years, Product: New smartphone, Target Audience: 20-30 years old, Tech-savvy"
[1025] Feedback from the virtual customer is displayed on the smart glasses with questions such as:
[1026] "How does this smartphone's camera performance compare to other products?"
[1027] In this way, users can improve the quality of their presentations and sales training by repeating virtual simulations.
[1028] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1029] Step 1:
[1030] A user launches the application and enters profile and audience information.
[1031] This information includes the user's job title, experience, products they work on, and target audience demographics (age range, interests, skill level, etc.).
[1032] The server receives these input data and stores them in a database.
[1033] Step 2:
[1034] A user uploads a presentation or product description.
[1035] These materials are provided in formats such as PDF and PowerPoint files.
[1036] The server receives the material and uses text analysis software (e.g., Apache Tika) to extract text and image information from the material.
[1037] The extracted information is stored in a database on the server.
[1038] Step 3:
[1039] The server generates reactions using a generative AI model based on uploaded materials, user information, and audience information.
[1040] This process generates questions and comments from the virtual audience or customer based on the content of the material.
[1041] The generated reactions are temporarily stored on the server.
[1042] Step 4:
[1043] The user puts on smart glasses (e.g., Google Glass, Vuzix Blade), and the application establishes a connection with the smart glasses.
[1044] The server transmits the generated reaction data to the smart glasses.
[1045] The smart glasses display the received data in real time in the user's field of vision.
[1046] Step 5:
[1047] The user operates the application and starts a presentation simulation or a sales training simulation.
[1048] Smart glasses display the reactions of virtual audiences and virtual customers.
[1049] Users receive real-time feedback as they progress through their presentations and product explanations.
[1050] Step 6:
[1051] As part of a presentation simulation or sales training, a user answers questions posed by a virtual audience or virtual customers.
[1052] The user's answers are sent to the server via the smart glasses.
[1053] The server analyzes the user's answers and generates the next reaction or feedback.
[1054] The generated feedback is then sent back to the smart glasses and displayed to the user.
[1055] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1056] The present invention relates to a system for providing more advanced feedback by combining a system for users to practice presentations with an emotion engine that recognizes the emotions of the users. Specific embodiments of the present invention will be described below.
[1057] System configuration
[1058] Entering user profile and audience information
[1059] A user launches the application and logs in. The user enters their profile information (e.g., name, job title, experience) and also sets audience information (e.g., age, occupation, areas of interest) for the presentation.
[1060] Upload presentation materials
[1061] Users upload presentation materials (e.g., PowerPoint or PDF files) from their PC or smartphone through the application. These materials are then sent to the server, where they are analyzed. The analysis includes extracting text and image information for each slide.
[1062] Creating a reaction
[1063] The server uses an AI model to generate reactions based on the uploaded presentation materials, user profile information, and audience information. Specifically, it determines what comments and questions the virtual audience will make, as well as the timing of their reactions.
[1064] Linking the Emotion Engine
[1065] When a user puts on the smart glasses and starts the simulation mode in the application, the emotion engine analyzes the user's facial expressions and tone of voice to recognize the user's emotional state in real time. For example, if the user is nervous, it will provide advice on how to relax.
[1066] Start of presentation simulation
[1067] The user interacts with the application to enter a presentation simulation mode, where the smart glasses display the reactions of the virtual audience. The user advances through the slides, presenting, and receives feedback from the virtual audience.
[1068] Q&A practice
[1069] As part of the simulation, users respond to questions posed by a virtual audience, and their responses are sent via smart glasses to a server that analyzes their responses and generates subsequent reactions and feedback.
[1070] Emotional data feedback and adjustment
[1071] The emotion engine analyzes the user's emotional data as they make their presentation and generates feedback based on their emotional state. For example, if the user feels anxious, the engine will advise them to explain the key points more simply.
[1072] Specific examples
[1073] For example, if a user wants to give a presentation on the theme of "innovative technology," the user first enters their profile information (e.g., experience as an engineer, field of expertise) into the application, and sets the audience attributes as "corporate technical personnel, aged 30-40, interested in technological innovation." Then, the user uploads presentation materials (PowerPoint files).
[1074] When a user puts on the smart glasses and starts the simulation mode, the emotion engine analyzes the user's facial expressions and tone of voice, and displays feedback such as, "You're feeling a little nervous on this slide," while the smart glasses also display advice such as, "Take a deep breath to relax."
[1075] In the presentation simulation, a virtual audience asks the user, "Please tell us some specific examples of how your technology can be applied," and the user responds. The response is sent to the server, where it is analyzed and the results are provided as feedback. For example, the user might receive feedback such as, "The explanation of the application example was easy to understand. However, I would like more specific examples."
[1076] In this way, by linking with the emotion engine, users can receive more detailed and useful feedback, significantly improving the quality of their presentations and giving them more confidence when they take the stage.
[1077] The processing flow will be explained below.
[1078] Step 1:
[1079] A user launches the application and logs in. The user enters their profile information (name, job title, experience, etc.) and also sets audience information (age, occupation, areas of interest, etc.) for the presentation.
[1080] Step 2:
[1081] Users select and upload presentation materials (PowerPoint or PDF files) from their PC or smartphone through the application, which are then sent to the server.
[1082] Step 3:
[1083] The server receives and analyzes the uploaded presentation materials, extracting text and image information from each slide, and using natural language processing (NLP) and image recognition technology to identify key points and locations within the materials.
[1084] Step 4:
[1085] The server uses an AI model to generate reactions based on the user's profile information and audience information. Specifically, it determines what comments and questions the virtual audience will make, as well as the timing of their reactions. The emotion engine also prepares to acquire the user's emotional data.
[1086] Step 5:
[1087] The user puts on the smart glasses and starts the simulation mode in the application. The device connects to the smart glasses and performs data communication with the server. The smart glasses receive reaction data from the server and prepare to display it.
[1088] Step 6:
[1089] The server confirms the start of the simulation mode and sends the analysis results of the presentation materials and reaction data to the smart glasses. As the user advances through the slides, the smart glasses display appropriate reactions. The emotion engine also analyzes the user's facial expressions and tone of voice in real time.
[1090] Step 7:
[1091] The device sends the user's operational information (slide progress and comments) to the server. The server uses this information to generate the next reaction or question. At the same time, the emotion engine analyzes the user's mental state based on the acquired emotional data and creates corresponding feedback.
[1092] Step 8:
[1093] The user can see the reactions of the virtual audience (e.g., "Interesting!", questions like, "What are some applications of this technology?") through the smart glasses and respond appropriately. The user's answers are sent to the server via the device.
[1094] Step 9:
[1095] The server analyzes the user's answers and emotional data to generate follow-up questions and feedback. This information is then sent back to the smart glasses and displayed to the user. For example, if the user is nervous, the system will suggest, "Try taking a deep breath to relax."
[1096] Step 10:
[1097] Once the user has finished the simulation, the server analyzes the overall performance data (e.g., the evaluation of each slide, the quality and timing of the user's responses, overall progress, and emotional data) and generates an evaluation result.
[1098] Step 11:
[1099] The user checks the practice results through the application and refers to the feedback provided by the server to improve the presentation materials and delivery method.
[1100] Step 12:
[1101] Users can re-run the simulation and practice as needed. By linking the emotion engine and reaction generation system, users can significantly improve the quality of their presentations, giving them the confidence to take the stage in real life.
[1102] Example 2
[1103] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1104] Previous presentation practice systems lacked the ability to analyze users' emotional states in real time and provide feedback based on that analysis. As a result, users were unable to adequately manage their own emotions and it was difficult to improve the quality of their presentations. It was also difficult to obtain detailed feedback from a virtual audience. This could lead to users losing confidence in their actual presentations and potentially resulting in poor performance.
[1105] The specification processing by the specification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for analyzing uploaded materials and generating reactions of the virtual audience, means for generating evaluation results and feedback based on the progress of the presentation and the user's emotional data, and an emotion analysis engine for analyzing the user's facial expressions and voice in real time and generating feedback. This allows the user to receive feedback according to their own emotional state in real time, improving the quality of their presentation and giving them confidence in their actual presentation.
[1106] A "user" is an individual who uses the system to practice a presentation through the application.
[1107] "Profile information" is information about personal attributes including the user's name, job title, experience, etc.
[1108] "Audience information" refers to attribute information such as age, occupation, and areas of interest of the virtual audience to whom the presentation is directed.
[1109] "Presentation materials" are presentation content such as slides and PDF files uploaded by users.
[1110] The "server" is a computer system that performs processes such as analyzing uploaded presentation materials, generating reactions from a virtual audience, and analyzing user emotional data.
[1111] A "terminal" is a device that allows a user to send and receive data to and from a server and operate a wearable device.
[1112] A "wearable device" is a wearable device (e.g., smart glasses) that can be worn by the user to display reactions and perform emotional analysis.
[1113] "Reactions" are responses such as comments, feedback, and questions from the virtual audience.
[1114] An "emotion analysis engine" is software that analyzes a user's facial expressions, voice, etc. in real time, recognizes the user's emotional state, and generates feedback.
[1115] "Feedback" is response information that includes evaluations and advice regarding the user's behavior, emotional state, etc.
[1116] A "virtual audience" is an imagined audience that makes comments and asks questions about a user's presentation generated by the system.
[1117] "Real-time" refers to the analysis and feedback process occurring instantly while the presentation is in progress.
[1118] The present invention provides a system for providing more advanced feedback by combining a system for users to practice presentations with an emotion analysis engine. Specific embodiments of this system are described below.
[1119] First, the user launches a dedicated application and logs in. After logging in, the user enters profile information (e.g., name, job title, experience) and audience information for the presentation (e.g., age, occupation, areas of interest). Next, the user uploads presentation materials (e.g., PowerPoint or PDF files). These materials are sent to the server, where they are analyzed. This analysis includes extracting text and image information for each slide.
[1120] The server uses a generative AI model to generate reactions from a virtual audience based on the uploaded presentation materials and user input. These reactions include comments and questions, and the timing of their occurrence is also determined. For example, the server inputs a prompt statement such as "Generate questions that the audience would ask during a presentation on technological innovations" into the generative AI model, and uses the responses.
[1121] Next, the user puts on the wearable device (e.g., smart glasses) and starts the application's simulation mode. In this mode, the emotion analysis engine analyzes the user's facial expressions and voice tone in real time. Based on the analysis results, real-time feedback is provided according to the user's emotional state (e.g., nervousness, anxiety, confidence). For example, the feedback "This slide makes you a little nervous" is displayed along with advice such as "Take a deep breath to relax."
[1122] During the simulation, the user gives a presentation and receives reactions from the virtual audience. When the user responds to the comments and questions of the virtual audience, the response is sent to the server via the wearable device. The server analyzes the user's response and generates the next reaction or feedback based on it. For example, the server may provide feedback such as, "The explanation of the application example was easy to understand, but please add more concrete examples."
[1123] This system allows users to receive detailed and useful feedback based on their emotional state in real time, which will improve the quality of their presentations and increase their confidence in their presentations. This confidence is expected to contribute to the success of their presentations.
[1124] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1125] Step 1:
[1126] The user launches the application and logs in.
[1127] Specific behavior:
[1128] The user taps the application icon to launch it, and enters their user ID and password on the login screen that appears. If they are successfully authenticated, the user profile screen is displayed.
[1129] Input and Output:
[1130] Input: User ID, Password
[1131] Data processing / calculation: The authentication server checks the ID and password
[1132] Output: Notification of successful or failed login
[1133] Step 2:
[1134] The user enters profile and audience information.
[1135] Specific behavior:
[1136] The user enters profile information (name, job title, experience) and audience information (age, occupation, areas of interest) into the form and clicks the "Next" button.
[1137] Input and Output:
[1138] Input: Profile information, audience information
[1139] Data processing / calculation: Input data is saved on the server as a user profile
[1140] Output: Confirmation of input completion and moving to the next step
[1141] Step 3:
[1142] A user uploads a presentation.
[1143] Specific behavior:
[1144] The user clicks the "Upload Document" button and selects a presentation file (e.g., PowerPoint, PDF) from the file selection dialog. The selected file is sent to the server.
[1145] Input and Output:
[1146] Input: Selected presentation file
[1147] Data processing / calculation: File upload and file storage on the server
[1148] Output: File upload completion notification
[1149] Step 4:
[1150] The server analyzes the uploaded presentation materials.
[1151] Specific behavior:
[1152] The server analyzes the received presentation file and extracts the text and image information of each slide, which is then stored in a database.
[1153] Input and Output:
[1154] Input: Uploaded presentation file
[1155] Data processing / calculation: Extraction of text and image information, storage in database
[1156] Output: Analysis data saved
[1157] Step 5:
[1158] The server uses a generative AI model to generate reactions from the virtual audience.
[1159] Specific behavior:
[1160] Based on the extracted data and input information, the server sends prompts to the generative AI model to generate reactions from the virtual audience, such as "Generate questions that the audience would ask during a presentation about technological innovations."
[1161] Input and Output:
[1162] Input: Analytics data, profile information, audience information
[1163] Data processing / calculation: Reaction generation using generative AI models
[1164] Output: Virtual audience comments, feedback, and questions
[1165] Step 6:
[1166] The user puts on the wearable device and starts the simulation mode.
[1167] Specific behavior:
[1168] The user puts on the smart glasses and clicks the "Start Simulation Mode" button in the application, which activates the wearable device and displays the reactions of the virtual audience.
[1169] Input and Output:
[1170] Input: Instruction to start the wearable device
[1171] Data processing / calculation: Connection and data transmission between terminal and wearable device
[1172] Output: Start of simulation mode, display of virtual audience
[1173] Step 7:
[1174] The emotion analysis engine analyzes the user's facial expressions and tone of voice in real time.
[1175] Specific behavior:
[1176] The emotion analysis engine uses the smart glasses' built-in camera and microphone to capture and analyze the user's facial expressions and tone of voice, and the analysis data is fed back to the user in real time.
[1177] Input and Output:
[1178] Input: User's facial expression data, voice data
[1179] Data processing / calculation: Data analysis and feedback generation using a sentiment analysis engine
[1180] Output: Real-time emotional feedback
[1181] Step 8:
[1182] Get reactions from a virtual audience as you deliver your presentation.
[1183] Specific behavior:
[1184] As users advance through the slides, the virtual audience displays comments and questions at designated times, to which users can respond.
[1185] Input and Output:
[1186] Input: User's slide progress data
[1187] Data processing / calculation: Reaction display timing control, reaction display
[1188] Output: Display of virtual audience comments and questions
[1189] Step 9:
[1190] Users practice their answers to questions from a virtual audience.
[1191] Specific behavior:
[1192] The user answers questions posed by the virtual audience, and the answers are sent to a server via smart glasses, where they are analyzed and used to generate the next reaction or feedback.
[1193] Input and Output:
[1194] Input: User's response audio data
[1195] Data processing / calculation: Analysis of voice data by server, generation of next reaction
[1196] Output: Analysis results and feedback
[1197] Step 10:
[1198] The emotion analysis engine provides feedback on the user's emotion data during the presentation.
[1199] Specific behavior:
[1200] The emotion analysis engine analyzes the data as it goes through the presentation and provides emotional feedback to the user, such as "I'm a little nervous about the next slide," allowing the user to make appropriate adjustments (e.g., take a deep breath, slow down speaking speed).
[1201] Input and Output:
[1202] Input: Real-time facial expression data, voice data
[1203] Data processing / calculation: Analysis by emotion analysis engine and feedback generation
[1204] Output: Real-time feedback according to emotional state
[1205] (Application example 2)
[1206] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1207] In employee training, it is difficult to improve skills by simulating actual work situations and customer interactions and receiving appropriate feedback in real time. There is also a lack of technology that can analyze employees' emotional states and provide specific advice accordingly. Therefore, there is a need for a system that can effectively improve employees' work performance before they actually begin work.
[1208] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1209] In this invention, the server includes means for a user to input profile information and subject information, means for uploading training materials, means for the server to analyze the uploaded materials and generate feedback, means for a terminal to transmit and receive data between the user and the server and operate the smart device, means for the smart device to display the feedback and perform a training simulation, and means for the server to generate evaluation results and feedback based on the progress of the training. This allows employees to receive feedback from virtual customers in real time while undergoing job training, and can provide specific advice according to the employee's emotional state, thereby improving their skills and providing a sense of security.
[1210] "Profile information" is basic personal data such as the user's name, job title, experience, etc.
[1211] "Target information" is data such as the age, occupation, and areas of interest of a hypothetical customer or audience.
[1212] "Training materials" are presentation materials and simulation documents used by users.
[1213] "Uploading means" refers to a method or function that allows a user to send training materials to a server via a terminal.
[1214] "Means for analyzing and generating feedback" refers to the function that enables the server to analyze uploaded materials and provide appropriate comments and advice.
[1215] A "terminal" is a device used by a user, such as a smartphone or tablet.
[1216] A "smart device" is a device for displaying feedback to a user, such as smart glasses or a head-mounted display.
[1217] The "means for performing training simulation" is a function that allows users to simulate actual work through dialogue with virtual customers and scenarios.
[1218] The "means for generating evaluation results and feedback" refers to a method by which the server evaluates the user's performance based on the progress of the training and generates appropriate advice and comments.
[1219] This invention relates to a system for effectively conducting employee training. The system begins when a user enters their profile information and target information and uploads training materials. The uploaded materials are analyzed by a server, and feedback is generated. A terminal transmits and receives data between the user and the server and operates a smart device. The smart device displays the feedback and functions as a means for conducting training simulations.
[1220] Hardware and software used
[1221] Smart devices (smart glasses, head-mounted displays, etc.)
[1222] These devices are used to give users real-time feedback during a simulation.
[1223] Devices (smartphones, tablets)
[1224] Used by users to upload training materials and enter profile and audience information.
[1225] server
[1226] The server installed in the cloud environment is responsible for analyzing materials, generating feedback, sending and receiving data, etc. Specifically, a server with an AI model implemented (e.g., a machine learning model using TensorFlow or PyTorch) is used.
[1227] Cloud services (e.g. AWS, Azure)
[1228] It is used to assist in centralized data management and calculation processing.
[1229] System Operation
[1230] 1. User configuration
[1231] The user uses a terminal to enter their profile information (e.g., job title, experience, etc.) and target information (e.g., age and occupation of the expected customers) according to the training scenario, and then uploads presentation and training materials.
[1232] 2. Server analysis
[1233] The server analyzes the uploaded materials and uses a generative AI model to generate hypothetical customer reactions and questions, taking into account the user's past training and performance data to generate evaluation results and feedback.
[1234] 3. Real-time feedback
[1235] When a user puts on a smart device and starts training, feedback is displayed on the smart device via the terminal from the server. For example, if the user feels tense, advice on how to relax or specific advice on questions from the customer is provided in real time.
[1236] Specific examples
[1237] For example, if a user is conducting training on the theme of "new smartphone products," they can enter "new staff member with two months' experience" as their profile information and set "male in his 30s, technically knowledgeable customer" as the target scenario. Next, they select a training scenario and questions from the virtual customer are displayed.
[1238] Example prompt sentence:
[1239] A customer asks, "What are the camera features on this phone?" Observe how your employees react and ask the question, and provide appropriate feedback. If your employees seem nervous, include tips to help them relax.
[1240] The system allows employees to receive real-time feedback from virtual customers while training on the job, improving their skills and providing peace of mind.
[1241] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1242] Step 1:
[1243] Operation: A user uses a terminal to enter their profile information and target information.
[1244] Input: User profile information (e.g., name, job title, experience), audience information (e.g., age, occupation, areas of interest).
[1245] Output: The entered information is sent from the terminal to the server.
[1246] Step 2:
[1247] How it works: A user uploads presentation or training materials via a terminal.
[1248] Input: Presentation materials (e.g. PowerPoint files or PDFs).
[1249] Output: Uploaded materials are sent from the device to the server and saved.
[1250] Step 3:
[1251] How it works: The server receives the uploaded material and analyzes it.
[1252] Input: Presentation materials, user profile information, audience information.
[1253] Data processing / calculation: Extract text and image information from documents using generative AI models.
[1254] Output: Generate virtual customer reactions and questions based on the extracted information.
[1255] Step 4:
[1256] Operation: The server generates feedback for the virtual customer based on the user's input data and the analysis results.
[1257] Input: User profile information, privacy data, analytics results.
[1258] Data processing / calculation: Generate prompt statements and generate feedback from virtual customers using AI models.
[1259] Output: The generated feedback data is sent from the server to the terminal.
[1260] Step 5:
[1261] Actions: A user puts on a smart device and starts a training simulation.
[1262] Input: Generated feedback data, user voice, and facial expression data.
[1263] Output: Virtual customer feedback is displayed in real time on a smart device.
[1264] Step 6:
[1265] Operation: The server receives the user's voice and facial expression data during the simulation and analyzes it using the emotion engine.
[1266] Input: User's voice and facial expression data.
[1267] Data processing / calculation: Analyzed by the emotion engine to evaluate the user's stress level and tension.
[1268] Output: As a result of the analysis, feedback based on the user's emotional state is generated.
[1269] Step 7:
[1270] How it works: Based on the analysis results of the emotion engine, the server generates advice for the user in real time and sends it to the smart device.
[1271] Input: Analysis results of the emotion engine, feedback from virtual customers.
[1272] Output: Specific advice to the user (e.g., "Take deep breaths to relax") is displayed on the smart device.
[1273] Step 8:
[1274] Action: The user answers questions posed by the virtual customer.
[1275] Input: User response data.
[1276] Output: The user's answer data is sent to the server and analyzed.
[1277] Step 9:
[1278] How it works: The server analyzes the user's answers and generates the following feedback and rating results:
[1279] Input: User response data, past training data.
[1280] Data processing / calculation: Analyzed by AI model.
[1281] Output: As a result of the analysis, feedback and evaluation results from the next virtual customer are generated and sent to the smart device.
[1282] Step 10:
[1283] Operation: After the training is completed, the server generates an evaluation result of the entire training and provides it to the user.
[1284] Input: All training data, emotion data.
[1285] Data processing / calculation: Conduct an overall evaluation of training performance.
[1286] Output: The final evaluation results are displayed on the device along with feedback on future improvements.
[1287] These steps allow employees to virtually experience real-world work scenarios and receive appropriate feedback in real time, resulting in improved work skills and a sense of psychological security.
[1288] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1289] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1290] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1291] [Fourth embodiment]
[1292] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1293] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1294] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1295] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1296] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1297] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1298] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1299] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1300] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1301] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1302] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1303] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1304] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1305] The present invention relates to a system for allowing a user to practice a presentation. Specific embodiments of the present invention will be described below.
[1306] The system begins with the user entering their profile information and audience information and uploading presentation materials. Next, the server analyzes the materials and generates reactions. The device sends and receives data between the user and the server and operates the smart glasses. Finally, the smart glasses display the reactions and simulate the presentation.
[1307] Program processing
[1308] 1. Enter your user profile and audience information
[1309] A user launches the application and enters their profile information and the audience information for the presentation, allowing the system to build an appropriate simulation scenario based on the characteristics of the user and audience.
[1310] 2. Upload your presentation materials
[1311] Users upload presentation materials (e.g., PowerPoint or PDF files) through the application. These materials are sent to the server and analyzed there. Text and image information for each slide is extracted from the analyzed materials.
[1312] 3. Generating Reactions
[1313] The server uses an AI model to generate reactions based on the uploaded presentation materials, user profile information, and audience information, including comments, reactions, and questions from the virtual audience.
[1314] 4. Connecting to smart glasses
[1315] The user wears the smart glasses, and the application establishes a connection with the smart glasses, which are devices that receive data from the server and display reactions.
[1316] 5. Start the presentation simulation
[1317] The user operates the application to enter a presentation simulation mode, and the reactions of the virtual audience are displayed through the smart glasses. The user advances through the slides, presents, and receives feedback from the virtual audience.
[1318] 6. Practice Q&A
[1319] As part of the simulation, users respond to questions posed by a virtual audience, and their responses are sent via smart glasses to a server that analyzes their responses and generates subsequent reactions and feedback.
[1320] Specific examples
[1321] For example, if a user wants to give a presentation on the topic of "agile development," they first enter their profile information, such as their job title and experience, into the application, and set the audience attributes as "engineers, aged 30-40, interested in agile development." They then upload their PowerPoint presentation materials through the application, and the server analyzes them.
[1322] Once a user puts on the smart glasses and enters the simulation mode, a question is displayed from the virtual audience: "What are the main challenges in adopting agile development?" The user responds to this question, and the response is sent to the server for analysis. The server then generates appropriate feedback and additional reactions, which are then displayed to the user.
[1323] In this way, users can repeatedly simulate their presentation with a virtual audience in advance, improving the quality of their presentation and giving them confidence when they actually take the stage.
[1324] The processing flow will be explained below.
[1325] Step 1:
[1326] A user launches the application and logs in. The user enters their profile information and sets up audience information, including name, job title, experience, audience age, occupation, and areas of interest.
[1327] Step 2:
[1328] Users select presentation materials and upload them from their PC or smartphone through the application, which can be in PowerPoint or PDF format.
[1329] Step 3:
[1330] The server receives and stores the uploaded presentation materials. The server analyzes the materials and extracts text and image information for each slide. The analysis uses natural language processing and image recognition technology.
[1331] Step 4:
[1332] The server uses an AI model to generate reactions based on the user's profile information and audience information, specifically determining what comments and questions the virtual audience will make and when they will react.
[1333] Step 5:
[1334] The user wears the smart glasses and starts the simulation mode in the application. The device connects to the smart glasses and establishes data communication with the server.
[1335] Step 6:
[1336] The server confirms the start of the simulation mode and sends the analysis results of the presentation materials and reaction data to the smart glasses. As the user advances through the slides, the smart glasses display the appropriate reactions.
[1337] Step 7:
[1338] The device sends user operation information to the server, including the user's actions as they advance through the slides and the progress of the presentation using voice recognition. The server then uses this information to generate the next reaction or question.
[1339] Step 8:
[1340] The user can see the reactions of the virtual audience (e.g., "Interesting!", questions like, "What are some applications of this technology?") through the smart glasses and respond appropriately. The user's responses are sent to the server via the device.
[1341] Step 9:
[1342] The server analyzes the user's answers and generates follow-up questions and feedback, which are then sent back to the smart glasses and displayed to the user.
[1343] Step 10:
[1344] Once the user has finished the simulation, the server analyzes their overall performance and generates results, including a grade for each slide, the quality and timing of the user's responses, and their overall progress.
[1345] Step 11:
[1346] The user checks the practice results through the application and refers to the feedback provided by the server to improve the presentation materials and delivery method.
[1347] Step 12:
[1348] Users can re-run the simulation and practice as needed, improving the quality of their presentation and building confidence in speaking.
[1349] Example 1
[1350] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1351] Traditional presentation practice methods often require a live audience or offer limited simulations, meaning users miss out on opportunities to effectively improve their presentation skills. Additionally, it's difficult to immediately identify areas for improvement through real-time feedback and Q&A sessions.
[1352] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1353] In this invention, the server includes a means for a user to input profile information and student information, a means for uploading presentation materials, and a means for an information processing device to analyze the uploaded materials and generate responses, thereby enabling the user to receive real-time feedback and questions from a virtual audience and to practice presentations effectively.
[1354] "User" refers to an individual or group who uses the system to practice a presentation.
[1355] "Profile information" refers to detailed information about a user, such as the user's job title, years of experience, and area of expertise.
[1356] "Audience information" refers to information such as the attributes and areas of interest of the audience to whom the presentation is directed.
[1357] "Presentation materials" refers to data such as slides and documents that a user uses to give a presentation.
[1358] An "information processing device" refers to a server or computer used to analyze presentation materials and generate responses.
[1359] "Responses" refers to interactive feedback such as comments, reactions, and questions from the virtual audience.
[1360] A "terminal" is a device (e.g., a smartphone, tablet, or PC) that transmits and receives data between a user and an information processing device and operates a video display device.
[1361] A "video display device" is a display device (e.g., smart glasses, VR headset, etc.) that allows users to visually check reactions.
[1362] "Practice presentation" refers to an activity in which a user prepares for an actual presentation by simulating a presentation in front of a virtual audience.
[1363] The present invention relates to a system for users to practice presentations. The system begins with the user entering profile information and participant information and uploading presentation materials. Next, an information processing device analyzes the materials and generates responses. A terminal transmits and receives data between the user and the information processing device and operates a video display device. Finally, the video display device displays the responses, and the presentation is practiced.
[1364] In a specific embodiment of the system, a program is generated and processing is performed in the following procedure.
[1365] First, the user launches the application, logs in, and enters their profile information (e.g., name, job title, years of experience, etc.). Then, they enter information about the target audience (e.g., industry, age group, areas of interest, etc.). The system then creates an appropriate simulation scenario based on the characteristics of the user and the audience.
[1366] Next, the user uploads presentation materials (e.g., PowerPoint or PDF files) to the application. These materials are sent to an information processing device via the terminal. The information processing device analyzes the uploaded materials and extracts text and image information for each slide. This analysis is performed using text analysis software and image recognition software.
[1367] The information processing device uses a generative AI model to generate responses based on the extracted material content, user profile information, and participant information. The generated responses include comments, reactions, and questions from the virtual audience. For example, a question might be generated: "What are the main challenges in introducing agile development?"
[1368] The user wears a video display device (e.g., smart glasses) and selects the "Video Display Device Connection" option within the application. The terminal establishes a connection with the video display device using a connection method such as Bluetooth or Wi-Fi. The video display device receives the response data from the information processing device and prepares to display it.
[1369] When the user selects "Start Simulation," the presentation begins, and the virtual audience's reactions and questions are displayed on the video display device. The user then verbally responds to the questions, which are then transmitted to the information processing device via the video display device. The information processing device then uses voice analysis software to analyze the user's responses and generate appropriate feedback or additional responses.
[1370] For example, if a user is giving a presentation on the topic of "Agile Development," the steps would be:
[1371] 1. The user enters profile information such as their job title and experience, and sets the audience attributes as "Engineer, 30-40 years old, interested in agile development."
[1372] 2. Upload PowerPoint presentation materials through the application, and the information processing device analyzes the materials.
[1373] 3. Once the user puts on the video display device and starts the simulation mode, a question is displayed from the virtual audience: "What are the main challenges in adopting agile development?"
[1374] 4. The user answers the question, and the answer is sent to the information processing device and analyzed.
[1375] 5. The information processing device generates further appropriate feedback and additional reactions and displays them to the user.
[1376] Examples of prompts include: "The user's profile information is an engineer in his 30s with five years of experience in agile development," "The audience is engineers aged 30-40 who are interested in agile development," and "The theme of the presentation is 'Introduction to agile development and its challenges.' Please generate questions and feedback from the virtual audience."
[1377] In this way, users can repeatedly simulate their presentation with a virtual audience in advance, improving the quality of their presentation and giving them confidence when they actually take the stage.
[1378] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1379] Step 1:
[1380] Entering user profile and audience information
[1381] Input: User profile information (e.g., name, job title, years of experience) and audience information (e.g., industry, age group, interests)
[1382] Processing: The user starts the application and authenticates by entering their user ID and password on the login screen. After logging in, they enter their profile information and audience information in the form that appears. The server receives this information and stores it in a database. It also designs an appropriate simulation scenario based on the user and audience information.
[1383] Output: User and audience information stored in a database, and simulation scenario designs
[1384] Step 2:
[1385] Upload presentation materials
[1386] Input: Presentation materials (e.g. PowerPoint files, PDF files)
[1387] Processing: The user clicks the "File Upload" button in the application and selects presentation materials from a selection screen. The device reads the selected file and sends it to the server. The server analyzes the received file and extracts text and image data for each slide. This analysis is performed using text analysis software and image recognition software.
[1388] Output: Parsed text and image data for each slide
[1389] Step 3:
[1390] Creating a reaction
[1391] Input: Parsed presentation materials, user profile information, audience information
[1392] Processing: The server uses the generative AI model to generate comments, reactions, and questions from the virtual audience based on the analyzed content and the input user profile and audience information. For example, it dynamically generates questions and reactions related to the content of the slides.
[1393] Output: Generated virtual audience reaction data (comments, questions, responses, etc.)
[1394] Step 4:
[1395] Connecting to smart glasses
[1396] Input: Wi-Fi and Bluetooth information for connection
[1397] Processing: The user wears a video display device (e.g., smart glasses) and selects the "Connect to video display device" option within the application. The device establishes a connection with the video display device using the selected connection method. The device confirms that a connection with the video display device has been established.
[1398] Output: Established connection status with video display device
[1399] Step 5:
[1400] Start of presentation simulation
[1401] Input: Command to start the simulation
[1402] Processing: The user clicks the "Start Simulation" button in the application. The terminal sends the presentation materials and the generated reaction data to the video display device and prepares to start the simulation.
[1403] Output: Presentation materials and reaction data sent to a video display device
[1404] Step 6:
[1405] Q&A practice
[1406] Input: Virtual audience question displayed on a video display device
[1407] Processing: The video display device displays the question from the virtual audience. The user responds verbally to the question, and the response is captured using the video display device's audio input function and sent to the server. The server analyzes the audio data and generates feedback and additional reactions based on the user's response, which are then sent back to the video display device.
[1408] Output: Parsed user response data, generated new feedback and reactions
[1409] (Application example 1)
[1410] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1411] Conventional presentation practice systems primarily require users to practice on their own, lacking the ability to provide real-time feedback or reflect the reactions of a virtual audience. Furthermore, systems for effectively training salespeople in brick-and-mortar stores on product explanations are also inadequate, and there is a need for a system that can respond immediately to customer questions and reactions. The present invention aims to solve these problems and improve the quality of presentation practice and sales training.
[1412] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1413] In this invention, the server includes: a means for a user to input profile information and audience information; a means for uploading presentation materials; a means for the server to analyze the uploaded materials and generate reactions; a means for a terminal to transmit and receive data between the user and the server and operate the smart glasses; a means for the smart glasses to display reactions and perform a presentation simulation; a means for a salesperson to input profile information and customer information and upload product explanation materials; a means for the server to analyze the product explanation materials and generate questions and reactions of virtual customers; and a means for the smart glasses to display the responses and questions of the virtual customers in real time and perform a sales training simulation. This allows the user to make a presentation while receiving reactions from the virtual audience, and enables the salesperson to receive practical training through interactions with virtual customers.
[1414] "User profile information" is information that includes personal information, job information, experience, and the like about the user.
[1415] "Audience information" is attribute information about the audience to whom the presentation is given.
[1416] "Presentation materials" are presentation materials provided in the form of slides or documents.
[1417] A "server" is a computer system that receives data from users, analyzes it, and generates appropriate reactions.
[1418] "Means for generating reactions" is a function that uses an AI model to generate virtual comments and questions based on presentation materials and user information.
[1419] A "terminal" is a device that transmits and receives data between a user and a server.
[1420] "Smart glasses" are eyeglass-type devices that utilize augmented reality technology to display information to users in real time.
[1421] The "means for simulating a presentation" is a function that allows a user to use smart glasses to give a presentation while interacting with a virtual audience.
[1422] "Salesperson profile information" is information including job information, products in charge, experience, etc., related to the salesperson.
[1423] "Customer information" is attribute information about customers who are the target of sales.
[1424] "Product description materials" are materials used to explain the characteristics and functions of a product.
[1425] "Virtual Customer" means a fictitious customer generated using AI technology.
[1426] "Means for generating questions and reactions" refers to a function that uses an AI model to generate virtual customer questions and reactions based on salesperson profile information and product explanation materials.
[1427] The "means for conducting sales training simulations" is a function that allows salespeople to use smart glasses to conduct practical training through interactions with virtual customers.
[1428] The present invention relates to a comprehensive system for users to practice presentations and sales training. The system begins when the user inputs profile information and audience information and uploads presentation or product explanation materials. The server analyzes the uploaded materials and generates reactions. The terminal transmits and receives data between the user and the server and operates the smart glasses. The smart glasses display reactions and perform a presentation simulation or sales training simulation.
[1429] Specific steps
[1430] 1. Enter your profile and audience information
[1431] A user launches the application and enters their profile information (e.g., job title, experience, products) and audience or customer information for the presentation or sale (e.g., audience age range, interests, skill level).
[1432] 2. Upload presentation or product description materials
[1433] Users upload presentation materials (e.g., PowerPoint or PDF files) or product descriptions through the application. These materials are then sent to the server, where they are analyzed. During the analysis process, important text and image information is extracted from the materials.
[1434] 3. Generating Reactions
[1435] The server uses a generative AI model to generate reactions based on the uploaded materials and user and audience information, including comments, reactions, and questions from the virtual audience or virtual customers.
[1436] 4. Connecting to smart glasses
[1437] The user wears the smart glasses, and the application establishes a connection with the smart glasses, which receives data from the server and displays reactions.
[1438] 5. Start a presentation simulation or sales training
[1439] The user interacts with the application and enters the simulation mode, where the smart glasses display the reactions of a virtual audience and virtual customers, providing virtual feedback as the user progresses through slides and product presentations.
[1440] 6. Practice Q&A
[1441] As part of the simulation, users respond to questions posed by a virtual audience or customer, and their answers are sent through the smart glasses to a server that analyzes their responses and generates subsequent reactions and feedback.
[1442] Hardware and software used
[1443] Hardware used: Smart glasses (e.g., Google Glass, Vuzix Blade)
[1444] Software used: Flask (server-side framework), generative AI model (reaction generation), data analysis system
[1445] Specific examples
[1446] For example, when training a salesperson to explain a new smartphone, the following prompts are fed into the generative AI model:
[1447] "User: Salesperson, Experience: 3 years, Product: New smartphone, Target Audience: 20-30 years old, Tech-savvy"
[1448] Feedback from the virtual customer is displayed on the smart glasses with questions such as:
[1449] "How does this smartphone's camera performance compare to other products?"
[1450] In this way, users can improve the quality of their presentations and sales training by repeating virtual simulations.
[1451] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1452] Step 1:
[1453] A user launches the application and enters profile and audience information.
[1454] This information includes the user's job title, experience, products they work on, and target audience demographics (age range, interests, skill level, etc.).
[1455] The server receives these input data and stores them in a database.
[1456] Step 2:
[1457] A user uploads a presentation or product description.
[1458] These materials are provided in formats such as PDF and PowerPoint files.
[1459] The server receives the material and uses text analysis software (e.g., Apache Tika) to extract text and image information from the material.
[1460] The extracted information is stored in a database on the server.
[1461] Step 3:
[1462] The server generates reactions using a generative AI model based on uploaded materials, user information, and audience information.
[1463] This process generates questions and comments from the virtual audience or customer based on the content of the material.
[1464] The generated reactions are temporarily stored on the server.
[1465] Step 4:
[1466] The user puts on smart glasses (e.g., Google Glass, Vuzix Blade), and the application establishes a connection with the smart glasses.
[1467] The server transmits the generated reaction data to the smart glasses.
[1468] The smart glasses display the received data in real time in the user's field of vision.
[1469] Step 5:
[1470] The user operates the application and starts a presentation simulation or a sales training simulation.
[1471] Smart glasses display the reactions of virtual audiences and virtual customers.
[1472] Users receive real-time feedback as they progress through their presentations and product explanations.
[1473] Step 6:
[1474] As part of a presentation simulation or sales training, a user answers questions posed by a virtual audience or virtual customers.
[1475] The user's answers are sent to the server via the smart glasses.
[1476] The server analyzes the user's answers and generates the next reaction or feedback.
[1477] The generated feedback is then sent back to the smart glasses and displayed to the user.
[1478] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1479] The present invention relates to a system for providing more advanced feedback by combining a system for users to practice presentations with an emotion engine that recognizes the emotions of the users. Specific embodiments of the present invention will be described below.
[1480] System configuration
[1481] Entering user profile and audience information
[1482] A user launches the application and logs in. The user enters their profile information (e.g., name, job title, experience) and also sets audience information (e.g., age, occupation, areas of interest) for the presentation.
[1483] Upload presentation materials
[1484] Users upload presentation materials (e.g., PowerPoint or PDF files) from their PC or smartphone through the application. These materials are then sent to the server, where they are analyzed. The analysis includes extracting text and image information for each slide.
[1485] Creating a reaction
[1486] The server uses an AI model to generate reactions based on the uploaded presentation materials, user profile information, and audience information. Specifically, it determines what comments and questions the virtual audience will make, as well as the timing of their reactions.
[1487] Linking the Emotion Engine
[1488] When a user puts on the smart glasses and starts the simulation mode in the application, the emotion engine analyzes the user's facial expressions and tone of voice to recognize the user's emotional state in real time. For example, if the user is nervous, it will provide advice on how to relax.
[1489] Start of presentation simulation
[1490] The user interacts with the application to enter a presentation simulation mode, where the smart glasses display the reactions of the virtual audience. The user advances through the slides, presenting, and receives feedback from the virtual audience.
[1491] Q&A practice
[1492] As part of the simulation, users respond to questions posed by a virtual audience, and their responses are sent via smart glasses to a server that analyzes their responses and generates subsequent reactions and feedback.
[1493] Emotional data feedback and adjustment
[1494] The emotion engine analyzes the user's emotional data as they make their presentation and generates feedback based on their emotional state. For example, if the user feels anxious, the engine will advise them to explain the key points more simply.
[1495] Specific examples
[1496] For example, if a user wants to give a presentation on the theme of "innovative technology," the user first enters their profile information (e.g., experience as an engineer, field of expertise) into the application, and sets the audience attributes as "corporate technical personnel, aged 30-40, interested in technological innovation." Then, the user uploads presentation materials (PowerPoint files).
[1497] When a user puts on the smart glasses and starts the simulation mode, the emotion engine analyzes the user's facial expressions and tone of voice, and displays feedback such as, "You're feeling a little nervous on this slide," while the smart glasses also display advice such as, "Take a deep breath to relax."
[1498] In the presentation simulation, a virtual audience asks the user, "Please tell us some specific examples of how your technology can be applied," and the user responds. The response is sent to the server, where it is analyzed and the results are provided as feedback. For example, the user might receive feedback such as, "The explanation of the application example was easy to understand. However, I would like more specific examples."
[1499] In this way, by linking with the emotion engine, users can receive more detailed and useful feedback, significantly improving the quality of their presentations and giving them more confidence when they take the stage.
[1500] The processing flow will be explained below.
[1501] Step 1:
[1502] A user launches the application and logs in. The user enters their profile information (name, job title, experience, etc.) and also sets audience information (age, occupation, areas of interest, etc.) for the presentation.
[1503] Step 2:
[1504] Users select and upload presentation materials (PowerPoint or PDF files) from their PC or smartphone through the application, which are then sent to the server.
[1505] Step 3:
[1506] The server receives and analyzes the uploaded presentation materials, extracting text and image information from each slide, and using natural language processing (NLP) and image recognition technology to identify key points and locations within the materials.
[1507] Step 4:
[1508] The server uses an AI model to generate reactions based on the user's profile information and audience information. Specifically, it determines what comments and questions the virtual audience will make, as well as the timing of their reactions. The emotion engine also prepares to acquire the user's emotional data.
[1509] Step 5:
[1510] The user puts on the smart glasses and starts the simulation mode in the application. The device connects to the smart glasses and performs data communication with the server. The smart glasses receive reaction data from the server and prepare to display it.
[1511] Step 6:
[1512] The server confirms the start of the simulation mode and sends the analysis results of the presentation materials and reaction data to the smart glasses. As the user advances through the slides, the smart glasses display appropriate reactions. The emotion engine also analyzes the user's facial expressions and tone of voice in real time.
[1513] Step 7:
[1514] The device sends the user's operational information (slide progress and comments) to the server. The server uses this information to generate the next reaction or question. At the same time, the emotion engine analyzes the user's mental state based on the acquired emotional data and creates corresponding feedback.
[1515] Step 8:
[1516] The user can see the reactions of the virtual audience (e.g., "Interesting!", questions like, "What are some applications of this technology?") through the smart glasses and respond appropriately. The user's answers are sent to the server via the device.
[1517] Step 9:
[1518] The server analyzes the user's answers and emotional data to generate follow-up questions and feedback. This information is then sent back to the smart glasses and displayed to the user. For example, if the user is nervous, the system will suggest, "Try taking a deep breath to relax."
[1519] Step 10:
[1520] Once the user has finished the simulation, the server analyzes the overall performance data (e.g., the evaluation of each slide, the quality and timing of the user's responses, overall progress, and emotional data) and generates an evaluation result.
[1521] Step 11:
[1522] The user checks the practice results through the application and refers to the feedback provided by the server to improve the presentation materials and delivery method.
[1523] Step 12:
[1524] Users can re-run the simulation and practice as needed. By linking the emotion engine and reaction generation system, users can significantly improve the quality of their presentations, giving them the confidence to take the stage in real life.
[1525] Example 2
[1526] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1527] Previous presentation practice systems lacked the ability to analyze users' emotional states in real time and provide feedback based on that analysis. As a result, users were unable to adequately manage their own emotions and it was difficult to improve the quality of their presentations. It was also difficult to obtain detailed feedback from a virtual audience. This could lead to users losing confidence in their actual presentations and potentially resulting in poor performance.
[1528] The specification processing by the specification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for analyzing uploaded materials and generating reactions of the virtual audience, means for generating evaluation results and feedback based on the progress of the presentation and the user's emotional data, and an emotion analysis engine for analyzing the user's facial expressions and voice in real time and generating feedback. This allows the user to receive feedback according to their own emotional state in real time, improving the quality of their presentation and giving them confidence in their actual presentation.
[1529] A "user" is an individual who uses the system to practice a presentation through the application.
[1530] "Profile information" is information about personal attributes including the user's name, job title, experience, etc.
[1531] "Audience information" refers to attribute information such as age, occupation, and areas of interest of the virtual audience to whom the presentation is directed.
[1532] "Presentation materials" are presentation content such as slides and PDF files uploaded by users.
[1533] The "server" is a computer system that performs processes such as analyzing uploaded presentation materials, generating reactions from a virtual audience, and analyzing user emotional data.
[1534] A "terminal" is a device that allows a user to send and receive data to and from a server and operate a wearable device.
[1535] A "wearable device" is a wearable device (e.g., smart glasses) that can be worn by the user to display reactions and perform emotional analysis.
[1536] "Reactions" are responses such as comments, feedback, and questions from the virtual audience.
[1537] An "emotion analysis engine" is software that analyzes a user's facial expressions, voice, etc. in real time, recognizes the user's emotional state, and generates feedback.
[1538] "Feedback" is response information that includes evaluations and advice regarding the user's behavior, emotional state, etc.
[1539] A "virtual audience" is an imagined audience that makes comments and asks questions about a user's presentation generated by the system.
[1540] "Real-time" refers to the analysis and feedback process occurring instantly while the presentation is in progress.
[1541] The present invention provides a system for providing more advanced feedback by combining a system for users to practice presentations with an emotion analysis engine. Specific embodiments of this system are described below.
[1542] First, the user launches a dedicated application and logs in. After logging in, the user enters profile information (e.g., name, job title, experience) and audience information for the presentation (e.g., age, occupation, areas of interest). Next, the user uploads presentation materials (e.g., PowerPoint or PDF files). These materials are sent to the server, where they are analyzed. This analysis includes extracting text and image information for each slide.
[1543] The server uses a generative AI model to generate reactions from a virtual audience based on the uploaded presentation materials and user input. These reactions include comments and questions, and the timing of their occurrence is also determined. For example, the server inputs a prompt statement such as "Generate questions that the audience would ask during a presentation on technological innovations" into the generative AI model, and uses the responses.
[1544] Next, the user puts on the wearable device (e.g., smart glasses) and starts the application's simulation mode. In this mode, the emotion analysis engine analyzes the user's facial expressions and voice tone in real time. Based on the analysis results, real-time feedback is provided according to the user's emotional state (e.g., nervousness, anxiety, confidence). For example, the feedback "This slide makes you a little nervous" is displayed along with advice such as "Take a deep breath to relax."
[1545] During the simulation, the user gives a presentation and receives reactions from the virtual audience. When the user responds to the comments and questions of the virtual audience, the response is sent to the server via the wearable device. The server analyzes the user's response and generates the next reaction or feedback based on it. For example, the server may provide feedback such as, "The explanation of the application example was easy to understand, but please add more concrete examples."
[1546] This system allows users to receive detailed and useful feedback based on their emotional state in real time, which will improve the quality of their presentations and increase their confidence in their presentations. This confidence is expected to contribute to the success of their presentations.
[1547] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1548] Step 1:
[1549] The user launches the application and logs in.
[1550] Specific behavior:
[1551] The user taps the application icon to launch it, and enters their user ID and password on the login screen that appears. If they are successfully authenticated, the user profile screen is displayed.
[1552] Input and Output:
[1553] Input: User ID, Password
[1554] Data processing / calculation: The authentication server checks the ID and password
[1555] Output: Notification of successful or failed login
[1556] Step 2:
[1557] The user enters profile and audience information.
[1558] Specific behavior:
[1559] The user enters profile information (name, job title, experience) and audience information (age, occupation, areas of interest) into the form and clicks the "Next" button.
[1560] Input and Output:
[1561] Input: Profile information, audience information
[1562] Data processing / calculation: Input data is saved on the server as a user profile
[1563] Output: Confirmation of input completion and moving to the next step
[1564] Step 3:
[1565] A user uploads a presentation.
[1566] Specific behavior:
[1567] The user clicks the "Upload Document" button and selects a presentation file (e.g., PowerPoint, PDF) from the file selection dialog. The selected file is sent to the server.
[1568] Input and Output:
[1569] Input: Selected presentation file
[1570] Data processing / calculation: File upload and file storage on the server
[1571] Output: File upload completion notification
[1572] Step 4:
[1573] The server analyzes the uploaded presentation materials.
[1574] Specific behavior:
[1575] The server analyzes the received presentation file and extracts the text and image information of each slide, which is then stored in a database.
[1576] Input and Output:
[1577] Input: Uploaded presentation file
[1578] Data processing / calculation: Extraction of text and image information, storage in database
[1579] Output: Analysis data saved
[1580] Step 5:
[1581] The server uses a generative AI model to generate reactions from the virtual audience.
[1582] Specific behavior:
[1583] Based on the extracted data and input information, the server sends prompts to the generative AI model to generate reactions from the virtual audience, such as "Generate questions that the audience would ask during a presentation about technological innovations."
[1584] Input and Output:
[1585] Input: Analytics data, profile information, audience information
[1586] Data processing / calculation: Reaction generation using generative AI models
[1587] Output: Virtual audience comments, feedback, and questions
[1588] Step 6:
[1589] The user puts on the wearable device and starts the simulation mode.
[1590] Specific behavior:
[1591] The user puts on the smart glasses and clicks the "Start Simulation Mode" button in the application, which activates the wearable device and displays the reactions of the virtual audience.
[1592] Input and Output:
[1593] Input: Instruction to start the wearable device
[1594] Data processing / calculation: Connection and data transmission between terminal and wearable device
[1595] Output: Start of simulation mode, display of virtual audience
[1596] Step 7:
[1597] The emotion analysis engine analyzes the user's facial expressions and tone of voice in real time.
[1598] Specific behavior:
[1599] The emotion analysis engine uses the smart glasses' built-in camera and microphone to capture and analyze the user's facial expressions and tone of voice, and the analysis data is fed back to the user in real time.
[1600] Input and Output:
[1601] Input: User's facial expression data, voice data
[1602] Data processing / calculation: Data analysis and feedback generation using a sentiment analysis engine
[1603] Output: Real-time emotional feedback
[1604] Step 8:
[1605] Get reactions from a virtual audience as you deliver your presentation.
[1606] Specific behavior:
[1607] As users advance through the slides, the virtual audience displays comments and questions at designated times, to which users can respond.
[1608] Input and Output:
[1609] Input: User's slide progress data
[1610] Data processing / calculation: Reaction display timing control, reaction display
[1611] Output: Display of virtual audience comments and questions
[1612] Step 9:
[1613] Users practice their answers to questions from a virtual audience.
[1614] Specific behavior:
[1615] The user answers questions posed by the virtual audience, and the answers are sent to a server via smart glasses, where they are analyzed and used to generate the next reaction or feedback.
[1616] Input and Output:
[1617] Input: User's response audio data
[1618] Data processing / calculation: Analysis of voice data by server, generation of next reaction
[1619] Output: Analysis results and feedback
[1620] Step 10:
[1621] The emotion analysis engine provides feedback on the user's emotion data during the presentation.
[1622] Specific behavior:
[1623] The emotion analysis engine analyzes the data as it goes through the presentation and provides emotional feedback to the user, such as "I'm a little nervous about the next slide," allowing the user to make appropriate adjustments (e.g., take a deep breath, slow down speaking speed).
[1624] Input and Output:
[1625] Input: Real-time facial expression data, voice data
[1626] Data processing / calculation: Analysis by emotion analysis engine and feedback generation
[1627] Output: Real-time feedback according to emotional state
[1628] (Application example 2)
[1629] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1630] In employee training, it is difficult to improve skills by simulating actual work situations and customer interactions and receiving appropriate feedback in real time. There is also a lack of technology that can analyze employees' emotional states and provide specific advice accordingly. Therefore, there is a need for a system that can effectively improve employees' work performance before they actually begin work.
[1631] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1632] In this invention, the server includes means for a user to input profile information and subject information, means for uploading training materials, means for the server to analyze the uploaded materials and generate feedback, means for a terminal to transmit and receive data between the user and the server and operate the smart device, means for the smart device to display the feedback and perform a training simulation, and means for the server to generate evaluation results and feedback based on the progress of the training. This allows employees to receive feedback from virtual customers in real time while undergoing job training, and can provide specific advice according to the employee's emotional state, thereby improving their skills and providing a sense of security.
[1633] "Profile information" is basic personal data such as the user's name, job title, experience, etc.
[1634] "Target information" is data such as the age, occupation, and areas of interest of a hypothetical customer or audience.
[1635] "Training materials" are presentation materials and simulation documents used by users.
[1636] "Uploading means" refers to a method or function that allows a user to send training materials to a server via a terminal.
[1637] "Means for analyzing and generating feedback" refers to the function that enables the server to analyze uploaded materials and provide appropriate comments and advice.
[1638] A "terminal" is a device used by a user, such as a smartphone or tablet.
[1639] A "smart device" is a device for displaying feedback to a user, such as smart glasses or a head-mounted display.
[1640] The "means for performing training simulation" is a function that allows users to simulate actual work through dialogue with virtual customers and scenarios.
[1641] The "means for generating evaluation results and feedback" refers to a method by which the server evaluates the user's performance based on the progress of the training and generates appropriate advice and comments.
[1642] This invention relates to a system for effectively conducting employee training. The system begins when a user enters their profile information and target information and uploads training materials. The uploaded materials are analyzed by a server, and feedback is generated. A terminal transmits and receives data between the user and the server and operates a smart device. The smart device displays the feedback and functions as a means for conducting training simulations.
[1643] Hardware and software used
[1644] Smart devices (smart glasses, head-mounted displays, etc.)
[1645] These devices are used to give users real-time feedback during a simulation.
[1646] Devices (smartphones, tablets)
[1647] Used by users to upload training materials and enter profile and audience information.
[1648] server
[1649] The server installed in the cloud environment is responsible for analyzing materials, generating feedback, sending and receiving data, etc. Specifically, a server with an AI model implemented (e.g., a machine learning model using TensorFlow or PyTorch) is used.
[1650] Cloud services (e.g. AWS, Azure)
[1651] It is used to assist in centralized data management and calculation processing.
[1652] System Operation
[1653] 1. User configuration
[1654] The user uses a terminal to enter their profile information (e.g., job title, experience, etc.) and target information (e.g., age and occupation of the expected customers) according to the training scenario, and then uploads presentation and training materials.
[1655] 2. Server analysis
[1656] The server analyzes the uploaded materials and uses a generative AI model to generate hypothetical customer reactions and questions, taking into account the user's past training and performance data to generate evaluation results and feedback.
[1657] 3. Real-time feedback
[1658] When a user puts on a smart device and starts training, feedback is displayed on the smart device via the terminal from the server. For example, if the user feels tense, advice on how to relax or specific advice on questions from the customer is provided in real time.
[1659] Specific examples
[1660] For example, if a user is conducting training on the theme of "new smartphone products," they can enter "new staff member with two months' experience" as their profile information and set "male in his 30s, technically knowledgeable customer" as the target scenario. Next, they select a training scenario and questions from the virtual customer are displayed.
[1661] Example prompt sentence:
[1662] A customer asks, "What are the camera features on this phone?" Observe how your employees react and ask the question, and provide appropriate feedback. If your employees seem nervous, include tips to help them relax.
[1663] The system allows employees to receive real-time feedback from virtual customers while training on the job, improving their skills and providing peace of mind.
[1664] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1665] Step 1:
[1666] Operation: A user uses a terminal to enter their profile information and target information.
[1667] Input: User profile information (e.g., name, job title, experience), audience information (e.g., age, occupation, areas of interest).
[1668] Output: The entered information is sent from the terminal to the server.
[1669] Step 2:
[1670] How it works: A user uploads presentation or training materials via a terminal.
[1671] Input: Presentation materials (e.g. PowerPoint files or PDFs).
[1672] Output: Uploaded materials are sent from the device to the server and saved.
[1673] Step 3:
[1674] How it works: The server receives the uploaded material and analyzes it.
[1675] Input: Presentation materials, user profile information, audience information.
[1676] Data processing / calculation: Extract text and image information from documents using generative AI models.
[1677] Output: Generate virtual customer reactions and questions based on the extracted information.
[1678] Step 4:
[1679] Operation: The server generates feedback for the virtual customer based on the user's input data and the analysis results.
[1680] Input: User profile information, privacy data, analytics results.
[1681] Data processing / calculation: Generate prompt statements and generate feedback from virtual customers using AI models.
[1682] Output: The generated feedback data is sent from the server to the terminal.
[1683] Step 5:
[1684] Actions: A user puts on a smart device and starts a training simulation.
[1685] Input: Generated feedback data, user voice, and facial expression data.
[1686] Output: Virtual customer feedback is displayed in real time on a smart device.
[1687] Step 6:
[1688] Operation: The server receives the user's voice and facial expression data during the simulation and analyzes it using the emotion engine.
[1689] Input: User's voice and facial expression data.
[1690] Data processing / calculation: Analyzed by the emotion engine to evaluate the user's stress level and tension.
[1691] Output: As a result of the analysis, feedback based on the user's emotional state is generated.
[1692] Step 7:
[1693] How it works: Based on the analysis results of the emotion engine, the server generates advice for the user in real time and sends it to the smart device.
[1694] Input: Analysis results of the emotion engine, feedback from virtual customers.
[1695] Output: Specific advice to the user (e.g., "Take deep breaths to relax") is displayed on the smart device.
[1696] Step 8:
[1697] Action: The user answers questions posed by the virtual customer.
[1698] Input: User response data.
[1699] Output: The user's answer data is sent to the server and analyzed.
[1700] Step 9:
[1701] How it works: The server analyzes the user's answers and generates the following feedback and rating results:
[1702] Input: User response data, past training data.
[1703] Data processing / calculation: Analyzed by AI model.
[1704] Output: As a result of the analysis, feedback and evaluation results from the next virtual customer are generated and sent to the smart device.
[1705] Step 10:
[1706] Operation: After the training is completed, the server generates an evaluation result of the entire training and provides it to the user.
[1707] Input: All training data, emotion data.
[1708] Data processing / calculation: Conduct an overall evaluation of training performance.
[1709] Output: The final evaluation results are displayed on the device along with feedback on future improvements.
[1710] These steps allow employees to virtually experience real-world work scenarios and receive appropriate feedback in real time, resulting in improved work skills and a sense of psychological security.
[1711] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1712] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1713] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1714] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1715] FIG. 9 illustrates an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and behaviors arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1716] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1717] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1718] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1719] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1720] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1721] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1722] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1723] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1724] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1725] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1726] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1727] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1728] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1729] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1730] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1731] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1732] The following is further disclosed regarding the above embodiment.
[1733] (Claim 1)
[1734] a means for users to input profile information and audience information;
[1735] A means to upload presentation materials;
[1736] a means for the server to analyze the uploaded material and generate reactions;
[1737] A means for the terminal to transmit and receive data between the user and the server and operate the smart glasses;
[1738] The system includes a means for smart glasses to display reactions and simulate presentations.
[1739] (Claim 2)
[1740] 10. The system of claim 1, wherein the reactions include virtual comments, reactions, and questions from the audience.
[1741] (Claim 3)
[1742] 10. The system of claim 1, wherein the server comprises means for generating evaluation results and feedback based on the progress of the presentation.
[1743] (Claim 4)
[1744] The system according to claim 1, wherein the terminal has means for acquiring the user's response through the smart glasses and transmitting the response to the server.
[1745] "Example 1"
[1746] (Claim 1)
[1747] a means for a user to input profile information and student information;
[1748] A means to upload presentation materials;
[1749] A means for the information processing device to analyze the uploaded materials and generate a response;
[1750] a means for transmitting and receiving data between a user and an information processing device in the terminal and for operating a video display device;
[1751] The system includes a video display device that displays responses and a means for practicing presentations.
[1752] (Claim 2)
[1753] 10. The system of claim 1, wherein the responses include virtual comments, reactions, and questions of the student.
[1754] (Claim 3)
[1755] 2. The system according to claim 1, wherein the information processing device has means for generating evaluation results and feedback based on the progress of the presentation.
[1756] "Application Example 1"
[1757] (Claim 1)
[1758] a means for users to input profile information and audience information;
[1759] A means to upload presentation materials;
[1760] a means for the server to analyze the uploaded material and generate reactions;
[1761] A means for the terminal to transmit and receive data between the user and the server and operate the smart glasses;
[1762] Smart glasses will display reactions and provide a means for presentation simulations.
[1763] means for sales associates to enter profile information, customer information, and upload product description materials;
[1764] A means for the server to analyze product description materials and generate questions and reactions of virtual customers;
[1765] How smart glasses can display real-time virtual customer responses and questions for sales training simulations
[1766] A system including:
[1767] (Claim 2)
[1768] 10. The system of claim 1, wherein the reactions include virtual comments, reactions, and questions from the audience.
[1769] (Claim 3)
[1770] 10. The system of claim 1, wherein the server comprises means for generating evaluation results and feedback based on the progress of the presentation.
[1771] "Example 2: Combining Emotion Engines"
[1772] (Claim 1)
[1773] a means for users to input profile information and audience information;
[1774] A means to upload presentation materials;
[1775] a means for the server to analyze the uploaded material and generate virtual audience reactions;
[1776] A terminal is a device for transmitting and receiving data between a user and a server and for operating a wearable device;
[1777] A wearable device displays reactions and simulates presentations.
[1778] A means for an emotion analysis engine to analyze a user's facial expressions and voice in real time and generate feedback
[1779] A system including:
[1780] (Claim 2)
[1781] 10. The system of claim 1, wherein the reactions include comments, reactions, and questions of the virtual audience.
[1782] (Claim 3)
[1783] 2. The system according to claim 1, wherein the server has means for generating evaluation results and feedback based on the progress of the presentation and the user's emotional data.
[1784] "Application example 2 when combining emotion engines"
[1785] (Claim 1)
[1786] a means for a user to input profile information and target information;
[1787] a means for uploading training materials;
[1788] means for the server to analyze the uploaded material and generate feedback;
[1789] A means for the terminal to transmit and receive data between the user and the server and operate the smart device;
[1790] The system includes a means for the smart device to display feedback and conduct training simulations.
[1791] (Claim 2)
[1792] 10. The system of claim 1, wherein the feedback includes virtual comments, reactions, and questions of the subject.
[1793] (Claim 3)
[1794] 10. The system of claim 1, wherein the server comprises means for generating evaluation results and feedback based on the progress of the training. [Explanation of symbols]
[1795] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. a means for users to input profile information and audience information; A means to upload presentation materials; a means for the server to analyze the uploaded material and generate reactions; A means for the terminal to transmit and receive data between the user and the server and operate the smart glasses; The system includes a means for smart glasses to display reactions and simulate presentations.
2. The system of claim 1 , wherein the reactions include virtual comments, reactions, and questions from the audience.
3. 2. The system of claim 1, wherein the server comprises means for generating evaluation results and feedback based on the progress of the presentation.
4. The system according to claim 1 , wherein the terminal has means for acquiring a user's response through the smart glasses and transmitting the response to the server.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A