System

The system provides comprehensive feedback on voice, facial expressions, and slide materials to enhance presentation skills by integrating data analysis and question generation, addressing the limitations of existing technologies.

JP2026022547APending Publication Date: 2026-02-12SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024124064
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-30
Publication Date
2026-02-12

AI Technical Summary

Technical Problem

Existing systems lack the ability to comprehensively analyze a user's voice, facial expressions, gestures, and slides to provide comprehensive feedback, making it difficult for users to effectively improve their presentation skills.

Method used

A system that includes a user interface for inputting voice, facial expressions, and slide materials, data transmission to a server, audio and video analysis for evaluating speaking and presentation composure, slide material analysis for content consistency, and feedback generation to provide detailed improvements, along with question generation and additional feedback based on user responses.

Benefits of technology

Enables users to receive comprehensive and detailed feedback, allowing them to practice presentations in a realistic environment and improve their skills effectively.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026022547000001_ABST
    Figure 2026022547000001_ABST
Patent Text Reader

Abstract

To provide a system capable of effectively training a presentation in an environment close to a real world.SOLUTION: The specification processing unit 290 of the data processing device 12 in the system receives the voice, the facial expression, the gesture, and the slide material input by the user, converts the received voice data into text, evaluates the speaking speed, the volume, and the pause, analyzes the facial expression, the line of sight, and the gesture of the user from the received video data, evaluates the degree of calmness and the degree of confidence of the presentation, analyzes the slide material, evaluates the consistency of the content and the visual effect, generates feedback to the user based on these evaluation results, and transmits the feedback to the terminal of the user.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] In the real world, presentations are required in a wide variety of situations, including job interviews, reports and proposals at meetings, and speaking in front of large audiences. However, there are few ways to efficiently practice and train presentations, making it difficult to deliver high-quality presentations. As a result, it becomes difficult to gain audience understanding and acceptance, reducing the effectiveness of the presentation. This invention aims to solve the above problem by providing a system that utilizes generative AI to enable users to receive feedback in real time and effectively practice presentations in an environment that closely resembles the real world. [Means for solving the problem]

[0005] The present invention provides a system including the following means: a user interface means for a user to input voice, facial expressions, gestures, and slide materials; a data transmission means for transmitting the data input from the user interface means to a server; an audio analysis means for converting the voice data received by the server into text and evaluating the speaking speed, volume, and pauses; a video analysis means for analyzing the user's facial expressions, gaze, and gestures from video data received by the server and evaluating the presentation's composure and confidence; a slide material analysis means for analyzing the slide materials received by the server and evaluating the consistency of the content and visual effect; a feedback generation means for generating feedback to the user based on the evaluation results of the voice analysis means, video analysis means, and slide material analysis means; and a feedback transmission means for transmitting the feedback generated by the feedback generation means to the user's terminal. The system of the present invention further includes an authentication management means for the server to manage user authentication information and store each user's presentation history and feedback results. Furthermore, the feedback generating means further includes a question generating means for generating questions based on the content of the user's presentation and presenting them to the user, and an additional feedback generating means for analyzing the answers given by the user to the presented questions and generating additional feedback based on the evaluation results, thereby allowing the user to practice presentations in a realistic environment and improve their skills.

[0006] The "user interface means" is an interface through which the user inputs voice, facial expressions, gestures, and slide materials.

[0007] The "data transmission means" is a means for transmitting data input from the user interface means to the server.

[0008] The "voice analysis means" is a means for converting voice data received by the server into text and evaluating speaking speed, volume, and pauses.

[0009] The "video analysis means" is a means for analyzing the user's facial expressions, gaze, and gestures from the video data received by the server, and evaluating the user's level of composure and confidence in the presentation.

[0010] The "slide material analysis means" is a means for analyzing the slide materials received by the server and evaluating the consistency of their contents and their visual effects.

[0011] The "feedback generation means" is a means for generating feedback to the user based on the evaluation results of the audio analysis means, the video analysis means, and the slide material analysis means.

[0012] The "feedback transmitting means" is a means for transmitting the feedback generated by the feedback generating means to the user's terminal.

[0013] The "authentication management means" is a means by which the server manages user authentication information and stores each user's presentation history and feedback results.

[0014] The "question generation means" is a means for generating questions based on the content of a user's presentation and presenting them to the user.

[0015] The "additional feedback generating means" is a means for analyzing the answers given by the user to the questions presented, and generating additional feedback based on the evaluation results. [Brief explanation of the drawings]

[0016] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4]FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0017] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0018] First, the terms used in the following description will be explained.

[0019] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0020] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0021] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0022] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0023] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0024] [First embodiment]

[0025] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0026] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0027] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0028] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0029] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0030] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0031] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0032] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0033] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0034] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0035] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0036] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0037] The system for improving presentation skills according to the present invention can be implemented as follows.

[0038] Basic configuration

[0039] The system consists of a user, a terminal, and a server. The user is the person giving the presentation and interfaces with the system through the terminal. The terminal collects the user's voice, facial expressions, gestures, and slides and transmits them to the server. The server analyzes the received data and provides feedback to the user.

[0040] User Interface Means

[0041] A user gives a presentation using a terminal. The terminal is equipped with an audio input device (microphone), a video input device (camera), and a screen for displaying slides. This allows the user to input voice, facial expressions, gestures, and slides into the system.

[0042] Data transmission method

[0043] The terminal has the function of transmitting the audio, video, and slides input by the user to the server. This process is performed in real time, and the data is buffered and sent to the server.

[0044] Voice analysis methods

[0045] The server analyzes the received audio data. Specifically, it uses speech recognition technology to convert the audio into text, and then evaluates the speaking speed, volume, and pauses based on the text data. For example, if the user speaks too fast, this will be reflected in the feedback.

[0046] Video analysis methods

[0047] The server analyzes the received video data, using computer vision technology to analyze the user's facial expressions, eye contact, and gestures, and evaluates the presentation's poise and confidence. For example, if the user is not making eye contact with the audience, this will be included in the feedback.

[0048] Slide analysis tools

[0049] The server analyzes the slides it receives, using optical character recognition (OCR) technology to read their content and evaluate whether they are consistent with the presentation theme and whether the visual effect is appropriate. For example, if a slide contains a typographical error, that will be reflected in the feedback.

[0050] Feedback generation means and feedback transmission means

[0051] The server generates feedback to be provided to the user based on the evaluation results obtained by the audio analysis means, video analysis means, and slide analysis means. The generated feedback is sent to the terminal and displayed to the user. The feedback may include adjustments to speaking speed and volume, improvements to facial expressions and gestures, and corrections to slides.

[0052] Authentication Management Methods

[0053] The server manages user authentication information and stores the user's presentation history and feedback results, allowing users to check their progress and refer to past feedback.

[0054] Question generation and additional feedback generation

[0055] The server generates questions based on the user's presentation and presents them to the user. When the user answers these questions, the answers are also analyzed and additional feedback is generated. For example, if the user does not answer a question accurately, that point is reflected in the additional feedback.

[0056] Specific examples

[0057] For example, if a user gives a presentation on the theme of "introducing a new product," the system functions as follows: The user gives the presentation using a device, and the server uses its audio analysis means to evaluate that the user's voice is "quiet." The video analysis means determines that the user's eyes are always looking at the slides, and the slide analysis means analyzes that there are typos on some of the slides. The server generates feedback based on these evaluations and provides the user with specific advice such as "increase the volume of your voice when speaking," "make eye contact with the audience," and "correct typos in the slides."

[0058] In this way, the system of the present invention functions as an effective tool for users to efficiently improve their presentation skills.

[0059] The processing flow will be explained below.

[0060] Step 1:

[0061] A user logs into the system using a terminal, enters a username and password, and submits the authentication information.

[0062] Step 2:

[0063] The device sends the entered authentication information to the server. The server compares it with the database and returns the authentication result to the device. If authentication is successful, the user can start the presentation.

[0064] Step 3:

[0065] The user begins preparing for a presentation. They upload presentation materials (slides) using their device. The device then sends the uploaded materials to the server.

[0066] Step 4:

[0067] The server stores the received presentation materials in an appropriate format.

[0068] Step 5:

[0069] The user presses the "Start Presentation" button. The device begins recording the user's voice and video (facial expressions and gestures) using the built-in camera and microphone. The recorded data is buffered.

[0070] Step 6:

[0071] The terminal transmits the buffered data to the server in real time. The data for the audio, video, and slide materials are transmitted to the server.

[0072] Step 7:

[0073] The server analyzes the received voice data, converts it into text using a speech recognition engine, and evaluates speaking speed, volume, and pauses.

[0074] Step 8:

[0075] The server analyzes the received video data and uses computer vision technology to analyze the user's facial expressions, eye movements, and gestures to evaluate the presentation's poise and confidence.

[0076] Step 9:

[0077] The server analyzes the slides it receives, using optical character recognition (OCR) technology to read their content and evaluate whether they are consistent with the presentation theme and whether their visual impact is appropriate.

[0078] Step 10:

[0079] The server integrates the results of the audio, video, and slide analysis and generates feedback for the user. Based on each evaluation result, it summarizes specific improvements and advice in written form.

[0080] Step 11:

[0081] The server sends the generated feedback to the user's terminal, which receives the feedback and displays it to the user.

[0082] Step 12:

[0083] Users can review the feedback and incorporate improvements for their next presentation. If desired, a simulated Q&A session can also be conducted.

[0084] Step 13:

[0085] If a question and answer simulation is desired, the server generates questions based on the content of the presentation and presents them to the user.

[0086] Step 14:

[0087] The user answers the questions, and the device sends the answer data to the server.

[0088] Step 15:

[0089] The server analyzes the user's answers and generates additional feedback, and reflects suggestions for improvement in the advice based on the evaluation results.

[0090] Step 16:

[0091] The server sends additional feedback to the user's device, which the user receives and uses as a reference for their next presentation.

[0092] This series of processing steps allows users to improve their presentation skills efficiently and effectively.

[0093] Example 1

[0094] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0095] Conventional presentation support systems lack the ability to comprehensively analyze a user's voice, facial expressions, gestures, and slides to provide comprehensive feedback. They also lack the ability to generate questions based on the user's presentation content, analyze the answers, and provide additional feedback. This makes it difficult for users to effectively improve their presentation skills.

[0096] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0097] In this invention, the server includes a user interface means for a user to input voice, facial expressions, gestures, and slide materials, a data transmission means for transmitting the data input from the user interface means to the server, an audio analysis means for converting the voice data received by the server into text and evaluating the speaking speed, volume, and pauses, a video analysis means for analyzing the user's facial expressions, gaze, and gestures from the video data received by the server and evaluating the poise and confidence of the presentation, a slide material analysis means for analyzing the slide materials received by the server and evaluating the consistency and visual effect of the content, a feedback generation means for generating feedback to the user based on the evaluation results of the audio analysis means, the video analysis means, and the slide material analysis means, a feedback transmission means for transmitting the feedback generated by the feedback generation means to the user's terminal, a question generation means for the server to generate questions based on the content of the user's presentation and present them to the user, and an additional feedback generation means for analyzing the user's answers to the questions posed via the user interface means and generating additional feedback based on the evaluation results. This allows the user to receive comprehensive and detailed feedback, thereby effectively improving their presentation skills.

[0098] "User interface means" refers to means by which a user inputs voice, facial expressions, gestures, and slide materials.

[0099] The "data transmission means" is a means for transmitting data input from the user interface means to the server.

[0100] The "voice analysis means" is a means for converting voice data received by the server into text and evaluating speaking speed, volume, and pauses.

[0101] The "video analysis means" is a means for analyzing the user's facial expressions, gaze, and gestures from the video data received by the server, and evaluating the user's level of composure and confidence in the presentation.

[0102] The "slide material analysis means" is a means for analyzing the slide materials received by the server and evaluating the consistency of the content and the visual effect.

[0103] The "feedback generation means" is a means for generating feedback to the user based on the evaluation results of the audio analysis means, the video analysis means, and the slide material analysis means.

[0104] The "feedback transmitting means" is a means for transmitting the feedback generated by the feedback generating means to the user's terminal.

[0105] The "question generation means" is a means by which the server generates a question based on the content of the user's presentation and presents it to the user.

[0106] The "additional feedback generating means" is a means for analyzing the answers given by the user to questions presented via the user interface means and generating additional feedback based on the evaluation results.

[0107] The "authentication management means" is a means by which the server manages user authentication information and stores each user's presentation history and feedback results.

[0108] The system for improving presentation skills according to the present invention comprises a user, a terminal, and a server. By using this system, the user can effectively analyze their voice, facial expressions, gestures, and slides, and receive comprehensive feedback.

[0109] User Interface Means

[0110] When giving a presentation, a user uses a terminal equipped with a microphone, a camera, and a screen for displaying slides. This allows the user to input voice, facial expressions, gestures, and slides.

[0111] Data transmission method

[0112] The terminal transmits the audio, video, and slides input by the user to the server in real time. During this process, the data is buffered, allowing for smooth data transfer in a stable communication environment.

[0113] Voice analysis methods

[0114] The server analyzes the received voice data. Specifically, it converts the voice into text using speech recognition technology (e.g., Google Cloud Speech-to-Text API), and evaluates the speaking speed, volume, and pauses based on the text data. For example, if a user says "introducing a new product," if they speak too quickly, this will be reflected in the feedback.

[0115] Video analysis methods

[0116] The server analyzes the received video data. Using computer vision techniques (e.g., OpenCV and TensorFlow), it analyzes the user's facial expressions, eye movements, and gestures to assess the presentation's poise and confidence. For example, if the user's eyes are fixed on the slides and they are not making eye contact with the audience, this will be included in the feedback.

[0117] Slide analysis tools

[0118] The server analyzes the slides it receives. It uses optical character recognition (OCR) technology (e.g., Tesseract OCR) to read the content of the slides and evaluates whether they are consistent with the presentation theme and whether the visual effect is appropriate. For example, if a slide contains a typographical error, that error is reflected in the feedback.

[0119] Feedback Generation Method

[0120] The server generates feedback to be provided to the user based on the evaluation results obtained by the audio analysis means, video analysis means, and slide analysis means. The generated feedback includes adjustments to speaking speed and volume, improvements to facial expressions and gestures, and corrections to slides.

[0121] Feedback sending method

[0122] The server sends the generated feedback to the device, which then displays it to the user, who receives real-time feedback and can immediately see how to improve their presentation.

[0123] Question generation means

[0124] The server generates questions based on the content of the user's presentation and presents them to the user via the terminal, encouraging the user to consider further the presentation.

[0125] Additional Feedback Generation Means

[0126] The server analyzes the answers given by the user to the questions posed and generates additional feedback based on the evaluation results, allowing the user to receive feedback at a deeper level.

[0127] Authentication Management Methods

[0128] The server manages user authentication information and stores presentation history and past feedback results, allowing users to check their progress and continuously improve their presentation skills.

[0129] Specific examples

[0130] For example, if a user is giving a presentation on the theme of "introducing a new product," the system functions as follows: The user gives the presentation using a terminal, and the server uses the audio analysis means to evaluate that the user's voice is quiet. The video analysis means determines that the user's eyes are always looking at the slides, and the slide analysis means analyzes that there are typos in some of the slides. The server generates feedback based on these evaluations and provides the user with specific advice, such as speaking louder, making eye contact with the audience, or correcting typos in the slides. In this way, the system of the present invention functions as an effective tool for users to efficiently improve their presentation skills.

[0131] Prompt Sentence Examples

[0132] For example, if you are giving a presentation on the topic of "Introducing a New Product," you might use the following prompt:

[0133] "You will be asked to give a presentation on the topic of introducing a new product. A microphone and camera will be used to collect data, and analysis will be performed using speech recognition, computer vision, and optical character recognition (OCR). Feedback will include an evaluation of speaking speed, volume, facial expressions, eye contact, gestures, and the content of your slides."

[0134] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0135] Step 1:

[0136] The user starts the application on the device, sets the theme to "New Product Introduction," and starts a presentation. The user can input voice, facial expressions, gestures, and slides.

[0137] Input: Set presentation theme, start presentation

[0138] Output:Presentation settings information

[0139] Step 2:

[0140] The device's microphone collects the user's voice, and the camera records the user's facial expressions and gestures, as well as capturing the slides the user uses.

[0141] Input: User's voice, facial expressions, gestures, slides

[0142] Output: Collected audio data, video data, and slide data

[0143] Step 3:

[0144] The devices transmit the collected data to the server in real time, a process in which the data is buffered and temporarily stored before being transmitted over the network.

[0145] Input: Audio data, video data, slide data

[0146] Output: Data sent

[0147] Step 4:

[0148] The server analyzes the received voice data, converts the voice into text using speech recognition technology (e.g., Google Cloud Speech-to-Text API), and evaluates the speaking speed, volume, and pauses based on the text data.

[0149] Input: Audio data

[0150] Output: Text data and its evaluation results

[0151] Specific operation: The server uses voice recognition technology to convert audio such as "introducing a new product" into text, and analyzes the speaking speed, how many characters per second, whether the voice volume is appropriate, etc.

[0152] Step 5:

[0153] The server analyzes the received video data and uses computer vision techniques (e.g., OpenCV and TensorFlow) to analyze the user's facial expressions, gaze, and gestures to evaluate the presentation's poise and confidence.

[0154] Input: Video data

[0155] Output: Analysis results

[0156] Specific operation: The server uses video analysis technology to evaluate whether the user's eyes are fixed on the slide, whether they are making eye contact with the audience, whether their gestures are natural, etc.

[0157] Step 6:

[0158] The server analyzes the received slides, using optical character recognition (OCR) technology (e.g., Tesseract OCR) to read the content of the slides and evaluate whether they are consistent with the presentation theme and have appropriate visual effects.

[0159] Input: Slide data

[0160] Output: Analysis results

[0161] How it works: The server uses OCR technology to analyze the slides and identify typos and inconsistencies in content, such as when "new product" is misspelled as "new store."

[0162] Step 7:

[0163] The server generates feedback to provide to the user based on the results of the audio, video, and slide analysis, including suggestions for adjusting speaking speed and volume, improving facial expressions and gestures, and correcting slides.

[0164] Input: Audio analysis results, video analysis results, slide analysis results

[0165] Output: Generated feedback

[0166] Specific operation: The server combines the results of each analysis and generates feedback on specific areas for improvement, such as "speaking too fast," "not making enough eye contact with the audience," and "there are typos on the slides."

[0167] Step 8:

[0168] The server sends the generated feedback to the device, which then displays it to the user, who receives real-time feedback and can immediately see how to improve their presentation.

[0169] Input: Generated feedback

[0170] Output: Feedback displayed to the user

[0171] What it does: The device will display the feedback as a pop-up notification for the user to view immediately.

[0172] Step 9:

[0173] The server manages user authentication information, such as authentication when a user logs into the system, and stores presentation history and past feedback results for future reference.

[0174] Input: User authentication information, presentation history, feedback results

[0175] Output: Authenticated user information, saved history and feedback

[0176] Specific operation: The server authenticates the user ID and password and stores each user's presentation history and feedback results in a database.

[0177] Step 10:

[0178] The server generates questions based on the user's presentation and presents them to the user via the terminal. The user answers the questions, and the answers are also analyzed to provide more detailed feedback.

[0179] Input: Presentation content, user responses

[0180] Output: Generated questions, analysis results and additional feedback

[0181] Specific operation: The server generates questions such as "What are the advantages of the new product?" and presents them to the user through the terminal. If the user answers "The advantage is multifunctionality," the answer is analyzed and additional feedback is generated.

[0182] (Application example 1)

[0183] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0184] Previous systems for improving presentation skills only provided feedback specific to presentations, which meant they were not particularly practical for customer service. Customer service skills also depend on a variety of important factors, such as voice, facial expressions, gestures, eye contact, and slides, but there was no system that could comprehensively analyze these and provide specific feedback to customer service staff in real time. Furthermore, while improving the ability to respond to questions in customer service is also important, previous technologies were inadequate in this regard as well.

[0185] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0186] In this invention, the server includes a user interface means for allowing a user to input voice, facial expressions, gestures, slides, and eye contact; a data transmission means for transmitting data input from the user interface means to the server; an audio analysis means for converting the voice data received by the server into text and evaluating the user's speaking speed, volume, and pauses; a video analysis means for analyzing the user's facial expressions, eye contact, and gestures from the video data received by the server and evaluating their level of composure and confidence; a slide analysis means for analyzing the slides received by the server and evaluating the consistency of content, visual effect, and typos; a feedback generation means for generating feedback to the user based on the evaluation results of the audio analysis means, video analysis means, and slide analysis means; a feedback transmission means and a visual output means for transmitting the feedback generated by the feedback generation means to the user's terminal; and a question generation means for generating questions based on the user's customer service content and presenting them to the user. This allows customer service staff to receive specific feedback in real time on various aspects of customer service, such as voice, facial expressions, gestures, eye contact, and consistency of content. The question generation means also contributes to improving customer service skills.

[0187] "User interface means" refers to the functions of devices and software that allow the wait staff to input voice, facial expressions, gestures, slides, and eye movements.

[0188] The "data transmission means" is a mechanism for transmitting data input from the user interface means to the server.

[0189] The "voice analysis means" is a function that converts the voice data received by the server into text and evaluates the speaking speed, volume, and pauses.

[0190] The "video analysis means" is a function that analyzes the facial expressions, gaze, and gestures of the customer service staff from the video data received by the server, and evaluates their level of composure and confidence.

[0191] The "slide material analysis means" is a function that analyzes the slide materials received by the server and evaluates the consistency of the content, the visual effect, and typographical errors.

[0192] The "feedback generation means" is a mechanism for generating feedback to the customer service staff based on the evaluation results of the audio analysis means, video analysis means, and slide data analysis means.

[0193] The "feedback sending means" is a function that sends the generated feedback to the terminal of the customer service staff member.

[0194] "Visual output means" refers to a device or function that visually displays feedback on a terminal.

[0195] The "question generation means" is a function in which the server generates questions based on the content of customer service and presents them to the customer service staff.

[0196] The system for improving customer service skills according to the present invention can be implemented as follows. The system is composed of a customer service staff member, a terminal, and a server. The customer service staff member is a trainee and interfaces with the system through the terminal. The terminal serves to collect the customer service staff member's voice, facial expressions, gestures, eye movements, and slide materials and transmit them to the server. The server analyzes the received data and provides feedback to the customer service staff member.

[0197] Basic configuration

[0198] The system includes the following basic means:

[0199] 1. User interface means: Devices and software that allow wait staff to input voice, facial expressions, gestures, slides, and eye movements. Equipped with a voice input device (microphone), a video input device (camera), and an eye movement detector, the wait staff can input voice, facial expressions, gestures, eye movements, and slides into the system.

[0200] 2. Data transmission means: The terminal has the function to transmit the audio, video and slides input by the customer service staff to the server. This process is performed in real time, and the data is buffered and sent to the server.

[0201] 3. Speech analysis: The server analyzes the received voice data. Specifically, it converts the voice into text using Google Cloud Speech-to-Text, and evaluates the speaking speed, volume, and pauses based on the text data. For example, if a customer service staff member speaks too quickly, this will be reflected in the feedback.

[0202] 4. Video analysis: The server analyzes the received video data. Using computer vision technologies such as OpenCV and dlib, it analyzes the facial expressions, gaze, and gestures of the wait staff, and evaluates their level of composure and confidence based on the analysis. For example, if the wait staff is not making eye contact with the viewer, this will be included in the feedback.

[0203] 5. Slide Analysis: The server analyzes the received slides. It uses optical character recognition (OCR) technology to read the content of the slides and evaluates whether they are consistent, visually appropriate, and free of typos. For example, if a slide contains a typographical error, this is reflected in the feedback.

[0204] 6. Feedback generation means and feedback transmission means: The server generates feedback to be provided to the customer service staff based on the evaluation results obtained by the audio analysis means, video analysis means, and slide analysis means. The generated feedback is sent to the terminal and visually displayed to the customer service staff. This feedback may include adjustments to speaking speed and volume, improvements to facial expressions and gestures, and corrections to slides.

[0205] 7. Question generation means and additional feedback generation means: The server generates questions based on the customer service content and presents them to the customer service staff. When the customer service staff answers these questions, the answers are also analyzed and additional feedback is generated. Questions are generated using a generative AI model such as OpenAI's GPT-3, and the customer service staff's answers are analyzed. For example, questions are generated on the theme of "explaining new products," and the necessary feedback is provided to the customer service staff when they answer them.

[0206] Specific examples

[0207] For example, if a customer service staff member practices on the theme of "explaining a new product," the system functions as follows: The customer service staff member practices using a terminal, and the server uses its audio analysis means to evaluate that "their voice is too quiet." It also uses its video analysis means to determine that "their eyes are always looking toward the product," and its slide analysis means to analyze that "there are typos in some of the slides." The server generates feedback based on these evaluations and provides the customer service staff with specific advice such as "increase the volume of your voice when speaking," "make eye contact with the customer," and "correct typos in the slides."

[0208] Example prompt sentence:

[0209] "Our customer service staff are explaining a new product. What advice would be effective? Please also give us specific examples of areas where our staff can improve."

[0210] In this way, the system of the present invention functions as an effective tool for customer service staff to efficiently improve their customer service skills.

[0211] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0212] Step 1:

[0213] The customer service staff uses the microphone, camera, and gaze detector of the terminal to emit voice and input facial expressions, gestures, gaze, and slides. The input data includes voice data, video data, gaze data, and slides.

[0214] Step 2:

[0215] The terminal buffers the voice data, video data, eye movement data, and slide material input by the customer service staff in real time and transmits them to the server using the data transmission means, whereby the voice, video, eye movement, and slide material data are input to the server.

[0216] Step 3:

[0217] The server converts the received voice data into text data using Google Cloud Speech-to-Text. The input voice data is analyzed and evaluated for speaking speed, volume, and pauses. For example, if the voice data indicates that the speaker is speaking too fast, an evaluation result will be generated.

[0218] Step 4:

[0219] The server uses OpenCV and dlib to analyze facial expressions, eye movements, and gestures from the received video data. It receives the video data as evaluation input and evaluates the customer's level of composure and confidence. For example, it generates an evaluation result such as "the customer's eyes are always directed toward the products."

[0220] Step 5:

[0221] The server uses optical character recognition (OCR) technology to read the content of the received slides and evaluates the consistency of the content, visual effect, and typographical errors. It receives the slides as input data and generates the evaluation result: "Some slides contain typographical errors."

[0222] Step 6:

[0223] The server generates feedback for the customer service staff based on the evaluation results obtained from the audio analysis means, video analysis means, and slide analysis means. For example, the generated feedback includes specific advice such as "increase your speaking volume," "make eye contact with the customer," and "correct typos in the slides."

[0224] Step 7:

[0225] The feedback generated by the feedback generating means is sent to the terminal via the feedback transmitting means and the visual output means. As feedback, information such as "increase your speaking volume" and "make eye contact with the customer" is displayed to the wait staff.

[0226] Step 8:

[0227] The server uses a generative AI model such as OpenAI's GPT-3 to generate questions based on the customer service content. The server generates appropriate questions based on the scenario the customer service staff is practicing and sends them to the device. For example, a question might be generated such as, "I'm explaining a new product. What kind of advice would be effective?"

[0228] Step 9:

[0229] The customer service staff responds to questions posed by the system and sends the answers to the server via their terminal. The server analyzes the received answers and generates additional feedback based on the evaluation results.

[0230] Step 10:

[0231] The server transmits the additional feedback generated by the additional feedback generating means to the terminal using the feedback transmitting means and the visual output means, thereby providing further specific advice to the wait staff.

[0232] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0233] The system for improving presentation skills according to the present invention can provide more advanced feedback by incorporating an emotion engine that analyzes the user's emotions. Specific embodiments will be described below.

[0234] Basic configuration

[0235] The system is composed of a user, a terminal, and a server. The user is the person giving the presentation and interfaces with the system through the terminal. The terminal collects the user's voice, facial expressions, gestures, and slides and transmits them to the server. The server analyzes the received data and provides feedback to the user. In addition, this embodiment includes an emotion engine that recognizes the user's emotions.

[0236] User Interface Means

[0237] A user gives a presentation using a terminal. The terminal is equipped with an audio input device (microphone), a video input device (camera), and a screen for displaying slides. An emotion engine is also installed, which analyzes the user's emotions.

[0238] Data transmission method

[0239] The device has the ability to transmit audio, video, and slides input by the user to the server. This process is performed in real time, and the data is buffered and sent to the server. The user's emotional data is also transmitted in the same way.

[0240] Voice analysis methods

[0241] The server analyzes the received audio data, converts it into text using speech recognition technology, and evaluates the speed, volume, and pauses of speech based on the text data. For example, if the user speaks too quickly, this will be reflected in the feedback.

[0242] Video analysis methods

[0243] The server analyzes the received video data, using computer vision technology to analyze the user's facial expressions, eye contact, and gestures, and evaluates the presentation's poise and confidence. For example, if the user is not making eye contact with the audience, this will be included in the feedback.

[0244] Slide analysis tools

[0245] The server analyzes the slides it receives, using optical character recognition (OCR) technology to read their content and evaluate whether they are consistent with the presentation theme and whether their visual impact is appropriate. For example, if a slide contains a typographical error, that error is reflected in the feedback.

[0246] Emotion Engine

[0247] The server uses an emotion engine to analyze the user's emotional data. Specifically, it analyzes the user's facial expressions and voice to evaluate the fluctuations and stability of their emotions during the presentation. For example, if the user is nervous during the presentation, that emotional data will be included in the feedback.

[0248] Feedback generation means and feedback transmission means

[0249] The server integrates the evaluation results of the audio analysis means, video analysis means, slide analysis means, and emotion engine to generate feedback for the user. The generated feedback is sent to the terminal and displayed to the user. The feedback includes advice on adjusting speaking speed and volume, improving facial expressions and gestures, correcting slides, and emotional advice.

[0250] Authentication Management Methods

[0251] The server manages user authentication information and stores the user's presentation history and feedback results, allowing users to check their progress and refer to past feedback.

[0252] Question generation and additional feedback generation

[0253] The server generates questions based on the user's presentation and presents them to the user. When the user answers these questions, the answers are also analyzed and additional feedback is generated. For example, if the user does not answer a question accurately, that point is reflected in the additional feedback.

[0254] Specific examples

[0255] For example, if a user gives a presentation on the theme of "introducing a new product," the system functions as follows: The user gives the presentation using a device, and the server uses its audio analysis means to evaluate the user's voice as "quiet." The video analysis means determines that the user's eyes are always looking at the slides, and the slide analysis means analyzes that there are typos on some of the slides. The emotion engine then determines that the user is nervous. The server generates feedback based on these evaluations and provides the user with specific advice such as "increase the volume of your voice," "make eye contact with the audience," "correct typos in the slides," and "regulate your breathing to relieve tension."

[0256] In this way, the system of the present invention functions as an effective tool for users to efficiently improve their presentation skills. In particular, the introduction of an emotion engine makes it possible to provide comprehensive feedback that takes into account the user's emotional state.

[0257] The processing flow will be explained below.

[0258] Step 1:

[0259] A user logs into the system using a terminal, enters a username and password, and submits the authentication information.

[0260] Step 2:

[0261] The device sends the entered authentication information to the server. The server compares it with the database and returns the authentication result to the device. If authentication is successful, the user can start the presentation.

[0262] Step 3:

[0263] The user begins preparing for a presentation. They upload presentation materials (slides) using their device. The device then sends the uploaded materials to the server.

[0264] Step 4:

[0265] The server stores the received presentation materials in an appropriate format.

[0266] Step 5:

[0267] The user presses the "Start Presentation" button. The device begins recording the user's voice and video (facial expressions and gestures) using the built-in camera and microphone. The recorded data is buffered.

[0268] Step 6:

[0269] The terminal transmits the buffered data to the server in real time, including audio, video, slides, and emotion data.

[0270] Step 7:

[0271] The server analyzes the received voice data, converts the voice data into text using a speech recognition engine, evaluates speaking speed, volume, and pauses, and generates text data.

[0272] Step 8:

[0273] The server analyzes the received video data and uses computer vision technology to analyze the user's facial expressions, eye movements, and gestures to evaluate the presentation's poise and confidence.

[0274] Step 9:

[0275] The server analyzes the slides received, using optical character recognition (OCR) technology to read the content of the slides and evaluate their consistency and visual impact.

[0276] Step 10:

[0277] The server uses an emotion engine to analyze the user's emotional data, identifying emotions from facial expressions and vocal tones, and assessing emotional stability and fluctuations.

[0278] Step 11:

[0279] The server integrates the evaluation results obtained from the audio analysis means, video analysis means, slide analysis means, and emotion engine, and generates feedback for the user, including advice on speaking speed, volume, naturalness of gestures, corrections to slides, and emotional advice.

[0280] Step 12:

[0281] The server sends the generated feedback to the user's terminal, which receives the feedback and displays it to the user.

[0282] Step 13:

[0283] Users can review the feedback and incorporate improvements for their next presentation. Users can also request a Q&A simulation if necessary.

[0284] Step 14:

[0285] If a question and answer simulation is desired, the server generates questions based on the content of the presentation and presents them to the user.

[0286] Step 15:

[0287] The user answers the questions, and the device sends the answer data to the server.

[0288] Step 16:

[0289] The server analyzes the user's answers and generates additional feedback, providing further improvements based on the evaluation results.

[0290] Step 17:

[0291] The server sends additional feedback to the user's device, which the user receives and uses as a reference for their next presentation.

[0292] This series of processing steps allows users to improve their presentation skills efficiently and effectively. In addition, the introduction of an emotion engine makes it possible to provide comprehensive feedback that takes emotional aspects into consideration.

[0293] Example 2

[0294] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0295] Current presentation skill improvement systems have difficulty accurately analyzing users' emotions and reflecting the analysis results in presentation feedback. Furthermore, they lack feedback on not only the technical aspects of the presentation but also the emotional and psychological aspects, making it difficult for users to comprehensively improve their skills. Furthermore, existing systems lack sufficient functionality for generating questions and providing additional feedback, making it difficult for users to self-evaluate.

[0296] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0297] In this invention, the server includes an audio analysis means, a video analysis means, a slide analysis means, an emotion engine, a feedback generation means, and a feedback transmission means. This allows for comprehensive analysis of a user's voice, facial expressions, gestures, slides, and even emotions, making it possible to provide detailed and accurate feedback. Furthermore, by including a question generation means and an additional feedback generation means, it is possible to automatically generate questions based on the content of the user's presentation and evaluate the answers to those questions, facilitating self-evaluation.

[0298] "User interface means" refers to devices and software that allow a user to input voice, facial expressions, gestures, and slide materials.

[0299] "Data transmission means" refers to a device or software for transmitting data input from the user interface means to the server.

[0300] "Voice analysis means" refers to devices or software that convert the voice data received by the server into text and evaluate speaking speed, volume, and pauses.

[0301] "Video analysis means" refers to devices or software that analyze the user's facial expressions, gaze, and gestures from the video data received by the server, and evaluate the presentation's level of composure and confidence.

[0302] "Slide material analysis means" refers to a device or software that analyzes the slide materials received by the server and evaluates the consistency of the content and the visual effect.

[0303] "Emotion engine" refers to a device or software that allows the server to analyze emotion data and evaluate the fluctuations and stability of a user's emotions.

[0304] "Feedback generation means" refers to a device or software for generating feedback to a user based on the evaluation results of the audio analysis means, video analysis means, slide data analysis means, and emotion engine.

[0305] The "feedback sending means" refers to a device or software for sending the feedback generated by the feedback generating means to the user's terminal.

[0306] "Authentication management means" refers to a device or software for managing user authentication information and storing each user's presentation history and feedback results.

[0307] The "question generation means" refers to a device or software that generates questions based on the content of a user's presentation and presents them to the user.

[0308] The "additional feedback generating means" refers to a device or software that analyzes the answers given by the user to the questions presented and generates additional feedback based on the evaluation results.

[0309] The system for improving presentation skills according to the present invention provides more advanced feedback by incorporating an emotion engine that analyzes the user's emotions. Specific embodiments will be described below.

[0310] Basic configuration

[0311] The system is composed of a user, a terminal, and a server. The user is the person giving the presentation and interfaces with the system through the terminal. The terminal collects the user's voice, facial expressions, gestures, and slides and transmits them to the server. The server analyzes the received data and provides feedback to the user. In addition, this embodiment includes an emotion engine that recognizes the user's emotions.

[0312] User Interface Means

[0313] A user gives a presentation using a terminal. The terminal is equipped with an audio input device (microphone), a video input device (camera), and a screen for displaying slides. An emotion engine is also installed, which analyzes the user's emotions.

[0314] Data transmission method

[0315] The device has the ability to transmit audio, video, and slides input by the user to the server. This process is performed in real time, and the data is buffered and sent to the server. The user's emotional data is also transmitted in the same way.

[0316] Voice analysis methods

[0317] The server analyzes the received audio data, converts it into text using speech recognition technology, and evaluates speaking speed, volume, and pauses based on the text data. The specific software used is the Google Cloud Speech-to-Text API. For example, if the user speaks too quickly, this is reflected in the feedback.

[0318] Video analysis methods

[0319] The server analyzes the received video data. It uses computer vision technology to analyze the user's facial expressions, eye movements, and gestures, and evaluates the presentation's poise and confidence based on the results. The OpenCV library is used for this software. For example, if the user is not making eye contact with the audience, this will be included in the feedback.

[0320] Slide analysis tools

[0321] The server analyzes the received slides. It uses optical character recognition (OCR) technology to read the content of the slides and evaluate whether they are consistent with the presentation theme and have an appropriate visual effect. Specific software includes Tesseract OCR. For example, if a slide contains a typographical error, that error is reflected in the feedback.

[0322] Emotion Engine

[0323] The server uses an emotion engine to analyze the user's emotional data. Specifically, it analyzes the user's facial expressions and voice to evaluate the fluctuations and stability of their emotions during the presentation. For example, if the user is nervous during the presentation, that emotional data will be included in the feedback.

[0324] Feedback generation means and feedback transmission means

[0325] The server integrates the evaluation results of the audio analysis means, video analysis means, slide analysis means, and emotion engine to generate feedback for the user. The generated feedback is sent to the terminal and displayed to the user. The feedback includes advice on adjusting speaking speed and volume, improving facial expressions and gestures, correcting slides, and emotional advice.

[0326] Authentication Management Methods

[0327] The server manages user authentication information and stores the user's presentation history and feedback results, allowing users to check their progress and refer to past feedback.

[0328] Question generation and additional feedback generation

[0329] The server generates questions based on the user's presentation and presents them to the user. When the user answers these questions, the answers are also analyzed and additional feedback is generated. For example, if the user does not answer a question accurately, that point is reflected in the additional feedback.

[0330] Specific examples

[0331] For example, if a user is giving a presentation on the theme of "introducing a new product," the system functions as follows: The user gives the presentation using a terminal, and the server uses the audio analysis means to evaluate the user's voice as "quiet." The video analysis means determines that the user's eyes are always directed toward the slides, and the slide analysis means analyzes that there are typos on some slides. The emotion engine then determines that the user is nervous. Based on these evaluations, the server generates feedback and provides the user with specific advice, such as "increase your speaking volume," "make eye contact with the audience," "correct typos in the slides," or "regulate your breathing to relieve tension." In this way, the system of the present invention functions as an effective tool for users to efficiently improve their presentation skills. In particular, the introduction of the emotion engine makes it possible to provide comprehensive feedback that takes the user's emotional state into consideration.

[0332] Prompt Sentence Examples

[0333] Here are some example prompts to input to a generative AI model:

[0334] A user will give a presentation introducing a new product. Staff will generate feedback based on the analysis of audio, video, and slides. Also, analyze the user's emotional data and include it in the feedback. Examples of feedback include "your voice is too quiet," "you should make eye contact with the audience," "there is a typo on the slide," and "you seem nervous."

[0335] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0336] Specific processing flow of the program

[0337] Step 1: Start your presentation

[0338] The user clicks a button on the device to start the presentation. The device verifies the user's authentication information and sends a request to the server to start the session.

[0339] Input: User credentials, session initiation request

[0340] Processing: The server verifies the authentication information and generates a token to start the session.

[0341] Output: Session initiation token

[0342] Step 2: Data collection

[0343] The user gives a presentation, and the device collects audio, video, and slides via a microphone, camera, and screen display.

[0344] Input: User's audio, video, and slides

[0345] Processing: The device buffers the data in real time and prepares it for transfer to the server.

[0346] Output: Buffered data

[0347] Step 3: Send data

[0348] The terminal transmits the collected data (audio, video, slides, and emotional data) to the server in real time.

[0349] Input: Buffered audio, video, slide deck, emotion data

[0350] Processing: The device divides the data into packets and sends them over the network to the server.

[0351] Output: Data packets arriving at the server

[0352] Step 4: Audio data analysis

[0353] The server analyzes the received voice data, converts the voice into text using speech recognition technology, and evaluates the speaking speed, volume, and pauses based on the text data.

[0354] Input: Audio data

[0355] Processing: Converts audio to text using the Google Cloud Speech-to-Text API and analyzes data on speaking rate, volume, and pauses.

[0356] Output: Text data, speaking speed, volume, and interval evaluation data

[0357] Step 5: Video data analysis

[0358] The server analyzes the received video data, using facial recognition technology and computer vision to analyze the user's facial expressions, gaze, and gestures, and evaluates the presentation's poise and confidence based on that information.

[0359] Input: Video data

[0360] Processing: Using the OpenCV library, we analyze the user's facial expressions, gaze, and gestures to assess their calmness and confidence.

[0361] Output: Facial expression, gaze, gesture analysis data, calmness and confidence evaluation data

[0362] Step 6: Slide analysis

[0363] The server analyzes the received slides, using optical character recognition (OCR) technology to read the slide content and evaluate whether it is consistent with the presentation theme and whether the visual effects are appropriate.

[0364] Input: Slide data

[0365] Processing: Extract text from slides using Tesseract OCR and evaluate content consistency and visual impact.

[0366] Output: Slide text data, consistency and visual effect evaluation data

[0367] Step 7: Sentiment Data Analysis

[0368] The server uses an emotion engine to analyze the user's emotion data. Specifically, it analyzes emotions from the user's facial expressions and voice and evaluates the fluctuations and stability of emotions during the presentation.

[0369] Input: Emotion data (facial expressions, voice)

[0370] Processing: An emotion engine is used to analyze facial and voice data to assess emotional variability and stability.

[0371] Output: Emotional evaluation data

[0372] Step 8: Feedback Generation

[0373] The server integrates the results of the audio analysis, video analysis, slide analysis, and emotion data analysis, and generates feedback for the user.

[0374] Input: Audio analysis data, video analysis data, slide analysis data, emotion analysis data

[0375] Processing: Comprehensive feedback is generated based on the results of each assessment. The feedback is provided in text and graphical format.

[0376] Output: Feedback data (text, graph)

[0377] Step 9: Send and view feedback

[0378] The server sends the generated feedback to the terminal, which displays it to the user.

[0379] Input: Feedback data

[0380] Processing: The server sends the feedback data to the terminal, which displays it to the user. The feedback includes specific advice.

[0381] Output: Feedback that is displayed to the user

[0382] Step 10: Question generation and additional feedback

[0383] The server generates relevant questions based on the user's presentation and presents them to the user via the terminal. When the user answers the questions, the server analyzes them again and generates additional feedback.

[0384] Input: Presentation content data, user questions and answers

[0385] Processing: A question generation algorithm is used to generate questions based on the presentation content, and the answers are analyzed and evaluated. Additional feedback is generated based on the evaluation results.

[0386] Output: Additional feedback data

[0387] Step 11: Authentication Management and History Storage

[0388] The server manages user authentication information and feedback history, and stores them in a database after the presentation session ends.

[0389] Input: User authentication information, feedback history data

[0390] Processing: Manages authentication information and stores feedback history in a database.

[0391] Output: Updated database

[0392] Through this process, users receive feedback to comprehensively improve their presentation skills.

[0393] (Application example 2)

[0394] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0395] Conventional presentation skill improvement systems are primarily limited to analyzing voice, facial expressions, and slides, and are not capable of improving specialized skills tailored to specific environments or situations. For example, to improve a factory operator's skills so that they can issue appropriate commands to a robot, analysis and feedback that takes into account the appropriateness of the robot's operation instructions and the operator's emotional state is necessary. A system that can address this issue is needed.

[0396] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a user interface means for a user to input voice, facial expressions, gestures, and slide materials, a data transmission means for transmitting the data input from the user interface means to the server, a voice analysis means for converting voice data received by the server into text and evaluating the speaking speed, volume, and pauses, a video analysis means for analyzing the user's facial expressions, line of sight, and gestures from video data received by the server and evaluating the calmness and confidence of the presentation, a slide material analysis means for analyzing the slide materials received by the server and evaluating the consistency of the content and the visual effect, a feedback generation means for generating feedback to the user based on the evaluation results of the voice analysis means, the video analysis means, and the slide material analysis means, a feedback transmission means for transmitting the feedback generated by the feedback generation means to the user's terminal, a robot operation instruction evaluation means for evaluating the appropriateness of instructions given by the operator to the robot, and an emotion evaluation means for analyzing the operator's emotional state and evaluating the mental state that affects the appropriateness of the instructions. This will enable factory operators to improve their skills in operating robots properly.

[0397] "User interface means" refers to devices and software that allow a user to input voice, facial expressions, gestures, and slide materials.

[0398] The "data transmission means" is a communication means for transmitting data input from the user interface means to the server in real time.

[0399] The "voice analysis means" is a technology that converts the voice data received by the server into text and evaluates the speaking speed, volume, and pauses.

[0400] The "video analysis means" is a technology that analyzes the user's facial expressions, gaze, and gestures from the video data received by the server, and evaluates the presentation's level of composure and confidence.

[0401] The "slide material analysis means" is a technology for analyzing the slide materials received by the server and evaluating the consistency of the content and the visual effect.

[0402] The "feedback generation means" is a technology that generates feedback indicating areas for improvement to the user based on the evaluation results of the audio analysis means, video analysis means, and slide data analysis means.

[0403] The "feedback transmitting means" is a means for transmitting the feedback generated by the feedback generating means to the user's terminal.

[0404] The "robot operation instruction evaluation means" is a technology for evaluating the appropriateness of instructions given by an operator to a robot.

[0405] "Emotion evaluation means" is a technology that analyzes the emotional state of an operator and evaluates the mental state that affects the appropriateness of instructions.

[0406] "Authentication management means" is a technology that enables the server to manage user authentication information and store each user's presentation history, instruction history, and feedback results.

[0407] The "question generation means" is a technology for generating questions based on the content of the user's presentation and instructions for operating the robot, and presenting the questions to the user.

[0408] The "additional feedback generating means" is a technology that analyzes the answers given by the user to the questions presented and generates additional feedback based on the evaluation results.

[0409] The system according to the present invention is intended to improve the skills of factory operators to properly operate robots. This system is composed of a user (operator), a terminal, and a server. Specific embodiments are described below.

[0410] Basic configuration

[0411] The system includes a user interface for inputting voice, facial expressions, gestures, and instructions from the user's terminal. These data are transmitted in real time from the terminal to a server, which analyzes the received data and provides feedback to the user.

[0412] Voice analysis methods

[0413] The server first receives the user's voice data and converts it into text using speech recognition software (for example, Google's speech recognition API). The server then evaluates the user's speaking speed, volume, and pauses. Any inappropriate parts are reflected in the feedback.

[0414] Video analysis methods

[0415] The server then analyzes the user's video data. Using video analysis technology (e.g., OpenCV), it detects and evaluates the user's facial expressions, eye movements, and gestures. The server evaluates the user's poise and confidence during the presentation, and provides feedback on any shortcomings.

[0416] Slide analysis tools

[0417] The server analyzes the slide deck it receives and uses optical character recognition (OCR) technology to review the slide content and evaluate its consistency and visual effectiveness. For example, typographical errors and visual flaws are flagged.

[0418] Emotion assessment tools

[0419] The server then uses an emotion engine to analyze the user's emotional data. It extracts emotions from the user's facial expressions and voice, and evaluates the emotional fluctuations and stability during the presentation. If tension or stress is detected, this is also reflected in the feedback.

[0420] Robot operation instruction evaluation means

[0421] In a specific application, the server evaluates the appropriateness of the instructions given by the operator to the robot, for example, whether the instructions are clear and unambiguous when given by voice or gesture.

[0422] Feedback generation and transmission methods

[0423] The server then aggregates all of these analysis results and generates specific feedback for the user, suggesting areas for improvement. The generated feedback is then sent to the user's device and displayed to the user.

[0424] Authentication Management Methods

[0425] The server also manages user authentication information and stores each user's presentation history, instruction history, and feedback results, allowing users to improve themselves based on past data.

[0426] Question generation and additional feedback generation means

[0427] The server generates questions based on the user's presentation and robot operation instructions, and presents them to the user. When the user answers these questions, the server analyzes the answers and generates additional feedback.

[0428] Specific examples

[0429] For example, imagine a scene in which an operator is giving verbal instructions to a robot in a factory. The scene is recorded using the user interface means, and the audio and video data are sent to the server. The server analyzes the audio and judges it as "too quiet," and analyzes the video and judges it as "tension is evident in the operator's facial expression." Furthermore, the robot operation instruction evaluation means judges it as "ambiguous." Based on these evaluations, the server generates specific feedback and provides the user with advice such as "speak louder" or "give clearer instructions."

[0430] Prompt Sentence Examples

[0431] "Analyze sample videos, evaluate the operator's emotional state and the appropriateness of their voice instructions, and provide feedback."

[0432] This system allows factory operators to effectively improve their robot operation skills, aiming to efficiently improve productivity.

[0433] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0434] Step 1: Data entry

[0435] The user inputs voice, facial expressions, gestures, and instructions. Specifically, data is collected using user interface means (camera, microphone, touch panel, etc.). The inputs in this step are the user's voice data, video data, and instructions. The output is that these data are stored in the terminal.

[0436] Step 2: Send data

[0437] The terminal transmits the voice data, video data, and instruction contents input from the user interface means to the server. Specifically, the data is transmitted to the server in real time via the Internet. The input of this step is the user's voice data, video data, and instruction contents, and the output is the server receiving these data.

[0438] Step 3: Audio analysis

[0439] The server analyzes the received voice data. Specifically, it uses speech recognition software (such as Google's speech recognition API) to convert the speech into text and evaluates the speaking speed, volume, and pauses. The input of this step is the received voice data, and the output is the analyzed text data and the evaluation results.

[0440] Step 4: Video analysis

[0441] The server analyzes the received video data. Specifically, it uses video analysis technology (such as OpenCV) to detect the user's facial expressions, eye movements, and gestures, and evaluates their level of composure and confidence during the presentation. The input to this step is the received video data, and the output is the analyzed facial expression data, eye movement data, and gesture data, along with the evaluation results.

[0442] Step 5: Slide analysis

[0443] The server analyzes the received slides. Specifically, it uses optical character recognition (OCR) technology to review the content of the slides and evaluate their consistency and visual effectiveness. The input for this step is the received slides, and the output is the analyzed text data and the evaluation results.

[0444] Step 6: Emotional assessment

[0445] The server evaluates emotions using the received video and audio data. Specifically, it uses emotion recognition technology (such as Face++ or Microsoft Azure Face API) to evaluate the user's emotional state and evaluate the fluctuations and stability of emotions during the presentation. The input for this step is facial expression data and audio data, and the output is analyzed emotional data and its evaluation results.

[0446] Step 7: Evaluate robot operation instructions

[0447] The server evaluates the operator's instruction data. Specifically, it uses voice recognition and video analysis to evaluate the appropriateness of the operator's instructions to the robot. The input for this step is the operator's voice data and gesture data, and the output is the analysis result.

[0448] Step 8: Feedback Generation

[0449] The server integrates all the above analysis results and generates specific feedback for the user indicating areas for improvement. Feedback is generated based on the overall evaluation results using a feedback generation means. The inputs for this step are the audio analysis results, video analysis results, slide analysis results, emotion evaluation results, and robot operation instruction evaluation results, and the output is the generated feedback information.

[0450] Step 9: Submit your feedback

[0451] The server sends the generated feedback to the user's terminal, specifically, provides the feedback information to the user via the Internet. The input of this step is the generated feedback information, and the output is the feedback being displayed on the user's terminal.

[0452] Step 10: Question generation and additional feedback generation

[0453] The server generates additional questions based on the user's presentation and instructions and presents them to the user. The server further analyzes the user's answers and generates additional feedback, which helps the user identify areas for further improvement. The input of this step is the user's answer data, and the output is the additional feedback information.

[0454] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0455] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0456] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0457] [Second embodiment]

[0458] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0459] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0460] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0461] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0462] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0463] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0464] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0465] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0466] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0467] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0468] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0469] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0470] The system for improving presentation skills according to the present invention can be implemented as follows.

[0471] Basic configuration

[0472] The system consists of a user, a terminal, and a server. The user is the person giving the presentation and interfaces with the system through the terminal. The terminal collects the user's voice, facial expressions, gestures, and slides and transmits them to the server. The server analyzes the received data and provides feedback to the user.

[0473] User Interface Means

[0474] A user gives a presentation using a terminal. The terminal is equipped with an audio input device (microphone), a video input device (camera), and a screen for displaying slides. This allows the user to input voice, facial expressions, gestures, and slides into the system.

[0475] Data transmission method

[0476] The terminal has the function of transmitting the audio, video, and slides input by the user to the server. This process is performed in real time, and the data is buffered and sent to the server.

[0477] Voice analysis methods

[0478] The server analyzes the received audio data. Specifically, it uses speech recognition technology to convert the audio into text, and then evaluates the speaking speed, volume, and pauses based on the text data. For example, if the user speaks too fast, this will be reflected in the feedback.

[0479] Video analysis methods

[0480] The server analyzes the received video data, using computer vision technology to analyze the user's facial expressions, eye contact, and gestures, and evaluates the presentation's poise and confidence. For example, if the user is not making eye contact with the audience, this will be included in the feedback.

[0481] Slide analysis tools

[0482] The server analyzes the slides it receives, using optical character recognition (OCR) technology to read their content and evaluate whether they are consistent with the presentation theme and whether the visual effect is appropriate. For example, if a slide contains a typographical error, that will be reflected in the feedback.

[0483] Feedback generation means and feedback transmission means

[0484] The server generates feedback to be provided to the user based on the evaluation results obtained by the audio analysis means, video analysis means, and slide analysis means. The generated feedback is sent to the terminal and displayed to the user. The feedback may include adjustments to speaking speed and volume, improvements to facial expressions and gestures, and corrections to slides.

[0485] Authentication Management Methods

[0486] The server manages user authentication information and stores the user's presentation history and feedback results, allowing users to check their progress and refer to past feedback.

[0487] Question generation and additional feedback generation

[0488] The server generates questions based on the user's presentation and presents them to the user. When the user answers these questions, the answers are also analyzed and additional feedback is generated. For example, if the user does not answer a question accurately, that point is reflected in the additional feedback.

[0489] Specific examples

[0490] For example, if a user gives a presentation on the theme of "introducing a new product," the system functions as follows: The user gives the presentation using a device, and the server uses its audio analysis means to evaluate that the user's voice is "quiet." The video analysis means determines that the user's eyes are always looking at the slides, and the slide analysis means analyzes that there are typos on some of the slides. The server generates feedback based on these evaluations and provides the user with specific advice such as "increase the volume of your voice when speaking," "make eye contact with the audience," and "correct typos in the slides."

[0491] In this way, the system of the present invention functions as an effective tool for users to efficiently improve their presentation skills.

[0492] The processing flow will be explained below.

[0493] Step 1:

[0494] A user logs into the system using a terminal, enters a username and password, and submits the authentication information.

[0495] Step 2:

[0496] The device sends the entered authentication information to the server. The server compares it with the database and returns the authentication result to the device. If authentication is successful, the user can start the presentation.

[0497] Step 3:

[0498] The user begins preparing for a presentation. They upload presentation materials (slides) using their device. The device then sends the uploaded materials to the server.

[0499] Step 4:

[0500] The server stores the received presentation materials in an appropriate format.

[0501] Step 5:

[0502] The user presses the "Start Presentation" button. The device begins recording the user's voice and video (facial expressions and gestures) using the built-in camera and microphone. The recorded data is buffered.

[0503] Step 6:

[0504] The terminal transmits the buffered data to the server in real time. The data for the audio, video, and slide materials are transmitted to the server.

[0505] Step 7:

[0506] The server analyzes the received voice data, converts it into text using a speech recognition engine, and evaluates speaking speed, volume, and pauses.

[0507] Step 8:

[0508] The server analyzes the received video data and uses computer vision technology to analyze the user's facial expressions, eye movements, and gestures to evaluate the presentation's poise and confidence.

[0509] Step 9:

[0510] The server analyzes the slides it receives, using optical character recognition (OCR) technology to read their content and evaluate whether they are consistent with the presentation theme and whether their visual impact is appropriate.

[0511] Step 10:

[0512] The server integrates the results of the audio, video, and slide analysis and generates feedback for the user. Based on each evaluation result, it summarizes specific improvements and advice in written form.

[0513] Step 11:

[0514] The server sends the generated feedback to the user's terminal, which receives the feedback and displays it to the user.

[0515] Step 12:

[0516] Users can review the feedback and incorporate improvements for their next presentation. If desired, a simulated Q&A session can also be conducted.

[0517] Step 13:

[0518] If a question and answer simulation is desired, the server generates questions based on the content of the presentation and presents them to the user.

[0519] Step 14:

[0520] The user answers the questions, and the device sends the answer data to the server.

[0521] Step 15:

[0522] The server analyzes the user's answers and generates additional feedback, and reflects suggestions for improvement in the advice based on the evaluation results.

[0523] Step 16:

[0524] The server sends additional feedback to the user's device, which the user receives and uses as a reference for their next presentation.

[0525] This series of processing steps allows users to improve their presentation skills efficiently and effectively.

[0526] Example 1

[0527] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0528] Conventional presentation support systems lack the ability to comprehensively analyze a user's voice, facial expressions, gestures, and slides to provide comprehensive feedback. They also lack the ability to generate questions based on the user's presentation content, analyze the answers, and provide additional feedback. This makes it difficult for users to effectively improve their presentation skills.

[0529] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0530] In this invention, the server includes a user interface means for a user to input voice, facial expressions, gestures, and slide materials, a data transmission means for transmitting the data input from the user interface means to the server, an audio analysis means for converting the voice data received by the server into text and evaluating the speaking speed, volume, and pauses, a video analysis means for analyzing the user's facial expressions, gaze, and gestures from the video data received by the server and evaluating the poise and confidence of the presentation, a slide material analysis means for analyzing the slide materials received by the server and evaluating the consistency and visual effect of the content, a feedback generation means for generating feedback to the user based on the evaluation results of the audio analysis means, the video analysis means, and the slide material analysis means, a feedback transmission means for transmitting the feedback generated by the feedback generation means to the user's terminal, a question generation means for the server to generate questions based on the content of the user's presentation and present them to the user, and an additional feedback generation means for analyzing the user's answers to the questions posed via the user interface means and generating additional feedback based on the evaluation results. This allows the user to receive comprehensive and detailed feedback, thereby effectively improving their presentation skills.

[0531] "User interface means" refers to means by which a user inputs voice, facial expressions, gestures, and slide materials.

[0532] The "data transmission means" is a means for transmitting data input from the user interface means to the server.

[0533] The "voice analysis means" is a means for converting voice data received by the server into text and evaluating speaking speed, volume, and pauses.

[0534] The "video analysis means" is a means for analyzing the user's facial expressions, gaze, and gestures from the video data received by the server, and evaluating the user's level of composure and confidence in the presentation.

[0535] The "slide material analysis means" is a means for analyzing the slide materials received by the server and evaluating the consistency of the content and the visual effect.

[0536] The "feedback generation means" is a means for generating feedback to the user based on the evaluation results of the audio analysis means, the video analysis means, and the slide material analysis means.

[0537] The "feedback transmitting means" is a means for transmitting the feedback generated by the feedback generating means to the user's terminal.

[0538] The "question generation means" is a means by which the server generates a question based on the content of the user's presentation and presents it to the user.

[0539] The "additional feedback generating means" is a means for analyzing the answers given by the user to questions presented via the user interface means and generating additional feedback based on the evaluation results.

[0540] The "authentication management means" is a means by which the server manages user authentication information and stores each user's presentation history and feedback results.

[0541] The system for improving presentation skills according to the present invention comprises a user, a terminal, and a server. By using this system, the user can effectively analyze their voice, facial expressions, gestures, and slides, and receive comprehensive feedback.

[0542] User Interface Means

[0543] When giving a presentation, a user uses a terminal equipped with a microphone, a camera, and a screen for displaying slides. This allows the user to input voice, facial expressions, gestures, and slides.

[0544] Data transmission method

[0545] The terminal transmits the audio, video, and slides input by the user to the server in real time. During this process, the data is buffered, allowing for smooth data transfer in a stable communication environment.

[0546] Voice analysis methods

[0547] The server analyzes the received voice data. Specifically, it converts the voice into text using speech recognition technology (e.g., Google Cloud Speech-to-Text API), and evaluates the speaking speed, volume, and pauses based on the text data. For example, if a user says "introducing a new product," if they speak too quickly, this will be reflected in the feedback.

[0548] Video analysis methods

[0549] The server analyzes the received video data. Using computer vision techniques (e.g., OpenCV and TensorFlow), it analyzes the user's facial expressions, eye movements, and gestures to assess the presentation's poise and confidence. For example, if the user's eyes are fixed on the slides and they are not making eye contact with the audience, this will be included in the feedback.

[0550] Slide analysis tools

[0551] The server analyzes the slides it receives. It uses optical character recognition (OCR) technology (e.g., Tesseract OCR) to read the content of the slides and evaluates whether they are consistent with the presentation theme and whether the visual effect is appropriate. For example, if a slide contains a typographical error, that error is reflected in the feedback.

[0552] Feedback Generation Method

[0553] The server generates feedback to be provided to the user based on the evaluation results obtained by the audio analysis means, video analysis means, and slide analysis means. The generated feedback includes adjustments to speaking speed and volume, improvements to facial expressions and gestures, and corrections to slides.

[0554] Feedback sending method

[0555] The server sends the generated feedback to the device, which then displays it to the user, who receives real-time feedback and can immediately see how to improve their presentation.

[0556] Question generation means

[0557] The server generates questions based on the content of the user's presentation and presents them to the user via the terminal, encouraging the user to consider further the presentation.

[0558] Additional Feedback Generation Means

[0559] The server analyzes the answers given by the user to the questions posed and generates additional feedback based on the evaluation results, allowing the user to receive feedback at a deeper level.

[0560] Authentication Management Methods

[0561] The server manages user authentication information and stores presentation history and past feedback results, allowing users to check their progress and continuously improve their presentation skills.

[0562] Specific examples

[0563] For example, if a user is giving a presentation on the theme of "introducing a new product," the system functions as follows: The user gives the presentation using a terminal, and the server uses the audio analysis means to evaluate that the user's voice is quiet. The video analysis means determines that the user's eyes are always looking at the slides, and the slide analysis means analyzes that there are typos in some of the slides. The server generates feedback based on these evaluations and provides the user with specific advice, such as speaking louder, making eye contact with the audience, or correcting typos in the slides. In this way, the system of the present invention functions as an effective tool for users to efficiently improve their presentation skills.

[0564] Prompt Sentence Examples

[0565] For example, if you are giving a presentation on the topic of "Introducing a New Product," you might use the following prompt:

[0566] "You will be asked to give a presentation on the topic of introducing a new product. A microphone and camera will be used to collect data, and analysis will be performed using speech recognition, computer vision, and optical character recognition (OCR). Feedback will include an evaluation of speaking speed, volume, facial expressions, eye contact, gestures, and the content of your slides."

[0567] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0568] Step 1:

[0569] The user starts the application on the device, sets the theme to "New Product Introduction," and starts a presentation. The user can input voice, facial expressions, gestures, and slides.

[0570] Input: Set presentation theme, start presentation

[0571] Output:Presentation settings information

[0572] Step 2:

[0573] The device's microphone collects the user's voice, and the camera records the user's facial expressions and gestures, as well as capturing the slides the user uses.

[0574] Input: User's voice, facial expressions, gestures, slides

[0575] Output: Collected audio data, video data, and slide data

[0576] Step 3:

[0577] The devices transmit the collected data to the server in real time, a process in which the data is buffered and temporarily stored before being transmitted over the network.

[0578] Input: Audio data, video data, slide data

[0579] Output: Data sent

[0580] Step 4:

[0581] The server analyzes the received voice data, converts the voice into text using speech recognition technology (e.g., Google Cloud Speech-to-Text API), and evaluates the speaking speed, volume, and pauses based on the text data.

[0582] Input: Audio data

[0583] Output: Text data and its evaluation results

[0584] Specific operation: The server uses voice recognition technology to convert audio such as "introducing a new product" into text, and analyzes the speaking speed, how many characters per second, whether the voice volume is appropriate, etc.

[0585] Step 5:

[0586] The server analyzes the received video data and uses computer vision techniques (e.g., OpenCV and TensorFlow) to analyze the user's facial expressions, gaze, and gestures to evaluate the presentation's poise and confidence.

[0587] Input: Video data

[0588] Output: Analysis results

[0589] Specific operation: The server uses video analysis technology to evaluate whether the user's eyes are fixed on the slide, whether they are making eye contact with the audience, whether their gestures are natural, etc.

[0590] Step 6:

[0591] The server analyzes the received slides, using optical character recognition (OCR) technology (e.g., Tesseract OCR) to read the content of the slides and evaluate whether they are consistent with the presentation theme and have appropriate visual effects.

[0592] Input: Slide data

[0593] Output: Analysis results

[0594] How it works: The server uses OCR technology to analyze the slides and identify typos and inconsistencies in content, such as when "new product" is misspelled as "new store."

[0595] Step 7:

[0596] The server generates feedback to provide to the user based on the results of the audio, video, and slide analysis, including suggestions for adjusting speaking speed and volume, improving facial expressions and gestures, and correcting slides.

[0597] Input: Audio analysis results, video analysis results, slide analysis results

[0598] Output: Generated feedback

[0599] Specific operation: The server combines the results of each analysis and generates feedback on specific areas for improvement, such as "speaking too fast," "not making enough eye contact with the audience," and "there are typos on the slides."

[0600] Step 8:

[0601] The server sends the generated feedback to the device, which then displays it to the user, who receives real-time feedback and can immediately see how to improve their presentation.

[0602] Input: Generated feedback

[0603] Output: Feedback displayed to the user

[0604] What it does: The device will display the feedback as a pop-up notification for the user to view immediately.

[0605] Step 9:

[0606] The server manages user authentication information, such as authentication when a user logs into the system, and stores presentation history and past feedback results for future reference.

[0607] Input: User authentication information, presentation history, feedback results

[0608] Output: Authenticated user information, saved history and feedback

[0609] Specific operation: The server authenticates the user ID and password and stores each user's presentation history and feedback results in a database.

[0610] Step 10:

[0611] The server generates questions based on the user's presentation and presents them to the user via the terminal. The user answers the questions, and the answers are also analyzed to provide more detailed feedback.

[0612] Input: Presentation content, user responses

[0613] Output: Generated questions, analysis results and additional feedback

[0614] Specific operation: The server generates questions such as "What are the advantages of the new product?" and presents them to the user through the terminal. If the user answers "The advantage is multifunctionality," the answer is analyzed and additional feedback is generated.

[0615] (Application example 1)

[0616] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0617] Previous systems for improving presentation skills only provided feedback specific to presentations, which meant they were not particularly practical for customer service. Customer service skills also depend on a variety of important factors, such as voice, facial expressions, gestures, eye contact, and slides, but there was no system that could comprehensively analyze these and provide specific feedback to customer service staff in real time. Furthermore, while improving the ability to respond to questions in customer service is also important, previous technologies were inadequate in this regard as well.

[0618] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0619] In this invention, the server includes a user interface means for allowing a user to input voice, facial expressions, gestures, slides, and eye contact; a data transmission means for transmitting data input from the user interface means to the server; an audio analysis means for converting the voice data received by the server into text and evaluating the user's speaking speed, volume, and pauses; a video analysis means for analyzing the user's facial expressions, eye contact, and gestures from the video data received by the server and evaluating their level of composure and confidence; a slide analysis means for analyzing the slides received by the server and evaluating the consistency of content, visual effect, and typos; a feedback generation means for generating feedback to the user based on the evaluation results of the audio analysis means, video analysis means, and slide analysis means; a feedback transmission means and a visual output means for transmitting the feedback generated by the feedback generation means to the user's terminal; and a question generation means for generating questions based on the user's customer service content and presenting them to the user. This allows customer service staff to receive specific feedback in real time on various aspects of customer service, such as voice, facial expressions, gestures, eye contact, and consistency of content. The question generation means also contributes to improving customer service skills.

[0620] "User interface means" refers to the functions of devices and software that allow the wait staff to input voice, facial expressions, gestures, slides, and eye movements.

[0621] The "data transmission means" is a mechanism for transmitting data input from the user interface means to the server.

[0622] The "voice analysis means" is a function that converts the voice data received by the server into text and evaluates the speaking speed, volume, and pauses.

[0623] The "video analysis means" is a function that analyzes the facial expressions, gaze, and gestures of the customer service staff from the video data received by the server, and evaluates their level of composure and confidence.

[0624] The "slide material analysis means" is a function that analyzes the slide materials received by the server and evaluates the consistency of the content, the visual effect, and typographical errors.

[0625] The "feedback generation means" is a mechanism for generating feedback to the customer service staff based on the evaluation results of the audio analysis means, video analysis means, and slide data analysis means.

[0626] The "feedback sending means" is a function that sends the generated feedback to the terminal of the customer service staff member.

[0627] "Visual output means" refers to a device or function that visually displays feedback on a terminal.

[0628] The "question generation means" is a function in which the server generates questions based on the content of customer service and presents them to the customer service staff.

[0629] The system for improving customer service skills according to the present invention can be implemented as follows. The system is composed of a customer service staff member, a terminal, and a server. The customer service staff member is a trainee and interfaces with the system through the terminal. The terminal serves to collect the customer service staff member's voice, facial expressions, gestures, eye movements, and slide materials and transmit them to the server. The server analyzes the received data and provides feedback to the customer service staff member.

[0630] Basic configuration

[0631] The system includes the following basic means:

[0632] 1. User interface means: Devices and software that allow wait staff to input voice, facial expressions, gestures, slides, and eye movements. Equipped with a voice input device (microphone), a video input device (camera), and an eye movement detector, the wait staff can input voice, facial expressions, gestures, eye movements, and slides into the system.

[0633] 2. Data transmission means: The terminal has the function to transmit the audio, video and slides input by the customer service staff to the server. This process is performed in real time, and the data is buffered and sent to the server.

[0634] 3. Speech analysis: The server analyzes the received voice data. Specifically, it converts the voice into text using Google Cloud Speech-to-Text, and evaluates the speaking speed, volume, and pauses based on the text data. For example, if a customer service staff member speaks too quickly, this will be reflected in the feedback.

[0635] 4. Video analysis: The server analyzes the received video data. Using computer vision technologies such as OpenCV and dlib, it analyzes the facial expressions, gaze, and gestures of the wait staff, and evaluates their level of composure and confidence based on the analysis. For example, if the wait staff is not making eye contact with the viewer, this will be included in the feedback.

[0636] 5. Slide Analysis: The server analyzes the received slides. It uses optical character recognition (OCR) technology to read the content of the slides and evaluates whether they are consistent, visually appropriate, and free of typos. For example, if a slide contains a typographical error, this is reflected in the feedback.

[0637] 6. Feedback generation means and feedback transmission means: The server generates feedback to be provided to the customer service staff based on the evaluation results obtained by the audio analysis means, video analysis means, and slide analysis means. The generated feedback is sent to the terminal and visually displayed to the customer service staff. This feedback may include adjustments to speaking speed and volume, improvements to facial expressions and gestures, and corrections to slides.

[0638] 7. Question generation means and additional feedback generation means: The server generates questions based on the customer service content and presents them to the customer service staff. When the customer service staff answers these questions, the answers are also analyzed and additional feedback is generated. Questions are generated using a generative AI model such as OpenAI's GPT-3, and the customer service staff's answers are analyzed. For example, questions are generated on the theme of "explaining new products," and the necessary feedback is provided to the customer service staff when they answer them.

[0639] Specific examples

[0640] For example, if a customer service staff member practices on the theme of "explaining a new product," the system functions as follows: The customer service staff member practices using a terminal, and the server uses its audio analysis means to evaluate that "their voice is too quiet." It also uses its video analysis means to determine that "their eyes are always looking toward the product," and its slide analysis means to analyze that "there are typos in some of the slides." The server generates feedback based on these evaluations and provides the customer service staff with specific advice such as "increase the volume of your voice when speaking," "make eye contact with the customer," and "correct typos in the slides."

[0641] Example prompt sentence:

[0642] "Our customer service staff are explaining a new product. What advice would be effective? Please also give us specific examples of areas where our staff can improve."

[0643] In this way, the system of the present invention functions as an effective tool for customer service staff to efficiently improve their customer service skills.

[0644] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0645] Step 1:

[0646] The customer service staff uses the microphone, camera, and gaze detector of the terminal to emit voice and input facial expressions, gestures, gaze, and slides. The input data includes voice data, video data, gaze data, and slides.

[0647] Step 2:

[0648] The terminal buffers the voice data, video data, eye movement data, and slide material input by the customer service staff in real time and transmits them to the server using the data transmission means, whereby the voice, video, eye movement, and slide material data are input to the server.

[0649] Step 3:

[0650] The server converts the received voice data into text data using Google Cloud Speech-to-Text. The input voice data is analyzed and evaluated for speaking speed, volume, and pauses. For example, if the voice data indicates that the speaker is speaking too fast, an evaluation result will be generated.

[0651] Step 4:

[0652] The server uses OpenCV and dlib to analyze facial expressions, eye movements, and gestures from the received video data. It receives the video data as evaluation input and evaluates the customer's level of composure and confidence. For example, it generates an evaluation result such as "the customer's eyes are always directed toward the products."

[0653] Step 5:

[0654] The server uses optical character recognition (OCR) technology to read the content of the received slides and evaluates the consistency of the content, visual effect, and typographical errors. It receives the slides as input data and generates the evaluation result: "Some slides contain typographical errors."

[0655] Step 6:

[0656] The server generates feedback for the customer service staff based on the evaluation results obtained from the audio analysis means, video analysis means, and slide analysis means. For example, the generated feedback includes specific advice such as "increase your speaking volume," "make eye contact with the customer," and "correct typos in the slides."

[0657] Step 7:

[0658] The feedback generated by the feedback generating means is sent to the terminal via the feedback transmitting means and the visual output means. As feedback, information such as "increase your speaking volume" and "make eye contact with the customer" is displayed to the wait staff.

[0659] Step 8:

[0660] The server uses a generative AI model such as OpenAI's GPT-3 to generate questions based on the customer service content. The server generates appropriate questions based on the scenario the customer service staff is practicing and sends them to the device. For example, a question might be generated such as, "I'm explaining a new product. What kind of advice would be effective?"

[0661] Step 9:

[0662] The customer service staff responds to questions posed by the system and sends the answers to the server via their terminal. The server analyzes the received answers and generates additional feedback based on the evaluation results.

[0663] Step 10:

[0664] The server transmits the additional feedback generated by the additional feedback generating means to the terminal using the feedback transmitting means and the visual output means, thereby providing further specific advice to the wait staff.

[0665] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0666] The system for improving presentation skills according to the present invention can provide more advanced feedback by incorporating an emotion engine that analyzes the user's emotions. Specific embodiments will be described below.

[0667] Basic configuration

[0668] The system is composed of a user, a terminal, and a server. The user is the person giving the presentation and interfaces with the system through the terminal. The terminal collects the user's voice, facial expressions, gestures, and slides and transmits them to the server. The server analyzes the received data and provides feedback to the user. In addition, this embodiment includes an emotion engine that recognizes the user's emotions.

[0669] User Interface Means

[0670] A user gives a presentation using a terminal. The terminal is equipped with an audio input device (microphone), a video input device (camera), and a screen for displaying slides. An emotion engine is also installed, which analyzes the user's emotions.

[0671] Data transmission method

[0672] The device has the ability to transmit audio, video, and slides input by the user to the server. This process is performed in real time, and the data is buffered and sent to the server. The user's emotional data is also transmitted in the same way.

[0673] Voice analysis methods

[0674] The server analyzes the received audio data, converts it into text using speech recognition technology, and evaluates the speed, volume, and pauses of speech based on the text data. For example, if the user speaks too quickly, this will be reflected in the feedback.

[0675] Video analysis methods

[0676] The server analyzes the received video data, using computer vision technology to analyze the user's facial expressions, eye contact, and gestures, and evaluates the presentation's poise and confidence. For example, if the user is not making eye contact with the audience, this will be included in the feedback.

[0677] Slide analysis tools

[0678] The server analyzes the slides it receives, using optical character recognition (OCR) technology to read their content and evaluate whether they are consistent with the presentation theme and whether their visual impact is appropriate. For example, if a slide contains a typographical error, that error is reflected in the feedback.

[0679] Emotion Engine

[0680] The server uses an emotion engine to analyze the user's emotional data. Specifically, it analyzes the user's facial expressions and voice to evaluate the fluctuations and stability of their emotions during the presentation. For example, if the user is nervous during the presentation, that emotional data will be included in the feedback.

[0681] Feedback generation means and feedback transmission means

[0682] The server integrates the evaluation results of the audio analysis means, video analysis means, slide analysis means, and emotion engine to generate feedback for the user. The generated feedback is sent to the terminal and displayed to the user. The feedback includes advice on adjusting speaking speed and volume, improving facial expressions and gestures, correcting slides, and emotional advice.

[0683] Authentication Management Methods

[0684] The server manages user authentication information and stores the user's presentation history and feedback results, allowing users to check their progress and refer to past feedback.

[0685] Question generation and additional feedback generation

[0686] The server generates questions based on the user's presentation and presents them to the user. When the user answers these questions, the answers are also analyzed and additional feedback is generated. For example, if the user does not answer a question accurately, that point is reflected in the additional feedback.

[0687] Specific examples

[0688] For example, if a user gives a presentation on the theme of "introducing a new product," the system functions as follows: The user gives the presentation using a device, and the server uses its audio analysis means to evaluate the user's voice as "quiet." The video analysis means determines that the user's eyes are always looking at the slides, and the slide analysis means analyzes that there are typos on some of the slides. The emotion engine then determines that the user is nervous. The server generates feedback based on these evaluations and provides the user with specific advice such as "increase the volume of your voice," "make eye contact with the audience," "correct typos in the slides," and "regulate your breathing to relieve tension."

[0689] In this way, the system of the present invention functions as an effective tool for users to efficiently improve their presentation skills. In particular, the introduction of an emotion engine makes it possible to provide comprehensive feedback that takes into account the user's emotional state.

[0690] The processing flow will be explained below.

[0691] Step 1:

[0692] A user logs into the system using a terminal, enters a username and password, and submits the authentication information.

[0693] Step 2:

[0694] The device sends the entered authentication information to the server. The server compares it with the database and returns the authentication result to the device. If authentication is successful, the user can start the presentation.

[0695] Step 3:

[0696] The user begins preparing for a presentation. They upload presentation materials (slides) using their device. The device then sends the uploaded materials to the server.

[0697] Step 4:

[0698] The server stores the received presentation materials in an appropriate format.

[0699] Step 5:

[0700] The user presses the "Start Presentation" button. The device begins recording the user's voice and video (facial expressions and gestures) using the built-in camera and microphone. The recorded data is buffered.

[0701] Step 6:

[0702] The terminal transmits the buffered data to the server in real time, including audio, video, slides, and emotion data.

[0703] Step 7:

[0704] The server analyzes the received voice data, converts the voice data into text using a speech recognition engine, evaluates speaking speed, volume, and pauses, and generates text data.

[0705] Step 8:

[0706] The server analyzes the received video data and uses computer vision technology to analyze the user's facial expressions, eye movements, and gestures to evaluate the presentation's poise and confidence.

[0707] Step 9:

[0708] The server analyzes the slides received, using optical character recognition (OCR) technology to read the content of the slides and evaluate their consistency and visual impact.

[0709] Step 10:

[0710] The server uses an emotion engine to analyze the user's emotional data, identifying emotions from facial expressions and vocal tones, and assessing emotional stability and fluctuations.

[0711] Step 11:

[0712] The server integrates the evaluation results obtained from the audio analysis means, video analysis means, slide analysis means, and emotion engine, and generates feedback for the user, including advice on speaking speed, volume, naturalness of gestures, corrections to slides, and emotional advice.

[0713] Step 12:

[0714] The server sends the generated feedback to the user's terminal, which receives the feedback and displays it to the user.

[0715] Step 13:

[0716] Users can review the feedback and incorporate improvements for their next presentation. Users can also request a Q&A simulation if necessary.

[0717] Step 14:

[0718] If a question and answer simulation is desired, the server generates questions based on the content of the presentation and presents them to the user.

[0719] Step 15:

[0720] The user answers the questions, and the device sends the answer data to the server.

[0721] Step 16:

[0722] The server analyzes the user's answers and generates additional feedback, providing further improvements based on the evaluation results.

[0723] Step 17:

[0724] The server sends additional feedback to the user's device, which the user receives and uses as a reference for their next presentation.

[0725] This series of processing steps allows users to improve their presentation skills efficiently and effectively. In addition, the introduction of an emotion engine makes it possible to provide comprehensive feedback that takes emotional aspects into consideration.

[0726] Example 2

[0727] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0728] Current presentation skill improvement systems have difficulty accurately analyzing users' emotions and reflecting the analysis results in presentation feedback. Furthermore, they lack feedback on not only the technical aspects of the presentation but also the emotional and psychological aspects, making it difficult for users to comprehensively improve their skills. Furthermore, existing systems lack sufficient functionality for generating questions and providing additional feedback, making it difficult for users to self-evaluate.

[0729] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0730] In this invention, the server includes an audio analysis means, a video analysis means, a slide analysis means, an emotion engine, a feedback generation means, and a feedback transmission means. This allows for comprehensive analysis of a user's voice, facial expressions, gestures, slides, and even emotions, making it possible to provide detailed and accurate feedback. Furthermore, by including a question generation means and an additional feedback generation means, it is possible to automatically generate questions based on the content of the user's presentation and evaluate the answers to those questions, facilitating self-evaluation.

[0731] "User interface means" refers to devices and software that allow a user to input voice, facial expressions, gestures, and slide materials.

[0732] "Data transmission means" refers to a device or software for transmitting data input from the user interface means to the server.

[0733] "Voice analysis means" refers to devices or software that convert the voice data received by the server into text and evaluate speaking speed, volume, and pauses.

[0734] "Video analysis means" refers to devices or software that analyze the user's facial expressions, gaze, and gestures from the video data received by the server, and evaluate the presentation's level of composure and confidence.

[0735] "Slide material analysis means" refers to a device or software that analyzes the slide materials received by the server and evaluates the consistency of the content and the visual effect.

[0736] "Emotion engine" refers to a device or software that allows the server to analyze emotion data and evaluate the fluctuations and stability of a user's emotions.

[0737] "Feedback generation means" refers to a device or software for generating feedback to a user based on the evaluation results of the audio analysis means, video analysis means, slide data analysis means, and emotion engine.

[0738] The "feedback sending means" refers to a device or software for sending the feedback generated by the feedback generating means to the user's terminal.

[0739] "Authentication management means" refers to a device or software for managing user authentication information and storing each user's presentation history and feedback results.

[0740] The "question generation means" refers to a device or software that generates questions based on the content of a user's presentation and presents them to the user.

[0741] The "additional feedback generating means" refers to a device or software that analyzes the answers given by the user to the questions presented and generates additional feedback based on the evaluation results.

[0742] The system for improving presentation skills according to the present invention provides more advanced feedback by incorporating an emotion engine that analyzes the user's emotions. Specific embodiments will be described below.

[0743] Basic configuration

[0744] The system is composed of a user, a terminal, and a server. The user is the person giving the presentation and interfaces with the system through the terminal. The terminal collects the user's voice, facial expressions, gestures, and slides and transmits them to the server. The server analyzes the received data and provides feedback to the user. In addition, this embodiment includes an emotion engine that recognizes the user's emotions.

[0745] User Interface Means

[0746] A user gives a presentation using a terminal. The terminal is equipped with an audio input device (microphone), a video input device (camera), and a screen for displaying slides. An emotion engine is also installed, which analyzes the user's emotions.

[0747] Data transmission method

[0748] The device has the ability to transmit audio, video, and slides input by the user to the server. This process is performed in real time, and the data is buffered and sent to the server. The user's emotional data is also transmitted in the same way.

[0749] Voice analysis methods

[0750] The server analyzes the received audio data, converts it into text using speech recognition technology, and evaluates speaking speed, volume, and pauses based on the text data. The specific software used is the Google Cloud Speech-to-Text API. For example, if the user speaks too quickly, this is reflected in the feedback.

[0751] Video analysis methods

[0752] The server analyzes the received video data. It uses computer vision technology to analyze the user's facial expressions, eye movements, and gestures, and evaluates the presentation's poise and confidence based on the results. The OpenCV library is used for this software. For example, if the user is not making eye contact with the audience, this will be included in the feedback.

[0753] Slide analysis tools

[0754] The server analyzes the received slides. It uses optical character recognition (OCR) technology to read the content of the slides and evaluate whether they are consistent with the presentation theme and have an appropriate visual effect. Specific software includes Tesseract OCR. For example, if a slide contains a typographical error, that error is reflected in the feedback.

[0755] Emotion Engine

[0756] The server uses an emotion engine to analyze the user's emotional data. Specifically, it analyzes the user's facial expressions and voice to evaluate the fluctuations and stability of their emotions during the presentation. For example, if the user is nervous during the presentation, that emotional data will be included in the feedback.

[0757] Feedback generation means and feedback transmission means

[0758] The server integrates the evaluation results of the audio analysis means, video analysis means, slide analysis means, and emotion engine to generate feedback for the user. The generated feedback is sent to the terminal and displayed to the user. The feedback includes advice on adjusting speaking speed and volume, improving facial expressions and gestures, correcting slides, and emotional advice.

[0759] Authentication Management Methods

[0760] The server manages user authentication information and stores the user's presentation history and feedback results, allowing users to check their progress and refer to past feedback.

[0761] Question generation and additional feedback generation

[0762] The server generates questions based on the user's presentation and presents them to the user. When the user answers these questions, the answers are also analyzed and additional feedback is generated. For example, if the user does not answer a question accurately, that point is reflected in the additional feedback.

[0763] Specific examples

[0764] For example, if a user is giving a presentation on the theme of "introducing a new product," the system functions as follows: The user gives the presentation using a terminal, and the server uses the audio analysis means to evaluate the user's voice as "quiet." The video analysis means determines that the user's eyes are always directed toward the slides, and the slide analysis means analyzes that there are typos on some slides. The emotion engine then determines that the user is nervous. Based on these evaluations, the server generates feedback and provides the user with specific advice, such as "increase your speaking volume," "make eye contact with the audience," "correct typos in the slides," or "regulate your breathing to relieve tension." In this way, the system of the present invention functions as an effective tool for users to efficiently improve their presentation skills. In particular, the introduction of the emotion engine makes it possible to provide comprehensive feedback that takes the user's emotional state into consideration.

[0765] Prompt Sentence Examples

[0766] Here are some example prompts to input to a generative AI model:

[0767] A user will give a presentation introducing a new product. Staff will generate feedback based on the analysis of audio, video, and slides. Also, analyze the user's emotional data and include it in the feedback. Examples of feedback include "your voice is too quiet," "you should make eye contact with the audience," "there is a typo on the slide," and "you seem nervous."

[0768] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0769] Specific processing flow of the program

[0770] Step 1: Start your presentation

[0771] The user clicks a button on the device to start the presentation. The device verifies the user's authentication information and sends a request to the server to start the session.

[0772] Input: User credentials, session initiation request

[0773] Processing: The server verifies the authentication information and generates a token to start the session.

[0774] Output: Session initiation token

[0775] Step 2: Data collection

[0776] The user gives a presentation, and the device collects audio, video, and slides via a microphone, camera, and screen display.

[0777] Input: User's audio, video, and slides

[0778] Processing: The device buffers the data in real time and prepares it for transfer to the server.

[0779] Output: Buffered data

[0780] Step 3: Send data

[0781] The terminal transmits the collected data (audio, video, slides, and emotional data) to the server in real time.

[0782] Input: Buffered audio, video, slide deck, emotion data

[0783] Processing: The device divides the data into packets and sends them over the network to the server.

[0784] Output: Data packets arriving at the server

[0785] Step 4: Audio data analysis

[0786] The server analyzes the received voice data, converts the voice into text using speech recognition technology, and evaluates the speaking speed, volume, and pauses based on the text data.

[0787] Input: Audio data

[0788] Processing: Converts audio to text using the Google Cloud Speech-to-Text API and analyzes data on speaking rate, volume, and pauses.

[0789] Output: Text data, speaking speed, volume, and interval evaluation data

[0790] Step 5: Video data analysis

[0791] The server analyzes the received video data, using facial recognition technology and computer vision to analyze the user's facial expressions, gaze, and gestures, and evaluates the presentation's poise and confidence based on that information.

[0792] Input: Video data

[0793] Processing: Using the OpenCV library, we analyze the user's facial expressions, gaze, and gestures to assess their calmness and confidence.

[0794] Output: Facial expression, gaze, gesture analysis data, calmness and confidence evaluation data

[0795] Step 6: Slide analysis

[0796] The server analyzes the received slides, using optical character recognition (OCR) technology to read the slide content and evaluate whether it is consistent with the presentation theme and whether the visual effects are appropriate.

[0797] Input: Slide data

[0798] Processing: Extract text from slides using Tesseract OCR and evaluate content consistency and visual impact.

[0799] Output: Slide text data, consistency and visual effect evaluation data

[0800] Step 7: Sentiment Data Analysis

[0801] The server uses an emotion engine to analyze the user's emotion data. Specifically, it analyzes emotions from the user's facial expressions and voice and evaluates the fluctuations and stability of emotions during the presentation.

[0802] Input: Emotion data (facial expressions, voice)

[0803] Processing: An emotion engine is used to analyze facial and voice data to assess emotional variability and stability.

[0804] Output: Emotional evaluation data

[0805] Step 8: Feedback Generation

[0806] The server integrates the results of the audio analysis, video analysis, slide analysis, and emotion data analysis, and generates feedback for the user.

[0807] Input: Audio analysis data, video analysis data, slide analysis data, emotion analysis data

[0808] Processing: Comprehensive feedback is generated based on the results of each assessment. The feedback is provided in text and graphical format.

[0809] Output: Feedback data (text, graph)

[0810] Step 9: Send and view feedback

[0811] The server sends the generated feedback to the terminal, which displays it to the user.

[0812] Input: Feedback data

[0813] Processing: The server sends the feedback data to the terminal, which displays it to the user. The feedback includes specific advice.

[0814] Output: Feedback that is displayed to the user

[0815] Step 10: Question generation and additional feedback

[0816] The server generates relevant questions based on the user's presentation and presents them to the user via the terminal. When the user answers the questions, the server analyzes them again and generates additional feedback.

[0817] Input: Presentation content data, user questions and answers

[0818] Processing: A question generation algorithm is used to generate questions based on the presentation content, and the answers are analyzed and evaluated. Additional feedback is generated based on the evaluation results.

[0819] Output: Additional feedback data

[0820] Step 11: Authentication Management and History Storage

[0821] The server manages user authentication information and feedback history, and stores them in a database after the presentation session ends.

[0822] Input: User authentication information, feedback history data

[0823] Processing: Manages authentication information and stores feedback history in a database.

[0824] Output: Updated database

[0825] Through this process, users receive feedback to comprehensively improve their presentation skills.

[0826] (Application example 2)

[0827] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0828] Conventional presentation skill improvement systems are primarily limited to analyzing voice, facial expressions, and slides, and are not capable of improving specialized skills tailored to specific environments or situations. For example, to improve a factory operator's skills so that they can issue appropriate commands to a robot, analysis and feedback that takes into account the appropriateness of the robot's operation instructions and the operator's emotional state is necessary. A system that can address this issue is needed.

[0829] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a user interface means for a user to input voice, facial expressions, gestures, and slide materials, a data transmission means for transmitting the data input from the user interface means to the server, a voice analysis means for converting voice data received by the server into text and evaluating the speaking speed, volume, and pauses, a video analysis means for analyzing the user's facial expressions, line of sight, and gestures from video data received by the server and evaluating the calmness and confidence of the presentation, a slide material analysis means for analyzing the slide materials received by the server and evaluating the consistency of the content and the visual effect, a feedback generation means for generating feedback to the user based on the evaluation results of the voice analysis means, the video analysis means, and the slide material analysis means, a feedback transmission means for transmitting the feedback generated by the feedback generation means to the user's terminal, a robot operation instruction evaluation means for evaluating the appropriateness of instructions given by the operator to the robot, and an emotion evaluation means for analyzing the operator's emotional state and evaluating the mental state that affects the appropriateness of the instructions. This will enable factory operators to improve their skills in operating robots properly.

[0830] "User interface means" refers to devices and software that allow a user to input voice, facial expressions, gestures, and slide materials.

[0831] The "data transmission means" is a communication means for transmitting data input from the user interface means to the server in real time.

[0832] The "voice analysis means" is a technology that converts the voice data received by the server into text and evaluates the speaking speed, volume, and pauses.

[0833] The "video analysis means" is a technology that analyzes the user's facial expressions, gaze, and gestures from the video data received by the server, and evaluates the presentation's level of composure and confidence.

[0834] The "slide material analysis means" is a technology for analyzing the slide materials received by the server and evaluating the consistency of the content and the visual effect.

[0835] The "feedback generation means" is a technology that generates feedback indicating areas for improvement to the user based on the evaluation results of the audio analysis means, video analysis means, and slide data analysis means.

[0836] The "feedback transmitting means" is a means for transmitting the feedback generated by the feedback generating means to the user's terminal.

[0837] The "robot operation instruction evaluation means" is a technology for evaluating the appropriateness of instructions given by an operator to a robot.

[0838] "Emotion evaluation means" is a technology that analyzes the emotional state of an operator and evaluates the mental state that affects the appropriateness of instructions.

[0839] "Authentication management means" is a technology that enables the server to manage user authentication information and store each user's presentation history, instruction history, and feedback results.

[0840] The "question generation means" is a technology for generating questions based on the content of the user's presentation and instructions for operating the robot, and presenting the questions to the user.

[0841] The "additional feedback generating means" is a technology that analyzes the answers given by the user to the questions presented and generates additional feedback based on the evaluation results.

[0842] The system according to the present invention is intended to improve the skills of factory operators to properly operate robots. This system is composed of a user (operator), a terminal, and a server. Specific embodiments are described below.

[0843] Basic configuration

[0844] The system includes a user interface for inputting voice, facial expressions, gestures, and instructions from the user's terminal. These data are transmitted in real time from the terminal to a server, which analyzes the received data and provides feedback to the user.

[0845] Voice analysis methods

[0846] The server first receives the user's voice data and converts it into text using speech recognition software (for example, Google's speech recognition API). The server then evaluates the user's speaking speed, volume, and pauses. Any inappropriate parts are reflected in the feedback.

[0847] Video analysis methods

[0848] The server then analyzes the user's video data. Using video analysis technology (e.g., OpenCV), it detects and evaluates the user's facial expressions, eye movements, and gestures. The server evaluates the user's poise and confidence during the presentation, and provides feedback on any shortcomings.

[0849] Slide analysis tools

[0850] The server analyzes the slide deck it receives and uses optical character recognition (OCR) technology to review the slide content and evaluate its consistency and visual effectiveness. For example, typographical errors and visual flaws are flagged.

[0851] Emotion assessment tools

[0852] The server then uses an emotion engine to analyze the user's emotional data. It extracts emotions from the user's facial expressions and voice, and evaluates the emotional fluctuations and stability during the presentation. If tension or stress is detected, this is also reflected in the feedback.

[0853] Robot operation instruction evaluation means

[0854] In a specific application, the server evaluates the appropriateness of the instructions given by the operator to the robot, for example, whether the instructions are clear and unambiguous when given by voice or gesture.

[0855] Feedback generation and transmission methods

[0856] The server then aggregates all of these analysis results and generates specific feedback for the user, suggesting areas for improvement. The generated feedback is then sent to the user's device and displayed to the user.

[0857] Authentication Management Methods

[0858] The server also manages user authentication information and stores each user's presentation history, instruction history, and feedback results, allowing users to improve themselves based on past data.

[0859] Question generation and additional feedback generation means

[0860] The server generates questions based on the user's presentation and robot operation instructions, and presents them to the user. When the user answers these questions, the server analyzes the answers and generates additional feedback.

[0861] Specific examples

[0862] For example, imagine a scene in which an operator is giving verbal instructions to a robot in a factory. The scene is recorded using the user interface means, and the audio and video data are sent to the server. The server analyzes the audio and judges it as "too quiet," and analyzes the video and judges it as "tension is evident in the operator's facial expression." Furthermore, the robot operation instruction evaluation means judges it as "ambiguous." Based on these evaluations, the server generates specific feedback and provides the user with advice such as "speak louder" or "give clearer instructions."

[0863] Prompt Sentence Examples

[0864] "Analyze sample videos, evaluate the operator's emotional state and the appropriateness of their voice instructions, and provide feedback."

[0865] This system allows factory operators to effectively improve their robot operation skills, aiming to efficiently improve productivity.

[0866] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0867] Step 1: Data entry

[0868] The user inputs voice, facial expressions, gestures, and instructions. Specifically, data is collected using user interface means (camera, microphone, touch panel, etc.). The inputs in this step are the user's voice data, video data, and instructions. The output is that these data are stored in the terminal.

[0869] Step 2: Send data

[0870] The terminal transmits the voice data, video data, and instruction contents input from the user interface means to the server. Specifically, the data is transmitted to the server in real time via the Internet. The input of this step is the user's voice data, video data, and instruction contents, and the output is the server receiving these data.

[0871] Step 3: Audio analysis

[0872] The server analyzes the received voice data. Specifically, it uses speech recognition software (such as Google's speech recognition API) to convert the speech into text and evaluates the speaking speed, volume, and pauses. The input of this step is the received voice data, and the output is the analyzed text data and the evaluation results.

[0873] Step 4: Video analysis

[0874] The server analyzes the received video data. Specifically, it uses video analysis technology (such as OpenCV) to detect the user's facial expressions, eye movements, and gestures, and evaluates their level of composure and confidence during the presentation. The input to this step is the received video data, and the output is the analyzed facial expression data, eye movement data, and gesture data, along with the evaluation results.

[0875] Step 5: Slide analysis

[0876] The server analyzes the received slides. Specifically, it uses optical character recognition (OCR) technology to review the content of the slides and evaluate their consistency and visual effectiveness. The input for this step is the received slides, and the output is the analyzed text data and the evaluation results.

[0877] Step 6: Emotional assessment

[0878] The server evaluates emotions using the received video and audio data. Specifically, it uses emotion recognition technology (such as Face++ or Microsoft Azure Face API) to evaluate the user's emotional state and evaluate the fluctuations and stability of emotions during the presentation. The input for this step is facial expression data and audio data, and the output is analyzed emotional data and its evaluation results.

[0879] Step 7: Evaluate robot operation instructions

[0880] The server evaluates the operator's instruction data. Specifically, it uses voice recognition and video analysis to evaluate the appropriateness of the operator's instructions to the robot. The input for this step is the operator's voice data and gesture data, and the output is the analysis result.

[0881] Step 8: Feedback Generation

[0882] The server integrates all the above analysis results and generates specific feedback for the user indicating areas for improvement. Feedback is generated based on the overall evaluation results using a feedback generation means. The inputs for this step are the audio analysis results, video analysis results, slide analysis results, emotion evaluation results, and robot operation instruction evaluation results, and the output is the generated feedback information.

[0883] Step 9: Submit your feedback

[0884] The server sends the generated feedback to the user's terminal, specifically, provides the feedback information to the user via the Internet. The input of this step is the generated feedback information, and the output is the feedback being displayed on the user's terminal.

[0885] Step 10: Question generation and additional feedback generation

[0886] The server generates additional questions based on the user's presentation and instructions and presents them to the user. The server further analyzes the user's answers and generates additional feedback, which helps the user identify areas for further improvement. The input of this step is the user's answer data, and the output is the additional feedback information.

[0887] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0888] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0889] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0890] [Third embodiment]

[0891] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0892] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[0893] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0894] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0895] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0896] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0897] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0898] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0899] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0900] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0901] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0902] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0903] The system for improving presentation skills according to the present invention can be implemented as follows.

[0904] Basic configuration

[0905] The system consists of a user, a terminal, and a server. The user is the person giving the presentation and interfaces with the system through the terminal. The terminal collects the user's voice, facial expressions, gestures, and slides and transmits them to the server. The server analyzes the received data and provides feedback to the user.

[0906] User Interface Means

[0907] A user gives a presentation using a terminal. The terminal is equipped with an audio input device (microphone), a video input device (camera), and a screen for displaying slides. This allows the user to input voice, facial expressions, gestures, and slides into the system.

[0908] Data transmission method

[0909] The terminal has the function of transmitting the audio, video, and slides input by the user to the server. This process is performed in real time, and the data is buffered and sent to the server.

[0910] Voice analysis methods

[0911] The server analyzes the received audio data. Specifically, it uses speech recognition technology to convert the audio into text, and then evaluates the speaking speed, volume, and pauses based on the text data. For example, if the user speaks too fast, this will be reflected in the feedback.

[0912] Video analysis methods

[0913] The server analyzes the received video data, using computer vision technology to analyze the user's facial expressions, eye contact, and gestures, and evaluates the presentation's poise and confidence. For example, if the user is not making eye contact with the audience, this will be included in the feedback.

[0914] Slide analysis tools

[0915] The server analyzes the slides it receives, using optical character recognition (OCR) technology to read their content and evaluate whether they are consistent with the presentation theme and whether the visual effect is appropriate. For example, if a slide contains a typographical error, that will be reflected in the feedback.

[0916] Feedback generation means and feedback transmission means

[0917] The server generates feedback to be provided to the user based on the evaluation results obtained by the audio analysis means, video analysis means, and slide analysis means. The generated feedback is sent to the terminal and displayed to the user. The feedback may include adjustments to speaking speed and volume, improvements to facial expressions and gestures, and corrections to slides.

[0918] Authentication Management Methods

[0919] The server manages user authentication information and stores the user's presentation history and feedback results, allowing users to check their progress and refer to past feedback.

[0920] Question generation and additional feedback generation

[0921] The server generates questions based on the user's presentation and presents them to the user. When the user answers these questions, the answers are also analyzed and additional feedback is generated. For example, if the user does not answer a question accurately, that point is reflected in the additional feedback.

[0922] Specific examples

[0923] For example, if a user gives a presentation on the theme of "introducing a new product," the system functions as follows: The user gives the presentation using a device, and the server uses its audio analysis means to evaluate that the user's voice is "quiet." The video analysis means determines that the user's eyes are always looking at the slides, and the slide analysis means analyzes that there are typos on some of the slides. The server generates feedback based on these evaluations and provides the user with specific advice such as "increase the volume of your voice when speaking," "make eye contact with the audience," and "correct typos in the slides."

[0924] In this way, the system of the present invention functions as an effective tool for users to efficiently improve their presentation skills.

[0925] The processing flow will be explained below.

[0926] Step 1:

[0927] A user logs into the system using a terminal, enters a username and password, and submits the authentication information.

[0928] Step 2:

[0929] The device sends the entered authentication information to the server. The server compares it with the database and returns the authentication result to the device. If authentication is successful, the user can start the presentation.

[0930] Step 3:

[0931] The user begins preparing for a presentation. They upload presentation materials (slides) using their device. The device then sends the uploaded materials to the server.

[0932] Step 4:

[0933] The server stores the received presentation materials in an appropriate format.

[0934] Step 5:

[0935] The user presses the "Start Presentation" button. The device begins recording the user's voice and video (facial expressions and gestures) using the built-in camera and microphone. The recorded data is buffered.

[0936] Step 6:

[0937] The terminal transmits the buffered data to the server in real time. The data for the audio, video, and slide materials are transmitted to the server.

[0938] Step 7:

[0939] The server analyzes the received voice data, converts it into text using a speech recognition engine, and evaluates speaking speed, volume, and pauses.

[0940] Step 8:

[0941] The server analyzes the received video data and uses computer vision technology to analyze the user's facial expressions, eye movements, and gestures to evaluate the presentation's poise and confidence.

[0942] Step 9:

[0943] The server analyzes the slides it receives, using optical character recognition (OCR) technology to read their content and evaluate whether they are consistent with the presentation theme and whether their visual impact is appropriate.

[0944] Step 10:

[0945] The server integrates the results of the audio, video, and slide analysis and generates feedback for the user. Based on each evaluation result, it summarizes specific improvements and advice in written form.

[0946] Step 11:

[0947] The server sends the generated feedback to the user's terminal, which receives the feedback and displays it to the user.

[0948] Step 12:

[0949] Users can review the feedback and incorporate improvements for their next presentation. If desired, a simulated Q&A session can also be conducted.

[0950] Step 13:

[0951] If a question and answer simulation is desired, the server generates questions based on the content of the presentation and presents them to the user.

[0952] Step 14:

[0953] The user answers the questions, and the device sends the answer data to the server.

[0954] Step 15:

[0955] The server analyzes the user's answers and generates additional feedback, and reflects suggestions for improvement in the advice based on the evaluation results.

[0956] Step 16:

[0957] The server sends additional feedback to the user's device, which the user receives and uses as a reference for their next presentation.

[0958] This series of processing steps allows users to improve their presentation skills efficiently and effectively.

[0959] Example 1

[0960] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0961] Conventional presentation support systems lack the ability to comprehensively analyze a user's voice, facial expressions, gestures, and slides to provide comprehensive feedback. They also lack the ability to generate questions based on the user's presentation content, analyze the answers, and provide additional feedback. This makes it difficult for users to effectively improve their presentation skills.

[0962] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0963] In this invention, the server includes a user interface means for a user to input voice, facial expressions, gestures, and slide materials, a data transmission means for transmitting the data input from the user interface means to the server, an audio analysis means for converting the voice data received by the server into text and evaluating the speaking speed, volume, and pauses, a video analysis means for analyzing the user's facial expressions, gaze, and gestures from the video data received by the server and evaluating the poise and confidence of the presentation, a slide material analysis means for analyzing the slide materials received by the server and evaluating the consistency and visual effect of the content, a feedback generation means for generating feedback to the user based on the evaluation results of the audio analysis means, the video analysis means, and the slide material analysis means, a feedback transmission means for transmitting the feedback generated by the feedback generation means to the user's terminal, a question generation means for the server to generate questions based on the content of the user's presentation and present them to the user, and an additional feedback generation means for analyzing the user's answers to the questions posed via the user interface means and generating additional feedback based on the evaluation results. This allows the user to receive comprehensive and detailed feedback, thereby effectively improving their presentation skills.

[0964] "User interface means" refers to means by which a user inputs voice, facial expressions, gestures, and slide materials.

[0965] The "data transmission means" is a means for transmitting data input from the user interface means to the server.

[0966] The "voice analysis means" is a means for converting voice data received by the server into text and evaluating speaking speed, volume, and pauses.

[0967] The "video analysis means" is a means for analyzing the user's facial expressions, gaze, and gestures from the video data received by the server, and evaluating the user's level of composure and confidence in the presentation.

[0968] The "slide material analysis means" is a means for analyzing the slide materials received by the server and evaluating the consistency of the content and the visual effect.

[0969] The "feedback generation means" is a means for generating feedback to the user based on the evaluation results of the audio analysis means, the video analysis means, and the slide material analysis means.

[0970] The "feedback transmitting means" is a means for transmitting the feedback generated by the feedback generating means to the user's terminal.

[0971] The "question generation means" is a means by which the server generates a question based on the content of the user's presentation and presents it to the user.

[0972] The "additional feedback generating means" is a means for analyzing the answers given by the user to questions presented via the user interface means and generating additional feedback based on the evaluation results.

[0973] The "authentication management means" is a means by which the server manages user authentication information and stores each user's presentation history and feedback results.

[0974] The system for improving presentation skills according to the present invention comprises a user, a terminal, and a server. By using this system, the user can effectively analyze their voice, facial expressions, gestures, and slides, and receive comprehensive feedback.

[0975] User Interface Means

[0976] When giving a presentation, a user uses a terminal equipped with a microphone, a camera, and a screen for displaying slides. This allows the user to input voice, facial expressions, gestures, and slides.

[0977] Data transmission method

[0978] The terminal transmits the audio, video, and slides input by the user to the server in real time. During this process, the data is buffered, allowing for smooth data transfer in a stable communication environment.

[0979] Voice analysis methods

[0980] The server analyzes the received voice data. Specifically, it converts the voice into text using speech recognition technology (e.g., Google Cloud Speech-to-Text API), and evaluates the speaking speed, volume, and pauses based on the text data. For example, if a user says "introducing a new product," if they speak too quickly, this will be reflected in the feedback.

[0981] Video analysis methods

[0982] The server analyzes the received video data. Using computer vision techniques (e.g., OpenCV and TensorFlow), it analyzes the user's facial expressions, eye movements, and gestures to assess the presentation's poise and confidence. For example, if the user's eyes are fixed on the slides and they are not making eye contact with the audience, this will be included in the feedback.

[0983] Slide analysis tools

[0984] The server analyzes the slides it receives. It uses optical character recognition (OCR) technology (e.g., Tesseract OCR) to read the content of the slides and evaluates whether they are consistent with the presentation theme and whether the visual effect is appropriate. For example, if a slide contains a typographical error, that error is reflected in the feedback.

[0985] Feedback Generation Method

[0986] The server generates feedback to be provided to the user based on the evaluation results obtained by the audio analysis means, video analysis means, and slide analysis means. The generated feedback includes adjustments to speaking speed and volume, improvements to facial expressions and gestures, and corrections to slides.

[0987] Feedback sending method

[0988] The server sends the generated feedback to the device, which then displays it to the user, who receives real-time feedback and can immediately see how to improve their presentation.

[0989] Question generation means

[0990] The server generates questions based on the content of the user's presentation and presents them to the user via the terminal, encouraging the user to consider further the presentation.

[0991] Additional Feedback Generation Means

[0992] The server analyzes the answers given by the user to the questions posed and generates additional feedback based on the evaluation results, allowing the user to receive feedback at a deeper level.

[0993] Authentication Management Methods

[0994] The server manages user authentication information and stores presentation history and past feedback results, allowing users to check their progress and continuously improve their presentation skills.

[0995] Specific examples

[0996] For example, if a user is giving a presentation on the theme of "introducing a new product," the system functions as follows: The user gives the presentation using a terminal, and the server uses the audio analysis means to evaluate that the user's voice is quiet. The video analysis means determines that the user's eyes are always looking at the slides, and the slide analysis means analyzes that there are typos in some of the slides. The server generates feedback based on these evaluations and provides the user with specific advice, such as speaking louder, making eye contact with the audience, or correcting typos in the slides. In this way, the system of the present invention functions as an effective tool for users to efficiently improve their presentation skills.

[0997] Prompt Sentence Examples

[0998] For example, if you are giving a presentation on the topic of "Introducing a New Product," you might use the following prompt:

[0999] "You will be asked to give a presentation on the topic of introducing a new product. A microphone and camera will be used to collect data, and analysis will be performed using speech recognition, computer vision, and optical character recognition (OCR). Feedback will include an evaluation of speaking speed, volume, facial expressions, eye contact, gestures, and the content of your slides."

[1000] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1001] Step 1:

[1002] The user starts the application on the device, sets the theme to "New Product Introduction," and starts a presentation. The user can input voice, facial expressions, gestures, and slides.

[1003] Input: Set presentation theme, start presentation

[1004] Output:Presentation settings information

[1005] Step 2:

[1006] The device's microphone collects the user's voice, and the camera records the user's facial expressions and gestures, as well as capturing the slides the user uses.

[1007] Input: User's voice, facial expressions, gestures, slides

[1008] Output: Collected audio data, video data, and slide data

[1009] Step 3:

[1010] The devices transmit the collected data to the server in real time, a process in which the data is buffered and temporarily stored before being transmitted over the network.

[1011] Input: Audio data, video data, slide data

[1012] Output: Data sent

[1013] Step 4:

[1014] The server analyzes the received voice data, converts the voice into text using speech recognition technology (e.g., Google Cloud Speech-to-Text API), and evaluates the speaking speed, volume, and pauses based on the text data.

[1015] Input: Audio data

[1016] Output: Text data and its evaluation results

[1017] Specific operation: The server uses voice recognition technology to convert audio such as "introducing a new product" into text, and analyzes the speaking speed, how many characters per second, whether the voice volume is appropriate, etc.

[1018] Step 5:

[1019] The server analyzes the received video data and uses computer vision techniques (e.g., OpenCV and TensorFlow) to analyze the user's facial expressions, gaze, and gestures to evaluate the presentation's poise and confidence.

[1020] Input: Video data

[1021] Output: Analysis results

[1022] Specific operation: The server uses video analysis technology to evaluate whether the user's eyes are fixed on the slide, whether they are making eye contact with the audience, whether their gestures are natural, etc.

[1023] Step 6:

[1024] The server analyzes the received slides, using optical character recognition (OCR) technology (e.g., Tesseract OCR) to read the content of the slides and evaluate whether they are consistent with the presentation theme and have appropriate visual effects.

[1025] Input: Slide data

[1026] Output: Analysis results

[1027] How it works: The server uses OCR technology to analyze the slides and identify typos and inconsistencies in content, such as when "new product" is misspelled as "new store."

[1028] Step 7:

[1029] The server generates feedback to provide to the user based on the results of the audio, video, and slide analysis, including suggestions for adjusting speaking speed and volume, improving facial expressions and gestures, and correcting slides.

[1030] Input: Audio analysis results, video analysis results, slide analysis results

[1031] Output: Generated feedback

[1032] Specific operation: The server combines the results of each analysis and generates feedback on specific areas for improvement, such as "speaking too fast," "not making enough eye contact with the audience," and "there are typos on the slides."

[1033] Step 8:

[1034] The server sends the generated feedback to the device, which then displays it to the user, who receives real-time feedback and can immediately see how to improve their presentation.

[1035] Input: Generated feedback

[1036] Output: Feedback displayed to the user

[1037] What it does: The device will display the feedback as a pop-up notification for the user to view immediately.

[1038] Step 9:

[1039] The server manages user authentication information, such as authentication when a user logs into the system, and stores presentation history and past feedback results for future reference.

[1040] Input: User authentication information, presentation history, feedback results

[1041] Output: Authenticated user information, saved history and feedback

[1042] Specific operation: The server authenticates the user ID and password and stores each user's presentation history and feedback results in a database.

[1043] Step 10:

[1044] The server generates questions based on the user's presentation and presents them to the user via the terminal. The user answers the questions, and the answers are also analyzed to provide more detailed feedback.

[1045] Input: Presentation content, user responses

[1046] Output: Generated questions, analysis results and additional feedback

[1047] Specific operation: The server generates questions such as "What are the advantages of the new product?" and presents them to the user through the terminal. If the user answers "The advantage is multifunctionality," the answer is analyzed and additional feedback is generated.

[1048] (Application example 1)

[1049] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1050] Previous systems for improving presentation skills only provided feedback specific to presentations, which meant they were not particularly practical for customer service. Customer service skills also depend on a variety of important factors, such as voice, facial expressions, gestures, eye contact, and slides, but there was no system that could comprehensively analyze these and provide specific feedback to customer service staff in real time. Furthermore, while improving the ability to respond to questions in customer service is also important, previous technologies were inadequate in this regard as well.

[1051] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1052] In this invention, the server includes a user interface means for allowing a user to input voice, facial expressions, gestures, slides, and eye contact; a data transmission means for transmitting data input from the user interface means to the server; an audio analysis means for converting the voice data received by the server into text and evaluating the user's speaking speed, volume, and pauses; a video analysis means for analyzing the user's facial expressions, eye contact, and gestures from the video data received by the server and evaluating their level of composure and confidence; a slide analysis means for analyzing the slides received by the server and evaluating the consistency of content, visual effect, and typos; a feedback generation means for generating feedback to the user based on the evaluation results of the audio analysis means, video analysis means, and slide analysis means; a feedback transmission means and a visual output means for transmitting the feedback generated by the feedback generation means to the user's terminal; and a question generation means for generating questions based on the user's customer service content and presenting them to the user. This allows customer service staff to receive specific feedback in real time on various aspects of customer service, such as voice, facial expressions, gestures, eye contact, and consistency of content. The question generation means also contributes to improving customer service skills.

[1053] "User interface means" refers to the functions of devices and software that allow the wait staff to input voice, facial expressions, gestures, slides, and eye movements.

[1054] The "data transmission means" is a mechanism for transmitting data input from the user interface means to the server.

[1055] The "voice analysis means" is a function that converts the voice data received by the server into text and evaluates the speaking speed, volume, and pauses.

[1056] The "video analysis means" is a function that analyzes the facial expressions, gaze, and gestures of the customer service staff from the video data received by the server, and evaluates their level of composure and confidence.

[1057] The "slide material analysis means" is a function that analyzes the slide materials received by the server and evaluates the consistency of the content, the visual effect, and typographical errors.

[1058] The "feedback generation means" is a mechanism for generating feedback to the customer service staff based on the evaluation results of the audio analysis means, video analysis means, and slide data analysis means.

[1059] The "feedback sending means" is a function that sends the generated feedback to the terminal of the customer service staff member.

[1060] "Visual output means" refers to a device or function that visually displays feedback on a terminal.

[1061] The "question generation means" is a function in which the server generates questions based on the content of customer service and presents them to the customer service staff.

[1062] The system for improving customer service skills according to the present invention can be implemented as follows. The system is composed of a customer service staff member, a terminal, and a server. The customer service staff member is a trainee and interfaces with the system through the terminal. The terminal serves to collect the customer service staff member's voice, facial expressions, gestures, eye movements, and slide materials and transmit them to the server. The server analyzes the received data and provides feedback to the customer service staff member.

[1063] Basic configuration

[1064] The system includes the following basic means:

[1065] 1. User interface means: Devices and software that allow wait staff to input voice, facial expressions, gestures, slides, and eye movements. Equipped with a voice input device (microphone), a video input device (camera), and an eye movement detector, the wait staff can input voice, facial expressions, gestures, eye movements, and slides into the system.

[1066] 2. Data transmission means: The terminal has the function to transmit the audio, video and slides input by the customer service staff to the server. This process is performed in real time, and the data is buffered and sent to the server.

[1067] 3. Speech analysis: The server analyzes the received voice data. Specifically, it converts the voice into text using Google Cloud Speech-to-Text, and evaluates the speaking speed, volume, and pauses based on the text data. For example, if a customer service staff member speaks too quickly, this will be reflected in the feedback.

[1068] 4. Video analysis: The server analyzes the received video data. Using computer vision technologies such as OpenCV and dlib, it analyzes the facial expressions, gaze, and gestures of the wait staff, and evaluates their level of composure and confidence based on the analysis. For example, if the wait staff is not making eye contact with the viewer, this will be included in the feedback.

[1069] 5. Slide Analysis: The server analyzes the received slides. It uses optical character recognition (OCR) technology to read the content of the slides and evaluates whether they are consistent, visually appropriate, and free of typos. For example, if a slide contains a typographical error, this is reflected in the feedback.

[1070] 6. Feedback generation means and feedback transmission means: The server generates feedback to be provided to the customer service staff based on the evaluation results obtained by the audio analysis means, video analysis means, and slide analysis means. The generated feedback is sent to the terminal and visually displayed to the customer service staff. This feedback may include adjustments to speaking speed and volume, improvements to facial expressions and gestures, and corrections to slides.

[1071] 7. Question generation means and additional feedback generation means: The server generates questions based on the customer service content and presents them to the customer service staff. When the customer service staff answers these questions, the answers are also analyzed and additional feedback is generated. Questions are generated using a generative AI model such as OpenAI's GPT-3, and the customer service staff's answers are analyzed. For example, questions are generated on the theme of "explaining new products," and the necessary feedback is provided to the customer service staff when they answer them.

[1072] Specific examples

[1073] For example, if a customer service staff member practices on the theme of "explaining a new product," the system functions as follows: The customer service staff member practices using a terminal, and the server uses its audio analysis means to evaluate that "their voice is too quiet." It also uses its video analysis means to determine that "their eyes are always looking toward the product," and its slide analysis means to analyze that "there are typos in some of the slides." The server generates feedback based on these evaluations and provides the customer service staff with specific advice such as "increase the volume of your voice when speaking," "make eye contact with the customer," and "correct typos in the slides."

[1074] Example prompt sentence:

[1075] "Our customer service staff are explaining a new product. What advice would be effective? Please also give us specific examples of areas where our staff can improve."

[1076] In this way, the system of the present invention functions as an effective tool for customer service staff to efficiently improve their customer service skills.

[1077] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1078] Step 1:

[1079] The customer service staff uses the microphone, camera, and gaze detector of the terminal to emit voice and input facial expressions, gestures, gaze, and slides. The input data includes voice data, video data, gaze data, and slides.

[1080] Step 2:

[1081] The terminal buffers the voice data, video data, eye movement data, and slide material input by the customer service staff in real time and transmits them to the server using the data transmission means, whereby the voice, video, eye movement, and slide material data are input to the server.

[1082] Step 3:

[1083] The server converts the received voice data into text data using Google Cloud Speech-to-Text. The input voice data is analyzed and evaluated for speaking speed, volume, and pauses. For example, if the voice data indicates that the speaker is speaking too fast, an evaluation result will be generated.

[1084] Step 4:

[1085] The server uses OpenCV and dlib to analyze facial expressions, eye movements, and gestures from the received video data. It receives the video data as evaluation input and evaluates the customer's level of composure and confidence. For example, it generates an evaluation result such as "the customer's eyes are always directed toward the products."

[1086] Step 5:

[1087] The server uses optical character recognition (OCR) technology to read the content of the received slides and evaluates the consistency of the content, visual effect, and typographical errors. It receives the slides as input data and generates the evaluation result: "Some slides contain typographical errors."

[1088] Step 6:

[1089] The server generates feedback for the customer service staff based on the evaluation results obtained from the audio analysis means, video analysis means, and slide analysis means. For example, the generated feedback includes specific advice such as "increase your speaking volume," "make eye contact with the customer," and "correct typos in the slides."

[1090] Step 7:

[1091] The feedback generated by the feedback generating means is sent to the terminal via the feedback transmitting means and the visual output means. As feedback, information such as "increase your speaking volume" and "make eye contact with the customer" is displayed to the wait staff.

[1092] Step 8:

[1093] The server uses a generative AI model such as OpenAI's GPT-3 to generate questions based on the customer service content. The server generates appropriate questions based on the scenario the customer service staff is practicing and sends them to the device. For example, a question might be generated such as, "I'm explaining a new product. What kind of advice would be effective?"

[1094] Step 9:

[1095] The customer service staff responds to questions posed by the system and sends the answers to the server via their terminal. The server analyzes the received answers and generates additional feedback based on the evaluation results.

[1096] Step 10:

[1097] The server transmits the additional feedback generated by the additional feedback generating means to the terminal using the feedback transmitting means and the visual output means, thereby providing further specific advice to the wait staff.

[1098] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1099] The system for improving presentation skills according to the present invention can provide more advanced feedback by incorporating an emotion engine that analyzes the user's emotions. Specific embodiments will be described below.

[1100] Basic configuration

[1101] The system is composed of a user, a terminal, and a server. The user is the person giving the presentation and interfaces with the system through the terminal. The terminal collects the user's voice, facial expressions, gestures, and slides and transmits them to the server. The server analyzes the received data and provides feedback to the user. In addition, this embodiment includes an emotion engine that recognizes the user's emotions.

[1102] User Interface Means

[1103] A user gives a presentation using a terminal. The terminal is equipped with an audio input device (microphone), a video input device (camera), and a screen for displaying slides. An emotion engine is also installed, which analyzes the user's emotions.

[1104] Data transmission method

[1105] The device has the ability to transmit audio, video, and slides input by the user to the server. This process is performed in real time, and the data is buffered and sent to the server. The user's emotional data is also transmitted in the same way.

[1106] Voice analysis methods

[1107] The server analyzes the received audio data, converts it into text using speech recognition technology, and evaluates the speed, volume, and pauses of speech based on the text data. For example, if the user speaks too quickly, this will be reflected in the feedback.

[1108] Video analysis methods

[1109] The server analyzes the received video data, using computer vision technology to analyze the user's facial expressions, eye contact, and gestures, and evaluates the presentation's poise and confidence. For example, if the user is not making eye contact with the audience, this will be included in the feedback.

[1110] Slide analysis tools

[1111] The server analyzes the slides it receives, using optical character recognition (OCR) technology to read their content and evaluate whether they are consistent with the presentation theme and whether their visual impact is appropriate. For example, if a slide contains a typographical error, that error is reflected in the feedback.

[1112] Emotion Engine

[1113] The server uses an emotion engine to analyze the user's emotional data. Specifically, it analyzes the user's facial expressions and voice to evaluate the fluctuations and stability of their emotions during the presentation. For example, if the user is nervous during the presentation, that emotional data will be included in the feedback.

[1114] Feedback generation means and feedback transmission means

[1115] The server integrates the evaluation results of the audio analysis means, video analysis means, slide analysis means, and emotion engine to generate feedback for the user. The generated feedback is sent to the terminal and displayed to the user. The feedback includes advice on adjusting speaking speed and volume, improving facial expressions and gestures, correcting slides, and emotional advice.

[1116] Authentication Management Methods

[1117] The server manages user authentication information and stores the user's presentation history and feedback results, allowing users to check their progress and refer to past feedback.

[1118] Question generation and additional feedback generation

[1119] The server generates questions based on the user's presentation and presents them to the user. When the user answers these questions, the answers are also analyzed and additional feedback is generated. For example, if the user does not answer a question accurately, that point is reflected in the additional feedback.

[1120] Specific examples

[1121] For example, if a user gives a presentation on the theme of "introducing a new product," the system functions as follows: The user gives the presentation using a device, and the server uses its audio analysis means to evaluate the user's voice as "quiet." The video analysis means determines that the user's eyes are always looking at the slides, and the slide analysis means analyzes that there are typos on some of the slides. The emotion engine then determines that the user is nervous. The server generates feedback based on these evaluations and provides the user with specific advice such as "increase the volume of your voice," "make eye contact with the audience," "correct typos in the slides," and "regulate your breathing to relieve tension."

[1122] In this way, the system of the present invention functions as an effective tool for users to efficiently improve their presentation skills. In particular, the introduction of an emotion engine makes it possible to provide comprehensive feedback that takes into account the user's emotional state.

[1123] The processing flow will be explained below.

[1124] Step 1:

[1125] A user logs into the system using a terminal, enters a username and password, and submits the authentication information.

[1126] Step 2:

[1127] The device sends the entered authentication information to the server. The server compares it with the database and returns the authentication result to the device. If authentication is successful, the user can start the presentation.

[1128] Step 3:

[1129] The user begins preparing for a presentation. They upload presentation materials (slides) using their device. The device then sends the uploaded materials to the server.

[1130] Step 4:

[1131] The server stores the received presentation materials in an appropriate format.

[1132] Step 5:

[1133] The user presses the "Start Presentation" button. The device begins recording the user's voice and video (facial expressions and gestures) using the built-in camera and microphone. The recorded data is buffered.

[1134] Step 6:

[1135] The terminal transmits the buffered data to the server in real time, including audio, video, slides, and emotion data.

[1136] Step 7:

[1137] The server analyzes the received voice data, converts the voice data into text using a speech recognition engine, evaluates speaking speed, volume, and pauses, and generates text data.

[1138] Step 8:

[1139] The server analyzes the received video data and uses computer vision technology to analyze the user's facial expressions, eye movements, and gestures to evaluate the presentation's poise and confidence.

[1140] Step 9:

[1141] The server analyzes the slides received, using optical character recognition (OCR) technology to read the content of the slides and evaluate their consistency and visual impact.

[1142] Step 10:

[1143] The server uses an emotion engine to analyze the user's emotional data, identifying emotions from facial expressions and vocal tones, and assessing emotional stability and fluctuations.

[1144] Step 11:

[1145] The server integrates the evaluation results obtained from the audio analysis means, video analysis means, slide analysis means, and emotion engine, and generates feedback for the user, including advice on speaking speed, volume, naturalness of gestures, corrections to slides, and emotional advice.

[1146] Step 12:

[1147] The server sends the generated feedback to the user's terminal, which receives the feedback and displays it to the user.

[1148] Step 13:

[1149] Users can review the feedback and incorporate improvements for their next presentation. Users can also request a Q&A simulation if necessary.

[1150] Step 14:

[1151] If a question and answer simulation is desired, the server generates questions based on the content of the presentation and presents them to the user.

[1152] Step 15:

[1153] The user answers the questions, and the device sends the answer data to the server.

[1154] Step 16:

[1155] The server analyzes the user's answers and generates additional feedback, providing further improvements based on the evaluation results.

[1156] Step 17:

[1157] The server sends additional feedback to the user's device, which the user receives and uses as a reference for their next presentation.

[1158] This series of processing steps allows users to improve their presentation skills efficiently and effectively. In addition, the introduction of an emotion engine makes it possible to provide comprehensive feedback that takes emotional aspects into consideration.

[1159] Example 2

[1160] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1161] Current presentation skill improvement systems have difficulty accurately analyzing users' emotions and reflecting the analysis results in presentation feedback. Furthermore, they lack feedback on not only the technical aspects of the presentation but also the emotional and psychological aspects, making it difficult for users to comprehensively improve their skills. Furthermore, existing systems lack sufficient functionality for generating questions and providing additional feedback, making it difficult for users to self-evaluate.

[1162] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1163] In this invention, the server includes an audio analysis means, a video analysis means, a slide analysis means, an emotion engine, a feedback generation means, and a feedback transmission means. This allows for comprehensive analysis of a user's voice, facial expressions, gestures, slides, and even emotions, making it possible to provide detailed and accurate feedback. Furthermore, by including a question generation means and an additional feedback generation means, it is possible to automatically generate questions based on the content of the user's presentation and evaluate the answers to those questions, facilitating self-evaluation.

[1164] "User interface means" refers to devices and software that allow a user to input voice, facial expressions, gestures, and slide materials.

[1165] "Data transmission means" refers to a device or software for transmitting data input from the user interface means to the server.

[1166] "Voice analysis means" refers to devices or software that convert the voice data received by the server into text and evaluate speaking speed, volume, and pauses.

[1167] "Video analysis means" refers to devices or software that analyze the user's facial expressions, gaze, and gestures from the video data received by the server, and evaluate the presentation's level of composure and confidence.

[1168] "Slide material analysis means" refers to a device or software that analyzes the slide materials received by the server and evaluates the consistency of the content and the visual effect.

[1169] "Emotion engine" refers to a device or software that allows the server to analyze emotion data and evaluate the fluctuations and stability of a user's emotions.

[1170] "Feedback generation means" refers to a device or software for generating feedback to a user based on the evaluation results of the audio analysis means, video analysis means, slide data analysis means, and emotion engine.

[1171] The "feedback sending means" refers to a device or software for sending the feedback generated by the feedback generating means to the user's terminal.

[1172] "Authentication management means" refers to a device or software for managing user authentication information and storing each user's presentation history and feedback results.

[1173] The "question generation means" refers to a device or software that generates questions based on the content of a user's presentation and presents them to the user.

[1174] The "additional feedback generating means" refers to a device or software that analyzes the answers given by the user to the questions presented and generates additional feedback based on the evaluation results.

[1175] The system for improving presentation skills according to the present invention provides more advanced feedback by incorporating an emotion engine that analyzes the user's emotions. Specific embodiments will be described below.

[1176] Basic configuration

[1177] The system is composed of a user, a terminal, and a server. The user is the person giving the presentation and interfaces with the system through the terminal. The terminal collects the user's voice, facial expressions, gestures, and slides and transmits them to the server. The server analyzes the received data and provides feedback to the user. In addition, this embodiment includes an emotion engine that recognizes the user's emotions.

[1178] User Interface Means

[1179] A user gives a presentation using a terminal. The terminal is equipped with an audio input device (microphone), a video input device (camera), and a screen for displaying slides. An emotion engine is also installed, which analyzes the user's emotions.

[1180] Data transmission method

[1181] The device has the ability to transmit audio, video, and slides input by the user to the server. This process is performed in real time, and the data is buffered and sent to the server. The user's emotional data is also transmitted in the same way.

[1182] Voice analysis methods

[1183] The server analyzes the received audio data, converts it into text using speech recognition technology, and evaluates speaking speed, volume, and pauses based on the text data. The specific software used is the Google Cloud Speech-to-Text API. For example, if the user speaks too quickly, this is reflected in the feedback.

[1184] Video analysis methods

[1185] The server analyzes the received video data. It uses computer vision technology to analyze the user's facial expressions, eye movements, and gestures, and evaluates the presentation's poise and confidence based on the results. The OpenCV library is used for this software. For example, if the user is not making eye contact with the audience, this will be included in the feedback.

[1186] Slide analysis tools

[1187] The server analyzes the received slides. It uses optical character recognition (OCR) technology to read the content of the slides and evaluate whether they are consistent with the presentation theme and have an appropriate visual effect. Specific software includes Tesseract OCR. For example, if a slide contains a typographical error, that error is reflected in the feedback.

[1188] Emotion Engine

[1189] The server uses an emotion engine to analyze the user's emotional data. Specifically, it analyzes the user's facial expressions and voice to evaluate the fluctuations and stability of their emotions during the presentation. For example, if the user is nervous during the presentation, that emotional data will be included in the feedback.

[1190] Feedback generation means and feedback transmission means

[1191] The server integrates the evaluation results of the audio analysis means, video analysis means, slide analysis means, and emotion engine to generate feedback for the user. The generated feedback is sent to the terminal and displayed to the user. The feedback includes advice on adjusting speaking speed and volume, improving facial expressions and gestures, correcting slides, and emotional advice.

[1192] Authentication Management Methods

[1193] The server manages user authentication information and stores the user's presentation history and feedback results, allowing users to check their progress and refer to past feedback.

[1194] Question generation and additional feedback generation

[1195] The server generates questions based on the user's presentation and presents them to the user. When the user answers these questions, the answers are also analyzed and additional feedback is generated. For example, if the user does not answer a question accurately, that point is reflected in the additional feedback.

[1196] Specific examples

[1197] For example, if a user is giving a presentation on the theme of "introducing a new product," the system functions as follows: The user gives the presentation using a terminal, and the server uses the audio analysis means to evaluate the user's voice as "quiet." The video analysis means determines that the user's eyes are always directed toward the slides, and the slide analysis means analyzes that there are typos on some slides. The emotion engine then determines that the user is nervous. Based on these evaluations, the server generates feedback and provides the user with specific advice, such as "increase your speaking volume," "make eye contact with the audience," "correct typos in the slides," or "regulate your breathing to relieve tension." In this way, the system of the present invention functions as an effective tool for users to efficiently improve their presentation skills. In particular, the introduction of the emotion engine makes it possible to provide comprehensive feedback that takes the user's emotional state into consideration.

[1198] Prompt Sentence Examples

[1199] Here are some example prompts to input to a generative AI model:

[1200] A user will give a presentation introducing a new product. Staff will generate feedback based on the analysis of audio, video, and slides. Also, analyze the user's emotional data and include it in the feedback. Examples of feedback include "your voice is too quiet," "you should make eye contact with the audience," "there is a typo on the slide," and "you seem nervous."

[1201] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1202] Specific processing flow of the program

[1203] Step 1: Start your presentation

[1204] The user clicks a button on the device to start the presentation. The device verifies the user's authentication information and sends a request to the server to start the session.

[1205] Input: User credentials, session initiation request

[1206] Processing: The server verifies the authentication information and generates a token to start the session.

[1207] Output: Session initiation token

[1208] Step 2: Data collection

[1209] The user gives a presentation, and the device collects audio, video, and slides via a microphone, camera, and screen display.

[1210] Input: User's audio, video, and slides

[1211] Processing: The device buffers the data in real time and prepares it for transfer to the server.

[1212] Output: Buffered data

[1213] Step 3: Send data

[1214] The terminal transmits the collected data (audio, video, slides, and emotional data) to the server in real time.

[1215] Input: Buffered audio, video, slide deck, emotion data

[1216] Processing: The device divides the data into packets and sends them over the network to the server.

[1217] Output: Data packets arriving at the server

[1218] Step 4: Audio data analysis

[1219] The server analyzes the received voice data, converts the voice into text using speech recognition technology, and evaluates the speaking speed, volume, and pauses based on the text data.

[1220] Input: Audio data

[1221] Processing: Converts audio to text using the Google Cloud Speech-to-Text API and analyzes data on speaking rate, volume, and pauses.

[1222] Output: Text data, speaking speed, volume, and interval evaluation data

[1223] Step 5: Video data analysis

[1224] The server analyzes the received video data, using facial recognition technology and computer vision to analyze the user's facial expressions, gaze, and gestures, and evaluates the presentation's poise and confidence based on that information.

[1225] Input: Video data

[1226] Processing: Using the OpenCV library, we analyze the user's facial expressions, gaze, and gestures to assess their calmness and confidence.

[1227] Output: Facial expression, gaze, gesture analysis data, calmness and confidence evaluation data

[1228] Step 6: Slide analysis

[1229] The server analyzes the received slides, using optical character recognition (OCR) technology to read the slide content and evaluate whether it is consistent with the presentation theme and whether the visual effects are appropriate.

[1230] Input: Slide data

[1231] Processing: Extract text from slides using Tesseract OCR and evaluate content consistency and visual impact.

[1232] Output: Slide text data, consistency and visual effect evaluation data

[1233] Step 7: Sentiment Data Analysis

[1234] The server uses an emotion engine to analyze the user's emotion data. Specifically, it analyzes emotions from the user's facial expressions and voice and evaluates the fluctuations and stability of emotions during the presentation.

[1235] Input: Emotion data (facial expressions, voice)

[1236] Processing: An emotion engine is used to analyze facial and voice data to assess emotional variability and stability.

[1237] Output: Emotional evaluation data

[1238] Step 8: Feedback Generation

[1239] The server integrates the results of the audio analysis, video analysis, slide analysis, and emotion data analysis, and generates feedback for the user.

[1240] Input: Audio analysis data, video analysis data, slide analysis data, emotion analysis data

[1241] Processing: Comprehensive feedback is generated based on the results of each assessment. The feedback is provided in text and graphical format.

[1242] Output: Feedback data (text, graph)

[1243] Step 9: Send and view feedback

[1244] The server sends the generated feedback to the terminal, which displays it to the user.

[1245] Input: Feedback data

[1246] Processing: The server sends the feedback data to the terminal, which displays it to the user. The feedback includes specific advice.

[1247] Output: Feedback that is displayed to the user

[1248] Step 10: Question generation and additional feedback

[1249] The server generates relevant questions based on the user's presentation and presents them to the user via the terminal. When the user answers the questions, the server analyzes them again and generates additional feedback.

[1250] Input: Presentation content data, user questions and answers

[1251] Processing: A question generation algorithm is used to generate questions based on the presentation content, and the answers are analyzed and evaluated. Additional feedback is generated based on the evaluation results.

[1252] Output: Additional feedback data

[1253] Step 11: Authentication Management and History Storage

[1254] The server manages user authentication information and feedback history, and stores them in a database after the presentation session ends.

[1255] Input: User authentication information, feedback history data

[1256] Processing: Manages authentication information and stores feedback history in a database.

[1257] Output: Updated database

[1258] Through this process, users receive feedback to comprehensively improve their presentation skills.

[1259] (Application example 2)

[1260] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1261] Conventional presentation skill improvement systems are primarily limited to analyzing voice, facial expressions, and slides, and are not capable of improving specialized skills tailored to specific environments or situations. For example, to improve a factory operator's skills so that they can issue appropriate commands to a robot, analysis and feedback that takes into account the appropriateness of the robot's operation instructions and the operator's emotional state is necessary. A system that can address this issue is needed.

[1262] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a user interface means for a user to input voice, facial expressions, gestures, and slide materials, a data transmission means for transmitting the data input from the user interface means to the server, a voice analysis means for converting voice data received by the server into text and evaluating the speaking speed, volume, and pauses, a video analysis means for analyzing the user's facial expressions, line of sight, and gestures from video data received by the server and evaluating the calmness and confidence of the presentation, a slide material analysis means for analyzing the slide materials received by the server and evaluating the consistency of the content and the visual effect, a feedback generation means for generating feedback to the user based on the evaluation results of the voice analysis means, the video analysis means, and the slide material analysis means, a feedback transmission means for transmitting the feedback generated by the feedback generation means to the user's terminal, a robot operation instruction evaluation means for evaluating the appropriateness of instructions given by the operator to the robot, and an emotion evaluation means for analyzing the operator's emotional state and evaluating the mental state that affects the appropriateness of the instructions. This will enable factory operators to improve their skills in operating robots properly.

[1263] "User interface means" refers to devices and software that allow a user to input voice, facial expressions, gestures, and slide materials.

[1264] The "data transmission means" is a communication means for transmitting data input from the user interface means to the server in real time.

[1265] The "voice analysis means" is a technology that converts the voice data received by the server into text and evaluates the speaking speed, volume, and pauses.

[1266] The "video analysis means" is a technology that analyzes the user's facial expressions, gaze, and gestures from the video data received by the server, and evaluates the presentation's level of composure and confidence.

[1267] The "slide material analysis means" is a technology for analyzing the slide materials received by the server and evaluating the consistency of the content and the visual effect.

[1268] The "feedback generation means" is a technology that generates feedback indicating areas for improvement to the user based on the evaluation results of the audio analysis means, video analysis means, and slide data analysis means.

[1269] The "feedback transmitting means" is a means for transmitting the feedback generated by the feedback generating means to the user's terminal.

[1270] The "robot operation instruction evaluation means" is a technology for evaluating the appropriateness of instructions given by an operator to a robot.

[1271] "Emotion evaluation means" is a technology that analyzes the emotional state of an operator and evaluates the mental state that affects the appropriateness of instructions.

[1272] "Authentication management means" is a technology that enables the server to manage user authentication information and store each user's presentation history, instruction history, and feedback results.

[1273] The "question generation means" is a technology for generating questions based on the content of the user's presentation and instructions for operating the robot, and presenting the questions to the user.

[1274] The "additional feedback generating means" is a technology that analyzes the answers given by the user to the questions presented and generates additional feedback based on the evaluation results.

[1275] The system according to the present invention is intended to improve the skills of factory operators to properly operate robots. This system is composed of a user (operator), a terminal, and a server. Specific embodiments are described below.

[1276] Basic configuration

[1277] The system includes a user interface for inputting voice, facial expressions, gestures, and instructions from the user's terminal. These data are transmitted in real time from the terminal to a server, which analyzes the received data and provides feedback to the user.

[1278] Voice analysis methods

[1279] The server first receives the user's voice data and converts it into text using speech recognition software (for example, Google's speech recognition API). The server then evaluates the user's speaking speed, volume, and pauses. Any inappropriate parts are reflected in the feedback.

[1280] Video analysis methods

[1281] The server then analyzes the user's video data. Using video analysis technology (e.g., OpenCV), it detects and evaluates the user's facial expressions, eye movements, and gestures. The server evaluates the user's poise and confidence during the presentation, and provides feedback on any shortcomings.

[1282] Slide analysis tools

[1283] The server analyzes the slide deck it receives and uses optical character recognition (OCR) technology to review the slide content and evaluate its consistency and visual effectiveness. For example, typographical errors and visual flaws are flagged.

[1284] Emotion assessment tools

[1285] The server then uses an emotion engine to analyze the user's emotional data. It extracts emotions from the user's facial expressions and voice, and evaluates the emotional fluctuations and stability during the presentation. If tension or stress is detected, this is also reflected in the feedback.

[1286] Robot operation instruction evaluation means

[1287] In a specific application, the server evaluates the appropriateness of the instructions given by the operator to the robot, for example, whether the instructions are clear and unambiguous when given by voice or gesture.

[1288] Feedback generation and transmission methods

[1289] The server then aggregates all of these analysis results and generates specific feedback for the user, suggesting areas for improvement. The generated feedback is then sent to the user's device and displayed to the user.

[1290] Authentication Management Methods

[1291] The server also manages user authentication information and stores each user's presentation history, instruction history, and feedback results, allowing users to improve themselves based on past data.

[1292] Question generation and additional feedback generation means

[1293] The server generates questions based on the user's presentation and robot operation instructions, and presents them to the user. When the user answers these questions, the server analyzes the answers and generates additional feedback.

[1294] Specific examples

[1295] For example, imagine a scene in which an operator is giving verbal instructions to a robot in a factory. The scene is recorded using the user interface means, and the audio and video data are sent to the server. The server analyzes the audio and judges it as "too quiet," and analyzes the video and judges it as "tension is evident in the operator's facial expression." Furthermore, the robot operation instruction evaluation means judges it as "ambiguous." Based on these evaluations, the server generates specific feedback and provides the user with advice such as "speak louder" or "give clearer instructions."

[1296] Prompt Sentence Examples

[1297] "Analyze sample videos, evaluate the operator's emotional state and the appropriateness of their voice instructions, and provide feedback."

[1298] This system allows factory operators to effectively improve their robot operation skills, aiming to efficiently improve productivity.

[1299] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1300] Step 1: Data entry

[1301] The user inputs voice, facial expressions, gestures, and instructions. Specifically, data is collected using user interface means (camera, microphone, touch panel, etc.). The inputs in this step are the user's voice data, video data, and instructions. The output is that these data are stored in the terminal.

[1302] Step 2: Send data

[1303] The terminal transmits the voice data, video data, and instruction contents input from the user interface means to the server. Specifically, the data is transmitted to the server in real time via the Internet. The input of this step is the user's voice data, video data, and instruction contents, and the output is the server receiving these data.

[1304] Step 3: Audio analysis

[1305] The server analyzes the received voice data. Specifically, it uses speech recognition software (such as Google's speech recognition API) to convert the speech into text and evaluates the speaking speed, volume, and pauses. The input of this step is the received voice data, and the output is the analyzed text data and the evaluation results.

[1306] Step 4: Video analysis

[1307] The server analyzes the received video data. Specifically, it uses video analysis technology (such as OpenCV) to detect the user's facial expressions, eye movements, and gestures, and evaluates their level of composure and confidence during the presentation. The input to this step is the received video data, and the output is the analyzed facial expression data, eye movement data, and gesture data, along with the evaluation results.

[1308] Step 5: Slide analysis

[1309] The server analyzes the received slides. Specifically, it uses optical character recognition (OCR) technology to review the content of the slides and evaluate their consistency and visual effectiveness. The input for this step is the received slides, and the output is the analyzed text data and the evaluation results.

[1310] Step 6: Emotional assessment

[1311] The server evaluates emotions using the received video and audio data. Specifically, it uses emotion recognition technology (such as Face++ or Microsoft Azure Face API) to evaluate the user's emotional state and evaluate the fluctuations and stability of emotions during the presentation. The input for this step is facial expression data and audio data, and the output is analyzed emotional data and its evaluation results.

[1312] Step 7: Evaluate robot operation instructions

[1313] The server evaluates the operator's instruction data. Specifically, it uses voice recognition and video analysis to evaluate the appropriateness of the operator's instructions to the robot. The input for this step is the operator's voice data and gesture data, and the output is the analysis result.

[1314] Step 8: Feedback Generation

[1315] The server integrates all the above analysis results and generates specific feedback for the user indicating areas for improvement. Feedback is generated based on the overall evaluation results using a feedback generation means. The inputs for this step are the audio analysis results, video analysis results, slide analysis results, emotion evaluation results, and robot operation instruction evaluation results, and the output is the generated feedback information.

[1316] Step 9: Submit your feedback

[1317] The server sends the generated feedback to the user's terminal, specifically, provides the feedback information to the user via the Internet. The input of this step is the generated feedback information, and the output is the feedback being displayed on the user's terminal.

[1318] Step 10: Question generation and additional feedback generation

[1319] The server generates additional questions based on the user's presentation and instructions and presents them to the user. The server further analyzes the user's answers and generates additional feedback, which helps the user identify areas for further improvement. The input of this step is the user's answer data, and the output is the additional feedback information.

[1320] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1321] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1322] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1323] [Fourth embodiment]

[1324] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1325] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1326] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1327] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1328] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1329] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1330] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1331] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1332] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1333] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1334] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1335] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1336] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1337] The system for improving presentation skills according to the present invention can be implemented as follows.

[1338] Basic configuration

[1339] The system consists of a user, a terminal, and a server. The user is the person giving the presentation and interfaces with the system through the terminal. The terminal collects the user's voice, facial expressions, gestures, and slides and transmits them to the server. The server analyzes the received data and provides feedback to the user.

[1340] User Interface Means

[1341] A user gives a presentation using a terminal. The terminal is equipped with an audio input device (microphone), a video input device (camera), and a screen for displaying slides. This allows the user to input voice, facial expressions, gestures, and slides into the system.

[1342] Data transmission method

[1343] The terminal has the function of transmitting the audio, video, and slides input by the user to the server. This process is performed in real time, and the data is buffered and sent to the server.

[1344] Voice analysis methods

[1345] The server analyzes the received audio data. Specifically, it uses speech recognition technology to convert the audio into text, and then evaluates the speaking speed, volume, and pauses based on the text data. For example, if the user speaks too fast, this will be reflected in the feedback.

[1346] Video analysis methods

[1347] The server analyzes the received video data, using computer vision technology to analyze the user's facial expressions, eye contact, and gestures, and evaluates the presentation's poise and confidence. For example, if the user is not making eye contact with the audience, this will be included in the feedback.

[1348] Slide analysis tools

[1349] The server analyzes the slides it receives, using optical character recognition (OCR) technology to read their content and evaluate whether they are consistent with the presentation theme and whether the visual effect is appropriate. For example, if a slide contains a typographical error, that will be reflected in the feedback.

[1350] Feedback generation means and feedback transmission means

[1351] The server generates feedback to be provided to the user based on the evaluation results obtained by the audio analysis means, video analysis means, and slide analysis means. The generated feedback is sent to the terminal and displayed to the user. The feedback may include adjustments to speaking speed and volume, improvements to facial expressions and gestures, and corrections to slides.

[1352] Authentication Management Methods

[1353] The server manages user authentication information and stores the user's presentation history and feedback results, allowing users to check their progress and refer to past feedback.

[1354] Question generation and additional feedback generation

[1355] The server generates questions based on the user's presentation and presents them to the user. When the user answers these questions, the answers are also analyzed and additional feedback is generated. For example, if the user does not answer a question accurately, that point is reflected in the additional feedback.

[1356] Specific examples

[1357] For example, if a user gives a presentation on the theme of "introducing a new product," the system functions as follows: The user gives the presentation using a device, and the server uses its audio analysis means to evaluate that the user's voice is "quiet." The video analysis means determines that the user's eyes are always looking at the slides, and the slide analysis means analyzes that there are typos on some of the slides. The server generates feedback based on these evaluations and provides the user with specific advice such as "increase the volume of your voice when speaking," "make eye contact with the audience," and "correct typos in the slides."

[1358] In this way, the system of the present invention functions as an effective tool for users to efficiently improve their presentation skills.

[1359] The processing flow will be explained below.

[1360] Step 1:

[1361] A user logs into the system using a terminal, enters a username and password, and submits the authentication information.

[1362] Step 2:

[1363] The device sends the entered authentication information to the server. The server compares it with the database and returns the authentication result to the device. If authentication is successful, the user can start the presentation.

[1364] Step 3:

[1365] The user begins preparing for a presentation. They upload presentation materials (slides) using their device. The device then sends the uploaded materials to the server.

[1366] Step 4:

[1367] The server stores the received presentation materials in an appropriate format.

[1368] Step 5:

[1369] The user presses the "Start Presentation" button. The device begins recording the user's voice and video (facial expressions and gestures) using the built-in camera and microphone. The recorded data is buffered.

[1370] Step 6:

[1371] The terminal transmits the buffered data to the server in real time. The data for the audio, video, and slide materials are transmitted to the server.

[1372] Step 7:

[1373] The server analyzes the received voice data, converts it into text using a speech recognition engine, and evaluates speaking speed, volume, and pauses.

[1374] Step 8:

[1375] The server analyzes the received video data and uses computer vision technology to analyze the user's facial expressions, eye movements, and gestures to evaluate the presentation's poise and confidence.

[1376] Step 9:

[1377] The server analyzes the slides it receives, using optical character recognition (OCR) technology to read their content and evaluate whether they are consistent with the presentation theme and whether their visual impact is appropriate.

[1378] Step 10:

[1379] The server integrates the results of the audio, video, and slide analysis and generates feedback for the user. Based on each evaluation result, it summarizes specific improvements and advice in written form.

[1380] Step 11:

[1381] The server sends the generated feedback to the user's terminal, which receives the feedback and displays it to the user.

[1382] Step 12:

[1383] Users can review the feedback and incorporate improvements for their next presentation. If desired, a simulated Q&A session can also be conducted.

[1384] Step 13:

[1385] If a question and answer simulation is desired, the server generates questions based on the content of the presentation and presents them to the user.

[1386] Step 14:

[1387] The user answers the questions, and the device sends the answer data to the server.

[1388] Step 15:

[1389] The server analyzes the user's answers and generates additional feedback, and reflects suggestions for improvement in the advice based on the evaluation results.

[1390] Step 16:

[1391] The server sends additional feedback to the user's device, which the user receives and uses as a reference for their next presentation.

[1392] This series of processing steps allows users to improve their presentation skills efficiently and effectively.

[1393] Example 1

[1394] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1395] Conventional presentation support systems lack the ability to comprehensively analyze a user's voice, facial expressions, gestures, and slides to provide comprehensive feedback. They also lack the ability to generate questions based on the user's presentation content, analyze the answers, and provide additional feedback. This makes it difficult for users to effectively improve their presentation skills.

[1396] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1397] In this invention, the server includes a user interface means for a user to input voice, facial expressions, gestures, and slide materials, a data transmission means for transmitting the data input from the user interface means to the server, an audio analysis means for converting the voice data received by the server into text and evaluating the speaking speed, volume, and pauses, a video analysis means for analyzing the user's facial expressions, gaze, and gestures from the video data received by the server and evaluating the poise and confidence of the presentation, a slide material analysis means for analyzing the slide materials received by the server and evaluating the consistency and visual effect of the content, a feedback generation means for generating feedback to the user based on the evaluation results of the audio analysis means, the video analysis means, and the slide material analysis means, a feedback transmission means for transmitting the feedback generated by the feedback generation means to the user's terminal, a question generation means for the server to generate questions based on the content of the user's presentation and present them to the user, and an additional feedback generation means for analyzing the user's answers to the questions posed via the user interface means and generating additional feedback based on the evaluation results. This allows the user to receive comprehensive and detailed feedback, thereby effectively improving their presentation skills.

[1398] "User interface means" refers to means by which a user inputs voice, facial expressions, gestures, and slide materials.

[1399] The "data transmission means" is a means for transmitting data input from the user interface means to the server.

[1400] The "voice analysis means" is a means for converting voice data received by the server into text and evaluating speaking speed, volume, and pauses.

[1401] The "video analysis means" is a means for analyzing the user's facial expressions, gaze, and gestures from the video data received by the server, and evaluating the user's level of composure and confidence in the presentation.

[1402] The "slide material analysis means" is a means for analyzing the slide materials received by the server and evaluating the consistency of the content and the visual effect.

[1403] The "feedback generation means" is a means for generating feedback to the user based on the evaluation results of the audio analysis means, the video analysis means, and the slide material analysis means.

[1404] The "feedback transmitting means" is a means for transmitting the feedback generated by the feedback generating means to the user's terminal.

[1405] The "question generation means" is a means by which the server generates a question based on the content of the user's presentation and presents it to the user.

[1406] The "additional feedback generating means" is a means for analyzing the answers given by the user to questions presented via the user interface means and generating additional feedback based on the evaluation results.

[1407] The "authentication management means" is a means by which the server manages user authentication information and stores each user's presentation history and feedback results.

[1408] The system for improving presentation skills according to the present invention comprises a user, a terminal, and a server. By using this system, the user can effectively analyze their voice, facial expressions, gestures, and slides, and receive comprehensive feedback.

[1409] User Interface Means

[1410] When giving a presentation, a user uses a terminal equipped with a microphone, a camera, and a screen for displaying slides. This allows the user to input voice, facial expressions, gestures, and slides.

[1411] Data transmission method

[1412] The terminal transmits the audio, video, and slides input by the user to the server in real time. During this process, the data is buffered, allowing for smooth data transfer in a stable communication environment.

[1413] Voice analysis methods

[1414] The server analyzes the received voice data. Specifically, it converts the voice into text using speech recognition technology (e.g., Google Cloud Speech-to-Text API), and evaluates the speaking speed, volume, and pauses based on the text data. For example, if a user says "introducing a new product," if they speak too quickly, this will be reflected in the feedback.

[1415] Video analysis methods

[1416] The server analyzes the received video data. Using computer vision techniques (e.g., OpenCV and TensorFlow), it analyzes the user's facial expressions, eye movements, and gestures to assess the presentation's poise and confidence. For example, if the user's eyes are fixed on the slides and they are not making eye contact with the audience, this will be included in the feedback.

[1417] Slide analysis tools

[1418] The server analyzes the slides it receives. It uses optical character recognition (OCR) technology (e.g., Tesseract OCR) to read the content of the slides and evaluates whether they are consistent with the presentation theme and whether the visual effect is appropriate. For example, if a slide contains a typographical error, that error is reflected in the feedback.

[1419] Feedback Generation Method

[1420] The server generates feedback to be provided to the user based on the evaluation results obtained by the audio analysis means, video analysis means, and slide analysis means. The generated feedback includes adjustments to speaking speed and volume, improvements to facial expressions and gestures, and corrections to slides.

[1421] Feedback sending method

[1422] The server sends the generated feedback to the device, which then displays it to the user, who receives real-time feedback and can immediately see how to improve their presentation.

[1423] Question generation means

[1424] The server generates questions based on the content of the user's presentation and presents them to the user via the terminal, encouraging the user to consider further the presentation.

[1425] Additional Feedback Generation Means

[1426] The server analyzes the answers given by the user to the questions posed and generates additional feedback based on the evaluation results, allowing the user to receive feedback at a deeper level.

[1427] Authentication Management Methods

[1428] The server manages user authentication information and stores presentation history and past feedback results, allowing users to check their progress and continuously improve their presentation skills.

[1429] Specific examples

[1430] For example, if a user is giving a presentation on the theme of "introducing a new product," the system functions as follows: The user gives the presentation using a terminal, and the server uses the audio analysis means to evaluate that the user's voice is quiet. The video analysis means determines that the user's eyes are always looking at the slides, and the slide analysis means analyzes that there are typos in some of the slides. The server generates feedback based on these evaluations and provides the user with specific advice, such as speaking louder, making eye contact with the audience, or correcting typos in the slides. In this way, the system of the present invention functions as an effective tool for users to efficiently improve their presentation skills.

[1431] Prompt Sentence Examples

[1432] For example, if you are giving a presentation on the topic of "Introducing a New Product," you might use the following prompt:

[1433] "You will be asked to give a presentation on the topic of introducing a new product. A microphone and camera will be used to collect data, and analysis will be performed using speech recognition, computer vision, and optical character recognition (OCR). Feedback will include an evaluation of speaking speed, volume, facial expressions, eye contact, gestures, and the content of your slides."

[1434] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1435] Step 1:

[1436] The user starts the application on the device, sets the theme to "New Product Introduction," and starts a presentation. The user can input voice, facial expressions, gestures, and slides.

[1437] Input: Set presentation theme, start presentation

[1438] Output:Presentation settings information

[1439] Step 2:

[1440] The device's microphone collects the user's voice, and the camera records the user's facial expressions and gestures, as well as capturing the slides the user uses.

[1441] Input: User's voice, facial expressions, gestures, slides

[1442] Output: Collected audio data, video data, and slide data

[1443] Step 3:

[1444] The devices transmit the collected data to the server in real time, a process in which the data is buffered and temporarily stored before being transmitted over the network.

[1445] Input: Audio data, video data, slide data

[1446] Output: Data sent

[1447] Step 4:

[1448] The server analyzes the received voice data, converts the voice into text using speech recognition technology (e.g., Google Cloud Speech-to-Text API), and evaluates the speaking speed, volume, and pauses based on the text data.

[1449] Input: Audio data

[1450] Output: Text data and its evaluation results

[1451] Specific operation: The server uses voice recognition technology to convert audio such as "introducing a new product" into text, and analyzes the speaking speed, how many characters per second, whether the voice volume is appropriate, etc.

[1452] Step 5:

[1453] The server analyzes the received video data and uses computer vision techniques (e.g., OpenCV and TensorFlow) to analyze the user's facial expressions, gaze, and gestures to evaluate the presentation's poise and confidence.

[1454] Input: Video data

[1455] Output: Analysis results

[1456] Specific operation: The server uses video analysis technology to evaluate whether the user's eyes are fixed on the slide, whether they are making eye contact with the audience, whether their gestures are natural, etc.

[1457] Step 6:

[1458] The server analyzes the received slides, using optical character recognition (OCR) technology (e.g., Tesseract OCR) to read the content of the slides and evaluate whether they are consistent with the presentation theme and have appropriate visual effects.

[1459] Input: Slide data

[1460] Output: Analysis results

[1461] How it works: The server uses OCR technology to analyze the slides and identify typos and inconsistencies in content, such as when "new product" is misspelled as "new store."

[1462] Step 7:

[1463] The server generates feedback to provide to the user based on the results of the audio, video, and slide analysis, including suggestions for adjusting speaking speed and volume, improving facial expressions and gestures, and correcting slides.

[1464] Input: Audio analysis results, video analysis results, slide analysis results

[1465] Output: Generated feedback

[1466] Specific operation: The server combines the results of each analysis and generates feedback on specific areas for improvement, such as "speaking too fast," "not making enough eye contact with the audience," and "there are typos on the slides."

[1467] Step 8:

[1468] The server sends the generated feedback to the device, which then displays it to the user, who receives real-time feedback and can immediately see how to improve their presentation.

[1469] Input: Generated feedback

[1470] Output: Feedback displayed to the user

[1471] What it does: The device will display the feedback as a pop-up notification for the user to view immediately.

[1472] Step 9:

[1473] The server manages user authentication information, such as authentication when a user logs into the system, and stores presentation history and past feedback results for future reference.

[1474] Input: User authentication information, presentation history, feedback results

[1475] Output: Authenticated user information, saved history and feedback

[1476] Specific operation: The server authenticates the user ID and password and stores each user's presentation history and feedback results in a database.

[1477] Step 10:

[1478] The server generates questions based on the user's presentation and presents them to the user via the terminal. The user answers the questions, and the answers are also analyzed to provide more detailed feedback.

[1479] Input: Presentation content, user responses

[1480] Output: Generated questions, analysis results and additional feedback

[1481] Specific operation: The server generates questions such as "What are the advantages of the new product?" and presents them to the user through the terminal. If the user answers "The advantage is multifunctionality," the answer is analyzed and additional feedback is generated.

[1482] (Application example 1)

[1483] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1484] Previous systems for improving presentation skills only provided feedback specific to presentations, which meant they were not particularly practical for customer service. Customer service skills also depend on a variety of important factors, such as voice, facial expressions, gestures, eye contact, and slides, but there was no system that could comprehensively analyze these and provide specific feedback to customer service staff in real time. Furthermore, while improving the ability to respond to questions in customer service is also important, previous technologies were inadequate in this regard as well.

[1485] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1486] In this invention, the server includes a user interface means for allowing a user to input voice, facial expressions, gestures, slides, and eye contact; a data transmission means for transmitting data input from the user interface means to the server; an audio analysis means for converting the voice data received by the server into text and evaluating the user's speaking speed, volume, and pauses; a video analysis means for analyzing the user's facial expressions, eye contact, and gestures from the video data received by the server and evaluating their level of composure and confidence; a slide analysis means for analyzing the slides received by the server and evaluating the consistency of content, visual effect, and typos; a feedback generation means for generating feedback to the user based on the evaluation results of the audio analysis means, video analysis means, and slide analysis means; a feedback transmission means and a visual output means for transmitting the feedback generated by the feedback generation means to the user's terminal; and a question generation means for generating questions based on the user's customer service content and presenting them to the user. This allows customer service staff to receive specific feedback in real time on various aspects of customer service, such as voice, facial expressions, gestures, eye contact, and consistency of content. The question generation means also contributes to improving customer service skills.

[1487] "User interface means" refers to the functions of devices and software that allow the wait staff to input voice, facial expressions, gestures, slides, and eye movements.

[1488] The "data transmission means" is a mechanism for transmitting data input from the user interface means to the server.

[1489] The "voice analysis means" is a function that converts the voice data received by the server into text and evaluates the speaking speed, volume, and pauses.

[1490] The "video analysis means" is a function that analyzes the facial expressions, gaze, and gestures of the customer service staff from the video data received by the server, and evaluates their level of composure and confidence.

[1491] The "slide material analysis means" is a function that analyzes the slide materials received by the server and evaluates the consistency of the content, the visual effect, and typographical errors.

[1492] The "feedback generation means" is a mechanism for generating feedback to the customer service staff based on the evaluation results of the audio analysis means, video analysis means, and slide data analysis means.

[1493] The "feedback sending means" is a function that sends the generated feedback to the terminal of the customer service staff member.

[1494] "Visual output means" refers to a device or function that visually displays feedback on a terminal.

[1495] The "question generation means" is a function in which the server generates questions based on the content of customer service and presents them to the customer service staff.

[1496] The system for improving customer service skills according to the present invention can be implemented as follows. The system is composed of a customer service staff member, a terminal, and a server. The customer service staff member is a trainee and interfaces with the system through the terminal. The terminal serves to collect the customer service staff member's voice, facial expressions, gestures, eye movements, and slide materials and transmit them to the server. The server analyzes the received data and provides feedback to the customer service staff member.

[1497] Basic configuration

[1498] The system includes the following basic means:

[1499] 1. User interface means: Devices and software that allow wait staff to input voice, facial expressions, gestures, slides, and eye movements. Equipped with a voice input device (microphone), a video input device (camera), and an eye movement detector, the wait staff can input voice, facial expressions, gestures, eye movements, and slides into the system.

[1500] 2. Data transmission means: The terminal has the function to transmit the audio, video and slides input by the customer service staff to the server. This process is performed in real time, and the data is buffered and sent to the server.

[1501] 3. Speech analysis: The server analyzes the received voice data. Specifically, it converts the voice into text using Google Cloud Speech-to-Text, and evaluates the speaking speed, volume, and pauses based on the text data. For example, if a customer service staff member speaks too quickly, this will be reflected in the feedback.

[1502] 4. Video analysis: The server analyzes the received video data. Using computer vision technologies such as OpenCV and dlib, it analyzes the facial expressions, gaze, and gestures of the wait staff, and evaluates their level of composure and confidence based on the analysis. For example, if the wait staff is not making eye contact with the viewer, this will be included in the feedback.

[1503] 5. Slide Analysis: The server analyzes the received slides. It uses optical character recognition (OCR) technology to read the content of the slides and evaluates whether they are consistent, visually appropriate, and free of typos. For example, if a slide contains a typographical error, this is reflected in the feedback.

[1504] 6. Feedback generation means and feedback transmission means: The server generates feedback to be provided to the customer service staff based on the evaluation results obtained by the audio analysis means, video analysis means, and slide analysis means. The generated feedback is sent to the terminal and visually displayed to the customer service staff. This feedback may include adjustments to speaking speed and volume, improvements to facial expressions and gestures, and corrections to slides.

[1505] 7. Question generation means and additional feedback generation means: The server generates questions based on the customer service content and presents them to the customer service staff. When the customer service staff answers these questions, the answers are also analyzed and additional feedback is generated. Questions are generated using a generative AI model such as OpenAI's GPT-3, and the customer service staff's answers are analyzed. For example, questions are generated on the theme of "explaining new products," and the necessary feedback is provided to the customer service staff when they answer them.

[1506] Specific examples

[1507] For example, if a customer service staff member practices on the theme of "explaining a new product," the system functions as follows: The customer service staff member practices using a terminal, and the server uses its audio analysis means to evaluate that "their voice is too quiet." It also uses its video analysis means to determine that "their eyes are always looking toward the product," and its slide analysis means to analyze that "there are typos in some of the slides." The server generates feedback based on these evaluations and provides the customer service staff with specific advice such as "increase the volume of your voice when speaking," "make eye contact with the customer," and "correct typos in the slides."

[1508] Example prompt sentence:

[1509] "Our customer service staff are explaining a new product. What advice would be effective? Please also give us specific examples of areas where our staff can improve."

[1510] In this way, the system of the present invention functions as an effective tool for customer service staff to efficiently improve their customer service skills.

[1511] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1512] Step 1:

[1513] The customer service staff uses the microphone, camera, and gaze detector of the terminal to emit voice and input facial expressions, gestures, gaze, and slides. The input data includes voice data, video data, gaze data, and slides.

[1514] Step 2:

[1515] The terminal buffers the voice data, video data, eye movement data, and slide material input by the customer service staff in real time and transmits them to the server using the data transmission means, whereby the voice, video, eye movement, and slide material data are input to the server.

[1516] Step 3:

[1517] The server converts the received voice data into text data using Google Cloud Speech-to-Text. The input voice data is analyzed and evaluated for speaking speed, volume, and pauses. For example, if the voice data indicates that the speaker is speaking too fast, an evaluation result will be generated.

[1518] Step 4:

[1519] The server uses OpenCV and dlib to analyze facial expressions, eye movements, and gestures from the received video data. It receives the video data as evaluation input and evaluates the customer's level of composure and confidence. For example, it generates an evaluation result such as "the customer's eyes are always directed toward the products."

[1520] Step 5:

[1521] The server uses optical character recognition (OCR) technology to read the content of the received slides and evaluates the consistency of the content, visual effect, and typographical errors. It receives the slides as input data and generates the evaluation result: "Some slides contain typographical errors."

[1522] Step 6:

[1523] The server generates feedback for the customer service staff based on the evaluation results obtained from the audio analysis means, video analysis means, and slide analysis means. For example, the generated feedback includes specific advice such as "increase your speaking volume," "make eye contact with the customer," and "correct typos in the slides."

[1524] Step 7:

[1525] The feedback generated by the feedback generating means is sent to the terminal via the feedback transmitting means and the visual output means. As feedback, information such as "increase your speaking volume" and "make eye contact with the customer" is displayed to the wait staff.

[1526] Step 8:

[1527] The server uses a generative AI model such as OpenAI's GPT-3 to generate questions based on the customer service content. The server generates appropriate questions based on the scenario the customer service staff is practicing and sends them to the device. For example, a question might be generated such as, "I'm explaining a new product. What kind of advice would be effective?"

[1528] Step 9:

[1529] The customer service staff responds to questions posed by the system and sends the answers to the server via their terminal. The server analyzes the received answers and generates additional feedback based on the evaluation results.

[1530] Step 10:

[1531] The server transmits the additional feedback generated by the additional feedback generating means to the terminal using the feedback transmitting means and the visual output means, thereby providing further specific advice to the wait staff.

[1532] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1533] The system for improving presentation skills according to the present invention can provide more advanced feedback by incorporating an emotion engine that analyzes the user's emotions. Specific embodiments will be described below.

[1534] Basic configuration

[1535] The system is composed of a user, a terminal, and a server. The user is the person giving the presentation and interfaces with the system through the terminal. The terminal collects the user's voice, facial expressions, gestures, and slides and transmits them to the server. The server analyzes the received data and provides feedback to the user. In addition, this embodiment includes an emotion engine that recognizes the user's emotions.

[1536] User Interface Means

[1537] A user gives a presentation using a terminal. The terminal is equipped with an audio input device (microphone), a video input device (camera), and a screen for displaying slides. An emotion engine is also installed, which analyzes the user's emotions.

[1538] Data transmission method

[1539] The device has the ability to transmit audio, video, and slides input by the user to the server. This process is performed in real time, and the data is buffered and sent to the server. The user's emotional data is also transmitted in the same way.

[1540] Voice analysis methods

[1541] The server analyzes the received audio data, converts it into text using speech recognition technology, and evaluates the speed, volume, and pauses of speech based on the text data. For example, if the user speaks too quickly, this will be reflected in the feedback.

[1542] Video analysis methods

[1543] The server analyzes the received video data, using computer vision technology to analyze the user's facial expressions, eye contact, and gestures, and evaluates the presentation's poise and confidence. For example, if the user is not making eye contact with the audience, this will be included in the feedback.

[1544] Slide analysis tools

[1545] The server analyzes the slides it receives, using optical character recognition (OCR) technology to read their content and evaluate whether they are consistent with the presentation theme and whether their visual impact is appropriate. For example, if a slide contains a typographical error, that error is reflected in the feedback.

[1546] Emotion Engine

[1547] The server uses an emotion engine to analyze the user's emotional data. Specifically, it analyzes the user's facial expressions and voice to evaluate the fluctuations and stability of their emotions during the presentation. For example, if the user is nervous during the presentation, that emotional data will be included in the feedback.

[1548] Feedback generation means and feedback transmission means

[1549] The server integrates the evaluation results of the audio analysis means, video analysis means, slide analysis means, and emotion engine to generate feedback for the user. The generated feedback is sent to the terminal and displayed to the user. The feedback includes advice on adjusting speaking speed and volume, improving facial expressions and gestures, correcting slides, and emotional advice.

[1550] Authentication Management Methods

[1551] The server manages user authentication information and stores the user's presentation history and feedback results, allowing users to check their progress and refer to past feedback.

[1552] Question generation and additional feedback generation

[1553] The server generates questions based on the user's presentation and presents them to the user. When the user answers these questions, the answers are also analyzed and additional feedback is generated. For example, if the user does not answer a question accurately, that point is reflected in the additional feedback.

[1554] Specific examples

[1555] For example, if a user gives a presentation on the theme of "introducing a new product," the system functions as follows: The user gives the presentation using a device, and the server uses its audio analysis means to evaluate the user's voice as "quiet." The video analysis means determines that the user's eyes are always looking at the slides, and the slide analysis means analyzes that there are typos on some of the slides. The emotion engine then determines that the user is nervous. The server generates feedback based on these evaluations and provides the user with specific advice such as "increase the volume of your voice," "make eye contact with the audience," "correct typos in the slides," and "regulate your breathing to relieve tension."

[1556] In this way, the system of the present invention functions as an effective tool for users to efficiently improve their presentation skills. In particular, the introduction of an emotion engine makes it possible to provide comprehensive feedback that takes into account the user's emotional state.

[1557] The processing flow will be explained below.

[1558] Step 1:

[1559] A user logs into the system using a terminal, enters a username and password, and submits the authentication information.

[1560] Step 2:

[1561] The device sends the entered authentication information to the server. The server compares it with the database and returns the authentication result to the device. If authentication is successful, the user can start the presentation.

[1562] Step 3:

[1563] The user begins preparing for a presentation. They upload presentation materials (slides) using their device. The device then sends the uploaded materials to the server.

[1564] Step 4:

[1565] The server stores the received presentation materials in an appropriate format.

[1566] Step 5:

[1567] The user presses the "Start Presentation" button. The device begins recording the user's voice and video (facial expressions and gestures) using the built-in camera and microphone. The recorded data is buffered.

[1568] Step 6:

[1569] The terminal transmits the buffered data to the server in real time, including audio, video, slides, and emotion data.

[1570] Step 7:

[1571] The server analyzes the received voice data, converts the voice data into text using a speech recognition engine, evaluates speaking speed, volume, and pauses, and generates text data.

[1572] Step 8:

[1573] The server analyzes the received video data and uses computer vision technology to analyze the user's facial expressions, eye movements, and gestures to evaluate the presentation's poise and confidence.

[1574] Step 9:

[1575] The server analyzes the slides received, using optical character recognition (OCR) technology to read the content of the slides and evaluate their consistency and visual impact.

[1576] Step 10:

[1577] The server uses an emotion engine to analyze the user's emotional data, identifying emotions from facial expressions and vocal tones, and assessing emotional stability and fluctuations.

[1578] Step 11:

[1579] The server integrates the evaluation results obtained from the audio analysis means, video analysis means, slide analysis means, and emotion engine, and generates feedback for the user, including advice on speaking speed, volume, naturalness of gestures, corrections to slides, and emotional advice.

[1580] Step 12:

[1581] The server sends the generated feedback to the user's terminal, which receives the feedback and displays it to the user.

[1582] Step 13:

[1583] Users can review the feedback and incorporate improvements for their next presentation. Users can also request a Q&A simulation if necessary.

[1584] Step 14:

[1585] If a question and answer simulation is desired, the server generates questions based on the content of the presentation and presents them to the user.

[1586] Step 15:

[1587] The user answers the questions, and the device sends the answer data to the server.

[1588] Step 16:

[1589] The server analyzes the user's answers and generates additional feedback, providing further improvements based on the evaluation results.

[1590] Step 17:

[1591] The server sends additional feedback to the user's device, which the user receives and uses as a reference for their next presentation.

[1592] This series of processing steps allows users to improve their presentation skills efficiently and effectively. In addition, the introduction of an emotion engine makes it possible to provide comprehensive feedback that takes emotional aspects into consideration.

[1593] Example 2

[1594] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1595] Current presentation skill improvement systems have difficulty accurately analyzing users' emotions and reflecting the analysis results in presentation feedback. Furthermore, they lack feedback on not only the technical aspects of the presentation but also the emotional and psychological aspects, making it difficult for users to comprehensively improve their skills. Furthermore, existing systems lack sufficient functionality for generating questions and providing additional feedback, making it difficult for users to self-evaluate.

[1596] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1597] In this invention, the server includes an audio analysis means, a video analysis means, a slide analysis means, an emotion engine, a feedback generation means, and a feedback transmission means. This allows for comprehensive analysis of a user's voice, facial expressions, gestures, slides, and even emotions, making it possible to provide detailed and accurate feedback. Furthermore, by including a question generation means and an additional feedback generation means, it is possible to automatically generate questions based on the content of the user's presentation and evaluate the answers to those questions, facilitating self-evaluation.

[1598] "User interface means" refers to devices and software that allow a user to input voice, facial expressions, gestures, and slide materials.

[1599] "Data transmission means" refers to a device or software for transmitting data input from the user interface means to the server.

[1600] "Voice analysis means" refers to devices or software that convert the voice data received by the server into text and evaluate speaking speed, volume, and pauses.

[1601] "Video analysis means" refers to devices or software that analyze the user's facial expressions, gaze, and gestures from the video data received by the server, and evaluate the presentation's level of composure and confidence.

[1602] "Slide material analysis means" refers to a device or software that analyzes the slide materials received by the server and evaluates the consistency of the content and the visual effect.

[1603] "Emotion engine" refers to a device or software that allows the server to analyze emotion data and evaluate the fluctuations and stability of a user's emotions.

[1604] "Feedback generation means" refers to a device or software for generating feedback to a user based on the evaluation results of the audio analysis means, video analysis means, slide data analysis means, and emotion engine.

[1605] The "feedback sending means" refers to a device or software for sending the feedback generated by the feedback generating means to the user's terminal.

[1606] "Authentication management means" refers to a device or software for managing user authentication information and storing each user's presentation history and feedback results.

[1607] The "question generation means" refers to a device or software that generates questions based on the content of a user's presentation and presents them to the user.

[1608] The "additional feedback generating means" refers to a device or software that analyzes the answers given by the user to the questions presented and generates additional feedback based on the evaluation results.

[1609] The system for improving presentation skills according to the present invention provides more advanced feedback by incorporating an emotion engine that analyzes the user's emotions. Specific embodiments will be described below.

[1610] Basic configuration

[1611] The system is composed of a user, a terminal, and a server. The user is the person giving the presentation and interfaces with the system through the terminal. The terminal collects the user's voice, facial expressions, gestures, and slides and transmits them to the server. The server analyzes the received data and provides feedback to the user. In addition, this embodiment includes an emotion engine that recognizes the user's emotions.

[1612] User Interface Means

[1613] A user gives a presentation using a terminal. The terminal is equipped with an audio input device (microphone), a video input device (camera), and a screen for displaying slides. An emotion engine is also installed, which analyzes the user's emotions.

[1614] Data transmission method

[1615] The device has the ability to transmit audio, video, and slides input by the user to the server. This process is performed in real time, and the data is buffered and sent to the server. The user's emotional data is also transmitted in the same way.

[1616] Voice analysis methods

[1617] The server analyzes the received audio data, converts it into text using speech recognition technology, and evaluates speaking speed, volume, and pauses based on the text data. The specific software used is the Google Cloud Speech-to-Text API. For example, if the user speaks too quickly, this is reflected in the feedback.

[1618] Video analysis methods

[1619] The server analyzes the received video data. It uses computer vision technology to analyze the user's facial expressions, eye movements, and gestures, and evaluates the presentation's poise and confidence based on the results. The OpenCV library is used for this software. For example, if the user is not making eye contact with the audience, this will be included in the feedback.

[1620] Slide analysis tools

[1621] The server analyzes the received slides. It uses optical character recognition (OCR) technology to read the content of the slides and evaluate whether they are consistent with the presentation theme and have an appropriate visual effect. Specific software includes Tesseract OCR. For example, if a slide contains a typographical error, that error is reflected in the feedback.

[1622] Emotion Engine

[1623] The server uses an emotion engine to analyze the user's emotional data. Specifically, it analyzes the user's facial expressions and voice to evaluate the fluctuations and stability of their emotions during the presentation. For example, if the user is nervous during the presentation, that emotional data will be included in the feedback.

[1624] Feedback generation means and feedback transmission means

[1625] The server integrates the evaluation results of the audio analysis means, video analysis means, slide analysis means, and emotion engine to generate feedback for the user. The generated feedback is sent to the terminal and displayed to the user. The feedback includes advice on adjusting speaking speed and volume, improving facial expressions and gestures, correcting slides, and emotional advice.

[1626] Authentication Management Methods

[1627] The server manages user authentication information and stores the user's presentation history and feedback results, allowing users to check their progress and refer to past feedback.

[1628] Question generation and additional feedback generation

[1629] The server generates questions based on the user's presentation and presents them to the user. When the user answers these questions, the answers are also analyzed and additional feedback is generated. For example, if the user does not answer a question accurately, that point is reflected in the additional feedback.

[1630] Specific examples

[1631] For example, if a user is giving a presentation on the theme of "introducing a new product," the system functions as follows: The user gives the presentation using a terminal, and the server uses the audio analysis means to evaluate the user's voice as "quiet." The video analysis means determines that the user's eyes are always directed toward the slides, and the slide analysis means analyzes that there are typos on some slides. The emotion engine then determines that the user is nervous. Based on these evaluations, the server generates feedback and provides the user with specific advice, such as "increase your speaking volume," "make eye contact with the audience," "correct typos in the slides," or "regulate your breathing to relieve tension." In this way, the system of the present invention functions as an effective tool for users to efficiently improve their presentation skills. In particular, the introduction of the emotion engine makes it possible to provide comprehensive feedback that takes the user's emotional state into consideration.

[1632] Prompt Sentence Examples

[1633] Here are some example prompts to input to a generative AI model:

[1634] A user will give a presentation introducing a new product. Staff will generate feedback based on the analysis of audio, video, and slides. Also, analyze the user's emotional data and include it in the feedback. Examples of feedback include "your voice is too quiet," "you should make eye contact with the audience," "there is a typo on the slide," and "you seem nervous."

[1635] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1636] Specific processing flow of the program

[1637] Step 1: Start your presentation

[1638] The user clicks a button on the device to start the presentation. The device verifies the user's authentication information and sends a request to the server to start the session.

[1639] Input: User credentials, session initiation request

[1640] Processing: The server verifies the authentication information and generates a token to start the session.

[1641] Output: Session initiation token

[1642] Step 2: Data collection

[1643] The user gives a presentation, and the device collects audio, video, and slides via a microphone, camera, and screen display. 【16...

Claims

1. a user interface means for a user to input voice, facial expressions, gestures, and slide materials; data transmission means for transmitting data input from the user interface means to a server; a voice analysis means for converting the voice data received by the server into text and evaluating the speaking speed, volume, and pauses; a video analysis means for analyzing the user's facial expressions, gaze, and gestures from the video data received by the server, and evaluating the user's level of composure and confidence in the presentation; a slide data analysis means for analyzing the slide data received by the server and evaluating the consistency of the content and the visual effect; a feedback generating means for generating feedback to a user based on the evaluation results of the audio analysis means, the video analysis means, and the slide data analysis means; a feedback transmission means for transmitting the feedback generated by the feedback generation means to a user terminal; A system including:

2. 2. The system according to claim 1, wherein the server further comprises an authentication management unit that manages user authentication information and stores the presentation history and feedback results of each user.

3. The system of claim 1, wherein the feedback generation means further includes a question generation means for generating questions based on the content of the user's presentation and presenting them to the user, and an additional feedback generation means for analyzing the answers given by the user to the presented questions and generating additional feedback based on the evaluation results.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A