system
The system addresses the challenge of quickly generating high-quality presentation content and improving speech skills by automatically creating materials and providing analytical feedback, allowing users to enhance their performance effectively and economically.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-16
- Publication Date
- 2026-04-28
AI Technical Summary
There is a lack of efficient means for individuals to generate high-quality presentation content quickly and improve their speech skills without expert guidance, leading to time and cost constraints in preparing for speeches and presentations.
A system that automatically generates presentation materials, speech scripts, and demo videos based on user input, analyzes practice videos for features like eye contact and voice tone, and provides instructional videos for self-improvement.
Enables users to prepare high-quality presentations efficiently and improve their speech skills by providing specific feedback, reducing the need for expert guidance and lowering costs.
Smart Images

Figure 2026070881000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, the method including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In preparing for a speech or presentation, there is a lack of means to generate high-quality content in a short time and improve one's own performance without receiving guidance from an expert. This problem does not meet the need to efficiently improve the skills that individual users should possess as the demand for speeches and presentations increases. Also, it is required to save high consulting fees and time constraints.
Means for Solving the Problems
[0005] This invention provides a system that automatically generates presentation materials, speech scripts, and demo videos quickly based on prompt input from the user. This allows users to prepare the framework for their presentations in a short amount of time. Furthermore, when a user uploads a practice video, the system analyzes the video using image analysis technology and evaluates eye contact, voice tone, posture, etc. In addition, by generating and providing instructional videos based on the analysis results, users can receive specific advice for self-improvement. In this way, this invention provides a system that supports the improvement of users' speech skills at a low cost.
[0006] A "prompt" is information that a user enters to provide instructions or topics to the system.
[0007] "Generation method" refers to a function or process that automatically generates content based on prompts provided by the user.
[0008] "Transmission means" refers to the part that has the functionality to send the generated content to the user's device.
[0009] "Analysis methods" refer to the technologies and processes used to analyze uploaded videos and evaluate the characteristics of the user's speech and posture.
[0010] "Instructional generation methods" refer to the technologies and processes used to create instructional videos that highlight areas for improvement based on analysis results.
[0011] "Presentation means" refers to the technology and functions that allow users to view the generated instructional videos. [Brief explanation of the drawing]
[0012] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] Shows an emotion map to which multiple emotions are mapped. [Figure 10] Shows an emotion map to which multiple emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when an emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when an emotion engine is combined.
Embodiments for Carrying Out the Invention
[0013] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0014] First, the language used in the following description will be explained.
[0015] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0016] In the following embodiments, the numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0017] In the following embodiments, the numbered storage is one or more non-volatile storage devices that store various programs, various parameters, and the like. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, and the like.
[0018] In the following embodiments, the numbered communication I / F (Interface) is an interface including a communication processor, an antenna, and the like. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0019] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0020] [First Embodiment]
[0021] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0022] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0023] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0024] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0025] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0026] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0027] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0028] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0029] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0030] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0031] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0032] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0033] This invention is a system that efficiently supports speeches and presentations through prompt input. The system mainly consists of a server and a terminal, and the user operates it via an interface. When the user inputs prompts regarding the theme and content of the presentation into the terminal, that information is sent to the server.
[0034] The server utilizes a generative AI model to automatically generate prompt-based presentation materials, speech scripts, and demo videos. In this process, the server collects relevant information from databases and external sources to construct specific and visually appealing materials tailored to the user's requirements. The generated content is then transmitted from the server to the terminal and provided to the user in a usable format.
[0035] Next, the user practices their speech using the generated content and records themselves. The practice video is uploaded to the server via the device, where it is analyzed. Image analysis technology is used to evaluate features such as the user's gaze, posture, and voice tone. The evaluation results are used to improve the user's speech performance.
[0036] Based on the analysis results, the server creates instructional videos that include specific improvement suggestions. These videos utilize synthesized speech and animation to provide users with visually and aurally easy-to-understand feedback. The instructional videos are sent to the user's device, and users can use them to improve their speech skills.
[0037] As a concrete example, consider preparing a presentation on the marketing strategy for a new product. When a user inputs the topic, the server generates presentation materials including product features and market information. Next, the user practices their speech, and feedback from the uploaded video indicates areas for improvement, such as how to emphasize certain points and how to direct eye contact. This allows users to deliver high-quality presentations in a short amount of time.
[0038] The following describes the processing flow.
[0039] Step 1:
[0040] The user enters prompts on the terminal regarding the presentation's theme and objectives. The entered prompts are formatted as digital data and sent to the server.
[0041] Step 2:
[0042] The server analyzes the received prompts and identifies the information needed to generate the corresponding presentation materials, speech scripts, and demo videos. The server searches for relevant content from databases and online resources and automatically generates content based on the collected data. A generation AI model is used to create materials in a format and expression appropriate to the prompt.
[0043] Step 3:
[0044] The server sends the generated presentation materials, speech scripts, and demo videos to the terminal. This content is then displayed on the terminal, allowing the user to review and utilize it.
[0045] Step 4:
[0046] The user practices their speech based on the materials and script they receive. The practice process is recorded with a camera and saved to the device. The recorded practice video is then uploaded to the server via the device.
[0047] Step 5:
[0048] The server analyzes the uploaded practice videos. Using image analysis technology, it evaluates the user's posture, gaze, voice tone, speed, etc. Based on the analysis results, it quantifies the user's performance and identifies areas that need improvement.
[0049] Step 6:
[0050] The server generates instructional videos based on the analysis results. These videos include specific improvement suggestions and advice, presented to the user visually and audibly. Synthesized speech and animation are used to provide feedback in an easily understandable format.
[0051] Step 7:
[0052] The server sends the generated instructional video to the user's device. The user watches the instructional video on their device and uses it to improve their speech skills. By practicing again based on the areas for improvement, it is possible to improve the quality of the presentation.
[0053] (Example 1)
[0054] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0055] Efficiently preparing for and improving speeches and presentations has traditionally required a great deal of time and effort. Furthermore, the lack of means to obtain concrete feedback through visual and auditory means hindered user growth. This invention proposes a system that provides concrete support for users to quickly prepare high-quality presentations and improve their skills.
[0056] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0057] In this invention, the server includes a device that receives information input from a user, including themes and content; a device that automatically creates materials, documents, and videos based on said information, and includes a device that collects relevant information from information sources; and a device that transmits the created data to the user's terminal. This allows users to efficiently create high-quality presentation materials and subsequently improve their skills based on performance analysis and specific feedback.
[0058] A "user" is an individual or group that uses this system to prepare speeches and presentations and to improve their own performance.
[0059] "Information including themes and content" refers to a collection of topics and specific information that users input to form the basis of their speeches or presentations.
[0060] A "device" is a combination of hardware and software that constitutes a system and performs a specific function.
[0061] "Automatically creating documents, texts, and videos" refers to the process of using computer programs based on input information to generate relevant content without user intervention.
[0062] "Collecting relevant information from sources" refers to the act of obtaining necessary data from internal databases or external information providers and forming the data set required for the generation process.
[0063] "Sending to the user's terminal" refers to the process of transmitting generated data to the user's terminal via the network, making the data available for the user to use.
[0064] "Exercise records" refer to data that is a recording or document of the content of a speech or presentation given by a user.
[0065] "Processing" refers to a set of procedures that involve analyzing data to achieve a specific purpose and generating results through calculations or operations.
[0066] "Eye movements, vocal characteristics, and posture" refer to the physical and vocal characteristics observed during a user's speech or presentation, and these are elements used to determine the quality of performance through evaluation.
[0067] "Educational videos" are video content created to provide users with feedback and instruction aimed at improving their skills.
[0068] Modes for carrying out the invention
[0069] This invention relates to a system for users to efficiently prepare and improve speeches and presentations. The system has multiple components, including a server, terminals, and interfaces.
[0070] User roles
[0071] The user first accesses the interface through a terminal and inputs prompts with themes and content related to their speech or presentation. These prompts are specific, such as "I would like to prepare materials to explain the importance of sustainable energy" or "I would like a presentation created that details the market impact of a new product."
[0072] Server Processing
[0073] The server receives prompts from users and automatically generates appropriate content using a generative AI model. This generative AI model incorporates natural language processing and machine learning technologies to generate materials, speeches, and video content that align with the user's intent. To gather relevant information, the server interacts with databases and publicly available online resources. The hardware used is a server equipped with a powerful processor.
[0074] Device usage and feedback
[0075] The generated content is sent from the server to the terminal and provided to the user in a format that the user can view. The user practices their speech based on this content and records their performance using the terminal's recording function. Afterwards, the user uploads the recorded practice session back to the server, where it uses image analysis technology to extract features and perform analysis. This analysis includes eye tracking, voice tone analysis, and posture monitoring.
[0076] Based on the analysis results, the server generates educational videos, visualizes feedback using synthesized speech and animation, and presents it to the user in an easy-to-understand format. This process gives users the opportunity to improve their speech skills in a concrete and efficient manner.
[0077] In this way, the invented system provides support for users to deliver high-quality speeches and presentations in a short amount of time.
[0078] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0079] Step 1:
[0080] The user enters a prompt message through the terminal. This prompt message includes the theme and content of the speech or presentation. The terminal sends this information to the server as input data.
[0081] Step 2:
[0082] The server analyzes the received prompt text and performs data processing using a generative AI model to automatically generate documents, text, and demo videos. It extracts relevant information from the input prompt text and retrieves information from databases and external sources to generate detailed and visual content. The generated content is obtained as the output of this process.
[0083] Step 3:
[0084] The server sends the generated content to the terminal, making it accessible to the user. The terminal then displays this received data in a format that the user can view and use.
[0085] Step 4:
[0086] Users practice their speeches using the generated materials and speech scripts on their devices. The practice sessions are recorded using the device's camera and microphone, and a file is created as a record of the practice.
[0087] Step 5:
[0088] Users upload their practice sessions to the server via their terminal. The server receives these recordings as input data and performs data processing using image analysis techniques to evaluate features such as gaze, posture, and voice tone. The output of the analysis is evaluation data regarding the user's performance.
[0089] Step 6:
[0090] Based on the analysis results, the server generates educational videos that include specific feedback. Utilizing a generation AI model and multimedia editing techniques, it uses synthesized speech and animation to create content that is easy for users to understand. The output of this process is the educational video.
[0091] Step 7:
[0092] The server sends the generated educational video to the terminal, allowing the user to watch it and provide feedback. The terminal plays the video and performs actions to enhance the user's viewing experience.
[0093] (Application Example 1)
[0094] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0095] The customer service skills of sales staff in physical stores significantly impact store sales and customer satisfaction. However, traditional training methods often lack sufficient specific feedback for individual sales staff, making efficient skill improvement difficult. In particular, the ability to effectively communicate information about new products and campaigns is crucial, and immediate and specific guidance is needed to enable sales staff to respond flexibly on the spot.
[0096] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0097] In this invention, the server includes an information input means for receiving instructions on a topic from the user, an automatic generation means, and an information transmission means. This allows sales staff to efficiently receive individualized customer service training. Furthermore, they can receive specific guidance to effectively explain new product information and improve customer interaction during actual sales activities in stores.
[0098] "Users" refer to salespeople and staff who use this system to improve their customer service skills.
[0099] "Theme-related instructions" refer to information and requests entered into the system regarding the products and services handled by the sales staff.
[0100] "Information input means" refers to devices or interfaces used by users to input instructions into a system.
[0101] "Automatic generation means" refers to devices or software that have the function of generating explanatory materials, speech scripts, and action videos based on a given theme.
[0102] "Information transmission means" refers to a system for transmitting generated content to a user's device.
[0103] "Training videos" refer to recorded videos of customer service skills practice that users upload to the system.
[0104] "Information analysis means" refers to devices or programs that have the function of analyzing exercise videos and evaluating important points such as the direction of gaze, pitch of voice, and posture.
[0105] "Educational generation means" refers to a system that creates instructional videos to improve users' skills based on analysis results.
[0106] "Display means" refers to devices or interfaces that present the generated instructional video to the user and enable them to view it.
[0107] "Training methods" refer to specific methods and programs for providing instruction aimed at improving customer service skills in physical stores.
[0108] To implement this invention, a device is first required for the user to input information, such as a smartphone or smart glasses. The user inputs instructions regarding products or services through this device. These instructions may include prompts such as, "Please create a presentation explaining the 100x zoom function of the latest smartphone."
[0109] The server receives input prompts and automatically generates the necessary content using a generative AI model. The generated content includes explanatory materials, presentation scripts, and even demonstration videos. This generation utilizes a Python backend program and OpenAI®'s GPT as the AI model. The generated content is then delivered to the user's device via a data transmission method.
[0110] Next, users practice customer service skills based on this content. This practice is recorded and uploaded to a server as practice video. The server analyzes this video using image analysis technologies such as OpenCV and TENSORFLOW®. Specifically, information such as gaze direction, voice pitch, and posture is extracted.
[0111] Based on the analysis results, the educational generation system generates instructional videos to support the user's skill improvement. Using synthesized speech and animation, specific areas for improvement are presented visually and audibly in an easy-to-understand manner. The instructional videos are displayed on the user's device, allowing sales staff to watch them and improve their customer service skills.
[0112] Furthermore, the system includes training methods that provide instruction on specific skills that users can immediately use in their in-store sales activities. This promotes effective customer service in stores.
[0113] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0114] Step 1:
[0115] The user enters a prompt using their device. The prompt is a topic-related instruction, such as, "Create a presentation explaining the 100x zoom feature of the latest smartphones." This input is structured as data sent from the device to the server.
[0116] Step 2:
[0117] The server uses a generative AI model to generate presentation materials, speech scripts, and action videos based on the received prompts. For data processing, the input prompts are analyzed using natural language processing, and relevant information is retrieved from databases and external sources. The output generates specific and visually appealing content tailored to the user's requests.
[0118] Step 3:
[0119] The generated content is transmitted from the server to the terminal via an information transmission mechanism. The user's terminal receives the data and begins practicing customer service skills based on it. The terminal's actions involve the appropriate display and demonstration of the generated content received from the server.
[0120] Step 4:
[0121] Users practice and record their sessions. The recorded practice videos are uploaded from the terminal to the server as data. This video is then input into the system and used as material for the next analysis.
[0122] Step 5:
[0123] The server receives the practice video and performs analysis using image analysis techniques. Specifically, it uses OpenCV and TensorFlow to extract features such as gaze, voice pitch, and posture. The input is the user's practice video, and the output is the analyzed feature data.
[0124] Step 6:
[0125] Based on the analysis results, the server generates instructional videos using educational generation tools. Synthesized speech and animation are used in the generation process to clearly and concretely demonstrate areas for improvement for the user. The output is an instructional video designed to promote skill development in the user.
[0126] Step 7:
[0127] The generated instructional video is sent from the server to the terminal. The user's terminal receives and displays the instructional video. The user can watch it and work on improving their skills based on the feedback. The terminal's operation includes playing the video and presenting it to the user.
[0128] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0129] This invention is a system that supports the improvement of users' speech skills, and in particular, recognizes the user's emotions using an emotion engine and provides feedback based on that information. The system mainly consists of a server, a terminal, and the user.
[0130] The user enters prompts related to the topic of their speech or presentation through their device. These prompts are sent from the device to a server. The server analyzes the received prompts, automatically collects relevant information, and generates presentation materials, speech scripts, and demo videos. This process is driven by a generative AI model, providing concrete and visually appealing content. The generated content is then sent to the user via their device and made available for use.
[0131] Users practice their speeches based on the generated content and record the process on video. The recorded practice videos are uploaded to the server via the device. The server analyzes the received practice videos using image and audio analysis technologies, evaluating not only the user's gaze, posture, and voice tone, but also their emotional state using an emotion engine. The emotion engine recognizes emotions from changes in facial expressions, voice tone, pitch, etc., and analyzes the type and intensity of those emotions.
[0132] The analysis results are used to generate instructional videos. The server creates instructional videos that include specific areas for improvement based on the user's expression and emotional state. These instructional videos provide feedback that reflects emotions in an easy-to-understand way for the user, using synthesized speech and animation. For example, if the video detects nervousness during a presentation, it will include suggestions on how to relax. These instructional videos are sent to the user's device, and the user watches them to guide further practice.
[0133] Through the operation of these systems, users can not only improve the quality of their presentations but also learn how to control their emotions and improve their overall public speaking skills.
[0134] The following describes the processing flow.
[0135] Step 1:
[0136] The user enters the presentation topic into the terminal. The entered prompt is digitized by the terminal and sent to the server.
[0137] Step 2:
[0138] The server receives a prompt and uses generation AI to search for relevant data. Based on the collected data, it automatically generates presentation materials, speech scripts, and demo videos. The created content is then sent from the server to the terminal.
[0139] Step 3:
[0140] The terminal provides the user with content received from the server. The user uses the provided materials and script to practice their speech.
[0141] Step 4:
[0142] Users record their practice sessions and save the videos to their devices. The completed practice videos are then uploaded from the device to the server.
[0143] Step 5:
[0144] The server analyzes the uploaded practice videos. Using image and audio analysis technologies, it evaluates the user's gaze, posture, voice tone, and emotions using an emotion engine. The emotion engine identifies the user's emotional state by analyzing facial expressions and changes in voice.
[0145] Step 6:
[0146] The server generates instructional videos based on the analysis results. These instructional videos include specific advice on areas for improvement and methods for controlling emotions. The videos utilize synthesized speech and animation to present information in a visual and auditory way that is easy for users to understand.
[0147] Step 7:
[0148] Instructional videos are sent to the user's device for viewing. Based on the feedback received, the user can use it to improve their performance in future speech practice sessions and actual presentations. By repeating this process, the user can refine their skills.
[0149] (Example 2)
[0150] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0151] A challenge exists in that users cannot effectively improve their emotional expression, visual information, and auditory information during speeches and presentations. Traditional methods limit opportunities for users to receive feedback and make self-assessment difficult, which can hinder skill improvement.
[0152] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0153] In this invention, the server includes means for receiving theme-based input information from the user, means for automatically generating material data and visual information using information processing technology, means for transferring the generated data to a terminal via communication, means for analyzing activity videos and evaluating visual and audio information, means for generating instructional information based on the analysis results using speech synthesis technology and visual technology, and means for providing the instructional information to the user visually and audibly. This enables the user to objectively improve their speech skills and emotional expression and enhance their abilities.
[0154] "Input information" refers to information provided by the user based on a given theme, and serves as the foundational data for the system to begin its analysis and generation processes.
[0155] "Information processing technology" refers to all technologies used to analyze received data and automatically generate document data and visual information.
[0156] "Document data" refers to automatically generated content for presentations and speeches, containing information that users can refer to.
[0157] "Visual information" refers to information that appeals to users visually, including demo videos and charts.
[0158] "Means of transfer via communication" refers to the infrastructure and protocols used to transmit generated data and information to a terminal.
[0159] "Activity videos" are videos recorded by users for speech practice, and their content is analyzed and used as material for evaluation.
[0160] "Means for evaluating visual and auditory information" refers to technologies for analyzing the content of activity videos and quantitatively and qualitatively evaluating the user's visual and auditory elements.
[0161] "Speech synthesis technology" is a technology for generating synthesized sounds that mimic human voices, and is used for voice feedback in instructional materials.
[0162] "Visual technology" refers to technologies that generate visual feedback and provide it to users in an easily understandable format.
[0163] "Instructional information" refers to feedback content generated based on analysis results, which includes suggestions and areas for improvement to enhance the user's skills.
[0164] This invention is an information processing system for supporting the improvement of speech skills. This system consists of three elements: a user, a terminal, and a server.
[0165] The user enters prompt text into the terminal based on the theme. For example, a prompt might be, "Please describe the elements of effective leadership." The terminal receives the user's input and sends it to the server via the internet.
[0166] The server utilizes a generative AI model to automatically generate document data and visual information based on this prompt. This process employs natural language processing techniques and machine learning algorithms. The materials include demo videos and presentation slides to visually represent specific information. The generated materials and data are returned to the terminal via a secure communication protocol.
[0167] The terminal displays data received from the server to the user. The user uses this data to practice their speech and records the process as an activity video. Recording is typically done using the terminal's built-in camera function.
[0168] Afterward, the user uploads the recorded activity video from their device to the server. The server uses image and audio analysis technologies to analyze this video and evaluate the user's visual and audio information. In addition to elements such as gaze, posture, and voice tone, an emotion analysis engine is also introduced to evaluate emotional expression from facial expressions.
[0169] Based on the analysis results, the server generates instructional information using speech synthesis and visual technologies. This instructional information is provided as feedback content, including specific areas for improvement and advice, with the aim of improving the user's skills. For example, if the server determines that the user is stressed, it will provide suggestions for relaxation techniques.
[0170] All data is transmitted to the device, and users can effectively improve their speech skills by viewing or using it. This system allows users to improve their skills and the quality of their presentations.
[0171] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0172] Step 1:
[0173] The user enters prompt sentences related to the topic of their speech or presentation into the terminal. These prompt sentences become data that instructs the system to analyze. After input, the terminal converts this data into a digital format and sends it to the server.
[0174] Step 2:
[0175] The server parses the received prompt message. This parsing uses a generative AI model, leveraging natural language processing techniques to understand the intent of the prompt. The AI model collects necessary information from a database to generate relevant data and visual information, creating presentation slides and demo videos. The output of this process is content in a user-friendly format.
[0176] Step 3:
[0177] The server sends the generated material data and visual information to the terminal. The terminal receives this and presents it to the user through an appropriate interface. The user practices their speech using the provided materials. Reviewing these materials helps improve understanding of the presentation content and self-expression.
[0178] Step 4:
[0179] The user records their speech practice using the device's camera function. The recorded content includes data on the user's visual and auditory expression. After recording, the video is converted to an appropriate format on the device and uploaded to the server.
[0180] Step 5:
[0181] The server receives uploaded practice videos and analyzes them using image and audio analysis technologies. This process evaluates data related to the user's gaze, posture, and voice tone. Additionally, an emotion analysis engine is used to extract emotional information from the user's facial expressions and voice. The output after analysis is evaluation information that can be used for instructional feedback.
[0182] Step 6:
[0183] The server generates instructional feedback based on the analysis results. Utilizing a generation AI model, it creates videos containing specific advice for the user using speech synthesis and visual technologies. For example, this feedback might include suggestions for relaxation techniques or tips for improving expressiveness.
[0184] Step 7:
[0185] The server sends instructional feedback to the device. The device receives this feedback and provides it to the user in a viewable format. Through this feedback video, the user can objectively evaluate their own performance and use it to further improve their skills.
[0186] (Application Example 2)
[0187] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0188] To improve the efficiency and safety of workers on site, a system is needed that can appropriately understand workers' emotional states and stress levels and provide feedback at the appropriate time. However, conventional technology has made it difficult to recognize workers' emotional states in real time and respond to changes in them. There is a need to provide a system that solves these problems and realizes a better working environment.
[0189] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0190] In this invention, the server includes an input device that receives prompts from the user, a generation device that automatically generates visual materials, audio scripts, and demo sound images based on the prompts, and a communication device that transmits the generated information to a terminal. This allows the worker's emotional state to be recognized and to receive feedback as needed.
[0191] A "prompt" is an instruction or request provided by the user through input.
[0192] An "input device" refers to a device used to receive data and instructions from a user.
[0193] A "generation device" is a device that has the function of processing information based on received instructions and creating the necessary content.
[0194] A "communication device" is a device used to transmit generated information to other devices or networks.
[0195] "Reference videos" refer to video content used to analyze information uploaded by users.
[0196] An "analysis device" is a device used to analyze the characteristics of an object based on collected data.
[0197] A "support structure" refers to a system that provides necessary responses and feedback based on the analysis results.
[0198] A "presentation device" is a device that provides generated feedback and instructional content to the user visually or audibly.
[0199] An "improvement suggestion" is a measure to improve the work environment or methods suggested by the system, based on the analysis results.
[0200] In a system for realizing this invention, the server receives prompts from the user via an input device and uses a generation AI model to create visual materials, audio scripts, and demo sound images using a generation device. This generated content is transmitted to a terminal via a communication device. The user operating the terminal can then begin actual work using the received content.
[0201] Users create reference videos during their work and upload them to the server via their terminal. An analysis device on the server analyzes the uploaded videos, evaluating the user's gaze, voice characteristics, posture, and emotional state. This analysis result is provided as feedback through a support structure that includes suggestions for necessary improvements and work environment enhancements.
[0202] As a concrete example, when a robot assists an operator in a factory, it can sense the operator's stress and suggest improvements such as, "Let's pause this task, have some water, and relax." An example of a prompt used in this process would be, "To improve work efficiency in the factory, the robot will recognize the operator's emotions and suggest providing appropriate feedback." This makes it possible to improve the work environment and enhance safety.
[0203] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0204] Step 1:
[0205] The user enters prompts using a terminal. The input device receives these prompts and prepares to send them to the generative AI model. Here, the input is the user's requests or instructions, and the output is the data sent to the generative AI model.
[0206] Step 2:
[0207] The server uses a generation AI model to generate visual materials, audio scripts, and demo sound images based on the received prompts. The generation device automatically collects relevant information based on the prompts and performs the process of generating content. The data processing performed here is information generation based on the prompts, and the output is the generated content.
[0208] Step 3:
[0209] The generated content is transmitted from the server to the terminal via a communication device. The server's role is to send and receive content; the input is the generated content, and the output is the data sent to the terminal.
[0210] Step 4:
[0211] The user performs tasks and exercises based on content received through the device. The actions in this step represent actual practice and work using the content, and indicate a specific feedback point. User evaluation audio and video data are obtained as output.
[0212] Step 5:
[0213] The user records a video of their work in progress and uploads it to the server via their device. The uploaded video becomes the input data. At this stage, the server prepares the received data for the next analysis step.
[0214] Step 6:
[0215] The server uses an analysis device to analyze uploaded videos and evaluate the user's gaze, posture, voice tone, and emotional state. The input is video data, and the output is the analyzed evaluation information. The process involves both video and audio analysis.
[0216] Step 7:
[0217] Based on the analysis results, the server generates necessary improvements and suggestions for improving the work environment through the support structure. The output is returned to the terminal as feedback information. The server utilizes a generation AI model to create specific improvement suggestions.
[0218] Step 8:
[0219] Users utilize feedback received on their devices to further improve their work and skills. This feedback includes specific suggestions tailored to the user's emotional state and serves as guidance for future work. The input is feedback information, and the output is the improved user behavior and environment.
[0220] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0221] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0222] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0223] [Second Embodiment]
[0224] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0225] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0226] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0227] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0228] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0229] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0230] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0231] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0232] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0233] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0234] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0235] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0236] This invention is a system that efficiently supports speeches and presentations through prompt input. The system mainly consists of a server and a terminal, and the user operates it via an interface. When the user inputs prompts regarding the theme and content of the presentation into the terminal, that information is sent to the server.
[0237] The server utilizes a generative AI model to automatically generate prompt-based presentation materials, speech scripts, and demo videos. In this process, the server collects relevant information from databases and external sources to construct specific and visually appealing materials tailored to the user's requirements. The generated content is then transmitted from the server to the terminal and provided to the user in a usable format.
[0238] Next, the user practices their speech using the generated content and records themselves. The practice video is uploaded to the server via the device, where it is analyzed. Image analysis technology is used to evaluate features such as the user's gaze, posture, and voice tone. The evaluation results are used to improve the user's speech performance.
[0239] Based on the analysis results, the server creates instructional videos that include specific improvement suggestions. These videos utilize synthesized speech and animation to provide users with visually and aurally easy-to-understand feedback. The instructional videos are sent to the user's device, and users can use them to improve their speech skills.
[0240] As a concrete example, consider preparing a presentation on the marketing strategy for a new product. When a user inputs the topic, the server generates presentation materials including product features and market information. Next, the user practices their speech, and feedback from the uploaded video indicates areas for improvement, such as how to emphasize certain points and how to direct eye contact. This allows users to deliver high-quality presentations in a short amount of time.
[0241] The following describes the processing flow.
[0242] Step 1:
[0243] The user enters prompts on the terminal regarding the presentation's theme and objectives. The entered prompts are formatted as digital data and sent to the server.
[0244] Step 2:
[0245] The server analyzes the received prompts and identifies the information needed to generate the corresponding presentation materials, speech scripts, and demo videos. The server searches for relevant content from databases and online resources and automatically generates content based on the collected data. A generation AI model is used to create materials in a format and expression appropriate to the prompt.
[0246] Step 3:
[0247] The server sends the generated presentation materials, speech scripts, and demo videos to the terminal. This content is then displayed on the terminal, allowing the user to review and utilize it.
[0248] Step 4:
[0249] The user practices their speech based on the materials and script they receive. The practice process is recorded with a camera and saved to the device. The recorded practice video is then uploaded to the server via the device.
[0250] Step 5:
[0251] The server analyzes the uploaded practice videos. Using image analysis technology, it evaluates the user's posture, gaze, voice tone, speed, etc. Based on the analysis results, it quantifies the user's performance and identifies areas that need improvement.
[0252] Step 6:
[0253] The server generates instructional videos based on the analysis results. These videos include specific improvement suggestions and advice, presented to the user visually and audibly. Synthesized speech and animation are used to provide feedback in an easily understandable format.
[0254] Step 7:
[0255] The server sends the generated instructional video to the user's device. The user watches the instructional video on their device and uses it to improve their speech skills. By practicing again based on the areas for improvement, it is possible to improve the quality of the presentation.
[0256] (Example 1)
[0257] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0258] Efficiently preparing for and improving speeches and presentations has traditionally required a great deal of time and effort. Furthermore, the lack of means to obtain concrete feedback through visual and auditory means hindered user growth. This invention proposes a system that provides concrete support for users to quickly prepare high-quality presentations and improve their skills.
[0259] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0260] In this invention, the server includes a device that receives information input from a user, including themes and content; a device that automatically creates materials, documents, and videos based on said information, and includes a device that collects relevant information from information sources; and a device that transmits the created data to the user's terminal. This allows users to efficiently create high-quality presentation materials and subsequently improve their skills based on performance analysis and specific feedback.
[0261] A "user" is an individual or group that uses this system to prepare speeches and presentations and to improve their own performance.
[0262] "Information including themes and content" refers to a collection of topics and specific information that users input to form the basis of their speeches or presentations.
[0263] A "device" is a combination of hardware and software that constitutes a system and performs a specific function.
[0264] "Automatically creating documents, texts, and videos" refers to the process of using computer programs based on input information to generate relevant content without user intervention.
[0265] "Collecting relevant information from sources" refers to the act of obtaining necessary data from internal databases or external information providers and forming the data set required for the generation process.
[0266] "Sending to the user's terminal" refers to the process of transmitting generated data to the user's terminal via the network, making the data available for the user to use.
[0267] "Exercise records" refer to data that is a recording or document of the content of a speech or presentation given by a user.
[0268] "Processing" refers to a set of procedures that involve analyzing data to achieve a specific purpose and generating results through calculations or operations.
[0269] "Eye movements, vocal characteristics, and posture" refer to the physical and vocal characteristics observed during a user's speech or presentation, and these are elements used to determine the quality of performance through evaluation.
[0270] "Educational videos" are video content created to provide users with feedback and instruction aimed at improving their skills.
[0271] Modes for carrying out the invention
[0272] This invention relates to a system for users to efficiently prepare and improve speeches and presentations. The system has multiple components, including a server, terminals, and interfaces.
[0273] User roles
[0274] The user first accesses the interface through a terminal and inputs prompts with themes and content related to their speech or presentation. These prompts are specific, such as "I would like to prepare materials to explain the importance of sustainable energy" or "I would like a presentation created that details the market impact of a new product."
[0275] Server Processing
[0276] The server receives prompts from users and automatically generates appropriate content using a generative AI model. This generative AI model incorporates natural language processing and machine learning technologies to generate materials, speeches, and video content that align with the user's intent. To gather relevant information, the server interacts with databases and publicly available online resources. The hardware used is a server equipped with a powerful processor.
[0277] Device usage and feedback
[0278] The generated content is sent from the server to the terminal and provided to the user in a format that the user can view. The user practices their speech based on this content and records their performance using the terminal's recording function. Afterwards, the user uploads the recorded practice session back to the server, where it uses image analysis technology to extract features and perform analysis. This analysis includes eye tracking, voice tone analysis, and posture monitoring.
[0279] Based on the analysis results, the server generates educational videos, visualizes feedback by utilizing synthesized voices and animations, and provides it to the user in an easy-to-understand form. Through this process, the user can obtain an opportunity to specifically and efficiently improve their speech skills.
[0280] In this way, the inventive system provides support for the user to perform high-quality speech and presentations in a short time.
[0281] The flow of specific processing in Example 1 will be described using FIG. 11.
[0282] Step 1:
[0283] The user inputs a prompt sentence through the terminal. This prompt sentence includes the theme and content of the speech or presentation. The terminal transmits this information to the server as input data.
[0284] Step 2:
[0285] The server analyzes the received prompt sentence and performs data processing to automatically generate materials, articles, and demo videos using the generation AI model. By extracting relevant information from the input prompt sentence and obtaining information from the database or external information sources, detailed and visual content is generated. As the output of this process, the generated content is obtained.
[0286] Step 3:
[0287] The server transmits the generated content to the terminal to make it accessible to the user. The terminal displays this received data in a form that the user can view and utilize.
[0288] Step 4:
[0289] Users practice their speeches using the generated materials and speech scripts on their devices. The practice sessions are recorded using the device's camera and microphone, and a file is created as a record of the practice.
[0290] Step 5:
[0291] Users upload their practice sessions to the server via their terminal. The server receives these recordings as input data and performs data processing using image analysis techniques to evaluate features such as gaze, posture, and voice tone. The output of the analysis is evaluation data regarding the user's performance.
[0292] Step 6:
[0293] Based on the analysis results, the server generates educational videos that include specific feedback. Utilizing a generation AI model and multimedia editing techniques, it uses synthesized speech and animation to create content that is easy for users to understand. The output of this process is the educational video.
[0294] Step 7:
[0295] The server sends the generated educational video to the terminal, allowing the user to watch it and provide feedback. The terminal plays the video and performs actions to enhance the user's viewing experience.
[0296] (Application Example 1)
[0297] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0298] The customer service skills of sales staff in physical stores significantly impact store sales and customer satisfaction. However, traditional training methods often lack sufficient specific feedback for individual sales staff, making efficient skill improvement difficult. In particular, the ability to effectively communicate information about new products and campaigns is crucial, and immediate and specific guidance is needed to enable sales staff to respond flexibly on the spot.
[0299] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0300] In this invention, the server includes an information input means for receiving instructions on a topic from the user, an automatic generation means, and an information transmission means. This allows sales staff to efficiently receive individualized customer service training. Furthermore, they can receive specific guidance to effectively explain new product information and improve customer interaction during actual sales activities in stores.
[0301] "Users" refer to salespeople and staff who use this system to improve their customer service skills.
[0302] "Theme-related instructions" refer to information and requests entered into the system regarding the products and services handled by the sales staff.
[0303] "Information input means" refers to devices or interfaces used by users to input instructions into a system.
[0304] "Automatic generation means" refers to devices or software that have the function of generating explanatory materials, speech scripts, and action videos based on a given theme.
[0305] "Information transmission means" refers to a system for transmitting generated content to a user's device.
[0306] "Practice video" refers to a recorded video of customer service skill practice uploaded by the user to the system.
[0307] "Information analysis means" refers to a device or program having a function for analyzing a practice video and evaluating important points such as the direction of the line of sight, the pitch of the voice, and the posture.
[0308] "Education generation means" refers to a system for creating an instructional video for improving the user's skills based on the analysis results.
[0309] "Display means" refers to a device or interface for presenting the generated instructional video to the user so that it can be viewed.
[0310] "Training means" refers to a specific method or program for implementing guidance aimed at improving customer service skills in a physical store.
[0311] To implement this invention, first, a device for the user to input information, such as a smartphone or smart glasses, is required. Through this, the user inputs instructions regarding products or services for sale. This instruction is, as a specific example, a prompt such as "Please create a presentation explaining the 100x zoom function of the latest smartphone."
[0312] Upon receiving the input prompt, the server automatically generates the necessary content using a generation AI model. The generated content includes explanatory materials, presentation manuscripts, and operation videos. For these generations, a backend program in Python and OpenAI's GPT as an AI model are used. The generated content is delivered to the user's device by the information transmission means.
[0313] Next, users practice customer service skills based on this content. This practice is recorded and uploaded to a server as practice video. The server analyzes this video using image analysis technologies such as OpenCV and TensorFlow. Specifically, information such as gaze direction, voice pitch, and posture is extracted.
[0314] Based on the analysis results, the educational generation system generates instructional videos to support the user's skill improvement. Using synthesized speech and animation, specific areas for improvement are presented visually and audibly in an easy-to-understand manner. The instructional videos are displayed on the user's device, allowing sales staff to watch them and improve their customer service skills.
[0315] Furthermore, the system includes training methods that provide instruction on specific skills that users can immediately use in their in-store sales activities. This promotes effective customer service in stores.
[0316] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0317] Step 1:
[0318] The user enters a prompt using their device. The prompt is a topic-related instruction, such as, "Create a presentation explaining the 100x zoom feature of the latest smartphones." This input is structured as data sent from the device to the server.
[0319] Step 2:
[0320] The server uses a generative AI model to generate presentation materials, speech scripts, and action videos based on the received prompts. For data processing, the input prompts are analyzed using natural language processing, and relevant information is retrieved from databases and external sources. The output generates specific and visually appealing content tailored to the user's requests.
[0321] Step 3:
[0322] The generated content is transmitted from the server to the terminal via an information transmission mechanism. The user's terminal receives the data and begins practicing customer service skills based on it. The terminal's actions involve the appropriate display and demonstration of the generated content received from the server.
[0323] Step 4:
[0324] Users practice and record their sessions. The recorded practice videos are uploaded from the terminal to the server as data. This video is then input into the system and used as material for the next analysis.
[0325] Step 5:
[0326] The server receives the practice video and performs analysis using image analysis techniques. Specifically, it uses OpenCV and TensorFlow to extract features such as gaze, voice pitch, and posture. The input is the user's practice video, and the output is the analyzed feature data.
[0327] Step 6:
[0328] Based on the analysis results, the server generates instructional videos using educational generation tools. Synthesized speech and animation are used in the generation process to clearly and concretely demonstrate areas for improvement for the user. The output is an instructional video designed to promote skill development in the user.
[0329] Step 7:
[0330] The generated instructional video is sent from the server to the terminal. The user's terminal receives and displays the instructional video. The user can watch it and work on improving their skills based on the feedback. The terminal's operation includes playing the video and presenting it to the user.
[0331] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0332] This invention is a system that supports the improvement of users' speech skills, and in particular, recognizes the user's emotions using an emotion engine and provides feedback based on that information. The system mainly consists of a server, a terminal, and the user.
[0333] The user enters prompts related to the topic of their speech or presentation through their device. These prompts are sent from the device to a server. The server analyzes the received prompts, automatically collects relevant information, and generates presentation materials, speech scripts, and demo videos. This process is driven by a generative AI model, providing concrete and visually appealing content. The generated content is then sent to the user via their device and made available for use.
[0334] Users practice their speeches based on the generated content and record the process on video. The recorded practice videos are uploaded to the server via the device. The server analyzes the received practice videos using image and audio analysis technologies, evaluating not only the user's gaze, posture, and voice tone, but also their emotional state using an emotion engine. The emotion engine recognizes emotions from changes in facial expressions, voice tone, pitch, etc., and analyzes the type and intensity of those emotions.
[0335] The analysis results are used to generate instructional videos. The server creates instructional videos that include specific areas for improvement based on the user's expression and emotional state. These instructional videos provide feedback that reflects emotions in an easy-to-understand way for the user, using synthesized speech and animation. For example, if the video detects nervousness during a presentation, it will include suggestions on how to relax. These instructional videos are sent to the user's device, and the user watches them to guide further practice.
[0336] Through the operation of these systems, users can not only improve the quality of their presentations but also learn how to control their emotions and improve their overall public speaking skills.
[0337] The following describes the processing flow.
[0338] Step 1:
[0339] The user enters the presentation topic into the terminal. The entered prompt is digitized by the terminal and sent to the server.
[0340] Step 2:
[0341] The server receives a prompt and uses generation AI to search for relevant data. Based on the collected data, it automatically generates presentation materials, speech scripts, and demo videos. The created content is then sent from the server to the terminal.
[0342] Step 3:
[0343] The terminal provides the user with content received from the server. The user uses the provided materials and script to practice their speech.
[0344] Step 4:
[0345] Users record their practice sessions and save the videos to their devices. The completed practice videos are then uploaded from the device to the server.
[0346] Step 5:
[0347] The server analyzes the uploaded practice videos. Using image and audio analysis technologies, it evaluates the user's gaze, posture, voice tone, and emotions using an emotion engine. The emotion engine identifies the user's emotional state by analyzing facial expressions and changes in voice.
[0348] Step 6:
[0349] The server generates instructional videos based on the analysis results. These instructional videos include specific advice on areas for improvement and methods for controlling emotions. The videos utilize synthesized speech and animation to present information in a visual and auditory way that is easy for users to understand.
[0350] Step 7:
[0351] Instructional videos are sent to the user's device for viewing. Based on the feedback received, the user can use it to improve their performance in future speech practice sessions and actual presentations. By repeating this process, the user can refine their skills.
[0352] (Example 2)
[0353] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0354] A challenge exists in that users cannot effectively improve their emotional expression, visual information, and auditory information during speeches and presentations. Traditional methods limit opportunities for users to receive feedback and make self-assessment difficult, which can hinder skill improvement.
[0355] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0356] In this invention, the server includes means for receiving theme-based input information from the user, means for automatically generating material data and visual information using information processing technology, means for transferring the generated data to a terminal via communication, means for analyzing activity videos and evaluating visual and audio information, means for generating instructional information based on the analysis results using speech synthesis technology and visual technology, and means for providing the instructional information to the user visually and audibly. This enables the user to objectively improve their speech skills and emotional expression and enhance their abilities.
[0357] "Input information" refers to information provided by the user based on a given theme, and serves as the foundational data for the system to begin its analysis and generation processes.
[0358] "Information processing technology" refers to all technologies used to analyze received data and automatically generate document data and visual information.
[0359] "Document data" refers to automatically generated content for presentations and speeches, containing information that users can refer to.
[0360] "Visual information" refers to information that appeals to users visually, including demo videos and charts.
[0361] "Means of transfer via communication" refers to the infrastructure and protocols used to transmit generated data and information to a terminal.
[0362] "Activity videos" are videos recorded by users for speech practice, and their content is analyzed and used as material for evaluation.
[0363] "Means for evaluating visual and auditory information" refers to technologies for analyzing the content of activity videos and quantitatively and qualitatively evaluating the user's visual and auditory elements.
[0364] "Speech synthesis technology" is a technology for generating synthesized sounds that mimic human voices, and is used for voice feedback in instructional materials.
[0365] "Visual technology" refers to technologies that generate visual feedback and provide it to users in an easily understandable format.
[0366] "Instructional information" refers to feedback content generated based on analysis results, which includes suggestions and areas for improvement to enhance the user's skills.
[0367] This invention is an information processing system for supporting the improvement of speech skills. This system consists of three elements: a user, a terminal, and a server.
[0368] The user enters prompt text into the terminal based on the theme. For example, a prompt might be, "Please describe the elements of effective leadership." The terminal receives the user's input and sends it to the server via the internet.
[0369] The server utilizes a generative AI model to automatically generate document data and visual information based on this prompt. This process employs natural language processing techniques and machine learning algorithms. The materials include demo videos and presentation slides to visually represent specific information. The generated materials and data are returned to the terminal via a secure communication protocol.
[0370] The terminal displays data received from the server to the user. The user uses this data to practice their speech and records the process as an activity video. Recording is typically done using the terminal's built-in camera function.
[0371] Afterward, the user uploads the recorded activity video from their device to the server. The server uses image and audio analysis technologies to analyze this video and evaluate the user's visual and audio information. In addition to elements such as gaze, posture, and voice tone, an emotion analysis engine is also introduced to evaluate emotional expression from facial expressions.
[0372] Based on the analysis results, the server generates instructional information using speech synthesis and visual technologies. This instructional information is provided as feedback content, including specific areas for improvement and advice, with the aim of improving the user's skills. For example, if the server determines that the user is stressed, it will provide suggestions for relaxation techniques.
[0373] All data is transmitted to the device, and users can effectively improve their speech skills by viewing or using it. This system allows users to improve their skills and the quality of their presentations.
[0374] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0375] Step 1:
[0376] The user enters prompt sentences related to the topic of their speech or presentation into the terminal. These prompt sentences become data that instructs the system to analyze. After input, the terminal converts this data into a digital format and sends it to the server.
[0377] Step 2:
[0378] The server parses the received prompt message. This parsing uses a generative AI model, leveraging natural language processing techniques to understand the intent of the prompt. The AI model collects necessary information from a database to generate relevant data and visual information, creating presentation slides and demo videos. The output of this process is content in a user-friendly format.
[0379] Step 3:
[0380] The server sends the generated material data and visual information to the terminal. The terminal receives this and presents it to the user through an appropriate interface. The user practices their speech using the provided materials. Reviewing these materials helps improve understanding of the presentation content and self-expression.
[0381] Step 4:
[0382] The user records their speech practice using the device's camera function. The recorded content includes data on the user's visual and auditory expression. After recording, the video is converted to an appropriate format on the device and uploaded to the server.
[0383] Step 5:
[0384] The server receives uploaded practice videos and analyzes them using image and audio analysis technologies. This process evaluates data related to the user's gaze, posture, and voice tone. Additionally, an emotion analysis engine is used to extract emotional information from the user's facial expressions and voice. The output after analysis is evaluation information that can be used for instructional feedback.
[0385] Step 6:
[0386] The server generates instructional feedback based on the analysis results. Utilizing a generation AI model, it creates videos containing specific advice for the user using speech synthesis and visual technologies. For example, this feedback might include suggestions for relaxation techniques or tips for improving expressiveness.
[0387] Step 7:
[0388] The server sends instructional feedback to the device. The device receives this feedback and provides it to the user in a viewable format. Through this feedback video, the user can objectively evaluate their own performance and use it to further improve their skills.
[0389] (Application Example 2)
[0390] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0391] To improve the efficiency and safety of workers on site, a system is needed that can appropriately understand workers' emotional states and stress levels and provide feedback at the appropriate time. However, conventional technology has made it difficult to recognize workers' emotional states in real time and respond to changes in them. There is a need to provide a system that solves these problems and realizes a better working environment.
[0392] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0393] In this invention, the server includes an input device that receives prompts from the user, a generation device that automatically generates visual materials, audio scripts, and demo sound images based on the prompts, and a communication device that transmits the generated information to a terminal. This allows the worker's emotional state to be recognized and to receive feedback as needed.
[0394] A "prompt" is an instruction or request provided by the user through input.
[0395] An "input device" refers to a device used to receive data and instructions from a user.
[0396] A "generation device" is a device that has the function of processing information based on received instructions and creating the necessary content.
[0397] A "communication device" is a device used to transmit generated information to other devices or networks.
[0398] "Reference videos" refer to video content used to analyze information uploaded by users.
[0399] An "analysis device" is a device used to analyze the characteristics of an object based on collected data.
[0400] A "support structure" refers to a system that provides necessary responses and feedback based on the analysis results.
[0401] A "presentation device" is a device that provides generated feedback and instructional content to the user visually or audibly.
[0402] An "improvement suggestion" is a measure to improve the work environment or methods suggested by the system, based on the analysis results.
[0403] In a system for realizing this invention, the server receives prompts from the user via an input device and uses a generation AI model to create visual materials, audio scripts, and demo sound images using a generation device. This generated content is transmitted to a terminal via a communication device. The user operating the terminal can then begin actual work using the received content.
[0404] Users create reference videos during their work and upload them to the server via their terminal. An analysis device on the server analyzes the uploaded videos, evaluating the user's gaze, voice characteristics, posture, and emotional state. This analysis result is provided as feedback through a support structure that includes suggestions for necessary improvements and work environment enhancements.
[0405] As a concrete example, when a robot assists an operator in a factory, it can sense the operator's stress and suggest improvements such as, "Let's pause this task, have some water, and relax." An example of a prompt used in this process would be, "To improve work efficiency in the factory, the robot will recognize the operator's emotions and suggest providing appropriate feedback." This makes it possible to improve the work environment and enhance safety.
[0406] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0407] Step 1:
[0408] The user enters prompts using a terminal. The input device receives these prompts and prepares to send them to the generative AI model. Here, the input is the user's requests or instructions, and the output is the data sent to the generative AI model.
[0409] Step 2:
[0410] The server uses a generation AI model to generate visual materials, audio scripts, and demo sound images based on the received prompts. The generation device automatically collects relevant information based on the prompts and performs the process of generating content. The data processing performed here is information generation based on the prompts, and the output is the generated content.
[0411] Step 3:
[0412] The generated content is transmitted from the server to the terminal via a communication device. The server's role is to send and receive content; the input is the generated content, and the output is the data sent to the terminal.
[0413] Step 4:
[0414] The user performs tasks and exercises based on content received through the device. The actions in this step represent actual practice and work using the content, and indicate a specific feedback point. User evaluation audio and video data are obtained as output.
[0415] Step 5:
[0416] The user records a video of their work in progress and uploads it to the server via their device. The uploaded video becomes the input data. At this stage, the server prepares the received data for the next analysis step.
[0417] Step 6:
[0418] The server uses an analysis device to analyze uploaded videos and evaluate the user's gaze, posture, voice tone, and emotional state. The input is video data, and the output is the analyzed evaluation information. The process involves both video and audio analysis.
[0419] Step 7:
[0420] Based on the analysis results, the server generates necessary improvements and suggestions for improving the work environment through the support structure. The output is returned to the terminal as feedback information. The server utilizes a generation AI model to create specific improvement suggestions.
[0421] Step 8:
[0422] Users utilize feedback received on their devices to further improve their work and skills. This feedback includes specific suggestions tailored to the user's emotional state and serves as guidance for future work. The input is feedback information, and the output is the improved user behavior and environment.
[0423] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0424] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0425] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0426] [Third Embodiment]
[0427] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0428] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0429] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0430] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0431] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0432] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0433] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0434] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0435] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0436] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0437] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0438] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0439] This invention is a system that efficiently supports speeches and presentations through prompt input. The system mainly consists of a server and a terminal, and the user operates it via an interface. When the user inputs prompts regarding the theme and content of the presentation into the terminal, that information is sent to the server.
[0440] The server utilizes a generative AI model to automatically generate prompt-based presentation materials, speech scripts, and demo videos. In this process, the server collects relevant information from databases and external sources to construct specific and visually appealing materials tailored to the user's requirements. The generated content is then transmitted from the server to the terminal and provided to the user in a usable format.
[0441] Next, the user practices their speech using the generated content and records themselves. The practice video is uploaded to the server via the device, where it is analyzed. Image analysis technology is used to evaluate features such as the user's gaze, posture, and voice tone. The evaluation results are used to improve the user's speech performance.
[0442] Based on the analysis results, the server creates instructional videos that include specific improvement suggestions. These videos utilize synthesized speech and animation to provide users with visually and aurally easy-to-understand feedback. The instructional videos are sent to the user's device, and users can use them to improve their speech skills.
[0443] As a concrete example, consider preparing a presentation on the marketing strategy for a new product. When a user inputs the topic, the server generates presentation materials including product features and market information. Next, the user practices their speech, and feedback from the uploaded video indicates areas for improvement, such as how to emphasize certain points and how to direct eye contact. This allows users to deliver high-quality presentations in a short amount of time.
[0444] The following describes the processing flow.
[0445] Step 1:
[0446] The user enters prompts on the terminal regarding the presentation's theme and objectives. The entered prompts are formatted as digital data and sent to the server.
[0447] Step 2:
[0448] The server analyzes the received prompts and identifies the information needed to generate the corresponding presentation materials, speech scripts, and demo videos. The server searches for relevant content from databases and online resources and automatically generates content based on the collected data. A generation AI model is used to create materials in a format and expression appropriate to the prompt.
[0449] Step 3:
[0450] The server sends the generated presentation materials, speech scripts, and demo videos to the terminal. This content is then displayed on the terminal, allowing the user to review and utilize it.
[0451] Step 4:
[0452] The user practices their speech based on the materials and script they receive. The practice process is recorded with a camera and saved to the device. The recorded practice video is then uploaded to the server via the device.
[0453] Step 5:
[0454] The server analyzes the uploaded practice videos. Using image analysis technology, it evaluates the user's posture, gaze, voice tone, speed, etc. Based on the analysis results, it quantifies the user's performance and identifies areas that need improvement.
[0455] Step 6:
[0456] The server generates instructional videos based on the analysis results. These videos include specific improvement suggestions and advice, presented to the user visually and audibly. Synthesized speech and animation are used to provide feedback in an easily understandable format.
[0457] Step 7:
[0458] The server sends the generated instructional video to the user's device. The user watches the instructional video on their device and uses it to improve their speech skills. By practicing again based on the areas for improvement, it is possible to improve the quality of the presentation.
[0459] (Example 1)
[0460] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0461] Efficiently preparing for and improving speeches and presentations has traditionally required a great deal of time and effort. Furthermore, the lack of means to obtain concrete feedback through visual and auditory means hindered user growth. This invention proposes a system that provides concrete support for users to quickly prepare high-quality presentations and improve their skills.
[0462] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0463] In this invention, the server includes a device that receives information input from a user, including themes and content; a device that automatically creates materials, documents, and videos based on said information, and includes a device that collects relevant information from information sources; and a device that transmits the created data to the user's terminal. This allows users to efficiently create high-quality presentation materials and subsequently improve their skills based on performance analysis and specific feedback.
[0464] A "user" is an individual or group that uses this system to prepare speeches and presentations and to improve their own performance.
[0465] "Information including themes and content" refers to a collection of topics and specific information that users input to form the basis of their speeches or presentations.
[0466] A "device" is a combination of hardware and software that constitutes a system and performs a specific function.
[0467] "Automatically creating documents, texts, and videos" refers to the process of using computer programs based on input information to generate relevant content without user intervention.
[0468] "Collecting relevant information from sources" refers to the act of obtaining necessary data from internal databases or external information providers and forming the data set required for the generation process.
[0469] "Sending to the user's terminal" refers to the process of transmitting generated data to the user's terminal via the network, making the data available for the user to use.
[0470] "Exercise records" refer to data that is a recording or document of the content of a speech or presentation given by a user.
[0471] "Processing" refers to a set of procedures that involve analyzing data to achieve a specific purpose and generating results through calculations or operations.
[0472] "Eye movements, vocal characteristics, and posture" refer to the physical and vocal characteristics observed during a user's speech or presentation, and these are elements used to determine the quality of performance through evaluation.
[0473] "Educational videos" are video content created to provide users with feedback and instruction aimed at improving their skills.
[0474] Modes for carrying out the invention
[0475] This invention relates to a system for users to efficiently prepare and improve speeches and presentations. The system has multiple components, including a server, terminals, and interfaces.
[0476] User roles
[0477] The user first accesses the interface through a terminal and inputs prompts with themes and content related to their speech or presentation. These prompts are specific, such as "I would like to prepare materials to explain the importance of sustainable energy" or "I would like a presentation created that details the market impact of a new product."
[0478] Server Processing
[0479] The server receives prompts from users and automatically generates appropriate content using a generative AI model. This generative AI model incorporates natural language processing and machine learning technologies to generate materials, speeches, and video content that align with the user's intent. To gather relevant information, the server interacts with databases and publicly available online resources. The hardware used is a server equipped with a powerful processor.
[0480] Device usage and feedback
[0481] The generated content is sent from the server to the terminal and provided to the user in a format that the user can view. The user practices their speech based on this content and records their performance using the terminal's recording function. Afterwards, the user uploads the recorded practice session back to the server, where it uses image analysis technology to extract features and perform analysis. This analysis includes eye tracking, voice tone analysis, and posture monitoring.
[0482] Based on the analysis results, the server generates educational videos, visualizes feedback using synthesized speech and animation, and presents it to the user in an easy-to-understand format. This process gives users the opportunity to improve their speech skills in a concrete and efficient manner.
[0483] In this way, the invented system provides support for users to deliver high-quality speeches and presentations in a short amount of time.
[0484] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0485] Step 1:
[0486] The user enters a prompt message through the terminal. This prompt message includes the theme and content of the speech or presentation. The terminal sends this information to the server as input data.
[0487] Step 2:
[0488] The server analyzes the received prompt text and performs data processing using a generative AI model to automatically generate documents, text, and demo videos. It extracts relevant information from the input prompt text and retrieves information from databases and external sources to generate detailed and visual content. The generated content is obtained as the output of this process.
[0489] Step 3:
[0490] The server sends the generated content to the terminal, making it accessible to the user. The terminal then displays this received data in a format that the user can view and use.
[0491] Step 4:
[0492] Users practice their speeches using the generated materials and speech scripts on their devices. The practice sessions are recorded using the device's camera and microphone, and a file is created as a record of the practice.
[0493] Step 5:
[0494] Users upload their practice sessions to the server via their terminal. The server receives these recordings as input data and performs data processing using image analysis techniques to evaluate features such as gaze, posture, and voice tone. The output of the analysis is evaluation data regarding the user's performance.
[0495] Step 6:
[0496] Based on the analysis results, the server generates educational videos that include specific feedback. Utilizing a generation AI model and multimedia editing techniques, it uses synthesized speech and animation to create content that is easy for users to understand. The output of this process is the educational video.
[0497] Step 7:
[0498] The server sends the generated educational video to the terminal, allowing the user to watch it and provide feedback. The terminal plays the video and performs actions to enhance the user's viewing experience.
[0499] (Application Example 1)
[0500] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0501] The customer service skills of sales staff in physical stores significantly impact store sales and customer satisfaction. However, traditional training methods often lack sufficient specific feedback for individual sales staff, making efficient skill improvement difficult. In particular, the ability to effectively communicate information about new products and campaigns is crucial, and immediate and specific guidance is needed to enable sales staff to respond flexibly on the spot.
[0502] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0503] In this invention, the server includes an information input means for receiving instructions on a topic from the user, an automatic generation means, and an information transmission means. This allows sales staff to efficiently receive individualized customer service training. Furthermore, they can receive specific guidance to effectively explain new product information and improve customer interaction during actual sales activities in stores.
[0504] "Users" refer to salespeople and staff who use this system to improve their customer service skills.
[0505] "Theme-related instructions" refer to information and requests entered into the system regarding the products and services handled by the sales staff.
[0506] "Information input means" refers to devices or interfaces used by users to input instructions into a system.
[0507] "Automatic generation means" refers to devices or software that have the function of generating explanatory materials, speech scripts, and action videos based on a given theme.
[0508] "Information transmission means" refers to a system for transmitting generated content to a user's device.
[0509] "Training videos" refer to recorded videos of customer service skills practice that users upload to the system.
[0510] "Information analysis means" refers to devices or programs that have the function of analyzing exercise videos and evaluating important points such as the direction of gaze, pitch of voice, and posture.
[0511] "Educational generation means" refers to a system that creates instructional videos to improve users' skills based on analysis results.
[0512] "Display means" refers to devices or interfaces that present the generated instructional video to the user and enable them to view it.
[0513] "Training methods" refer to specific methods and programs for providing instruction aimed at improving customer service skills in physical stores.
[0514] To implement this invention, a device is first required for the user to input information, such as a smartphone or smart glasses. The user inputs instructions regarding products or services through this device. These instructions may include prompts such as, "Please create a presentation explaining the 100x zoom function of the latest smartphone."
[0515] The server receives input prompts and automatically generates the necessary content using a generative AI model. The generated content includes explanatory materials, presentation scripts, and even demonstration videos. This generation utilizes a Python backend program and OpenAI's GPT AI model. The generated content is then delivered to the user's device via a data transmission method.
[0516] Next, users practice customer service skills based on this content. This practice is recorded and uploaded to a server as practice video. The server analyzes this video using image analysis technologies such as OpenCV and TensorFlow. Specifically, information such as gaze direction, voice pitch, and posture is extracted.
[0517] Based on the analysis results, the educational generation system generates instructional videos to support the user's skill improvement. Using synthesized speech and animation, specific areas for improvement are presented visually and audibly in an easy-to-understand manner. The instructional videos are displayed on the user's device, allowing sales staff to watch them and improve their customer service skills.
[0518] Furthermore, the system includes training methods that provide instruction on specific skills that users can immediately use in their in-store sales activities. This promotes effective customer service in stores.
[0519] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0520] Step 1:
[0521] The user enters a prompt using their device. The prompt is a topic-related instruction, such as, "Create a presentation explaining the 100x zoom feature of the latest smartphones." This input is structured as data sent from the device to the server.
[0522] Step 2:
[0523] The server uses a generative AI model to generate presentation materials, speech scripts, and action videos based on the received prompts. For data processing, the input prompts are analyzed using natural language processing, and relevant information is retrieved from databases and external sources. The output generates specific and visually appealing content tailored to the user's requests.
[0524] Step 3:
[0525] The generated content is transmitted from the server to the terminal via an information transmission mechanism. The user's terminal receives the data and begins practicing customer service skills based on it. The terminal's actions involve the appropriate display and demonstration of the generated content received from the server.
[0526] Step 4:
[0527] Users practice and record their sessions. The recorded practice videos are uploaded from the terminal to the server as data. This video is then input into the system and used as material for the next analysis.
[0528] Step 5:
[0529] The server receives the practice video and performs analysis using image analysis techniques. Specifically, it uses OpenCV and TensorFlow to extract features such as gaze, voice pitch, and posture. The input is the user's practice video, and the output is the analyzed feature data.
[0530] Step 6:
[0531] Based on the analysis results, the server generates instructional videos using educational generation tools. Synthesized speech and animation are used in the generation process to clearly and concretely demonstrate areas for improvement for the user. The output is an instructional video designed to promote skill development in the user.
[0532] Step 7:
[0533] The generated instructional video is sent from the server to the terminal. The user's terminal receives and displays the instructional video. The user can watch it and work on improving their skills based on the feedback. The terminal's operation includes playing the video and presenting it to the user.
[0534] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0535] This invention is a system that supports the improvement of users' speech skills, and in particular, recognizes the user's emotions using an emotion engine and provides feedback based on that information. The system mainly consists of a server, a terminal, and the user.
[0536] The user enters prompts related to the topic of their speech or presentation through their device. These prompts are sent from the device to a server. The server analyzes the received prompts, automatically collects relevant information, and generates presentation materials, speech scripts, and demo videos. This process is driven by a generative AI model, providing concrete and visually appealing content. The generated content is then sent to the user via their device and made available for use.
[0537] Users practice their speeches based on the generated content and record the process on video. The recorded practice videos are uploaded to the server via the device. The server analyzes the received practice videos using image and audio analysis technologies, evaluating not only the user's gaze, posture, and voice tone, but also their emotional state using an emotion engine. The emotion engine recognizes emotions from changes in facial expressions, voice tone, pitch, etc., and analyzes the type and intensity of those emotions.
[0538] The analysis results are used to generate instructional videos. The server creates instructional videos that include specific areas for improvement based on the user's expression and emotional state. These instructional videos provide feedback that reflects emotions in an easy-to-understand way for the user, using synthesized speech and animation. For example, if the video detects nervousness during a presentation, it will include suggestions on how to relax. These instructional videos are sent to the user's device, and the user watches them to guide further practice.
[0539] Through the operation of these systems, users can not only improve the quality of their presentations but also learn how to control their emotions and improve their overall public speaking skills.
[0540] The following describes the processing flow.
[0541] Step 1:
[0542] The user enters the presentation topic into the terminal. The entered prompt is digitized by the terminal and sent to the server.
[0543] Step 2:
[0544] The server receives a prompt and uses generation AI to search for relevant data. Based on the collected data, it automatically generates presentation materials, speech scripts, and demo videos. The created content is then sent from the server to the terminal.
[0545] Step 3:
[0546] The terminal provides the user with content received from the server. The user uses the provided materials and script to practice their speech.
[0547] Step 4:
[0548] Users record their practice sessions and save the videos to their devices. The completed practice videos are then uploaded from the device to the server.
[0549] Step 5:
[0550] The server analyzes the uploaded practice videos. Using image and audio analysis technologies, it evaluates the user's gaze, posture, voice tone, and emotions using an emotion engine. The emotion engine identifies the user's emotional state by analyzing facial expressions and changes in voice.
[0551] Step 6:
[0552] The server generates instructional videos based on the analysis results. These instructional videos include specific advice on areas for improvement and methods for controlling emotions. The videos utilize synthesized speech and animation to present information in a visual and auditory way that is easy for users to understand.
[0553] Step 7:
[0554] Instructional videos are sent to the user's device for viewing. Based on the feedback received, the user can use it to improve their performance in future speech practice sessions and actual presentations. By repeating this process, the user can refine their skills.
[0555] (Example 2)
[0556] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0557] A challenge exists in that users cannot effectively improve their emotional expression, visual information, and auditory information during speeches and presentations. Traditional methods limit opportunities for users to receive feedback and make self-assessment difficult, which can hinder skill improvement.
[0558] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0559] In this invention, the server includes means for receiving theme-based input information from the user, means for automatically generating material data and visual information using information processing technology, means for transferring the generated data to a terminal via communication, means for analyzing activity videos and evaluating visual and audio information, means for generating instructional information based on the analysis results using speech synthesis technology and visual technology, and means for providing the instructional information to the user visually and audibly. This enables the user to objectively improve their speech skills and emotional expression and enhance their abilities.
[0560] "Input information" refers to information provided by the user based on a given theme, and serves as the foundational data for the system to begin its analysis and generation processes.
[0561] "Information processing technology" refers to all technologies used to analyze received data and automatically generate document data and visual information.
[0562] "Document data" refers to automatically generated content for presentations and speeches, containing information that users can refer to.
[0563] "Visual information" refers to information that appeals to users visually, including demo videos and charts.
[0564] "Means of transfer via communication" refers to the infrastructure and protocols used to transmit generated data and information to a terminal.
[0565] "Activity videos" are videos recorded by users for speech practice, and their content is analyzed and used as material for evaluation.
[0566] "Means for evaluating visual and auditory information" refers to technologies for analyzing the content of activity videos and quantitatively and qualitatively evaluating the user's visual and auditory elements.
[0567] "Speech synthesis technology" is a technology for generating synthesized sounds that mimic human voices, and is used for voice feedback in instructional materials.
[0568] "Visual technology" refers to technologies that generate visual feedback and provide it to users in an easily understandable format.
[0569] "Instructional information" refers to feedback content generated based on analysis results, which includes suggestions and areas for improvement to enhance the user's skills.
[0570] This invention is an information processing system for supporting the improvement of speech skills. This system consists of three elements: a user, a terminal, and a server.
[0571] The user enters prompt text into the terminal based on the theme. For example, a prompt might be, "Please describe the elements of effective leadership." The terminal receives the user's input and sends it to the server via the internet.
[0572] The server utilizes a generative AI model to automatically generate document data and visual information based on this prompt. This process employs natural language processing techniques and machine learning algorithms. The materials include demo videos and presentation slides to visually represent specific information. The generated materials and data are returned to the terminal via a secure communication protocol.
[0573] The terminal displays data received from the server to the user. The user uses this data to practice their speech and records the process as an activity video. Recording is typically done using the terminal's built-in camera function.
[0574] Afterward, the user uploads the recorded activity video from their device to the server. The server uses image and audio analysis technologies to analyze this video and evaluate the user's visual and audio information. In addition to elements such as gaze, posture, and voice tone, an emotion analysis engine is also introduced to evaluate emotional expression from facial expressions.
[0575] Based on the analysis results, the server generates instructional information using speech synthesis and visual technologies. This instructional information is provided as feedback content, including specific areas for improvement and advice, with the aim of improving the user's skills. For example, if the server determines that the user is stressed, it will provide suggestions for relaxation techniques.
[0576] All data is transmitted to the device, and users can effectively improve their speech skills by viewing or using it. This system allows users to improve their skills and the quality of their presentations.
[0577] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0578] Step 1:
[0579] The user enters prompt sentences related to the topic of their speech or presentation into the terminal. These prompt sentences become data that instructs the system to analyze. After input, the terminal converts this data into a digital format and sends it to the server.
[0580] Step 2:
[0581] The server parses the received prompt message. This parsing uses a generative AI model, leveraging natural language processing techniques to understand the intent of the prompt. The AI model collects necessary information from a database to generate relevant data and visual information, creating presentation slides and demo videos. The output of this process is content in a user-friendly format.
[0582] Step 3:
[0583] The server sends the generated material data and visual information to the terminal. The terminal receives this and presents it to the user through an appropriate interface. The user practices their speech using the provided materials. Reviewing these materials helps improve understanding of the presentation content and self-expression.
[0584] Step 4:
[0585] The user records their speech practice using the device's camera function. The recorded content includes data on the user's visual and auditory expression. After recording, the video is converted to an appropriate format on the device and uploaded to the server.
[0586] Step 5:
[0587] The server receives uploaded practice videos and analyzes them using image and audio analysis technologies. This process evaluates data related to the user's gaze, posture, and voice tone. Additionally, an emotion analysis engine is used to extract emotional information from the user's facial expressions and voice. The output after analysis is evaluation information that can be used for instructional feedback.
[0588] Step 6:
[0589] The server generates instructional feedback based on the analysis results. Utilizing a generation AI model, it creates videos containing specific advice for the user using speech synthesis and visual technologies. For example, this feedback might include suggestions for relaxation techniques or tips for improving expressiveness.
[0590] Step 7:
[0591] The server sends instructional feedback to the device. The device receives this feedback and provides it to the user in a viewable format. Through this feedback video, the user can objectively evaluate their own performance and use it to further improve their skills.
[0592] (Application Example 2)
[0593] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0594] To improve the efficiency and safety of workers on site, a system is needed that can appropriately understand workers' emotional states and stress levels and provide feedback at the appropriate time. However, conventional technology has made it difficult to recognize workers' emotional states in real time and respond to changes in them. There is a need to provide a system that solves these problems and realizes a better working environment.
[0595] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0596] In this invention, the server includes an input device that receives prompts from the user, a generation device that automatically generates visual materials, audio scripts, and demo sound images based on the prompts, and a communication device that transmits the generated information to a terminal. This allows the worker's emotional state to be recognized and to receive feedback as needed.
[0597] A "prompt" is an instruction or request provided by the user through input.
[0598] An "input device" refers to a device used to receive data and instructions from a user.
[0599] A "generation device" is a device that has the function of processing information based on received instructions and creating the necessary content.
[0600] A "communication device" is a device used to transmit generated information to other devices or networks.
[0601] "Reference videos" refer to video content used to analyze information uploaded by users.
[0602] An "analysis device" is a device used to analyze the characteristics of an object based on collected data.
[0603] A "support structure" refers to a system that provides necessary responses and feedback based on the analysis results.
[0604] A "presentation device" is a device that provides generated feedback and instructional content to the user visually or audibly.
[0605] An "improvement suggestion" is a measure to improve the work environment or methods suggested by the system, based on the analysis results.
[0606] In a system for realizing this invention, the server receives prompts from the user via an input device and uses a generation AI model to create visual materials, audio scripts, and demo sound images using a generation device. This generated content is transmitted to a terminal via a communication device. The user operating the terminal can then begin actual work using the received content.
[0607] Users create reference videos during their work and upload them to the server via their terminal. An analysis device on the server analyzes the uploaded videos, evaluating the user's gaze, voice characteristics, posture, and emotional state. This analysis result is provided as feedback through a support structure that includes suggestions for necessary improvements and work environment enhancements.
[0608] As a concrete example, when a robot assists an operator in a factory, it can sense the operator's stress and suggest improvements such as, "Let's pause this task, have some water, and relax." An example of a prompt used in this process would be, "To improve work efficiency in the factory, the robot will recognize the operator's emotions and suggest providing appropriate feedback." This makes it possible to improve the work environment and enhance safety.
[0609] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0610] Step 1:
[0611] The user enters prompts using a terminal. The input device receives these prompts and prepares to send them to the generative AI model. Here, the input is the user's requests or instructions, and the output is the data sent to the generative AI model.
[0612] Step 2:
[0613] The server uses a generation AI model to generate visual materials, audio scripts, and demo sound images based on the received prompts. The generation device automatically collects relevant information based on the prompts and performs the process of generating content. The data processing performed here is information generation based on the prompts, and the output is the generated content.
[0614] Step 3:
[0615] The generated content is transmitted from the server to the terminal via a communication device. The server's role is to send and receive content; the input is the generated content, and the output is the data sent to the terminal.
[0616] Step 4:
[0617] The user performs tasks and exercises based on content received through the device. The actions in this step represent actual practice and work using the content, and indicate a specific feedback point. User evaluation audio and video data are obtained as output.
[0618] Step 5:
[0619] The user records a video of their work in progress and uploads it to the server via their device. The uploaded video becomes the input data. At this stage, the server prepares the received data for the next analysis step.
[0620] Step 6:
[0621] The server uses an analysis device to analyze uploaded videos and evaluate the user's gaze, posture, voice tone, and emotional state. The input is video data, and the output is the analyzed evaluation information. The process involves both video and audio analysis.
[0622] Step 7:
[0623] Based on the analysis results, the server generates necessary improvements and suggestions for improving the work environment through the support structure. The output is returned to the terminal as feedback information. The server utilizes a generation AI model to create specific improvement suggestions.
[0624] Step 8:
[0625] Users utilize feedback received on their devices to further improve their work and skills. This feedback includes specific suggestions tailored to the user's emotional state and serves as guidance for future work. The input is feedback information, and the output is the improved user behavior and environment.
[0626] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0627] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0628] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[0629] [Fourth Embodiment]
[0630] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0631] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0632] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0633] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0634] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0635] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0636] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0637] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0638] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0639] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0640] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0641] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0642] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0643] This invention is a system that efficiently supports speeches and presentations through prompt input. The system mainly consists of a server and a terminal, and the user operates it via an interface. When the user inputs prompts regarding the theme and content of the presentation into the terminal, that information is sent to the server.
[0644] The server utilizes a generative AI model to automatically generate prompt-based presentation materials, speech scripts, and demo videos. In this process, the server collects relevant information from databases and external sources to construct specific and visually appealing materials tailored to the user's requirements. The generated content is then transmitted from the server to the terminal and provided to the user in a usable format.
[0645] Next, the user practices their speech using the generated content and records themselves. The practice video is uploaded to the server via the device, where it is analyzed. Image analysis technology is used to evaluate features such as the user's gaze, posture, and voice tone. The evaluation results are used to improve the user's speech performance.
[0646] Based on the analysis results, the server creates instructional videos that include specific improvement suggestions. These videos utilize synthesized speech and animation to provide users with visually and aurally easy-to-understand feedback. The instructional videos are sent to the user's device, and users can use them to improve their speech skills.
[0647] As a concrete example, consider preparing a presentation on the marketing strategy for a new product. When a user inputs the topic, the server generates presentation materials including product features and market information. Next, the user practices their speech, and feedback from the uploaded video indicates areas for improvement, such as how to emphasize certain points and how to direct eye contact. This allows users to deliver high-quality presentations in a short amount of time.
[0648] The following describes the processing flow.
[0649] Step 1:
[0650] The user enters prompts on the terminal regarding the presentation's theme and objectives. The entered prompts are formatted as digital data and sent to the server.
[0651] Step 2:
[0652] The server analyzes the received prompts and identifies the information needed to generate the corresponding presentation materials, speech scripts, and demo videos. The server searches for relevant content from databases and online resources and automatically generates content based on the collected data. A generation AI model is used to create materials in a format and expression appropriate to the prompt.
[0653] Step 3:
[0654] The server sends the generated presentation materials, speech scripts, and demo videos to the terminal. This content is then displayed on the terminal, allowing the user to review and utilize it.
[0655] Step 4:
[0656] The user practices their speech based on the materials and script they receive. The practice process is recorded with a camera and saved to the device. The recorded practice video is then uploaded to the server via the device.
[0657] Step 5:
[0658] The server analyzes the uploaded practice videos. Using image analysis technology, it evaluates the user's posture, gaze, voice tone, speed, etc. Based on the analysis results, it quantifies the user's performance and identifies areas that need improvement.
[0659] Step 6:
[0660] The server generates instructional videos based on the analysis results. These videos include specific improvement suggestions and advice, presented to the user visually and audibly. Synthesized speech and animation are used to provide feedback in an easily understandable format.
[0661] Step 7:
[0662] The server sends the generated instructional video to the user's device. The user watches the instructional video on their device and uses it to improve their speech skills. By practicing again based on the areas for improvement, it is possible to improve the quality of the presentation.
[0663] (Example 1)
[0664] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0665] Efficiently preparing for and improving speeches and presentations has traditionally required a great deal of time and effort. Furthermore, the lack of means to obtain concrete feedback through visual and auditory means hindered user growth. This invention proposes a system that provides concrete support for users to quickly prepare high-quality presentations and improve their skills.
[0666] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0667] In this invention, the server includes a device that receives information input from a user, including themes and content; a device that automatically creates materials, documents, and videos based on said information, and includes a device that collects relevant information from information sources; and a device that transmits the created data to the user's terminal. This allows users to efficiently create high-quality presentation materials and subsequently improve their skills based on performance analysis and specific feedback.
[0668] A "user" is an individual or group that uses this system to prepare speeches and presentations and to improve their own performance.
[0669] "Information including themes and content" refers to a collection of topics and specific information that users input to form the basis of their speeches or presentations.
[0670] A "device" is a combination of hardware and software that constitutes a system and performs a specific function.
[0671] "Automatically creating documents, texts, and videos" refers to the process of using computer programs based on input information to generate relevant content without user intervention.
[0672] "Collecting relevant information from sources" refers to the act of obtaining necessary data from internal databases or external information providers and forming the data set required for the generation process.
[0673] "Sending to the user's terminal" refers to the process of transmitting generated data to the user's terminal via the network, making the data available for the user to use.
[0674] "Exercise records" refer to data that is a recording or document of the content of a speech or presentation given by a user.
[0675] "Processing" refers to a set of procedures that involve analyzing data to achieve a specific purpose and generating results through calculations or operations.
[0676] "Eye movements, vocal characteristics, and posture" refer to the physical and vocal characteristics observed during a user's speech or presentation, and these are elements used to determine the quality of performance through evaluation.
[0677] "Educational videos" are video content created to provide users with feedback and instruction aimed at improving their skills.
[0678] Modes for carrying out the invention
[0679] This invention relates to a system for users to efficiently prepare and improve speeches and presentations. The system has multiple components, including a server, terminals, and interfaces.
[0680] User roles
[0681] The user first accesses the interface through a terminal and inputs prompts with themes and content related to their speech or presentation. These prompts are specific, such as "I would like to prepare materials to explain the importance of sustainable energy" or "I would like a presentation created that details the market impact of a new product."
[0682] Server Processing
[0683] The server receives prompts from users and automatically generates appropriate content using a generative AI model. This generative AI model incorporates natural language processing and machine learning technologies to generate materials, speeches, and video content that align with the user's intent. To gather relevant information, the server interacts with databases and publicly available online resources. The hardware used is a server equipped with a powerful processor.
[0684] Device usage and feedback
[0685] The generated content is sent from the server to the terminal and provided to the user in a format that the user can view. The user practices their speech based on this content and records their performance using the terminal's recording function. Afterwards, the user uploads the recorded practice session back to the server, where it uses image analysis technology to extract features and perform analysis. This analysis includes eye tracking, voice tone analysis, and posture monitoring.
[0686] Based on the analysis results, the server generates educational videos, visualizes feedback using synthesized speech and animation, and presents it to the user in an easy-to-understand format. This process gives users the opportunity to improve their speech skills in a concrete and efficient manner.
[0687] In this way, the invented system provides support for users to deliver high-quality speeches and presentations in a short amount of time.
[0688] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0689] Step 1:
[0690] The user enters a prompt message through the terminal. This prompt message includes the theme and content of the speech or presentation. The terminal sends this information to the server as input data.
[0691] Step 2:
[0692] The server analyzes the received prompt text and performs data processing using a generative AI model to automatically generate documents, text, and demo videos. It extracts relevant information from the input prompt text and retrieves information from databases and external sources to generate detailed and visual content. The generated content is obtained as the output of this process.
[0693] Step 3:
[0694] The server sends the generated content to the terminal, making it accessible to the user. The terminal then displays this received data in a format that the user can view and use.
[0695] Step 4:
[0696] Users practice their speeches using the generated materials and speech scripts on their devices. The practice sessions are recorded using the device's camera and microphone, and a file is created as a record of the practice.
[0697] Step 5:
[0698] Users upload their practice sessions to the server via their terminal. The server receives these recordings as input data and performs data processing using image analysis techniques to evaluate features such as gaze, posture, and voice tone. The output of the analysis is evaluation data regarding the user's performance.
[0699] Step 6:
[0700] Based on the analysis results, the server generates educational videos that include specific feedback. Utilizing a generation AI model and multimedia editing techniques, it uses synthesized speech and animation to create content that is easy for users to understand. The output of this process is the educational video.
[0701] Step 7:
[0702] The server sends the generated educational video to the terminal, allowing the user to watch it and provide feedback. The terminal plays the video and performs actions to enhance the user's viewing experience.
[0703] (Application Example 1)
[0704] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0705] The customer service skills of sales staff in physical stores significantly impact store sales and customer satisfaction. However, traditional training methods often lack sufficient specific feedback for individual sales staff, making efficient skill improvement difficult. In particular, the ability to effectively communicate information about new products and campaigns is crucial, and immediate and specific guidance is needed to enable sales staff to respond flexibly on the spot.
[0706] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0707] In this invention, the server includes an information input means for receiving instructions on a topic from the user, an automatic generation means, and an information transmission means. This allows sales staff to efficiently receive individualized customer service training. Furthermore, they can receive specific guidance to effectively explain new product information and improve customer interaction during actual sales activities in stores.
[0708] "Users" refer to salespeople and staff who use this system to improve their customer service skills.
[0709] "Theme-related instructions" refer to information and requests entered into the system regarding the products and services handled by the sales staff.
[0710] "Information input means" refers to devices or interfaces used by users to input instructions into a system.
[0711] "Automatic generation means" refers to devices or software that have the function of generating explanatory materials, speech scripts, and action videos based on a given theme.
[0712] "Information transmission means" refers to a system for transmitting generated content to a user's device.
[0713] "Training videos" refer to recorded videos of customer service skills practice that users upload to the system.
[0714] "Information analysis means" refers to devices or programs that have the function of analyzing exercise videos and evaluating important points such as the direction of gaze, pitch of voice, and posture.
[0715] "Educational generation means" refers to a system that creates instructional videos to improve users' skills based on analysis results.
[0716] "Display means" refers to devices or interfaces that present the generated instructional video to the user and enable them to view it.
[0717] "Training methods" refer to specific methods and programs for providing instruction aimed at improving customer service skills in physical stores.
[0718] To implement this invention, a device is first required for the user to input information, such as a smartphone or smart glasses. The user inputs instructions regarding products or services through this device. These instructions may include prompts such as, "Please create a presentation explaining the 100x zoom function of the latest smartphone."
[0719] The server receives input prompts and automatically generates the necessary content using a generative AI model. The generated content includes explanatory materials, presentation scripts, and even demonstration videos. This generation utilizes a Python backend program and OpenAI's GPT AI model. The generated content is then delivered to the user's device via a data transmission method.
[0720] Next, users practice customer service skills based on this content. This practice is recorded and uploaded to a server as practice video. The server analyzes this video using image analysis technologies such as OpenCV and TensorFlow. Specifically, information such as gaze direction, voice pitch, and posture is extracted.
[0721] Based on the analysis results, the educational generation system generates instructional videos to support the user's skill improvement. Using synthesized speech and animation, specific areas for improvement are presented visually and audibly in an easy-to-understand manner. The instructional videos are displayed on the user's device, allowing sales staff to watch them and improve their customer service skills.
[0722] Furthermore, the system includes training methods that provide instruction on specific skills that users can immediately use in their in-store sales activities. This promotes effective customer service in stores.
[0723] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0724] Step 1:
[0725] The user enters a prompt using their device. The prompt is a topic-related instruction, such as, "Create a presentation explaining the 100x zoom feature of the latest smartphones." This input is structured as data sent from the device to the server.
[0726] Step 2:
[0727] The server uses a generative AI model to generate presentation materials, speech scripts, and action videos based on the received prompts. For data processing, the input prompts are analyzed using natural language processing, and relevant information is retrieved from databases and external sources. The output generates specific and visually appealing content tailored to the user's requests.
[0728] Step 3:
[0729] The generated content is transmitted from the server to the terminal via an information transmission mechanism. The user's terminal receives the data and begins practicing customer service skills based on it. The terminal's actions involve the appropriate display and demonstration of the generated content received from the server.
[0730] Step 4:
[0731] Users practice and record their sessions. The recorded practice videos are uploaded from the terminal to the server as data. This video is then input into the system and used as material for the next analysis.
[0732] Step 5:
[0733] The server receives the practice video and performs analysis using image analysis techniques. Specifically, it uses OpenCV and TensorFlow to extract features such as gaze, voice pitch, and posture. The input is the user's practice video, and the output is the analyzed feature data.
[0734] Step 6:
[0735] Based on the analysis results, the server generates instructional videos using educational generation tools. Synthesized speech and animation are used in the generation process to clearly and concretely demonstrate areas for improvement for the user. The output is an instructional video designed to promote skill development in the user.
[0736] Step 7:
[0737] The generated instructional video is sent from the server to the terminal. The user's terminal receives and displays the instructional video. The user can watch it and work on improving their skills based on the feedback. The terminal's operation includes playing the video and presenting it to the user.
[0738] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0739] This invention is a system that supports the improvement of users' speech skills, and in particular, recognizes the user's emotions using an emotion engine and provides feedback based on that information. The system mainly consists of a server, a terminal, and the user.
[0740] The user enters prompts related to the topic of their speech or presentation through their device. These prompts are sent from the device to a server. The server analyzes the received prompts, automatically collects relevant information, and generates presentation materials, speech scripts, and demo videos. This process is driven by a generative AI model, providing concrete and visually appealing content. The generated content is then sent to the user via their device and made available for use.
[0741] Users practice their speeches based on the generated content and record the process on video. The recorded practice videos are uploaded to the server via the device. The server analyzes the received practice videos using image and audio analysis technologies, evaluating not only the user's gaze, posture, and voice tone, but also their emotional state using an emotion engine. The emotion engine recognizes emotions from changes in facial expressions, voice tone, pitch, etc., and analyzes the type and intensity of those emotions.
[0742] The analysis results are used to generate instructional videos. The server creates instructional videos that include specific areas for improvement based on the user's expression and emotional state. These instructional videos provide feedback that reflects emotions in an easy-to-understand way for the user, using synthesized speech and animation. For example, if the video detects nervousness during a presentation, it will include suggestions on how to relax. These instructional videos are sent to the user's device, and the user watches them to guide further practice.
[0743] Through the operation of these systems, users can not only improve the quality of their presentations but also learn how to control their emotions and improve their overall public speaking skills.
[0744] The following describes the processing flow.
[0745] Step 1:
[0746] The user enters the presentation topic into the terminal. The entered prompt is digitized by the terminal and sent to the server.
[0747] Step 2:
[0748] The server receives a prompt and uses generation AI to search for relevant data. Based on the collected data, it automatically generates presentation materials, speech scripts, and demo videos. The created content is then sent from the server to the terminal.
[0749] Step 3:
[0750] The terminal provides the user with content received from the server. The user uses the provided materials and script to practice their speech.
[0751] Step 4:
[0752] Users record their practice sessions and save the videos to their devices. The completed practice videos are then uploaded from the device to the server.
[0753] Step 5:
[0754] The server analyzes the uploaded practice videos. Using image and audio analysis technologies, it evaluates the user's gaze, posture, voice tone, and emotions using an emotion engine. The emotion engine identifies the user's emotional state by analyzing facial expressions and changes in voice.
[0755] Step 6:
[0756] The server generates instructional videos based on the analysis results. These instructional videos include specific advice on areas for improvement and methods for controlling emotions. The videos utilize synthesized speech and animation to present information in a visual and auditory way that is easy for users to understand.
[0757] Step 7:
[0758] Instructional videos are sent to the user's device for viewing. Based on the feedback received, the user can use it to improve their performance in future speech practice sessions and actual presentations. By repeating this process, the user can refine their skills.
[0759] (Example 2)
[0760] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0761] A challenge exists in that users cannot effectively improve their emotional expression, visual information, and auditory information during speeches and presentations. Traditional methods limit opportunities for users to receive feedback and make self-assessment difficult, which can hinder skill improvement.
[0762] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0763] In this invention, the server includes means for receiving theme-based input information from the user, means for automatically generating material data and visual information using information processing technology, means for transferring the generated data to a terminal via communication, means for analyzing activity videos and evaluating visual and audio information, means for generating instructional information based on the analysis results using speech synthesis technology and visual technology, and means for providing the instructional information to the user visually and audibly. This enables the user to objectively improve their speech skills and emotional expression and enhance their abilities.
[0764] "Input information" refers to information provided by the user based on a given theme, and serves as the foundational data for the system to begin its analysis and generation processes.
[0765] "Information processing technology" refers to all technologies used to analyze received data and automatically generate document data and visual information.
[0766] "Document data" refers to automatically generated content for presentations and speeches, containing information that users can refer to.
[0767] "Visual information" refers to information that appeals to users visually, including demo videos and charts.
[0768] "Means of transfer via communication" refers to the infrastructure and protocols used to transmit generated data and information to a terminal.
[0769] "Activity videos" are videos recorded by users for speech practice, and their content is analyzed and used as material for evaluation.
[0770] "Means for evaluating visual and auditory information" refers to technologies for analyzing the content of activity videos and quantitatively and qualitatively evaluating the user's visual and auditory elements.
[0771] "Speech synthesis technology" is a technology for generating synthesized sounds that mimic human voices, and is used for voice feedback in instructional materials.
[0772] "Visual technology" refers to technologies that generate visual feedback and provide it to users in an easily understandable format.
[0773] "Instructional information" refers to feedback content generated based on analysis results, which includes suggestions and areas for improvement to enhance the user's skills.
[0774] This invention is an information processing system for supporting the improvement of speech skills. This system consists of three elements: a user, a terminal, and a server.
[0775] The user enters prompt text into the terminal based on the theme. For example, a prompt might be, "Please describe the elements of effective leadership." The terminal receives the user's input and sends it to the server via the internet.
[0776] The server utilizes a generative AI model to automatically generate document data and visual information based on this prompt. This process employs natural language processing techniques and machine learning algorithms. The materials include demo videos and presentation slides to visually represent specific information. The generated materials and data are returned to the terminal via a secure communication protocol.
[0777] The terminal displays data received from the server to the user. The user uses this data to practice their speech and records the process as an activity video. Recording is typically done using the terminal's built-in camera function.
[0778] Afterward, the user uploads the recorded activity video from their device to the server. The server uses image and audio analysis technologies to analyze this video and evaluate the user's visual and audio information. In addition to elements such as gaze, posture, and voice tone, an emotion analysis engine is also introduced to evaluate emotional expression from facial expressions.
[0779] Based on the analysis results, the server generates instructional information using speech synthesis and visual technologies. This instructional information is provided as feedback content, including specific areas for improvement and advice, with the aim of improving the user's skills. For example, if the server determines that the user is stressed, it will provide suggestions for relaxation techniques.
[0780] All data is transmitted to the device, and users can effectively improve their speech skills by viewing or using it. This system allows users to improve their skills and the quality of their presentations.
[0781] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0782] Step 1:
[0783] The user enters prompt sentences related to the topic of their speech or presentation into the terminal. These prompt sentences become data that instructs the system to analyze. After input, the terminal converts this data into a digital format and sends it to the server.
[0784] Step 2:
[0785] The server parses the received prompt message. This parsing uses a generative AI model, leveraging natural language processing techniques to understand the intent of the prompt. The AI model collects necessary information from a database to generate relevant data and visual information, creating presentation slides and demo videos. The output of this process is content in a user-friendly format.
[0786] Step 3:
[0787] The server sends the generated material data and visual information to the terminal. The terminal receives this and presents it to the user through an appropriate interface. The user practices their speech using the provided materials. Reviewing these materials helps improve understanding of the presentation content and self-expression.
[0788] Step 4:
[0789] The user records their speech practice using the device's camera function. The recorded content includes data on the user's visual and auditory expression. After recording, the video is converted to an appropriate format on the device and uploaded to the server.
[0790] Step 5:
[0791] The server receives uploaded practice videos and analyzes them using image and audio analysis technologies. This process evaluates data related to the user's gaze, posture, and voice tone. Additionally, an emotion analysis engine is used to extract emotional information from the user's facial expressions and voice. The output after analysis is evaluation information that can be used for instructional feedback.
[0792] Step 6:
[0793] The server generates instructional feedback based on the analysis results. Utilizing a generation AI model, it creates videos containing specific advice for the user using speech synthesis and visual technologies. For example, this feedback might include suggestions for relaxation techniques or tips for improving expressiveness.
[0794] Step 7:
[0795] The server sends instructional feedback to the device. The device receives this feedback and provides it to the user in a viewable format. Through this feedback video, the user can objectively evaluate their own performance and use it to further improve their skills.
[0796] (Application Example 2)
[0797] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0798] To improve the efficiency and safety of workers on site, a system is needed that can appropriately understand workers' emotional states and stress levels and provide feedback at the appropriate time. However, conventional technology has made it difficult to recognize workers' emotional states in real time and respond to changes in them. There is a need to provide a system that solves these problems and realizes a better working environment.
[0799] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0800] In this invention, the server includes an input device that receives prompts from the user, a generation device that automatically generates visual materials, audio scripts, and demo sound images based on the prompts, and a communication device that transmits the generated information to a terminal. This allows the worker's emotional state to be recognized and to receive feedback as needed.
[0801] A "prompt" is an instruction or request provided by the user through input.
[0802] An "input device" refers to a device used to receive data and instructions from a user.
[0803] A "generation device" is a device that has the function of processing information based on received instructions and creating the necessary content.
[0804] A "communication device" is a device used to transmit generated information to other devices or networks.
[0805] "Reference videos" refer to video content used to analyze information uploaded by users.
[0806] An "analysis device" is a device used to analyze the characteristics of an object based on collected data.
[0807] A "support structure" refers to a system that provides necessary responses and feedback based on the analysis results.
[0808] A "presentation device" is a device that provides generated feedback and instructional content to the user visually or audibly.
[0809] An "improvement suggestion" is a measure to improve the work environment or methods suggested by the system, based on the analysis results.
[0810] In a system for realizing this invention, the server receives prompts from the user via an input device and uses a generation AI model to create visual materials, audio scripts, and demo sound images using a generation device. This generated content is transmitted to a terminal via a communication device. The user operating the terminal can then begin actual work using the received content.
[0811] Users create reference videos during their work and upload them to the server via their terminal. An analysis device on the server analyzes the uploaded videos, evaluating the user's gaze, voice characteristics, posture, and emotional state. This analysis result is provided as feedback through a support structure that includes suggestions for necessary improvements and work environment enhancements.
[0812] As a concrete example, when a robot assists an operator in a factory, it can sense the operator's stress and suggest improvements such as, "Let's pause this task, have some water, and relax." An example of a prompt used in this process would be, "To improve work efficiency in the factory, the robot will recognize the operator's emotions and suggest providing appropriate feedback." This makes it possible to improve the work environment and enhance safety.
[0813] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0814] Step 1:
[0815] The user enters prompts using a terminal. The input device receives these prompts and prepares to send them to the generative AI model. Here, the input is the user's requests or instructions, and the output is the data sent to the generative AI model.
[0816] Step 2:
[0817] The server uses a generation AI model to generate visual materials, audio scripts, and demo sound images based on the received prompts. The generation device automatically collects relevant information based on the prompts and performs the process of generating content. The data processing performed here is information generation based on the prompts, and the output is the generated content.
[0818] Step 3:
[0819] The generated content is transmitted from the server to the terminal via a communication device. The server's role is to send and receive content; the input is the generated content, and the output is the data sent to the terminal.
[0820] Step 4:
[0821] The user performs tasks and exercises based on content received through the device. The actions in this step represent actual practice and work using the content, and indicate a specific feedback point. User evaluation audio and video data are obtained as output.
[0822] Step 5:
[0823] The user records a video of their work in progress and uploads it to the server via their device. The uploaded video becomes the input data. At this stage, the server prepares the received data for the next analysis step.
[0824] Step 6:
[0825] The server uses an analysis device to analyze uploaded videos and evaluate the user's gaze, posture, voice tone, and emotional state. The input is video data, and the output is the analyzed evaluation information. The process involves both video and audio analysis.
[0826] Step 7:
[0827] Based on the analysis results, the server generates necessary improvements and suggestions for improving the work environment through the support structure. The output is returned to the terminal as feedback information. The server utilizes a generation AI model to create specific improvement suggestions.
[0828] Step 8:
[0829] Users utilize feedback received on their devices to further improve their work and skills. This feedback includes specific suggestions tailored to the user's emotional state and serves as guidance for future work. The input is feedback information, and the output is the improved user behavior and environment.
[0830] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0831] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0832] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[0833] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0834] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0835] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0836] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0837] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0838] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0839] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0840] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[0841] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[0842] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0843] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0844] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0845] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0846] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0847] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0848] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0849] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0850] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[0851] The following is further disclosed regarding the embodiments described above.
[0852] (Claim 1)
[0853] An input means that accepts prompt input from the user,
[0854] A generation means that automatically generates presentation materials, speech scripts, and demo videos based on the prompt,
[0855] A transmission means for sending the generated content to the terminal,
[0856] An analysis method that analyzes practice videos uploaded by users and evaluates eye gaze, voice tone, posture, etc.
[0857] Instructional generation means for generating instructional videos based on analysis results,
[0858] A means of providing instructional videos to users,
[0859] A system that includes this.
[0860] (Claim 2)
[0861] The system according to claim 1, comprising means for visualizing and audibly presenting areas for improvement to the user using synthesized speech and animation when generating instructional videos.
[0862] (Claim 3)
[0863] The system according to claim 1, further comprising means for evaluating the richness of the user's emotional expression and interaction with the audience in the analysis of practice videos.
[0864] "Example 1"
[0865] (Claim 1)
[0866] A device that accepts information input from users, including themes and content,
[0867] A device that automatically creates documents, texts, and images based on said information, comprising a device that collects relevant information from an information source,
[0868] A device that transmits the created data to the user's terminal,
[0869] A device that processes exercise records uploaded by users and evaluates eye movements, vocal characteristics, posture, etc.
[0870] A device that generates educational videos based on processing results,
[0871] A device that provides educational videos to users,
[0872] A system that includes this.
[0873] (Claim 2)
[0874] The system according to claim 1, comprising a device that uses synthesized speech and visual effects to visualize and audibly present areas for improvement to the user when generating educational videos.
[0875] (Claim 3)
[0876] The system according to claim 1, further comprising a device for evaluating the diversity of user emotional expression and interaction with the audience in processing practice records.
[0877] "Application Example 1"
[0878] (Claim 1)
[0879] An information input means for receiving instructions from users regarding the theme,
[0880] An automatic generation means that generates explanatory materials, speech scripts, and action videos based on the said instructions,
[0881] Information transmission means for transmitting the generated content to a device,
[0882] An information analysis means that analyzes exercise videos submitted by users and evaluates gaze direction, voice pitch, and posture,
[0883] An educational generation means for generating instructional videos based on analysis results,
[0884] A display means for presenting instructional videos to the user,
[0885] Training methods for providing instruction aimed at improving customer service skills in sales settings,
[0886] A system that includes this.
[0887] (Claim 2)
[0888] The system according to claim 1, comprising means for displaying and allowing the user to hear improvements using synthesized speech and motion video when generating instructional videos.
[0889] (Claim 3)
[0890] The system according to claim 1, comprising means for evaluating the richness of the user's emotional expression and the interaction with the recipient in the analysis of the training video.
[0891] "Example 2 of combining an emotion engine"
[0892] (Claim 1)
[0893] A means of receiving theme-based input information from users,
[0894] A means for automatically generating data and visual information using information processing technology based on the input information,
[0895] A means of transferring the generated data to the terminal via communication,
[0896] A means for analyzing activity videos provided by users and evaluating visual and audio information,
[0897] A means for generating instructional information based on analysis results using speech synthesis technology and visual technology,
[0898] Means for providing instructional information to users visually and audibly,
[0899] A system that includes this.
[0900] (Claim 2)
[0901] The system according to claim 1, comprising means for presenting points for improvement to the user using speech synthesis technology and video technology when generating instructional information.
[0902] (Claim 3)
[0903] The system according to claim 1, comprising means for evaluating the diversity of user emotional expression and interactive operations in the analysis of activity videos.
[0904] "Application example 2 when combining with an emotional engine"
[0905] (Claim 1)
[0906] An input device that accepts prompts from the user,
[0907] A generation device that automatically generates visual materials, audio scripts, and demo sound images based on the prompt,
[0908] A communication device that transmits the generated information to a terminal,
[0909] An analysis device that analyzes user-uploaded video materials and evaluates eye gaze, voice characteristics, posture, etc.
[0910] Based on the analysis results, a support structure is provided to generate instructional videos, recognize the operator's emotional state, and provide necessary feedback.
[0911] A display device that provides instructional videos to operators,
[0912] A device that makes suggestions for improving the work environment according to the emotional state of the operator,
[0913] A system that includes this.
[0914] (Claim 2)
[0915] The system according to claim 1, comprising a device that visualizes and audibly presents areas for improvement to the operator using synthesized speech and animation, and provides feedback based on the operator's emotional state.
[0916] (Claim 3)
[0917] The system according to claim 1, further comprising a device for evaluating the diversity of the operator's emotional expression and their interaction with the work environment in the analysis of reference videos. [Explanation of Symbols]
[0918] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. An input means that accepts prompt input from the user, A generation means that automatically generates presentation materials, speech scripts, and demo videos based on the prompt, A transmission means for sending the generated content to the terminal, An analysis method that analyzes practice videos uploaded by users and evaluates eye gaze, voice tone, posture, etc. Instructional generation means for generating instructional videos based on analysis results, A means of providing instructional videos to users, A system that includes this.
2. The system according to claim 1, comprising means for visualizing and audibly presenting areas for improvement to the user using synthesized speech and animation when generating instructional videos.
3. The system according to claim 1, further comprising means for evaluating the richness of the user's emotional expression and interaction with the audience in the analysis of practice videos.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A