system
The system addresses the challenge of inadequate real-time instructions by using a generative model to provide actors with immediate, scenario-appropriate text and feedback, improving performance quality and training efficiency.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-16
- Publication Date
- 2026-04-28
AI Technical Summary
Performers face challenges in responding to improvisational scenes and unexpected scenario changes during performances, leading to a decline in performance quality due to inadequate real-time instructions and inefficient training processes.
A system utilizing a generative model provides actors with real-time text and instructions, collects performance data for feedback, and generates customized training programs based on past data to improve performance efficiency.
The system enhances performance quality by offering immediate, scenario-appropriate instructions and feedback, allowing actors to improve their skills systematically and efficiently.
Smart Images

Figure 2026070889000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, the method including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance as a response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] Performers, idols, and entertainers often face various challenges and troubles during performances. In conventional systems, it has been difficult to respond to such situations in real time, increasing the burden on the performers themselves. In particular, in improvisational scenes and unexpected scenario changes, appropriate instructions have not been obtained, often resulting in a decline in the quality of the performance. Also, although prior training and repeated practice are required, there has been a problem that it takes a lot of time and resources. Therefore, there is a need for a system that allows performers to improve their performances more efficiently in a shorter period of time.
Means for Solving the Problems
[0005] This invention provides a system that utilizes a generative model to provide actors with appropriate, real-time text and instructions. The generative model is operated based on user requests and helps actors adapt to unexpected situations. It can also collect actor performance data and generate feedback for improving future performance. Furthermore, it improves the efficiency of actor training by providing training programs to actors via a terminal and generating customized instructions based on past data. In this way, actors can achieve high-quality performance in a short period of time.
[0006] A "generative model" is a technology or system that automatically generates text or instructions in response to user requests or circumstances.
[0007] The term "performer" refers to an individual or group that engages in entertainment activities and provides performances for an audience.
[0008] "Real-time" refers to a time frame in which processing and responses are performed immediately without delay.
[0009] A "terminal" is an electronic device or equipment used by a user or performer to receive or display information.
[0010] "Feedback" refers to information provided for measuring effectiveness and making improvements based on the performance and results of an implemented task.
[0011] A "training program" is a set of practice and training sessions designed to help performers improve their skills and abilities.
[0012] "Customization" refers to a state that has been modified or adjusted to meet specific needs or conditions. [Brief explanation of the drawing]
[0013] [Figure 1]This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2, which incorporates an emotion engine. [Figure 14] This is a sequence diagram showing the processing flow of the data processing system in Application Example 2, which combines an emotion engine. [Modes for carrying out the invention]
[0014] Hereinafter, an example of an embodiment of the system relating to the technology of this disclosure will be described with reference to the attached drawings.
[0015] First, the terms used in the following description will be explained.
[0016] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0017] In the following embodiments, the numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0018] In the following embodiments, the numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, and the like.
[0019] In the following embodiments, the numbered communication I / F (Interface) is an interface including a communication processor and an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark), and the like.
[0020] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0021] [First Embodiment]
[0022] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0023] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0024] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0025] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0026] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0027] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0028] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0029] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0030] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0031] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0032] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0033] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0034] This invention is a system for providing actors with appropriate instructions in real time to improve the quality of their performance. This system utilizes a generative model to automatically generate situation-appropriate text and instructions, and provides feedback to actors, thereby enabling rapid skill improvement.
[0035] The server functions as the central management unit and hosts the generative model. The server operates the generative model in response to user requests, generating instructions and scenarios. It also collects past performance data from performers and uses this data to generate feedback for improving future performance.
[0036] As a prompt engineer, the user accesses the server and operates the system. The user requests necessary instructions from the server according to specific performances or events, and reviews and adjusts the generated information. Furthermore, the user reviews feedback from performers and optimizes the system settings for the next performance.
[0037] The terminal is a device or screen held by the performer, displaying instructions sent from the server in real time. Through this terminal, the performer can obtain necessary information and smoothly proceed with their performance. The terminal also displays a training program, supporting the performer in systematically improving their skills.
[0038] As a concrete example, in a scene where an actor is appearing on a music program, the user requests the generative model via the server to develop the scenario appropriately. The server immediately generates a script and displays it on the terminal. Based on this information, the actor can perform according to the script and respond quickly to any unexpected problems.
[0039] In this way, the present invention can significantly improve the quality of performers' performances and also contribute to the efficiency of entertainment production.
[0040] The following describes the processing flow.
[0041] Step 1:
[0042] Users input information about performances and events into the server and request the generation of necessary scenarios and instructions. Users provide information and specify specific conditions and requirements through the system interface.
[0043] Step 2:
[0044] The server receives requests from users and operates a generative model based on them. Scenarios and instructions are automatically generated and customized to match the user's requirements. This information is processed immediately, and instructions are generated.
[0045] Step 3:
[0046] The server sends the generated instructions to the terminal. The instructions are organized in a pre-specified order as information necessary for the performance to progress. The transmitted information is directed to the terminal used by the performer.
[0047] Step 4:
[0048] The terminal displays instructions received from the server to the performer in real time. The terminal visualizes the information and is designed to be intuitively easy for the performer to understand. Based on this information, the performer can decide on their actions in the actual performance.
[0049] Step 5:
[0050] Users can request further instructions from the server during performance or if an anomaly occurs. They can also modify system settings to prompt the server to generate more appropriate instructions.
[0051] Step 6:
[0052] The server analyzes the data collected after the performance and generates feedback to identify areas for improvement in the performer. This information is then used in training for the next performance.
[0053] Step 7:
[0054] The device displays the training program and provides the performer with a plan to improve their skills based on feedback. The performer can use this information to proceed with their training systematically.
[0055] (Example 1)
[0056] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0057] For performers to deliver high performances in stage and video content, they need to be provided with appropriate instructions and scenarios in real time. However, current systems often fail to adequately adapt the generated instructions to the performers' past performance data and real-time circumstances, resulting in a decline in performance quality. Furthermore, there is a lack of adequate mechanisms for performers to improve their next performance based on feedback.
[0058] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0059] In this invention, the server includes means for providing appropriate information to the performer using a generative AI model, means for analyzing the performer's past performance data and selecting a generative AI model, and means for collecting the performer's feedback and generating suggestions for improving the next performance. This ensures that the instructions the performer receives in real time are appropriately adjusted based on their individual performance history, and that effective feedback is provided to help them improve their next performance.
[0060] A "generative AI model" is an algorithm that uses machine learning techniques to automatically generate content and instructions that are appropriate for a given purpose from natural language.
[0061] "Past performance data" refers to records of performances an actor has given in the past, including the quality of the performance, audience reactions, and feedback.
[0062] "Feedback" refers to evaluations of a performer's performance and suggestions for improvement, and is important information used to improve future performances.
[0063] "Real-time" refers to the fact that information is processed and provided to the performer immediately at the moment it is requested, minimizing time lag and achieving immediacy.
[0064] "Display means" are devices or mechanisms used by performers to directly receive information, and they play a role in providing detailed instructions and feedback visually.
[0065] The server is the central component of this system, responsible for generating instructions and scenarios to be provided to the performers using a generative AI model. Specifically, the server utilizes machine learning libraries and runs the generative AI model using historical data related to the performers' performance. The generative AI model is based on natural language processing algorithms and generates optimal text information for the performers based on user prompts. For this to work, the server requires a computing system equipped with high-performance CPUs and GPUs.
[0066] The user acts as a prompt engineer, operating the system and inputting specific prompt messages to the server. Specifically, they input prompts related to performance themes or specific scenes. For example, they might send a prompt message to the server such as, "Generate a scenario for the next music performance that includes countermeasures for anticipated problems." Based on this input, the server uses a generative AI model to generate a scenario, which is then presented to the user for review.
[0067] The terminal is a device held by the performer, and it has the function of displaying instructions and scenarios generated by the server in real time. Through this terminal, the performer can instantly receive detailed instructions and feedback and incorporate them into their performance. For example, they can check the instructions for the next action on the terminal on stage and proceed with the performance smoothly. In order for the terminal to work, a mobile device or tablet PC with display capabilities is used. This allows performers to always perform based on the latest information.
[0068] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0069] Step 1:
[0070] The user inputs specific prompts into the server to generate instructions and scenarios necessary for the actors. For example, this might include instructions for acting in a specific scene or solutions for problems that may arise. This input includes the acting theme and specific requirements.
[0071] Step 2:
[0072] The server parses the prompt text received from the user. The parsed prompt is input into the generating AI model. Based on this input data, the server performs data calculations to generate appropriate instructions and scenarios from the model. As output, text data to be provided to the performer is derived.
[0073] Step 3:
[0074] The server temporarily stores the generated text data and notifies the user to check if the output is suitable for the actor. The user reviews it and makes adjustments as needed, such as adding parts of the scenario or adjusting the lines.
[0075] Step 4:
[0076] The server sends the final instructions to the terminal after confirmation. The terminal immediately displays the information received from the server, allowing the performer to refer to it in real time. This step utilizes high-speed data communication and the terminal's display capabilities.
[0077] Step 5:
[0078] The performer executes their performance based on the information displayed on the terminal. After the performance, they input feedback on areas for improvement and requests for the next performance into the terminal and send it to the server. This feedback data will be used to improve future performances.
[0079] (Application Example 1)
[0080] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0081] In conventional production facilities, workers often lack sufficient support to efficiently perform complex procedures, which contributes to decreased productivity and errors. There is a need to improve this situation and enable workers to receive appropriate instructions in real time, thereby increasing work efficiency and accuracy.
[0082] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0083] In this invention, the server includes means for providing appropriate information and instructions to an operator using a generative model, means for presenting information via an information processing device to support the operator's work in real time, and means for manipulating the generative model and generating instructions based on requests from the user. This enables the provision of immediate and effective work instructions in the production field.
[0084] A "generative model" is an algorithm that learns from past data to generate new information or instructions.
[0085] "Workers" refer to personnel or automated equipment that perform tasks in a factory or production facility.
[0086] An "information processing device" is an electronic device that receives and processes data to present appropriate information to the user.
[0087] "Presentation means" refers to a method or device for visually or audibly transmitting instructions or information generated by an information processing device to a worker.
[0088] "User" refers to a person or organization that can manipulate the generative model and manage the entire system.
[0089] "Feedback" refers to the process or information itself that returns to the system about the results of the work performed by an employee.
[0090] The system that realizes this invention utilizes a generative model server on the cloud to provide immediate and effective instructions to field workers. The server hosts generative AI models and analyzes past work data to generate information useful for future work. As a specific example of use, in a particular assembly task, the server generates the optimal procedure based on past assembly process data and provides it to the worker.
[0091] The terminal refers to an information processing device such as a smartphone or tablet held by the worker, which receives instructions from the server and displays them to the worker in real time. The terminal has an application built with React Native installed, and the HTTP protocol is used for communication with the server. Specifically, when a worker is assembling a product on the assembly line, the terminal has functions such as displaying a screen that visually confirms the next task to be performed.
[0092] Users are provided with an interface for manipulating the generated AI model. This allows for overall system management and adjustment, and improvement of next instructions based on worker feedback. Users can input prompts such as, "Instruct me on the most efficient way to join part A to part B in the next product assembly step," into the generated model, and generate new work procedures.
[0093] This system configuration allows workers to efficiently perform complex procedures, and is expected to improve both productivity and work accuracy. As described above, this embodiment of the invention provides a new method for supporting work in manufacturing sites.
[0094] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0095] Step 1:
[0096] The server retrieves past work data from the database. It receives past performance logs and work environment information of the worker as input and prepares to run the generated AI model based on this data. This data is analyzed to extract work efficiency and patterns.
[0097] Step 2:
[0098] The user inputs prompt statements to the generative model. For example, the user might send an instruction to the generative AI model such as, "Instruct me on the most efficient way to join part A to part B in the next product assembly step." Based on this prompt, the generative AI model executes and outputs the optimal work procedure.
[0099] Step 3:
[0100] The server receives the optimal work procedure as the output of the generated AI model. This output is specifically represented as text or a series of steps and processed for presentation to the worker. The processed data is later transferred to the terminal.
[0101] Step 4:
[0102] The terminal presents the worker with the optimal work procedure received from the server. The application installed on the terminal displays clear, sequential information to the worker in real time. Here, the work procedure is presented as a visualized diagram or a list of steps.
[0103] Step 5:
[0104] The worker performs the actual work based on the provided work procedures. Feedback obtained during the work and any newly generated data are sent to the server via the terminal.
[0105] Step 6:
[0106] The server records feedback received from workers and newly collected work data in a database. This data is used to improve future work instructions and adjust the generative model, thereby improving the overall system performance.
[0107] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0108] This invention provides a system that combines a generative model and an emotion engine to provide actors with real-time, appropriate, and emotion-based instructions. The system consists of a server, terminals, and users, and can further improve performance quality by utilizing user emotional feedback.
[0109] The server forms the core of the emotion engine and generative model. The emotion engine analyzes the user's emotions and provides that data to the generative model. Based on the emotion data, the generative model personalizes instructions and scenarios for the performers. This generates appropriate text that matches the user's current mental state.
[0110] The user provides emotional feedback to the server. The user's emotional state is continuously monitored and collected by the server through a dedicated emotion recognition device. The user can operate the system and update instructions as needed.
[0111] The terminal displays instructions and information sent from the server to the performer in real time. The terminal provides the performer with instructions adapted based on their emotions, helping them to deliver their best performance. Through this terminal, the performer can immediately receive information and take action to improve.
[0112] For example, if an actor is nervous on stage, the server senses the tension through its emotion engine, and a generative model generates instructions to promote relaxation. The terminal displays these instructions to the actor, who can then immediately take appropriate action. This process leads to optimal performance results.
[0113] Thus, the present invention enables the dynamic generation of instructions using emotional feedback and provides an excellent system for supporting performers in maximizing their abilities during performance.
[0114] The following describes the processing flow.
[0115] Step 1:
[0116] The user logs into the system and enters an outline of the scenario and instructions required for the event or performance into the server. They then put on an emotion recognition device and prepare to begin emotional feedback.
[0117] Step 2:
[0118] The server receives information provided by the user and inputs it into the generative model. At this point, the emotion engine begins collecting and analyzing real-time emotion data from the user.
[0119] Step 3:
[0120] The server passes emotional data obtained from the emotion engine to the generative model. Based on this information, the generative model generates appropriate instructions and scenarios in real time, according to the user's emotional state.
[0121] Step 4:
[0122] The terminal instantly displays instructions and scenarios sent from the server to the actor. The displayed content is adapted based on the user's emotions, allowing the actor to act most effectively.
[0123] Step 5:
[0124] Users provide real-time feedback to the server via the emotion engine to give additional instructions or corrections as needed during their performance. This allows the instructions to be adjusted accordingly.
[0125] Step 6:
[0126] The server analyzes user feedback and sentiment data to generate improvement suggestions for future performances and training. This data is then used to prepare subsequent performances.
[0127] Step 7:
[0128] The device displays generated improvement suggestions and training programs to the performer. Based on this, the performer can systematically improve their skills.
[0129] (Example 2)
[0130] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0131] In contemporary performance activities, it is difficult for individual performers to receive real-time, adaptable instructions. Furthermore, the lack of instructions that take into account the performer's emotional state makes it difficult for performances to consistently produce optimal results. Additionally, current practices do not fully utilize performers' past data, limiting the provision of dynamic, emotion-based feedback.
[0132] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0133] In this invention, the server includes means for providing emotion-based text and instructions to the performer using a generation module, means for acquiring the user's emotion data using an emotion recognition device and transmitting it to the server, and means for analyzing the emotion data using an emotion engine and identifying the emotional state. As a result, the performer can receive appropriate emotion-based instructions in real time and improve performance by utilizing past data.
[0134] A "generating module" is a program that automatically generates instructions and text for the performer based on the user's emotional state.
[0135] A "performer" is an individual or entity that performs based on emotionally driven instructions.
[0136] An "emotion recognition device" is a device that monitors the user's facial expressions, voice, heart rate, etc., to acquire emotional data.
[0137] A "server" is a central computing unit that receives emotional data, analyzes it, and then generates and transmits instructions from the generation module.
[0138] The "emotion engine" is a feature equipped with an algorithm that analyzes emotional data acquired from users and identifies their emotional state in real time.
[0139] "Text and instructions" refer to specific guidance and information created by the generative module to guide the performer's performance.
[0140] A "display device" is a terminal or device used to present generated instructions to the user in real time, either visually or audibly.
[0141] "Feedback" is information collected based on the performer's performance and emotional state, and is used to improve their performance in the future.
[0142] This invention provides a system for providing emotion-based instructions to an expresser in real time. This system mainly consists of a server, a user, and a terminal.
[0143] The server plays a central role in the system. The server receives emotional data sent by the user and analyzes it using an emotion engine. The emotion engine incorporates algorithms that identify emotional states based on facial expressions, vocal characteristics, and physical changes such as heart rate. Based on the identified emotional state, a generation module within the server automatically generates text and instructions appropriate to that emotion. This generation module can use a generation AI model to generate optimal instructions from prompt sentences. For example, in response to a prompt sentence such as "Generate specific instructions to encourage relaxation when the user is nervous on stage," it will generate instructions related to relaxation.
[0144] The user uses a dedicated emotion recognition device to send their emotional data to a server in real time. The emotion recognition device acquires data such as the user's facial expressions, voice, and heart rate via sensors and transmits this data to the server using wireless technology.
[0145] The terminal is a device that receives instructions sent from the server and displays the information necessary for the performer. Specifically, generated text and instructions are displayed on the terminal's screen, and in some cases, the instructions are read aloud via voice assistance. This terminal allows the performer to receive real-time feedback during their performance and adjust their performance on the spot.
[0146] In this way, the entire system works in conjunction to support performers in achieving their optimal performance.
[0147] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0148] Step 1:
[0149] The user collects their emotional data using an emotion recognition device. Inputs include facial expression data, voice information, and biometric information such as heart rate. This emotional data is immediately transmitted to a server via Bluetooth or Wi-Fi. Outputs are prepared for transfer to the server.
[0150] Step 2:
[0151] The server inputs the received emotional data into the emotion engine. This engine uses AI algorithms to analyze the data and identify the user's emotional state. Based on the emotional data as input, a machine learning model determines emotional categories such as anxiety, joy, and tension. As output, labels for the identified emotional states are generated.
[0152] Step 3:
[0153] The server passes a prompt for instruction generation to the generation module based on the already analyzed emotional state. A prompt sentence corresponding to the emotional state is prepared as input. This prompt is analyzed by the generation AI model and used as data necessary to form appropriate instructions. As output, specific instructions corresponding to the emotional state are created.
[0154] Step 4:
[0155] The server transfers the generated instructions to the terminal. The generated instruction data is received from the server as input. The instructions are sent to the terminal as output and displayed in real time on the performer's display device.
[0156] Step 5:
[0157] The device presents the received instructions to the user visually or audibly. For example, it might display "Take a deep breath and relax" on the screen, or read the instructions aloud using its voice assistance function. Based on the instruction data as input, the device generates display and audio output, and this information is delivered to the user as output.
[0158] This sequence of events allows the system to utilize the user's emotional state to provide immediate instructions and optimize the performer's performance.
[0159] (Application Example 2)
[0160] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0161] Traditional customer service in retail stores often involves providing a uniform service, making it difficult to offer appropriate customer service tailored to the customer's emotions. In particular, detecting changes in emotions in real time and dynamically adjusting responses accordingly is a significant burden for employees and poses an obstacle to providing optimal customer service.
[0162] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0163] In this invention, the server includes means for providing instructions via a display device to support the activities of role performers in real time using a generative model and an emotion engine; means for detecting the user's emotions using an emotion recognition device and dynamically generating instructions based on the results; and means for displaying guidance to optimize the actions of employees in the store based on emotional information. This enables detailed customer service that responds to the user's emotions.
[0164] A "generative model" is an artificial intelligence system that generates new content or instructions based on input data.
[0165] An "emotion engine" is a program or device that analyzes a user's emotions from their facial expressions, voice, etc., and processes the results.
[0166] A "role performer" refers to a person who fulfills a predetermined role in a specific place or situation, and may include employees who primarily perform customer service duties.
[0167] A "display device" is hardware used to visualize digital information, and includes devices such as smart glasses and smartphones.
[0168] A "means of providing instructions" is a system for delivering direct instructions or guidance to role performers in accordance with specific conditions or circumstances.
[0169] An "emotion recognition device" is a device that detects a user's emotional state from their facial expressions, voice, and other physiological responses, and transmits that information as data.
[0170] "Employee actions within the store" refers to the specific actions taken by employees in the store to perform tasks and provide services.
[0171] "A means of displaying guidance for optimization based on emotional information" refers to a system that visually provides guidance and instructions to adjust the work content of employees according to the user's emotions.
[0172] This invention is a system that improves the quality of service within a store through close cooperation between the server, terminal, and user.
[0173] The server first receives image and audio data from the emotion recognition device and analyzes the user's emotions in real time using an emotion engine. The analysis results are sent to a generative model, which generates appropriate instructions and information for the role performer. The generated instructions are immediately sent to the terminal.
[0174] The terminal visually displays personalized instructions sent from the server through a display device such as smart glasses or a mobile device worn by the role performer. This allows the role performer to respond quickly in accordance with the user's emotional state.
[0175] Users not only have emotional information collected by the emotion recognition device, but they can also provide feedback as needed. This feedback will be used to improve the service in the future.
[0176] As a concrete example, there is a scenario in which a cafe employee uses smart glasses to recognize the customer's emotions and, based on instructions from the emotion engine, suggests the most suitable drink to create a calm atmosphere. When the prompt "Please suggest a drink menu that can be served to the customer as quickly as possible to create a relaxed atmosphere" is input into the generative model, an appropriate menu is presented.
[0177] This system allows store employees to provide customized services tailored to the customer's emotions, thereby improving customer satisfaction.
[0178] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0179] Step 1:
[0180] The server receives image and audio data from the emotion recognition device. This image and audio data is used as input and is preprocessed to extract the features necessary for emotion recognition. The extracted features then become the input for estimating the emotional state in the next step.
[0181] Step 2:
[0182] The server uses an emotion engine to analyze the user's emotions based on pre-processed features. In this step, facial expressions and voice are analyzed based on the input features to estimate the emotional state (e.g., happiness, anger, sadness, etc.). The estimated emotional state is then output as input data for a generative AI model.
[0183] Step 3:
[0184] The server transfers emotional data obtained from the emotion engine to a generative AI model, which then generates appropriate instructions and information based on the prompt text. The input consists of emotional data and pre-prepared prompt texts, and the generative model generates theme-appropriate text and instructions based on this data. The generated instructions are stored on the server as output for display on the terminal.
[0185] Step 4:
[0186] Instructions generated by a generative model are sent from the server to the terminal. The terminal, such as smart glasses, displays personalized instructions based on the user's information. The input is instruction data sent from the server, which is converted into a user-friendly format and displayed. This allows role performers to respond appropriately to the user in real time.
[0187] Step 5:
[0188] Users provide feedback to the server via their terminal as needed. Feedback input includes verbal or selective comments from role performers. This feedback is collected by the server and used as data to improve future responses. The output is the updated feedback database.
[0189] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0190] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include those described above. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions shown by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0191] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0192] [Second Embodiment]
[0193] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0194] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0195] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0196] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0197] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0198] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0199] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0200] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0201] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0202] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0203] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0204] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0205] This invention is a system for providing actors with appropriate instructions in real time to improve the quality of their performance. This system utilizes a generative model to automatically generate situation-appropriate text and instructions, and provides feedback to actors, thereby enabling rapid skill improvement.
[0206] The server functions as the central management unit and hosts the generative model. The server operates the generative model in response to user requests, generating instructions and scenarios. It also collects past performance data from performers and uses this data to generate feedback for improving future performance.
[0207] As a prompt engineer, the user accesses the server and operates the system. The user requests necessary instructions from the server according to specific performances or events, and reviews and adjusts the generated information. Furthermore, the user reviews feedback from performers and optimizes the system settings for the next performance.
[0208] The terminal is a device or screen held by the performer, displaying instructions sent from the server in real time. Through this terminal, the performer can obtain necessary information and smoothly proceed with their performance. The terminal also displays a training program, supporting the performer in systematically improving their skills.
[0209] As a concrete example, in a scene where an actor is appearing on a music program, the user requests the generative model via the server to develop the scenario appropriately. The server immediately generates a script and displays it on the terminal. Based on this information, the actor can perform according to the script and respond quickly to any unexpected problems.
[0210] In this way, the present invention can significantly improve the quality of performers' performances and also contribute to the efficiency of entertainment production.
[0211] The following describes the processing flow.
[0212] Step 1:
[0213] Users input information about performances and events into the server and request the generation of necessary scenarios and instructions. Users provide information and specify specific conditions and requirements through the system interface.
[0214] Step 2:
[0215] The server receives requests from users and operates a generative model based on them. Scenarios and instructions are automatically generated and customized to match the user's requirements. This information is processed immediately, and instructions are generated.
[0216] Step 3:
[0217] The server sends the generated instructions to the terminal. The instructions are organized in a pre-specified order as information necessary for the performance to progress. The transmitted information is directed to the terminal used by the performer.
[0218] Step 4:
[0219] The terminal displays instructions received from the server to the performer in real time. The terminal visualizes the information and is designed to be intuitively easy for the performer to understand. Based on this information, the performer can decide on their actions in the actual performance.
[0220] Step 5:
[0221] Users can request further instructions from the server during performance or if an anomaly occurs. They can also modify system settings to prompt the server to generate more appropriate instructions.
[0222] Step 6:
[0223] The server analyzes the data collected after the performance and generates feedback to identify areas for improvement in the performer. This information is then used in training for the next performance.
[0224] Step 7:
[0225] The device displays the training program and provides the performer with a plan to improve their skills based on feedback. The performer can use this information to proceed with their training systematically.
[0226] (Example 1)
[0227] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0228] For performers to deliver high performances in stage and video content, they need to be provided with appropriate instructions and scenarios in real time. However, current systems often fail to adequately adapt the generated instructions to the performers' past performance data and real-time circumstances, resulting in a decline in performance quality. Furthermore, there is a lack of adequate mechanisms for performers to improve their next performance based on feedback.
[0229] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0230] In this invention, the server includes means for providing appropriate information to the performer using a generative AI model, means for analyzing the performer's past performance data and selecting a generative AI model, and means for collecting the performer's feedback and generating suggestions for improving the next performance. This ensures that the instructions the performer receives in real time are appropriately adjusted based on their individual performance history, and that effective feedback is provided to help them improve their next performance.
[0231] A "generative AI model" is an algorithm that uses machine learning techniques to automatically generate content and instructions that are appropriate for a given purpose from natural language.
[0232] "Past performance data" refers to records of performances an actor has given in the past, including the quality of the performance, audience reactions, and feedback.
[0233] "Feedback" refers to evaluations of a performer's performance and suggestions for improvement, and is important information used to improve future performances.
[0234] "Real-time" refers to the fact that information is processed and provided to the performer immediately at the moment it is requested, minimizing time lag and achieving immediacy.
[0235] "Display means" are devices or mechanisms used by performers to directly receive information, and they play a role in providing detailed instructions and feedback visually.
[0236] The server is the central component of this system, responsible for generating instructions and scenarios to be provided to the performers using a generative AI model. Specifically, the server utilizes machine learning libraries and runs the generative AI model using historical data related to the performers' performance. The generative AI model is based on natural language processing algorithms and generates optimal text information for the performers based on user prompts. For this to work, the server requires a computing system equipped with high-performance CPUs and GPUs.
[0237] The user acts as a prompt engineer, operating the system and inputting specific prompt messages to the server. Specifically, they input prompts related to performance themes or specific scenes. For example, they might send a prompt message to the server such as, "Generate a scenario for the next music performance that includes countermeasures for anticipated problems." Based on this input, the server uses a generative AI model to generate a scenario, which is then presented to the user for review.
[0238] The terminal is a device held by the performer, and it has the function of displaying instructions and scenarios generated by the server in real time. Through this terminal, the performer can instantly receive detailed instructions and feedback and incorporate them into their performance. For example, they can check the instructions for the next action on the terminal on stage and proceed with the performance smoothly. In order for the terminal to work, a mobile device or tablet PC with display capabilities is used. This allows performers to always perform based on the latest information.
[0239] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0240] Step 1:
[0241] The user inputs specific prompts into the server to generate instructions and scenarios necessary for the actors. For example, this might include instructions for acting in a specific scene or solutions for problems that may arise. This input includes the acting theme and specific requirements.
[0242] Step 2:
[0243] The server parses the prompt text received from the user. The parsed prompt is input into the generating AI model. Based on this input data, the server performs data calculations to generate appropriate instructions and scenarios from the model. As output, text data to be provided to the performer is derived.
[0244] Step 3:
[0245] The server temporarily stores the generated text data and notifies the user to check if the output is suitable for the actor. The user reviews it and makes adjustments as needed, such as adding parts of the scenario or adjusting the lines.
[0246] Step 4:
[0247] The server sends the final instructions to the terminal after confirmation. The terminal immediately displays the information received from the server, allowing the performer to refer to it in real time. This step utilizes high-speed data communication and the terminal's display capabilities.
[0248] Step 5:
[0249] The performer executes their performance based on the information displayed on the terminal. After the performance, they input feedback on areas for improvement and requests for the next performance into the terminal and send it to the server. This feedback data will be used to improve future performances.
[0250] (Application Example 1)
[0251] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0252] In conventional production facilities, workers often lack sufficient support to efficiently perform complex procedures, which contributes to decreased productivity and errors. There is a need to improve this situation and enable workers to receive appropriate instructions in real time, thereby increasing work efficiency and accuracy.
[0253] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0254] In this invention, the server includes means for providing appropriate information and instructions to an operator using a generative model, means for presenting information via an information processing device to support the operator's work in real time, and means for manipulating the generative model and generating instructions based on requests from the user. This enables the provision of immediate and effective work instructions in the production field.
[0255] A "generative model" is an algorithm that learns from past data to generate new information or instructions.
[0256] "Workers" refer to personnel or automated equipment that perform tasks in a factory or production facility.
[0257] An "information processing device" is an electronic device that receives and processes data to present appropriate information to the user.
[0258] "Presentation means" refers to a method or device for visually or audibly transmitting instructions or information generated by an information processing device to a worker.
[0259] "User" refers to a person or organization that can manipulate the generative model and manage the entire system.
[0260] "Feedback" refers to the process or information itself that returns to the system about the results of the work performed by an employee.
[0261] The system that realizes this invention utilizes a generative model server on the cloud to provide immediate and effective instructions to field workers. The server hosts generative AI models and analyzes past work data to generate information useful for future work. As a specific example of use, in a particular assembly task, the server generates the optimal procedure based on past assembly process data and provides it to the worker.
[0262] The terminal refers to an information processing device such as a smartphone or tablet held by the worker, which receives instructions from the server and displays them to the worker in real time. The terminal has an application built with React Native installed, and the HTTP protocol is used for communication with the server. Specifically, when a worker is assembling a product on the assembly line, the terminal has functions such as displaying a screen that visually confirms the next task to be performed.
[0263] Users are provided with an interface for manipulating the generated AI model. This allows for overall system management and adjustment, and improvement of next instructions based on worker feedback. Users can input prompts such as, "Instruct me on the most efficient way to join part A to part B in the next product assembly step," into the generated model, and generate new work procedures.
[0264] This system configuration allows workers to efficiently perform complex procedures, and is expected to improve both productivity and work accuracy. As described above, this embodiment of the invention provides a new method for supporting work in manufacturing sites.
[0265] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0266] Step 1:
[0267] The server retrieves past work data from the database. It receives past performance logs and work environment information of the worker as input and prepares to run the generated AI model based on this data. This data is analyzed to extract work efficiency and patterns.
[0268] Step 2:
[0269] The user inputs prompt statements to the generative model. For example, the user might send an instruction to the generative AI model such as, "Instruct me on the most efficient way to join part A to part B in the next product assembly step." Based on this prompt, the generative AI model executes and outputs the optimal work procedure.
[0270] Step 3:
[0271] The server receives the optimal work procedure as the output of the generated AI model. This output is specifically represented as text or a series of steps and processed for presentation to the worker. The processed data is later transferred to the terminal.
[0272] Step 4:
[0273] The terminal presents the worker with the optimal work procedure received from the server. The application installed on the terminal displays clear, sequential information to the worker in real time. Here, the work procedure is presented as a visualized diagram or a list of steps.
[0274] Step 5:
[0275] The worker performs the actual work based on the provided work procedures. Feedback obtained during the work and any newly generated data are sent to the server via the terminal.
[0276] Step 6:
[0277] The server records feedback received from workers and newly collected work data in a database. This data is used to improve future work instructions and adjust the generative model, thereby improving the overall system performance.
[0278] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0279] This invention provides a system that combines a generative model and an emotion engine to provide actors with real-time, appropriate, and emotion-based instructions. The system consists of a server, terminals, and users, and can further improve performance quality by utilizing user emotional feedback.
[0280] The server forms the core of the emotion engine and the generation model. The emotion engine analyzes the user's emotions and provides the data to the generation model. Based on the emotion data, the generation model personalizes the instructions and scenarios for the actor. As a result, appropriate text corresponding to the user's current mental state is generated.
[0281] The user provides emotion feedback to the server. The user's emotional state is continuously monitored through a dedicated emotion recognition device and collected by the server. The user can operate the system and update the instructions as needed.
[0282] The terminal displays the instructions and information sent from the server to the actor in real time. The terminal shows the actor the instructions adapted based on emotions, supporting the actor to provide an optimal performance. Through this terminal, the actor can immediately receive information and take actions for improvement.
[0283] For example, when the actor is nervous on stage, the server senses the nervousness through the emotion engine, and the generation model generates an instruction to promote relaxation. The terminal displays this to the actor, and the actor can immediately take countermeasures according to the instruction. Through this process, optimal performance results can be obtained.
[0284] In this way, the present invention enables dynamic instruction generation using emotion feedback and provides an excellent system to support the actor to maximize their ability during the performance.
[0285] The following describes the processing flow.
[0286] Step 1:
[0287] The user logs in to the system and inputs to the server the outline of the scenarios and instructions required for the event or program. The user wears the emotion recognition device and prepares for the start of emotion feedback.
[0288] Step 2:
[0289] The server receives information provided by the user and inputs it into the generative model. At this point, the emotion engine begins collecting and analyzing real-time emotion data from the user.
[0290] Step 3:
[0291] The server passes emotional data obtained from the emotion engine to the generative model. Based on this information, the generative model generates appropriate instructions and scenarios in real time, according to the user's emotional state.
[0292] Step 4:
[0293] The terminal instantly displays instructions and scenarios sent from the server to the actor. The displayed content is adapted based on the user's emotions, allowing the actor to act most effectively.
[0294] Step 5:
[0295] Users provide real-time feedback to the server via the emotion engine to give additional instructions or corrections as needed during their performance. This allows the instructions to be adjusted accordingly.
[0296] Step 6:
[0297] The server analyzes user feedback and sentiment data to generate improvement suggestions for future performances and training. This data is then used to prepare subsequent performances.
[0298] Step 7:
[0299] The device displays generated improvement suggestions and training programs to the performer. Based on this, the performer can systematically improve their skills.
[0300] (Example 2)
[0301] Next, Example 2 will be described. In the following description, the data processing device 12 is referred to as a "server", and the smart glasses 214 are referred to as a "terminal".
[0302] In modern performance activities, it is difficult for individual performers to receive instructions adapted in real time. Also, due to the lack of instructions considering the emotional state of the performer, it is difficult for the performance to consistently produce optimal results. Furthermore, at present, the past data of the performer cannot be fully utilized, and the provision of dynamic feedback based on emotions is limited.
[0303] The specific processing by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0304] In this invention, the server includes means for providing the performer with text and instructions based on emotions using a generation module, means for acquiring the user's emotional data by an emotion recognition device and transmitting it to the server, and means for analyzing the emotional data by an emotion engine and specifying the emotional state. As a result, the performer can receive appropriate instructions based on emotions in real time, and it becomes possible to improve the performance by making use of past data.
[0305] The "generation module" is a program that automatically generates instructions and text for the performer based on the emotional state of the user.
[0306] The "performer" is an individual or entity that performs based on instructions based on emotions.
[0307] The "emotion recognition device" is a device for monitoring the user's expression, voice, heart rate, etc. and acquiring emotional data.
[0308] The "server" is a central computing unit that receives emotional data, performs analysis, generates instructions from the generation module, and transmits them.
[0309] The "emotion engine" is a feature equipped with an algorithm that analyzes emotional data acquired from users and identifies their emotional state in real time.
[0310] "Text and instructions" refer to specific guidance and information created by the generative module to guide the performer's performance.
[0311] A "display device" is a terminal or device used to present generated instructions to the user in real time, either visually or audibly.
[0312] "Feedback" is information collected based on the performer's performance and emotional state, and is used to improve their performance in the future.
[0313] This invention provides a system for providing emotion-based instructions to an expresser in real time. This system mainly consists of a server, a user, and a terminal.
[0314] The server plays a central role in the system. The server receives emotional data sent by the user and analyzes it using an emotion engine. The emotion engine incorporates algorithms that identify emotional states based on facial expressions, vocal characteristics, and physical changes such as heart rate. Based on the identified emotional state, a generation module within the server automatically generates text and instructions appropriate to that emotion. This generation module can use a generation AI model to generate optimal instructions from prompt sentences. For example, in response to a prompt sentence such as "Generate specific instructions to encourage relaxation when the user is nervous on stage," it will generate instructions related to relaxation.
[0315] The user uses a dedicated emotion recognition device to send their emotional data to a server in real time. The emotion recognition device acquires data such as the user's facial expressions, voice, and heart rate via sensors and transmits this data to the server using wireless technology.
[0316] The terminal is a device that receives instructions sent from the server and displays the information necessary for the performer. Specifically, generated text and instructions are displayed on the terminal's screen, and in some cases, the instructions are read aloud via voice assistance. This terminal allows the performer to receive real-time feedback during their performance and adjust their performance on the spot.
[0317] In this way, the entire system works in conjunction to support performers in achieving their optimal performance.
[0318] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0319] Step 1:
[0320] The user collects their emotional data using an emotion recognition device. Inputs include facial expression data, voice information, and biometric information such as heart rate. This emotional data is immediately transmitted to a server via Bluetooth or Wi-Fi. Outputs are prepared for transfer to the server.
[0321] Step 2:
[0322] The server inputs the received emotional data into the emotion engine. This engine uses AI algorithms to analyze the data and identify the user's emotional state. Based on the emotional data as input, a machine learning model determines emotional categories such as anxiety, joy, and tension. As output, labels for the identified emotional states are generated.
[0323] Step 3:
[0324] The server passes a prompt for instruction generation to the generation module based on the already analyzed emotional state. A prompt sentence corresponding to the emotional state is prepared as input. This prompt is analyzed by the generation AI model and used as data necessary to form appropriate instructions. As output, specific instructions corresponding to the emotional state are created.
[0325] Step 4:
[0326] The server transfers the generated instructions to the terminal. The generated instruction data is received from the server as input. The instructions are sent to the terminal as output and displayed in real time on the performer's display device.
[0327] Step 5:
[0328] The device presents the received instructions to the user visually or audibly. For example, it might display "Take a deep breath and relax" on the screen, or read the instructions aloud using its voice assistance function. Based on the instruction data as input, the device generates display and audio output, and this information is delivered to the user as output.
[0329] This sequence of events allows the system to utilize the user's emotional state to provide immediate instructions and optimize the performer's performance.
[0330] (Application Example 2)
[0331] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0332] Traditional customer service in retail stores often involves providing a uniform service, making it difficult to offer appropriate customer service tailored to the customer's emotions. In particular, detecting changes in emotions in real time and dynamically adjusting responses accordingly is a significant burden for employees and poses an obstacle to providing optimal customer service.
[0333] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0334] In this invention, the server includes means for providing instructions via a display device to support the activities of role performers in real time using a generative model and an emotion engine; means for detecting the user's emotions using an emotion recognition device and dynamically generating instructions based on the results; and means for displaying guidance to optimize the actions of employees in the store based on emotional information. This enables detailed customer service that responds to the user's emotions.
[0335] A "generative model" is an artificial intelligence system that generates new content or instructions based on input data.
[0336] An "emotion engine" is a program or device that analyzes a user's emotions from their facial expressions, voice, etc., and processes the results.
[0337] A "role performer" refers to a person who fulfills a predetermined role in a specific place or situation, and may include employees who primarily perform customer service duties.
[0338] A "display device" is hardware used to visualize digital information, and includes devices such as smart glasses and smartphones.
[0339] A "means of providing instructions" is a system for delivering direct instructions or guidance to role performers in accordance with specific conditions or circumstances.
[0340] An "emotion recognition device" is a device that detects a user's emotional state from their facial expressions, voice, and other physiological responses, and transmits that information as data.
[0341] "Employee actions within the store" refers to the specific actions taken by employees in the store to perform tasks and provide services.
[0342] "A means of displaying guidance for optimization based on emotional information" refers to a system that visually provides guidance and instructions to adjust the work content of employees according to the user's emotions.
[0343] This invention is a system that improves the quality of service within a store through close cooperation between the server, terminal, and user.
[0344] The server first receives image and audio data from the emotion recognition device and analyzes the user's emotions in real time using an emotion engine. The analysis results are sent to a generative model, which generates appropriate instructions and information for the role performer. The generated instructions are immediately sent to the terminal.
[0345] The terminal visually displays personalized instructions sent from the server through a display device such as smart glasses or a mobile device worn by the role performer. This allows the role performer to respond quickly in accordance with the user's emotional state.
[0346] Users not only have emotional information collected by the emotion recognition device, but they can also provide feedback as needed. This feedback will be used to improve the service in the future.
[0347] As a concrete example, there is a scenario in which a cafe employee uses smart glasses to recognize the customer's emotions and, based on instructions from the emotion engine, suggests the most suitable drink to create a calm atmosphere. When the prompt "Please suggest a drink menu that can be served to the customer as quickly as possible to create a relaxed atmosphere" is input into the generative model, an appropriate menu is presented.
[0348] This system allows store employees to provide customized services tailored to the customer's emotions, thereby improving customer satisfaction.
[0349] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0350] Step 1:
[0351] The server receives image and audio data from the emotion recognition device. This image and audio data is used as input and is preprocessed to extract the features necessary for emotion recognition. The extracted features then become the input for estimating the emotional state in the next step.
[0352] Step 2:
[0353] The server uses an emotion engine to analyze the user's emotions based on pre-processed features. In this step, facial expressions and voice are analyzed based on the input features to estimate the emotional state (e.g., happiness, anger, sadness, etc.). The estimated emotional state is then output as input data for a generative AI model.
[0354] Step 3:
[0355] The server transfers emotional data obtained from the emotion engine to a generative AI model, which then generates appropriate instructions and information based on the prompt text. The input consists of emotional data and pre-prepared prompt texts, and the generative model generates theme-appropriate text and instructions based on this data. The generated instructions are stored on the server as output for display on the terminal.
[0356] Step 4:
[0357] Instructions generated by a generative model are sent from the server to the terminal. The terminal, such as smart glasses, displays personalized instructions based on the user's information. The input is instruction data sent from the server, which is converted into a user-friendly format and displayed. This allows role performers to respond appropriately to the user in real time.
[0358] Step 5:
[0359] Users provide feedback to the server via their terminal as needed. Feedback input includes verbal or selective comments from role performers. This feedback is collected by the server and used as data to improve future responses. The output is the updated feedback database.
[0360] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0361] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include those described above. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions shown by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0362] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0363] [Third Embodiment]
[0364] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0365] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0366] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0367] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0368] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0369] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0370] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0371] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0372] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0373] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0374] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0375] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0376] This invention is a system for providing actors with appropriate instructions in real time to improve the quality of their performance. This system utilizes a generative model to automatically generate situation-appropriate text and instructions, and provides feedback to actors, thereby enabling rapid skill improvement.
[0377] The server functions as the central management unit and hosts the generative model. The server operates the generative model in response to user requests, generating instructions and scenarios. It also collects past performance data from performers and uses this data to generate feedback for improving future performance.
[0378] As a prompt engineer, the user accesses the server and operates the system. The user requests necessary instructions from the server according to specific performances or events, and reviews and adjusts the generated information. Furthermore, the user reviews feedback from performers and optimizes the system settings for the next performance.
[0379] The terminal is a device or screen held by the performer, displaying instructions sent from the server in real time. Through this terminal, the performer can obtain necessary information and smoothly proceed with their performance. The terminal also displays a training program, supporting the performer in systematically improving their skills.
[0380] As a concrete example, in a scene where an actor is appearing on a music program, the user requests the generative model via the server to develop the scenario appropriately. The server immediately generates a script and displays it on the terminal. Based on this information, the actor can perform according to the script and respond quickly to any unexpected problems.
[0381] In this way, the present invention can significantly improve the quality of performers' performances and also contribute to the efficiency of entertainment production.
[0382] The following describes the processing flow.
[0383] Step 1:
[0384] Users input information about performances and events into the server and request the generation of necessary scenarios and instructions. Users provide information and specify specific conditions and requirements through the system interface.
[0385] Step 2:
[0386] The server receives requests from users and operates a generative model based on them. Scenarios and instructions are automatically generated and customized to match the user's requirements. This information is processed immediately, and instructions are generated.
[0387] Step 3:
[0388] The server sends the generated instructions to the terminal. The instructions are organized in a pre-specified order as information necessary for the performance to progress. The transmitted information is directed to the terminal used by the performer.
[0389] Step 4:
[0390] The terminal displays instructions received from the server to the performer in real time. The terminal visualizes the information and is designed to be intuitively easy for the performer to understand. Based on this information, the performer can decide on their actions in the actual performance.
[0391] Step 5:
[0392] Users can request further instructions from the server during performance or if an anomaly occurs. They can also modify system settings to prompt the server to generate more appropriate instructions.
[0393] Step 6:
[0394] The server analyzes the data collected after the performance and generates feedback to identify areas for improvement in the performer. This information is then used in training for the next performance.
[0395] Step 7:
[0396] The device displays the training program and provides the performer with a plan to improve their skills based on feedback. The performer can use this information to proceed with their training systematically.
[0397] (Example 1)
[0398] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0399] For performers to deliver high performances in stage and video content, they need to be provided with appropriate instructions and scenarios in real time. However, current systems often fail to adequately adapt the generated instructions to the performers' past performance data and real-time circumstances, resulting in a decline in performance quality. Furthermore, there is a lack of adequate mechanisms for performers to improve their next performance based on feedback.
[0400] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0401] In this invention, the server includes means for providing appropriate information to the performer using a generative AI model, means for analyzing the performer's past performance data and selecting a generative AI model, and means for collecting the performer's feedback and generating suggestions for improving the next performance. This ensures that the instructions the performer receives in real time are appropriately adjusted based on their individual performance history, and that effective feedback is provided to help them improve their next performance.
[0402] A "generative AI model" is an algorithm that uses machine learning techniques to automatically generate content and instructions that are appropriate for a given purpose from natural language.
[0403] "Past performance data" refers to records of performances an actor has given in the past, including the quality of the performance, audience reactions, and feedback.
[0404] "Feedback" refers to evaluations of a performer's performance and suggestions for improvement, and is important information used to improve future performances.
[0405] "Real-time" refers to the fact that information is processed and provided to the performer immediately at the moment it is requested, minimizing time lag and achieving immediacy.
[0406] "Display means" are devices or mechanisms used by performers to directly receive information, and they play a role in providing detailed instructions and feedback visually.
[0407] The server is the central component of this system, responsible for generating instructions and scenarios to be provided to the performers using a generative AI model. Specifically, the server utilizes machine learning libraries and runs the generative AI model using historical data related to the performers' performance. The generative AI model is based on natural language processing algorithms and generates optimal text information for the performers based on user prompts. For this to work, the server requires a computing system equipped with high-performance CPUs and GPUs.
[0408] The user acts as a prompt engineer, operating the system and inputting specific prompt messages to the server. Specifically, they input prompts related to performance themes or specific scenes. For example, they might send a prompt message to the server such as, "Generate a scenario for the next music performance that includes countermeasures for anticipated problems." Based on this input, the server uses a generative AI model to generate a scenario, which is then presented to the user for review.
[0409] The terminal is a device held by the performer, and it has the function of displaying instructions and scenarios generated by the server in real time. Through this terminal, the performer can instantly receive detailed instructions and feedback and incorporate them into their performance. For example, they can check the instructions for the next action on the terminal on stage and proceed with the performance smoothly. In order for the terminal to work, a mobile device or tablet PC with display capabilities is used. This allows performers to always perform based on the latest information.
[0410] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0411] Step 1:
[0412] The user inputs specific prompts into the server to generate instructions and scenarios necessary for the actors. For example, this might include instructions for acting in a specific scene or solutions for problems that may arise. This input includes the acting theme and specific requirements.
[0413] Step 2:
[0414] The server parses the prompt text received from the user. The parsed prompt is input into the generating AI model. Based on this input data, the server performs data calculations to generate appropriate instructions and scenarios from the model. As output, text data to be provided to the performer is derived.
[0415] Step 3:
[0416] The server temporarily stores the generated text data and notifies the user to check if the output is suitable for the actor. The user reviews it and makes adjustments as needed, such as adding parts of the scenario or adjusting the lines.
[0417] Step 4:
[0418] The server sends the final instructions to the terminal after confirmation. The terminal immediately displays the information received from the server, allowing the performer to refer to it in real time. This step utilizes high-speed data communication and the terminal's display capabilities.
[0419] Step 5:
[0420] The performer executes their performance based on the information displayed on the terminal. After the performance, they input feedback on areas for improvement and requests for the next performance into the terminal and send it to the server. This feedback data will be used to improve future performances.
[0421] (Application Example 1)
[0422] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0423] In conventional production facilities, workers often lack sufficient support to efficiently perform complex procedures, which contributes to decreased productivity and errors. There is a need to improve this situation and enable workers to receive appropriate instructions in real time, thereby increasing work efficiency and accuracy.
[0424] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0425] In this invention, the server includes means for providing appropriate information and instructions to an operator using a generative model, means for presenting information via an information processing device to support the operator's work in real time, and means for manipulating the generative model and generating instructions based on requests from the user. This enables the provision of immediate and effective work instructions in the production field.
[0426] A "generative model" is an algorithm that learns from past data to generate new information or instructions.
[0427] "Workers" refer to personnel or automated equipment that perform tasks in a factory or production facility.
[0428] An "information processing device" is an electronic device that receives and processes data to present appropriate information to the user.
[0429] "Presentation means" refers to a method or device for visually or audibly transmitting instructions or information generated by an information processing device to a worker.
[0430] "User" refers to a person or organization that can manipulate the generative model and manage the entire system.
[0431] "Feedback" refers to the process or information itself that returns to the system about the results of the work performed by an employee.
[0432] The system that realizes this invention utilizes a generative model server on the cloud to provide immediate and effective instructions to field workers. The server hosts generative AI models and analyzes past work data to generate information useful for future work. As a specific example of use, in a particular assembly task, the server generates the optimal procedure based on past assembly process data and provides it to the worker.
[0433] The terminal refers to an information processing device such as a smartphone or tablet held by the worker, which receives instructions from the server and displays them to the worker in real time. The terminal has an application built with React Native installed, and the HTTP protocol is used for communication with the server. Specifically, when a worker is assembling a product on the assembly line, the terminal has functions such as displaying a screen that visually confirms the next task to be performed.
[0434] Users are provided with an interface for manipulating the generated AI model. This allows for overall system management and adjustment, and improvement of next instructions based on worker feedback. Users can input prompts such as, "Instruct me on the most efficient way to join part A to part B in the next product assembly step," into the generated model, and generate new work procedures.
[0435] This system configuration allows workers to efficiently perform complex procedures, and is expected to improve both productivity and work accuracy. As described above, this embodiment of the invention provides a new method for supporting work in manufacturing sites.
[0436] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0437] Step 1:
[0438] The server retrieves past work data from the database. It receives past performance logs and work environment information of the worker as input and prepares to run the generated AI model based on this data. This data is analyzed to extract work efficiency and patterns.
[0439] Step 2:
[0440] The user inputs prompt statements to the generative model. For example, the user might send an instruction to the generative AI model such as, "Instruct me on the most efficient way to join part A to part B in the next product assembly step." Based on this prompt, the generative AI model executes and outputs the optimal work procedure.
[0441] Step 3:
[0442] The server receives the optimal work procedure as the output of the generated AI model. This output is specifically represented as text or a series of steps and processed for presentation to the worker. The processed data is later transferred to the terminal.
[0443] Step 4:
[0444] The terminal presents the worker with the optimal work procedure received from the server. The application installed on the terminal displays clear, sequential information to the worker in real time. Here, the work procedure is presented as a visualized diagram or a list of steps.
[0445] Step 5:
[0446] The worker performs the actual work based on the provided work procedures. Feedback obtained during the work and any newly generated data are sent to the server via the terminal.
[0447] Step 6:
[0448] The server records feedback received from workers and newly collected work data in a database. This data is used to improve future work instructions and adjust the generative model, thereby improving the overall system performance.
[0449] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0450] This invention provides a system that combines a generative model and an emotion engine to provide actors with real-time, appropriate, and emotion-based instructions. The system consists of a server, terminals, and users, and can further improve performance quality by utilizing user emotional feedback.
[0451] The server forms the core of the emotion engine and generative model. The emotion engine analyzes the user's emotions and provides that data to the generative model. Based on the emotion data, the generative model personalizes instructions and scenarios for the performers. This generates appropriate text that matches the user's current mental state.
[0452] The user provides emotional feedback to the server. The user's emotional state is continuously monitored and collected by the server through a dedicated emotion recognition device. The user can operate the system and update instructions as needed.
[0453] The terminal displays instructions and information sent from the server to the performer in real time. The terminal provides the performer with instructions adapted based on their emotions, helping them to deliver their best performance. Through this terminal, the performer can immediately receive information and take action to improve.
[0454] For example, if an actor is nervous on stage, the server senses the tension through its emotion engine, and a generative model generates instructions to promote relaxation. The terminal displays these instructions to the actor, who can then immediately take appropriate action. This process leads to optimal performance results.
[0455] Thus, the present invention enables the dynamic generation of instructions using emotional feedback and provides an excellent system for supporting performers in maximizing their abilities during performance.
[0456] The following describes the processing flow.
[0457] Step 1:
[0458] The user logs into the system and enters an outline of the scenario and instructions required for the event or performance into the server. They then put on an emotion recognition device and prepare to begin emotional feedback.
[0459] Step 2:
[0460] The server receives information provided by the user and inputs it into the generative model. At this point, the emotion engine begins collecting and analyzing real-time emotion data from the user.
[0461] Step 3:
[0462] The server passes emotional data obtained from the emotion engine to the generative model. Based on this information, the generative model generates appropriate instructions and scenarios in real time, according to the user's emotional state.
[0463] Step 4:
[0464] The terminal instantly displays instructions and scenarios sent from the server to the actor. The displayed content is adapted based on the user's emotions, allowing the actor to act most effectively.
[0465] Step 5:
[0466] Users provide real-time feedback to the server via the emotion engine to give additional instructions or corrections as needed during their performance. This allows the instructions to be adjusted accordingly.
[0467] Step 6:
[0468] The server analyzes user feedback and sentiment data to generate improvement suggestions for future performances and training. This data is then used to prepare subsequent performances.
[0469] Step 7:
[0470] The device displays generated improvement suggestions and training programs to the performer. Based on this, the performer can systematically improve their skills.
[0471] (Example 2)
[0472] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0473] In contemporary performance activities, it is difficult for individual performers to receive real-time, adaptable instructions. Furthermore, the lack of instructions that take into account the performer's emotional state makes it difficult for performances to consistently produce optimal results. Additionally, current practices do not fully utilize performers' past data, limiting the provision of dynamic, emotion-based feedback.
[0474] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0475] In this invention, the server includes means for providing emotion-based text and instructions to the performer using a generation module, means for acquiring the user's emotion data using an emotion recognition device and transmitting it to the server, and means for analyzing the emotion data using an emotion engine and identifying the emotional state. As a result, the performer can receive appropriate emotion-based instructions in real time and improve performance by utilizing past data.
[0476] A "generating module" is a program that automatically generates instructions and text for the performer based on the user's emotional state.
[0477] A "performer" is an individual or entity that performs based on emotionally driven instructions.
[0478] An "emotion recognition device" is a device that monitors the user's facial expressions, voice, heart rate, etc., to acquire emotional data.
[0479] A "server" is a central computing unit that receives emotional data, analyzes it, and then generates and transmits instructions from the generation module.
[0480] The "emotion engine" is a feature equipped with an algorithm that analyzes emotional data acquired from users and identifies their emotional state in real time.
[0481] "Text and instructions" refer to specific guidance and information created by the generative module to guide the performer's performance.
[0482] A "display device" is a terminal or device used to present generated instructions to the user in real time, either visually or audibly.
[0483] "Feedback" is information collected based on the performer's performance and emotional state, and is used to improve their performance in the future.
[0484] This invention provides a system for providing emotion-based instructions to an expresser in real time. This system mainly consists of a server, a user, and a terminal.
[0485] The server plays a central role in the system. The server receives emotional data sent by the user and analyzes it using an emotion engine. The emotion engine incorporates algorithms that identify emotional states based on facial expressions, vocal characteristics, and physical changes such as heart rate. Based on the identified emotional state, a generation module within the server automatically generates text and instructions appropriate to that emotion. This generation module can use a generation AI model to generate optimal instructions from prompt sentences. For example, in response to a prompt sentence such as "Generate specific instructions to encourage relaxation when the user is nervous on stage," it will generate instructions related to relaxation.
[0486] The user uses a dedicated emotion recognition device to send their emotional data to a server in real time. The emotion recognition device acquires data such as the user's facial expressions, voice, and heart rate via sensors and transmits this data to the server using wireless technology.
[0487] The terminal is a device that receives instructions sent from the server and displays the information necessary for the performer. Specifically, generated text and instructions are displayed on the terminal's screen, and in some cases, the instructions are read aloud via voice assistance. This terminal allows the performer to receive real-time feedback during their performance and adjust their performance on the spot.
[0488] In this way, the entire system works in conjunction to support performers in achieving their optimal performance.
[0489] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0490] Step 1:
[0491] The user collects their emotional data using an emotion recognition device. Inputs include facial expression data, voice information, and biometric information such as heart rate. This emotional data is immediately transmitted to a server via Bluetooth or Wi-Fi. Outputs are prepared for transfer to the server.
[0492] Step 2:
[0493] The server inputs the received emotional data into the emotion engine. This engine uses AI algorithms to analyze the data and identify the user's emotional state. Based on the emotional data as input, a machine learning model determines emotional categories such as anxiety, joy, and tension. As output, labels for the identified emotional states are generated.
[0494] Step 3:
[0495] The server passes a prompt for instruction generation to the generation module based on the already analyzed emotional state. A prompt sentence corresponding to the emotional state is prepared as input. This prompt is analyzed by the generation AI model and used as data necessary to form appropriate instructions. As output, specific instructions corresponding to the emotional state are created.
[0496] Step 4:
[0497] The server transfers the generated instructions to the terminal. The generated instruction data is received from the server as input. The instructions are sent to the terminal as output and displayed in real time on the performer's display device.
[0498] Step 5:
[0499] The device presents the received instructions to the user visually or audibly. For example, it might display "Take a deep breath and relax" on the screen, or read the instructions aloud using its voice assistance function. Based on the instruction data as input, the device generates display and audio output, and this information is delivered to the user as output.
[0500] This sequence of events allows the system to utilize the user's emotional state to provide immediate instructions and optimize the performer's performance.
[0501] (Application Example 2)
[0502] Next, we will explain Application Example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0503] Traditional customer service in retail stores often involves providing a uniform service, making it difficult to offer appropriate customer service tailored to the customer's emotions. In particular, detecting changes in emotions in real time and dynamically adjusting responses accordingly is a significant burden for employees and poses an obstacle to providing optimal customer service.
[0504] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0505] In this invention, the server includes means for providing instructions via a display device to support the activities of role performers in real time using a generative model and an emotion engine; means for detecting the user's emotions using an emotion recognition device and dynamically generating instructions based on the results; and means for displaying guidance to optimize the actions of employees in the store based on emotional information. This enables detailed customer service that responds to the user's emotions.
[0506] A "generative model" is an artificial intelligence system that generates new content or instructions based on input data.
[0507] An "emotion engine" is a program or device that analyzes a user's emotions from their facial expressions, voice, etc., and processes the results.
[0508] A "role performer" refers to a person who fulfills a predetermined role in a specific place or situation, and may include employees who primarily perform customer service duties.
[0509] A "display device" is hardware used to visualize digital information, and includes devices such as smart glasses and smartphones.
[0510] A "means of providing instructions" is a system for delivering direct instructions or guidance to role performers in accordance with specific conditions or circumstances.
[0511] An "emotion recognition device" is a device that detects a user's emotional state from their facial expressions, voice, and other physiological responses, and transmits that information as data.
[0512] "Employee actions within the store" refers to the specific actions taken by employees in the store to perform tasks and provide services.
[0513] "A means of displaying guidance for optimization based on emotional information" refers to a system that visually provides guidance and instructions to adjust the work content of employees according to the user's emotions.
[0514] This invention is a system that improves the quality of service within a store through close cooperation between the server, terminal, and user.
[0515] The server first receives image and audio data from the emotion recognition device and analyzes the user's emotions in real time using an emotion engine. The analysis results are sent to a generative model, which generates appropriate instructions and information for the role performer. The generated instructions are immediately sent to the terminal.
[0516] The terminal visually displays personalized instructions sent from the server through a display device such as smart glasses or a mobile device worn by the role performer. This allows the role performer to respond quickly in accordance with the user's emotional state.
[0517] Users not only have emotional information collected by the emotion recognition device, but they can also provide feedback as needed. This feedback will be used to improve the service in the future.
[0518] As a concrete example, there is a scenario in which a cafe employee uses smart glasses to recognize the customer's emotions and, based on instructions from the emotion engine, suggests the most suitable drink to create a calm atmosphere. When the prompt "Please suggest a drink menu that can be served to the customer as quickly as possible to create a relaxed atmosphere" is input into the generative model, an appropriate menu is presented.
[0519] This system allows store employees to provide customized services tailored to the customer's emotions, thereby improving customer satisfaction.
[0520] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0521] Step 1:
[0522] The server receives image and audio data from the emotion recognition device. This image and audio data is used as input and is preprocessed to extract the features necessary for emotion recognition. The extracted features then become the input for estimating the emotional state in the next step.
[0523] Step 2:
[0524] The server uses an emotion engine to analyze the user's emotions based on pre-processed features. In this step, facial expressions and voice are analyzed based on the input features to estimate the emotional state (e.g., happiness, anger, sadness, etc.). The estimated emotional state is then output as input data for a generative AI model.
[0525] Step 3:
[0526] The server transfers emotional data obtained from the emotion engine to a generative AI model, which then generates appropriate instructions and information based on the prompt text. The input consists of emotional data and pre-prepared prompt texts, and the generative model generates theme-appropriate text and instructions based on this data. The generated instructions are stored on the server as output for display on the terminal.
[0527] Step 4:
[0528] Instructions generated by a generative model are sent from the server to the terminal. The terminal, such as smart glasses, displays personalized instructions based on the user's information. The input is instruction data sent from the server, which is converted into a user-friendly format and displayed. This allows role performers to respond appropriately to the user in real time.
[0529] Step 5:
[0530] Users provide feedback to the server via their terminal as needed. Feedback input includes verbal or selective comments from role performers. This feedback is collected by the server and used as data to improve future responses. The output is the updated feedback database.
[0531] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0532] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include those described above. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions shown by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0533] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[0534] [Fourth Embodiment]
[0535] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0536] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0537] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0538] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0539] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0540] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0541] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0542] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0543] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0544] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0545] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0546] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0547] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0548] This invention is a system for providing actors with appropriate instructions in real time to improve the quality of their performance. This system utilizes a generative model to automatically generate situation-appropriate text and instructions, and provides feedback to actors, thereby enabling rapid skill improvement.
[0549] The server functions as the central management unit and hosts the generative model. The server operates the generative model in response to user requests, generating instructions and scenarios. It also collects past performance data from performers and uses this data to generate feedback for improving future performance.
[0550] As a prompt engineer, the user accesses the server and operates the system. The user requests necessary instructions from the server according to specific performances or events, and reviews and adjusts the generated information. Furthermore, the user reviews feedback from performers and optimizes the system settings for the next performance.
[0551] The terminal is a device or screen held by the performer, displaying instructions sent from the server in real time. Through this terminal, the performer can obtain necessary information and smoothly proceed with their performance. The terminal also displays a training program, supporting the performer in systematically improving their skills.
[0552] As a concrete example, in a scene where an actor is appearing on a music program, the user requests the generative model via the server to develop the scenario appropriately. The server immediately generates a script and displays it on the terminal. Based on this information, the actor can perform according to the script and respond quickly to any unexpected problems.
[0553] In this way, the present invention can significantly improve the quality of performers' performances and also contribute to the efficiency of entertainment production.
[0554] The following describes the processing flow.
[0555] Step 1:
[0556] Users input information about performances and events into the server and request the generation of necessary scenarios and instructions. Users provide information and specify specific conditions and requirements through the system interface.
[0557] Step 2:
[0558] The server receives requests from users and operates a generative model based on them. Scenarios and instructions are automatically generated and customized to match the user's requirements. This information is processed immediately, and instructions are generated.
[0559] Step 3:
[0560] The server sends the generated instructions to the terminal. The instructions are organized in a pre-specified order as information necessary for the performance to progress. The transmitted information is directed to the terminal used by the performer.
[0561] Step 4:
[0562] The terminal displays instructions received from the server to the performer in real time. The terminal visualizes the information and is designed to be intuitively easy for the performer to understand. Based on this information, the performer can decide on their actions in the actual performance.
[0563] Step 5:
[0564] Users can request further instructions from the server during performance or if an anomaly occurs. They can also modify system settings to prompt the server to generate more appropriate instructions.
[0565] Step 6:
[0566] The server analyzes the data collected after the performance and generates feedback to identify areas for improvement in the performer. This information is then used in training for the next performance.
[0567] Step 7:
[0568] The device displays the training program and provides the performer with a plan to improve their skills based on feedback. The performer can use this information to proceed with their training systematically.
[0569] (Example 1)
[0570] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0571] For performers to deliver high performances in stage and video content, they need to be provided with appropriate instructions and scenarios in real time. However, current systems often fail to adequately adapt the generated instructions to the performers' past performance data and real-time circumstances, resulting in a decline in performance quality. Furthermore, there is a lack of adequate mechanisms for performers to improve their next performance based on feedback.
[0572] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0573] In this invention, the server includes means for providing appropriate information to the performer using a generative AI model, means for analyzing the performer's past performance data and selecting a generative AI model, and means for collecting the performer's feedback and generating suggestions for improving the next performance. This ensures that the instructions the performer receives in real time are appropriately adjusted based on their individual performance history, and that effective feedback is provided to help them improve their next performance.
[0574] A "generative AI model" is an algorithm that uses machine learning techniques to automatically generate content and instructions that are appropriate for a given purpose from natural language.
[0575] "Past performance data" refers to records of performances an actor has given in the past, including the quality of the performance, audience reactions, and feedback.
[0576] "Feedback" refers to evaluations of a performer's performance and suggestions for improvement, and is important information used to improve future performances.
[0577] "Real-time" refers to the fact that information is processed and provided to the performer immediately at the moment it is requested, minimizing time lag and achieving immediacy.
[0578] "Display means" are devices or mechanisms used by performers to directly receive information, and they play a role in providing detailed instructions and feedback visually.
[0579] The server is the central component of this system, responsible for generating instructions and scenarios to be provided to the performers using a generative AI model. Specifically, the server utilizes machine learning libraries and runs the generative AI model using historical data related to the performers' performance. The generative AI model is based on natural language processing algorithms and generates optimal text information for the performers based on user prompts. For this to work, the server requires a computing system equipped with high-performance CPUs and GPUs.
[0580] The user acts as a prompt engineer, operating the system and inputting specific prompt messages to the server. Specifically, they input prompts related to performance themes or specific scenes. For example, they might send a prompt message to the server such as, "Generate a scenario for the next music performance that includes countermeasures for anticipated problems." Based on this input, the server uses a generative AI model to generate a scenario, which is then presented to the user for review.
[0581] The terminal is a device held by the performer, and it has the function of displaying instructions and scenarios generated by the server in real time. Through this terminal, the performer can instantly receive detailed instructions and feedback and incorporate them into their performance. For example, they can check the instructions for the next action on the terminal on stage and proceed with the performance smoothly. In order for the terminal to work, a mobile device or tablet PC with display capabilities is used. This allows performers to always perform based on the latest information.
[0582] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0583] Step 1:
[0584] The user inputs specific prompts into the server to generate instructions and scenarios necessary for the actors. For example, this might include instructions for acting in a specific scene or solutions for problems that may arise. This input includes the acting theme and specific requirements.
[0585] Step 2:
[0586] The server parses the prompt text received from the user. The parsed prompt is input into the generating AI model. Based on this input data, the server performs data calculations to generate appropriate instructions and scenarios from the model. As output, text data to be provided to the performer is derived.
[0587] Step 3:
[0588] The server temporarily stores the generated text data and notifies the user to check if the output is suitable for the actor. The user reviews it and makes adjustments as needed, such as adding parts of the scenario or adjusting the lines.
[0589] Step 4:
[0590] The server sends the final instructions to the terminal after confirmation. The terminal immediately displays the information received from the server, allowing the performer to refer to it in real time. This step utilizes high-speed data communication and the terminal's display capabilities.
[0591] Step 5:
[0592] The performer executes their performance based on the information displayed on the terminal. After the performance, they input feedback on areas for improvement and requests for the next performance into the terminal and send it to the server. This feedback data will be used to improve future performances.
[0593] (Application Example 1)
[0594] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0595] In conventional production facilities, workers often lack sufficient support to efficiently perform complex procedures, which contributes to decreased productivity and errors. There is a need to improve this situation and enable workers to receive appropriate instructions in real time, thereby increasing work efficiency and accuracy.
[0596] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0597] In this invention, the server includes means for providing appropriate information and instructions to an operator using a generative model, means for presenting information via an information processing device to support the operator's work in real time, and means for manipulating the generative model and generating instructions based on requests from the user. This enables the provision of immediate and effective work instructions in the production field.
[0598] A "generative model" is an algorithm that learns from past data to generate new information or instructions.
[0599] "Workers" refer to personnel or automated equipment that perform tasks in a factory or production facility.
[0600] An "information processing device" is an electronic device that receives and processes data to present appropriate information to the user.
[0601] "Presentation means" refers to a method or device for visually or audibly transmitting instructions or information generated by an information processing device to a worker.
[0602] "User" refers to a person or organization that can manipulate the generative model and manage the entire system.
[0603] "Feedback" refers to the process or information itself that returns to the system about the results of the work performed by an employee.
[0604] The system that realizes this invention utilizes a generative model server on the cloud to provide immediate and effective instructions to field workers. The server hosts generative AI models and analyzes past work data to generate information useful for future work. As a specific example of use, in a particular assembly task, the server generates the optimal procedure based on past assembly process data and provides it to the worker.
[0605] The terminal refers to an information processing device such as a smartphone or tablet held by the worker, which receives instructions from the server and displays them to the worker in real time. The terminal has an application built with React Native installed, and the HTTP protocol is used for communication with the server. Specifically, when a worker is assembling a product on the assembly line, the terminal has functions such as displaying a screen that visually confirms the next task to be performed.
[0606] Users are provided with an interface for manipulating the generated AI model. This allows for overall system management and adjustment, and improvement of next instructions based on worker feedback. Users can input prompts such as, "Instruct me on the most efficient way to join part A to part B in the next product assembly step," into the generated model, and generate new work procedures.
[0607] This system configuration allows workers to efficiently perform complex procedures, and is expected to improve both productivity and work accuracy. As described above, this embodiment of the invention provides a new method for supporting work in manufacturing sites.
[0608] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0609] Step 1:
[0610] The server retrieves past work data from the database. It receives past performance logs and work environment information of the worker as input and prepares to run the generated AI model based on this data. This data is analyzed to extract work efficiency and patterns.
[0611] Step 2:
[0612] The user inputs prompt statements to the generative model. For example, the user might send an instruction to the generative AI model such as, "Instruct me on the most efficient way to join part A to part B in the next product assembly step." Based on this prompt, the generative AI model executes and outputs the optimal work procedure.
[0613] Step 3:
[0614] The server receives the optimal work procedure as the output of the generated AI model. This output is specifically represented as text or a series of steps and processed for presentation to the worker. The processed data is later transferred to the terminal.
[0615] Step 4:
[0616] The terminal presents the worker with the optimal work procedure received from the server. The application installed on the terminal displays clear, sequential information to the worker in real time. Here, the work procedure is presented as a visualized diagram or a list of steps.
[0617] Step 5:
[0618] The worker performs the actual work based on the provided work procedures. Feedback obtained during the work and any newly generated data are sent to the server via the terminal.
[0619] Step 6:
[0620] The server records feedback received from workers and newly collected work data in a database. This data is used to improve future work instructions and adjust the generative model, thereby improving the overall system performance.
[0621] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0622] This invention provides a system that combines a generative model and an emotion engine to provide actors with real-time, appropriate, and emotion-based instructions. The system consists of a server, terminals, and users, and can further improve performance quality by utilizing user emotional feedback.
[0623] The server forms the core of the emotion engine and generative model. The emotion engine analyzes the user's emotions and provides that data to the generative model. Based on the emotion data, the generative model personalizes instructions and scenarios for the performers. This generates appropriate text that matches the user's current mental state.
[0624] The user provides emotional feedback to the server. The user's emotional state is continuously monitored and collected by the server through a dedicated emotion recognition device. The user can operate the system and update instructions as needed.
[0625] The terminal displays instructions and information sent from the server to the performer in real time. The terminal provides the performer with instructions adapted based on their emotions, helping them to deliver their best performance. Through this terminal, the performer can immediately receive information and take action to improve.
[0626] For example, if an actor is nervous on stage, the server senses the tension through its emotion engine, and a generative model generates instructions to promote relaxation. The terminal displays these instructions to the actor, who can then immediately take appropriate action. This process leads to optimal performance results.
[0627] Thus, the present invention enables the dynamic generation of instructions using emotional feedback and provides an excellent system for supporting performers in maximizing their abilities during performance.
[0628] The following describes the processing flow.
[0629] Step 1:
[0630] The user logs into the system and enters an outline of the scenario and instructions required for the event or performance into the server. They then put on an emotion recognition device and prepare to begin emotional feedback.
[0631] Step 2:
[0632] The server receives information provided by the user and inputs it into the generative model. At this point, the emotion engine begins collecting and analyzing real-time emotion data from the user.
[0633] Step 3:
[0634] The server passes emotional data obtained from the emotion engine to the generative model. Based on this information, the generative model generates appropriate instructions and scenarios in real time, according to the user's emotional state.
[0635] Step 4:
[0636] The terminal instantly displays instructions and scenarios sent from the server to the actor. The displayed content is adapted based on the user's emotions, allowing the actor to act most effectively.
[0637] Step 5:
[0638] Users provide real-time feedback to the server via the emotion engine to give additional instructions or corrections as needed during their performance. This allows the instructions to be adjusted accordingly.
[0639] Step 6:
[0640] The server analyzes user feedback and sentiment data to generate improvement suggestions for future performances and training. This data is then used to prepare subsequent performances.
[0641] Step 7:
[0642] The device displays generated improvement suggestions and training programs to the performer. Based on this, the performer can systematically improve their skills.
[0643] (Example 2)
[0644] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0645] In contemporary performance activities, it is difficult for individual performers to receive real-time, adaptable instructions. Furthermore, the lack of instructions that take into account the performer's emotional state makes it difficult for performances to consistently produce optimal results. Additionally, current practices do not fully utilize performers' past data, limiting the provision of dynamic, emotion-based feedback.
[0646] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0647] In this invention, the server includes means for providing emotion-based text and instructions to the performer using a generation module, means for acquiring the user's emotion data using an emotion recognition device and transmitting it to the server, and means for analyzing the emotion data using an emotion engine and identifying the emotional state. As a result, the performer can receive appropriate emotion-based instructions in real time and improve performance by utilizing past data.
[0648] A "generating module" is a program that automatically generates instructions and text for the performer based on the user's emotional state.
[0649] A "performer" is an individual or entity that performs based on emotionally driven instructions.
[0650] An "emotion recognition device" is a device that monitors the user's facial expressions, voice, heart rate, etc., to acquire emotional data.
[0651] A "server" is a central computing unit that receives emotional data, analyzes it, and then generates and transmits instructions from the generation module.
[0652] The "emotion engine" is a feature equipped with an algorithm that analyzes emotional data acquired from users and identifies their emotional state in real time.
[0653] "Text and instructions" refer to specific guidance and information created by the generative module to guide the performer's performance.
[0654] A "display device" is a terminal or device used to present generated instructions to the user in real time, either visually or audibly.
[0655] "Feedback" is information collected based on the performer's performance and emotional state, and is used to improve their performance in the future.
[0656] This invention provides a system for providing emotion-based instructions to an expresser in real time. This system mainly consists of a server, a user, and a terminal.
[0657] The server plays a central role in the system. The server receives emotional data sent by the user and analyzes it using an emotion engine. The emotion engine incorporates algorithms that identify emotional states based on facial expressions, vocal characteristics, and physical changes such as heart rate. Based on the identified emotional state, a generation module within the server automatically generates text and instructions appropriate to that emotion. This generation module can use a generation AI model to generate optimal instructions from prompt sentences. For example, in response to a prompt sentence such as "Generate specific instructions to encourage relaxation when the user is nervous on stage," it will generate instructions related to relaxation.
[0658] The user uses a dedicated emotion recognition device to send their emotional data to a server in real time. The emotion recognition device acquires data such as the user's facial expressions, voice, and heart rate via sensors and transmits this data to the server using wireless technology.
[0659] The terminal is a device that receives instructions sent from the server and displays the information necessary for the performer. Specifically, generated text and instructions are displayed on the terminal's screen, and in some cases, the instructions are read aloud via voice assistance. This terminal allows the performer to receive real-time feedback during their performance and adjust their performance on the spot.
[0660] In this way, the entire system works in conjunction to support performers in achieving their optimal performance.
[0661] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0662] Step 1:
[0663] The user collects their emotional data using an emotion recognition device. Inputs include facial expression data, voice information, and biometric information such as heart rate. This emotional data is immediately transmitted to a server via Bluetooth or Wi-Fi. Outputs are prepared for transfer to the server.
[0664] Step 2:
[0665] The server inputs the received emotional data into the emotion engine. This engine uses AI algorithms to analyze the data and identify the user's emotional state. Based on the emotional data as input, a machine learning model determines emotional categories such as anxiety, joy, and tension. As output, labels for the identified emotional states are generated.
[0666] Step 3:
[0667] The server passes a prompt for instruction generation to the generation module based on the already analyzed emotional state. A prompt sentence corresponding to the emotional state is prepared as input. This prompt is analyzed by the generation AI model and used as data necessary to form appropriate instructions. As output, specific instructions corresponding to the emotional state are created.
[0668] Step 4:
[0669] The server transfers the generated instructions to the terminal. The generated instruction data is received from the server as input. The instructions are sent to the terminal as output and displayed in real time on the performer's display device.
[0670] Step 5:
[0671] The device presents the received instructions to the user visually or audibly. For example, it might display "Take a deep breath and relax" on the screen, or read the instructions aloud using its voice assistance function. Based on the instruction data as input, the device generates display and audio output, and this information is delivered to the user as output.
[0672] This sequence of events allows the system to utilize the user's emotional state to provide immediate instructions and optimize the performer's performance.
[0673] (Application Example 2)
[0674] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0675] Traditional customer service in retail stores often involves providing a uniform service, making it difficult to offer appropriate customer service tailored to the customer's emotions. In particular, detecting changes in emotions in real time and dynamically adjusting responses accordingly is a significant burden for employees and poses an obstacle to providing optimal customer service.
[0676] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0677] In this invention, the server includes means for providing instructions via a display device to support the activities of role performers in real time using a generative model and an emotion engine; means for detecting the user's emotions using an emotion recognition device and dynamically generating instructions based on the results; and means for displaying guidance to optimize the actions of employees in the store based on emotional information. This enables detailed customer service that responds to the user's emotions.
[0678] A "generative model" is an artificial intelligence system that generates new content or instructions based on input data.
[0679] An "emotion engine" is a program or device that analyzes a user's emotions from their facial expressions, voice, etc., and processes the results.
[0680] A "role performer" refers to a person who fulfills a predetermined role in a specific place or situation, and may include employees who primarily perform customer service duties.
[0681] A "display device" is hardware used to visualize digital information, and includes devices such as smart glasses and smartphones.
[0682] A "means of providing instructions" is a system for delivering direct instructions or guidance to role performers in accordance with specific conditions or circumstances.
[0683] An "emotion recognition device" is a device that detects a user's emotional state from their facial expressions, voice, and other physiological responses, and transmits that information as data.
[0684] "Employee actions within the store" refers to the specific actions taken by employees in the store to perform tasks and provide services.
[0685] "A means of displaying guidance for optimization based on emotional information" refers to a system that visually provides guidance and instructions to adjust the work content of employees according to the user's emotions.
[0686] This invention is a system that improves the quality of service within a store through close cooperation between the server, terminal, and user.
[0687] The server first receives image and audio data from the emotion recognition device and analyzes the user's emotions in real time using an emotion engine. The analysis results are sent to a generative model, which generates appropriate instructions and information for the role performer. The generated instructions are immediately sent to the terminal.
[0688] The terminal visually displays personalized instructions sent from the server through a display device such as smart glasses or a mobile device worn by the role performer. This allows the role performer to respond quickly in accordance with the user's emotional state.
[0689] Users not only have emotional information collected by the emotion recognition device, but they can also provide feedback as needed. This feedback will be used to improve the service in the future.
[0690] As a concrete example, there is a scenario in which a cafe employee uses smart glasses to recognize the customer's emotions and, based on instructions from the emotion engine, suggests the most suitable drink to create a calm atmosphere. When the prompt "Please suggest a drink menu that can be served to the customer as quickly as possible to create a relaxed atmosphere" is input into the generative model, an appropriate menu is presented.
[0691] This system allows store employees to provide customized services tailored to the customer's emotions, thereby improving customer satisfaction.
[0692] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0693] Step 1:
[0694] The server receives image and audio data from the emotion recognition device. This image and audio data is used as input and is preprocessed to extract the features necessary for emotion recognition. The extracted features then become the input for estimating the emotional state in the next step.
[0695] Step 2:
[0696] The server uses an emotion engine to analyze the user's emotions based on pre-processed features. In this step, facial expressions and voice are analyzed based on the input features to estimate the emotional state (e.g., happiness, anger, sadness, etc.). The estimated emotional state is then output as input data for a generative AI model.
[0697] Step 3:
[0698] The server transfers emotional data obtained from the emotion engine to a generative AI model, which then generates appropriate instructions and information based on the prompt text. The input consists of emotional data and pre-prepared prompt texts, and the generative model generates theme-appropriate text and instructions based on this data. The generated instructions are stored on the server as output for display on the terminal.
[0699] Step 4:
[0700] Instructions generated by a generative model are sent from the server to the terminal. The terminal, such as smart glasses, displays personalized instructions based on the user's information. The input is instruction data sent from the server, which is converted into a user-friendly format and displayed. This allows role performers to respond appropriately to the user in real time.
[0701] Step 5:
[0702] Users provide feedback to the server via their terminal as needed. Feedback input includes verbal or selective comments from role performers. This feedback is collected by the server and used as data to improve future responses. The output is the updated feedback database.
[0703] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0704] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include those described above. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions shown by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0705] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[0706] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0707] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0708] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0709] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0710] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0711] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0712] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0713] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[0714] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[0715] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0716] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0717] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0718] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0719] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0720] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0721] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0722] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0723] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[0724] The following is further disclosed regarding the embodiments described above.
[0725] (Claim 1)
[0726] A means of providing actors with appropriate text and instructions using a generative model,
[0727] A display means via a terminal to support the performer's performance in real time,
[0728] Means for manipulating a generative model and generating instructions based on user requests,
[0729] A means of collecting feedback from performers and generating suggestions for improving their next performance,
[0730] A system that includes this.
[0731] (Claim 2)
[0732] The system according to claim 1, wherein the generative model customizes instructions to the performer using past performance data.
[0733] (Claim 3)
[0734] The system according to claim 1, comprising means for generating a training program for an actor and providing it to the actor via a terminal.
[0735] "Example 1"
[0736] (Claim 1)
[0737] A means of providing appropriate information to performers using a generative AI model,
[0738] A means of selecting a generation AI model by analyzing the performer's past performance data,
[0739] A display means via a device for assisting the performer's performance in real time,
[0740] A means of manipulating a generated AI model in response to user requests and generating instructions,
[0741] A means of collecting feedback from performers and generating suggestions for improving their next performance,
[0742] A system that includes this.
[0743] (Claim 2)
[0744] The system according to claim 1, wherein a generative AI model customizes instructions to performers using past performance data and optimizes text data.
[0745] (Claim 3)
[0746] The system according to claim 1, comprising means for generating a training program for an actor and providing it to the actor via a device.
[0747] "Application Example 1"
[0748] (Claim 1)
[0749] A means of providing appropriate information and instructions to workers using a generative model,
[0750] A presentation means via an information processing device to support the worker's work in real time,
[0751] Means for manipulating a generative model and generating instructions based on user requests,
[0752] A means of collecting worker feedback and generating suggestions for improving work in the future,
[0753] A system that includes this.
[0754] (Claim 2)
[0755] The system according to claim 1, wherein the generative model customizes instructions to the worker using past work data.
[0756] (Claim 3)
[0757] The system according to claim 1, comprising means for generating a training program for workers and providing it to workers via an information processing device.
[0758] "Example 2 of combining an emotion engine"
[0759] (Claim 1)
[0760] A means of providing emotion-based text and instructions to creators using a generation module,
[0761] A means for acquiring user emotion data using an emotion recognition device and transmitting it to a server,
[0762] A means of analyzing emotional data using an emotion engine and identifying emotional states,
[0763] A means by which a generation module generates instructions based on the analyzed emotional state and displays them to the performer in real time via a display device,
[0764] A means of collecting feedback on the activities of performers and generating suggestions for improving their next performance,
[0765] A system that includes this.
[0766] (Claim 2)
[0767] The system according to claim 1, wherein the generation module customizes instructions to the performer using past activity data.
[0768] (Claim 3)
[0769] The system according to claim 1, comprising means for generating an educational program for an artist and providing it to the artist via a display device.
[0770] "Application example 2 when combining with an emotional engine"
[0771] (Claim 1)
[0772] A means for providing instructions via a display device to support the activities of role performers in real time using a generative model and an emotion engine,
[0773] A means for detecting the user's emotions using an emotion recognition device and dynamically generating instructions based on the results,
[0774] A means of displaying guidance to optimize employee behavior within a store based on emotional information,
[0775] A system that includes this.
[0776] (Claim 2)
[0777] The system according to claim 1, wherein the generative model adapts instructions to role performers using past activity information.
[0778] (Claim 3)
[0779] The system according to claim 1, comprising means for generating a training plan for a role performer and providing it to the role performer via a display device. [Explanation of Symbols]
[0780] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. A means of providing actors with appropriate text and instructions using a generative model, A display means via a terminal to support the performer's performance in real time, Means for manipulating a generative model and generating instructions based on user requests, A means of collecting feedback from performers and generating suggestions for improving their next performance, A system that includes this.
2. The system according to claim 1, wherein the generative model customizes instructions to the performer using past performance data.
3. The system according to claim 1, comprising means for generating a training program for an actor and providing it to the actor via a terminal.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A