system
The system enhances improvisational theater performances by using real-time data capture and analysis to automate stage direction, addressing flexibility and quality issues in small-scale productions.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-10
- Publication Date
- 2026-04-22
Smart Images

Figure 2026068332000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, and includes steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance that responds to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In improvisational theater, since the actions and lines of the actors change on the spot, high-level production techniques are required. However, it is difficult for small-scale theater troupes with limited budgets and personnel to sufficiently ensure the required flexibility and quality of the production. Also, it is a problem to maintain the story while dealing with the ad-libs and mistakes of the actors. Technologies that reduce these problems and realize high-level improvisational performances are required.
Means for Solving the Problems
[0005] In this invention, an acquisition means is used to acquire action information and audio information from the play. Next, an analysis means analyzes this information and provides a control means for automating appropriate direction in real time. Furthermore, a generation means generates correction instructions that correspond to the performers' improvisational actions, thereby constructing a system that improves the flexibility and quality of stage direction. In addition, by acquiring audience reactions through an audience reaction acquisition means, it is possible to further dynamically adjust the direction.
[0006] "Acquisition means" refers to a device or method for acquiring action information and audio information from the play in real time.
[0007] "Analysis means" refers to a device or method for analyzing motion information and audio information obtained from acquisition means and deriving data related to performance.
[0008] "Control means" refers to a device or method for automatically adjusting the lighting and sound in a play based on data derived by the analysis means, and for performing the play.
[0009] "Generative means" refers to a device or method for creating necessary modification instructions to adapt to the performer's improvisational actions and for maintaining the performance.
[0010] "Audience reaction acquisition means" refers to a device or method for acquiring audience reactions during a play and providing them to an analysis means. [Brief explanation of the drawing]
[0011] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2, which incorporates an emotion engine. [Figure 14] This is a sequence diagram showing the processing flow of the data processing system in Application Example 2, which combines an emotion engine. [Modes for carrying out the invention]
[0012] Hereinafter, an example of an embodiment of the system relating to the technology of this disclosure will be described with reference to the attached drawings.
[0013] First, let's explain the terminology used in the following explanation.
[0014] In the following embodiments, the labeled processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0015] In the following embodiments, the labeled RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0016] In the following embodiments, the labeled storage is one or more non-volatile storage devices that store various programs, various parameters, and the like. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, and the like.
[0017] In the following embodiments, the labeled communication I / F (Interface) is an interface that includes a communication processor, an antenna, and the like. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0018] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0019] [First Embodiment]
[0020] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0021] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0022] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0023] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0024] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0025] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0026] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0027] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0028] As shown in Figure 2, in the data processing device 12, specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0029] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0030] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0031] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0032] The "improvisational AI direction" system according to the present invention is for dynamically adjusting the direction on a theatrical stage. This system consists of acquisition means, analysis means, control means, and generation means. The processing of the program in this system is described below in natural language.
[0033] First, the user (director) inputs the play's theme, story structure, and character settings into the system. This prepares the basic information for the play. Next, the terminal captures the actors' movements and conversations in real time through cameras and microphones installed on the stage. In addition, audience reactions are also captured by sensors.
[0034] This data is sent to a server and analyzed in detail by an analysis system. An AI model on the server analyzes the performers' movements and voice information to extract necessary insights regarding the performance. Based on these analysis results, the control system generates instructions to adjust lighting, sound, and other stage equipment in real time.
[0035] Furthermore, if an actor acts spontaneously or makes an unexpected move, the server uses a generation mechanism to instantly create appropriate corrective instructions. This allows for flexible responses that maintain the flow of the story while ensuring the performance is not interrupted.
[0036] For example, if loud laughter erupts from the audience, the server analyzes the reaction and instructs the control system to change the lighting to a warmer color and play upbeat background music at that point. Furthermore, if an actor makes a mistake with their lines, the generation system suggests alternative actions to correct the story in a natural flow. In this way, the present invention improves the quality of direction in improvisational theater.
[0037] The following describes the processing flow.
[0038] Step 1:
[0039] The user (director) uses a terminal to input the play's theme, storyline, character settings, and basic directorial guidelines into the system. The entered information is stored in the system's database and used by the AI for subsequent processing.
[0040] Step 2:
[0041] The devices (cameras and microphones installed on the stage) capture the performers' movements and lines in real time. They also capture audience reactions such as laughter and applause.
[0042] Step 3:
[0043] The server inputs the real-time data stream provided from the terminal into an analysis device, and the performer's actions and voice are analyzed using an AI model. Here, the performer's ad-libs and improvisational actions, the content of the lines, and audience reaction data are identified.
[0044] Step 4:
[0045] Based on the insights gained from the analysis results, the server generates staging instructions that are appropriate for the current stage situation. Specifically, it determines staging elements such as lighting patterns, sound adjustments, and background music selection.
[0046] Step 5:
[0047] The terminal receives performance instructions sent from the server and applies those instructions to the stage's lighting and sound systems. This allows for instantaneous adjustments to lighting color, brightness, and volume, as well as background music switching.
[0048] Step 6:
[0049] The user (stage operator) can monitor whether the automated effects provided by the system are being executed appropriately and make manual adjustments as needed.
[0050] Step 7:
[0051] The server collects feedback data from performers and audience members in response to their reactions, and stores it as points for improvement in the next session, in order to update future production strategies.
[0052] (Example 1)
[0053] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0054] In improvisational theater and live performances, there is a need to respond flexibly in real time to the performers' spontaneous actions and the audience's reactions, thereby improving the quality of the performance. However, conventional systems have the challenge of being unable to respond to changing situations and making appropriate adjustments to maintain the flow of the performance.
[0055] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0056] In this invention, the server includes means for acquiring operational information and audio information, means for performing analysis based on the operational information and audio information from the acquisition means, and means for controlling lighting and sound based on the results of the analysis means. This makes it possible to adjust the performance appropriately in real time even in response to improvisational situational changes.
[0057] "Motion information" refers to data that shows the physical movements of performers and people on stage, such as gestures, movements, and postures.
[0058] "Audio information" refers to data that represents acoustic signals, including sounds from the performers and their surroundings.
[0059] "Means of acquisition" refers to devices and technologies for collecting operational information and audio information.
[0060] "Analytical means" refers to the techniques and processes used to analyze acquired information and derive meaningful results.
[0061] "Control means" refers to devices and technologies used to operate stage equipment such as lighting and sound based on analysis results.
[0062] "Modification instructions" refer to instructions for appropriately changing the stage production in response to the performers' improvisational actions or the audience's reactions.
[0063] "Audience reaction" refers to information that shows the visual, auditory, and other feedback from the audience during a performance.
[0064] "Monitoring methods" refer to devices and technologies used to monitor audience reactions in real time and acquire data.
[0065] "Generative means" refers to the techniques and processes for generating staging and modifications that are appropriate to the improvisational situation.
[0066] The "improvisational AI direction" system of the present invention enables improvisational direction adjustments in theaters and on stage. The system comprises the following components:
[0067] The user inputs basic information into the system, such as the theme, story structure, and character settings of the play to be performed on stage. This information primarily forms the basis for determining the direction of the production. Once the input is complete, the system uses this information to set up the initial configuration.
[0068] The terminal uses cameras and microphones installed on the stage to capture the performers' movements and voices in real time. This data is sent to the server as movement and voice information. In addition, sensors are connected to the terminal to detect audience reactions, thereby capturing feedback information such as laughter and applause.
[0069] The server is responsible for analyzing the acquired data. An advanced generative AI model runs on the server, analyzing the performers' actions and audio information. This analysis provides insights necessary for the performance, generating appropriate direction based on the stage situation. Furthermore, the server can extract insights from audience reactions and incorporate them into the performance in real time.
[0070] For example, in a scene where the audience bursts into laughter, the server analyzes the reaction and generates instructions to change the lighting to a warmer tone and play upbeat music. Furthermore, by using a generative AI model, it is possible to immediately suggest new actions to maintain the flow of the story even if an actor makes a mistake in their lines.
[0071] An example of a prompt to the generative AI model in this system is the instruction, "When the performer makes an unexpected move, please suggest a new story development that is appropriate for the situation." This prompt allows the system to suggest an improvisational story development, helping to ensure smooth performance on stage.
[0072] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0073] Step 1:
[0074] The user inputs the play's theme, story structure, and character settings into the system. This input is stored in the system as basic data and used for subsequent processing. This sets the fundamental direction for the production.
[0075] Step 2:
[0076] The terminal uses cameras and microphones installed on the stage to acquire real-time information on the performers' movements and voices. The input data includes characteristics such as the performers' gestures, voice tone, and speed. This information is immediately transmitted to the server.
[0077] Step 3:
[0078] The device also acquires audience reactions. Sensors collect laughter, applause, and other reactions from the audience and send them to a server as reaction data. Based on this input, the system evaluates which scenes are effective for the audience.
[0079] Step 4:
[0080] The server receives motion information, audio information, and audience reactions transmitted from the terminals, and analyzes the data using a generative AI model. Based on the input data, it extracts insights into the performers' movements and audience feedback. This provides analytical results that allow for the determination of what kind of staging is necessary.
[0081] Step 5:
[0082] Based on the analysis results, the server generates instructions to adjust the stage equipment, including lighting and sound, using control devices. Specifically, for example, it might instruct the lighting to be dimmed and the music to be played softly during emotional scenes. This allows for the creation of an atmosphere appropriate to the scene.
[0083] Step 6:
[0084] If unexpected actions or improvisations occur, the server uses a generation mechanism to create appropriate corrective instructions. Based on the input data, new action plans are generated to naturally develop the scene while maintaining the continuity of the narrative.
[0085] Step 7:
[0086] Based on the generated results, users can fine-tune the performance as needed. This enables optimal performance tailored to real-time acting, enhancing the audience experience.
[0087] (Application Example 1)
[0088] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0089] Improving the customer experience within a store is difficult with static displays alone. Dynamically adjusting the displays in response to customer movements and reactions is essential. Furthermore, it's necessary to adapt to customer interest in specific products and emphasize their appeal.
[0090] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0091] In this invention, the server includes acquisition means for acquiring operational information and audio information, analysis means for performing analysis based on the information from the acquisition means, control means for controlling lighting and sound based on the analysis results, generation means for generating adaptive effects based on customer actions and audio information, and adjustment means for emphasizing product identification based on trends within the store. This makes it possible to dynamically enhance the store environment and the appeal of products in response to customer movements and reactions.
[0092] "Motion information" refers to data about the physical movements and actions of customers and spectators.
[0093] "Audio information" refers to data related to the voices of customers and spectators, as well as sounds within the store.
[0094] "Means of acquisition" refers to equipment and technology for collecting motion information and audio information.
[0095] "Analysis means" refers to equipment and software used to analyze acquired motion information and audio information and extract patterns and insights.
[0096] "Control means" refers to equipment and software used to adjust lighting and sound based on analysis results.
[0097] "Generation means" refers to equipment and software used to create adaptive presentations and instructions in response to customer actions and voice information.
[0098] "Adjustment measures" refer to technologies and devices used to highlight specific products in response to trends within a store or customer interest.
[0099] To implement this invention, a system for dynamically adjusting the in-store environment is required. This system consists of acquisition means, analysis means, control means, generation means, and adjustment means.
[0100] The server uses acquisition methods to collect motion and audio information through cameras and microphones installed in the store. For example, it captures actions such as customers stopping in front of product shelves and audio of conversations.
[0101] Next, the server uses analysis tools to perform a detailed analysis of the acquired behavioral and audio information. This analysis utilizes AI models and machine learning platforms such as TENSORFLOW®. This allows the server to understand customer behavior and reactions and extract necessary insights.
[0102] Based on the analysis results, the lighting and sound are adjusted in real time by the control system. For example, if many customers show interest in a particular product, the lighting will brighten and related music will play to highlight that product.
[0103] Furthermore, it is possible to create visuals that adapt to customer behavior using generation methods. This allows for the emphasis on specific products based on in-store trends, thereby improving the customer experience.
[0104] One concrete example is a store employee wearing smart glasses who walks around the store and provides information about specific products. This allows for immediate responses to customer questions and can increase their desire to purchase.
[0105] An example of a prompt message might be, "Which product is currently attracting the most attention in the store? Please suggest ways to highlight that product."
[0106] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0107] Step 1:
[0108] The terminal acquires motion and audio information through cameras and microphones installed in the store. Inputs include customer location, posture, voice volume, and frequency. Outputs are raw video and audio data. This allows for understanding customer movements within the store.
[0109] Step 2:
[0110] The server receives the acquired raw data and inputs it into the AI model using analysis tools. Data processing includes identifying customer attributes and behavioral patterns through image recognition and recognizing emotions through voice analysis. The output is detailed insights into customer trends and reactions.
[0111] Step 3:
[0112] The server's control system issues instructions to adjust the store's lighting and sound based on the analysis results. The input is insights obtained from the analysis system, and the output is settings for lighting color tone and brightness, and background music selection. For example, an instruction might be issued to brighten the area around a specific product.
[0113] Step 4:
[0114] The generation mechanism generates prompts to provide adaptive presentations and product information in response to changes in behavioral information and customer intentions. The inputs are the customer's current behavior and analysis results. The output is the information displayed to the store employee using smart glasses.
[0115] Step 5:
[0116] The user (store clerk) receives instructions from the server and provides appropriate guidance to customers through smart glasses. For example, based on a prompt such as, "Which product is currently attracting the most attention in the store? Please suggest ways to highlight that product," they can provide product descriptions to customers in real time.
[0117] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0118] This invention relates to an "improvisational AI performance" system that uses an emotion engine to recognize the emotions of performers and audience members in real time. By incorporating acquisition means, analysis means, control means, generation means, and an emotion engine, this system enables flexible adjustment of performances in improvisational theater. The program processing of this system will be explained in natural language below, along with specific examples.
[0119] First, the user (director) uses a terminal to input the play's theme, storyline, and character settings into the system. This information is stored in a database as basic data necessary for the play's progress. Next, the terminal uses cameras and microphones placed on the stage to capture the actors' movements and lines in real time. Simultaneously, an emotion engine analyzes this data and recognizes emotions from the actors' facial expressions and vocal characteristics.
[0120] This emotional information is received by a server and analyzed through an analysis system to determine the performance strategy based on the performer's current emotional state. The analysis results are interpreted by a control system, which generates instructions on how to effectively utilize the stage lighting and sound. For example, if a performer shows anger, the system can instruct the lighting to turn red to increase the sense of tension and play background music with emphasized bass.
[0121] The emotion engine can also capture audience reactions and amplify the performance effects in response to laughter, surprise, and other reactions. For example, when the audience laughs loudly, it can instruct the lighting to change to warmer colors and play positive background music.
[0122] Finally, the user (stage operator) can monitor whether the system's automatic adjustments are being performed accurately and make manual adjustments as needed. In this way, the present invention realizes sophisticated direction that is in line with the emotions of the performers and audience in improvisational theater, thereby improving the quality of the play.
[0123] The following describes the processing flow.
[0124] Step 1:
[0125] The user (director) uses a terminal to input the play's theme, story structure, and character settings into the system. This allows fundamental information for the play to be stored in a database.
[0126] Step 2:
[0127] The device captures the performers' movements and lines in real time through cameras and microphones installed on the stage. Simultaneously, audience responses are collected by voice sensors.
[0128] Step 3:
[0129] The emotion engine analyzes acquired video and audio data to recognize the emotional states of performers and audience members. For example, it infers emotions from changes in performers' facial expressions and tone of voice, and analyzes collective emotional trends from audience reactions.
[0130] Step 4:
[0131] The server thoroughly analyzes the emotional information transmitted from the emotion engine, along with the behavioral and audio information acquired by the terminal, using analysis tools. Based on the analysis results, it constructs a specific performance scenario.
[0132] Step 5:
[0133] The server uses a generation mechanism to create performance instructions based on the analysis results and transmits them to the control mechanism. These instructions include dynamic adjustments to lighting color tone and intensity, sound balance, and background music selection.
[0134] Step 6:
[0135] The terminal controls the equipment on stage and executes the specified effects according to instructions received from the server. This includes switching lighting and adjusting sound effects.
[0136] Step 7:
[0137] The user (stage operator) monitors whether the automated effects provided by the system are working correctly and makes manual adjustments if further fine-tuning is required.
[0138] This allows the system to create high-quality improvisational theater performances based on the emotions of both the performers and the audience.
[0139] (Example 2)
[0140] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0141] In theater, there is a challenge in improving the quality of improvisational plays by instantly grasping the actors' expressions and the audience's reactions, and adjusting the direction in real time. This challenge has been difficult to address with conventional technology, as it was challenging to accurately grasp the emotions of the actors and audience and reflect them in immediate changes to the direction.
[0142] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0143] In this invention, the server includes setting means for inputting the play's theme, storyline, and character settings; acquisition means for acquiring performer's movement information and voice information; emotion analysis means for analyzing emotional information; and generation means for generating control instructions for lighting and sound. This enables real-time adjustment of the performance based on the emotions of the performers and the audience.
[0144] The "setting method" refers to a function that allows performers to input themes, storylines, character settings, and other elements of an improvisational play into the system.
[0145] "Acquisition means" refers to a function for collecting information on performers' movements and voices on stage in real time.
[0146] "Emotional analysis means" refers to an analytical function that estimates the emotional state of a performer based on their actions and voice data.
[0147] The "generation means" refers to a function that generates control instructions for lighting and sound based on analyzed emotional information, and automatically adjusts the stage production.
[0148] The "audience reaction acquisition means" is a function that acquires audience reactions as data and transmits that information to an emotion analysis means for use in adjusting the performance.
[0149] This invention is a system that grasps the emotions of performers and audience in real time and dynamically adjusts the direction of improvisational theater. This system operates effectively by combining a server, terminals, and users.
[0150] First, the user (director or stage operator) inputs the play's theme, storyline, and character settings into the system via the terminal interface. This input information is stored in the server's database as basic data for the production.
[0151] The terminal uses sensors such as cameras and microphones installed on the stage to capture the performers' movements and voices in real time. This terminal has an emotion engine built in that analyzes the acquired raw data. Specifically, it uses image processing technology and voice analysis technology to read emotions from the performers' facial expressions and tone of voice. This makes it possible to determine what emotions the performers are currently expressing.
[0152] The server uses a generative AI model to calculate a performance strategy based on the performer's emotional information and basic data transmitted from the terminal. The server then generates automatic adjustment instructions for stage lighting and sound based on the generated strategy. For example, if a performer expresses anger, the server can issue instructions to turn on red lighting and play background music with emphasized bass.
[0153] The terminal also captures audience reactions using cameras and microphones, analyzing responses such as laughter and surprise. This information is sent to a server and used to fine-tune the performance. For example, if the audience is laughing hysterically, the server can instruct the system to change the lighting to a brighter, warmer tone and play positive background music.
[0154] The user (stage operator) monitors these automated adjustments and manually changes the settings as needed to ensure that the intended performance is being carried out.
[0155] As a concrete example, here is an example of a prompt sentence to be input to a generative AI model: "If the audience is crying during a sad scene for the protagonist, how should the lighting and background music be changed?"
[0156] This system enables timely and flexible direction that responds to the emotions of both performers and audience, thereby improving the quality of improvisational theater.
[0157] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0158] Step 1:
[0159] Users input initial data such as the play's theme, storyline, and character settings into the system via their terminal. This data is sent to the server and stored in a database as foundational data for the production. The input data serves as the basic criteria for future adjustments to the production.
[0160] Step 2:
[0161] The device uses cameras and microphones installed on the stage to capture the performers' movements and voices in real time. The cameras capture the performers' movements and facial expressions as video data, and the microphones collect dialogue and voice tones as audio data. This acquired data is temporarily stored within the device.
[0162] Step 3:
[0163] The device passes the acquired motion data and audio data to the emotion engine. The emotion engine processes the data using image processing algorithms and audio analysis techniques to estimate the performer's emotions. For example, changes in facial muscle movements and tone of voice can be used to classify the performer into a specific emotion such as "joy," "anger," or "sadness." This analysis result is returned to the device as emotion information.
[0164] Step 4:
[0165] The server receives emotional information transmitted from the terminal and uses a generative AI model to formulate a performance plan based on it. This plan includes specific instructions on what lighting and sound effects should be used. The generative AI model combines initial information stored in the database with the results of the emotional analysis to calculate the optimal performance and generate a set of instructions.
[0166] Step 5:
[0167] The server sends the generated performance instructions to the control system, which adjusts the stage lighting and sound in real time. For example, when an actor is expressing "anger," the stage lighting is changed to red, and a bass-heavy background music track is selected and played. This approach makes the actors' emotional expressions more effective for the audience.
[0168] Step 6:
[0169] The terminal further captures audience reactions using its camera and microphone, and performs emotion analysis based on this data. Laughter, surprise, applause, and other reactions are detected, and the results are sent back to the server. This audience emotion information is then used to fine-tune the stage production.
[0170] Step 7:
[0171] Users (stage operators) can monitor real-time performance adjustment information provided by the server and manually adjust lighting and sound as needed. This allows for further optimization of the performance in response to instantaneous audience reactions and changes in performers.
[0172] Through these steps, the entire system works in coordination, enabling flexible adjustments that respond to the emotions of both performers and audience members while maintaining high quality in the direction of improvisational theater.
[0173] (Application Example 2)
[0174] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0175] In recent years, physical stores have seen a growing demand for flexible environmental adjustments tailored to individual customer emotions in order to enhance the customer experience. However, current technology makes it difficult to accurately recognize customer emotions and automatically adjust store lighting and sound in real time based on those emotions. Therefore, a system that combines emotion recognition and environmental adjustment is necessary to improve the quality of the customer experience within a store.
[0176] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0177] In this invention, the server includes an acquisition means for acquiring information to estimate a person's emotional state, an analysis means for analyzing the emotional state based on the information obtained from the acquisition means, and a control means for controlling a light source and sound based on the emotional state from the analysis means. This enables dynamic environmental adjustments that are in line with the customer's emotions.
[0178] "A person's emotional state" refers to the mental or emotional state an individual is experiencing at a particular moment. This includes specific emotions such as joy, sadness, anger, and surprise.
[0179] "Means of acquisition" refers to means of obtaining information to estimate a person's emotional state. Specifically, this involves using input devices such as cameras and microphones.
[0180] "Analysis means" refers to methods for processing information to analyze a person's emotional state based on the acquired information. Image recognition technology and speech analysis technology may be used for this analysis.
[0181] "Control means" refers to methods for manipulating light sources and sounds based on the results of the analysis. This allows the living environment to be adjusted according to emotions.
[0182] "Generative means" refers to methods for generating specific instructions to adapt the environment to the customer's emotional state. This enables real-time environment configuration.
[0183] "Light sources and sound" refer to elements in physical space that affect vision and hearing, such as lighting equipment and speaker systems.
[0184] "Means of acquiring reactions" refers to means of detecting and acquiring a person's actions and reactions. This includes sensor and camera technologies.
[0185] The system that realizes this invention aims to identify a person's emotional state in real time and dynamically adjust the lighting and sound settings of a physical space, such as a store, accordingly. The system is centered around a server connected to a network.
[0186] The server acquires information such as facial expressions and voice tone of people inside the store through acquisition devices such as cameras and microphones. This acquisition method utilizes OpenCV, a video processing software, and Google's Cloud Speech-to-Text API for speech analysis. By integrating these devices and software, the emotional state of people is estimated, and data based on that estimation is collected.
[0187] The server then processes the acquired data using analysis tools to infer the emotional state. The emotion engine analyzes facial expressions and voice data to recognize specific emotions such as joy, sadness, and anger. The results of the analysis are interpreted by control tools within the server, which then generate environmental adjustment instructions corresponding to the emotional state.
[0188] The server controls the store's lighting and sound systems, adjusting the environmental settings. For example, if a customer smiles, the server instructs the sound system to change the lighting to a warmer color and play rhythmic music.
[0189] In this way, the server can adjust the environment in real time according to emotions. This invention is expected to further enrich the customer experience.
[0190] For example, when customers are in a good mood, the store can be filled with bright lighting and upbeat music to amplify their enjoyment. Another example of a prompt for a generative AI model would be: "Based on the customer's emotions, suggest the optimal lighting and background music settings. For example, if a customer is smiling, consider how to improve the store's atmosphere."
[0191] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0192] Step 1:
[0193] The server captures customers' facial expressions and voices through cameras and microphones placed within the store. The input consists of camera footage and microphone audio, which are collected as digital data. This data is temporarily stored in storage for subsequent analysis.
[0194] Step 2:
[0195] The server processes the collected video data using OpenCV to detect faces and estimate emotions from their expressions. The input is camera video data, and the output is emotional state (e.g., joy, surprise, anger). Emotions are quantified by analyzing facial shape and muscle movements.
[0196] Step 3:
[0197] The server uses the Google Cloud Speech-to-Text API to convert speech data into text and analyzes the tone of the speech. The input is microphone audio data, and the output is an estimate of the emotional state derived from the speech. Emotion is estimated based on volume and intonation.
[0198] Step 4:
[0199] The analyzed emotional information is integrated within the server to evaluate the overall emotional state. The input is the emotional estimation results from steps 2 and 3, and the output is the overall emotional state. This information is used to determine the next environmental adjustment step.
[0200] Step 5:
[0201] The server generates commands to adjust the lighting and sound settings within the store based on the integrated emotional state. The input is the overall emotional state, and the output is a specific control command. For example, if a positive emotion is detected, it recommends brighter lighting and upbeat music.
[0202] Step 6:
[0203] Upon receiving instructions from the server, the lighting and sound systems are automatically adjusted. The server sends control signals to these devices, ensuring that the environmental settings change instantly. Ultimately, a store environment that aligns with the customer's emotional needs is provided.
[0204] This series of steps enables dynamic environmental adjustments that respond to customer emotions.
[0205] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0206] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0207] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0208] [Second Embodiment]
[0209] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0210] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0211] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0212] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0213] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0214] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0215] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0216] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0217] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0218] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0219] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0220] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0221] The "improvisational AI direction" system according to the present invention is for dynamically adjusting the direction on a theatrical stage. This system consists of acquisition means, analysis means, control means, and generation means. The processing of the program in this system is described below in natural language.
[0222] First, the user (director) inputs the play's theme, story structure, and character settings into the system. This prepares the basic information for the play. Next, the terminal captures the actors' movements and conversations in real time through cameras and microphones installed on the stage. In addition, audience reactions are also captured by sensors.
[0223] This data is sent to a server and analyzed in detail by an analysis system. An AI model on the server analyzes the performers' movements and voice information to extract necessary insights regarding the performance. Based on these analysis results, the control system generates instructions to adjust lighting, sound, and other stage equipment in real time.
[0224] Furthermore, if an actor acts spontaneously or makes an unexpected move, the server uses a generation mechanism to instantly create appropriate corrective instructions. This allows for flexible responses that maintain the flow of the story while ensuring the performance is not interrupted.
[0225] For example, if loud laughter erupts from the audience, the server analyzes the reaction and instructs the control system to change the lighting to a warmer color and play upbeat background music at that point. Furthermore, if an actor makes a mistake with their lines, the generation system suggests alternative actions to correct the story in a natural flow. In this way, the present invention improves the quality of direction in improvisational theater.
[0226] The following describes the processing flow.
[0227] Step 1:
[0228] The user (director) uses a terminal to input the play's theme, storyline, character settings, and basic directorial guidelines into the system. The entered information is stored in the system's database and used by the AI for subsequent processing.
[0229] Step 2:
[0230] The devices (cameras and microphones installed on the stage) capture the performers' movements and lines in real time. They also capture audience reactions such as laughter and applause.
[0231] Step 3:
[0232] The server inputs the real-time data stream provided from the terminal into an analysis device, and the performer's actions and voice are analyzed using an AI model. Here, the performer's ad-libs and improvisational actions, the content of the lines, and audience reaction data are identified.
[0233] Step 4:
[0234] Based on the insights gained from the analysis results, the server generates staging instructions that are appropriate for the current stage situation. Specifically, it determines staging elements such as lighting patterns, sound adjustments, and background music selection.
[0235] Step 5:
[0236] The terminal receives performance instructions sent from the server and applies those instructions to the stage's lighting and sound systems. This allows for instantaneous adjustments to lighting color, brightness, and volume, as well as background music switching.
[0237] Step 6:
[0238] The user (stage operator) can monitor whether the automated effects provided by the system are being executed appropriately and make manual adjustments as needed.
[0239] Step 7:
[0240] The server collects feedback data from performers and audience members in response to their reactions, and stores it as points for improvement in the next session, in order to update future production strategies.
[0241] (Example 1)
[0242] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0243] In improvisational theater and live performances, there is a need to respond flexibly in real time to the performers' spontaneous actions and the audience's reactions, thereby improving the quality of the performance. However, conventional systems have the challenge of being unable to respond to changing situations and making appropriate adjustments to maintain the flow of the performance.
[0244] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0245] In this invention, the server includes means for acquiring operational information and audio information, means for performing analysis based on the operational information and audio information from the acquisition means, and means for controlling lighting and sound based on the results of the analysis means. This makes it possible to adjust the performance appropriately in real time even in response to improvisational situational changes.
[0246] "Motion information" refers to data that shows the physical movements of performers and people on stage, such as gestures, movements, and postures.
[0247] "Audio information" refers to data that represents acoustic signals, including sounds from the performers and their surroundings.
[0248] "Means of acquisition" refers to devices and technologies for collecting operational information and audio information.
[0249] "Analytical means" refers to the techniques and processes used to analyze acquired information and derive meaningful results.
[0250] "Control means" refers to devices and technologies used to operate stage equipment such as lighting and sound based on analysis results.
[0251] "Modification instructions" refer to instructions for appropriately changing the stage production in response to the performers' improvisational actions or the audience's reactions.
[0252] "Audience reaction" refers to information that shows the visual, auditory, and other feedback from the audience during a performance.
[0253] "Monitoring methods" refer to devices and technologies used to monitor audience reactions in real time and acquire data.
[0254] "Generative means" refers to the techniques and processes for generating staging and modifications that are appropriate to the improvisational situation.
[0255] The "improvisational AI direction" system of the present invention enables improvisational direction adjustments in theaters and on stage. The system comprises the following components:
[0256] The user inputs basic information into the system, such as the theme, story structure, and character settings of the play to be performed on stage. This information primarily forms the basis for determining the direction of the production. Once the input is complete, the system uses this information to set up the initial configuration.
[0257] The terminal uses cameras and microphones installed on the stage to capture the performers' movements and voices in real time. This data is sent to the server as movement and voice information. In addition, sensors are connected to the terminal to detect audience reactions, thereby capturing feedback information such as laughter and applause.
[0258] The server is responsible for analyzing the acquired data. An advanced generative AI model runs on the server, analyzing the performers' actions and audio information. This analysis provides insights necessary for the performance, generating appropriate direction based on the stage situation. Furthermore, the server can extract insights from audience reactions and incorporate them into the performance in real time.
[0259] For example, in a scene where the audience bursts into laughter, the server analyzes the reaction and generates instructions to change the lighting to a warmer tone and play upbeat music. Furthermore, by using a generative AI model, it is possible to immediately suggest new actions to maintain the flow of the story even if an actor makes a mistake in their lines.
[0260] An example of a prompt to the generative AI model in this system is the instruction, "When the performer makes an unexpected move, please suggest a new story development that is appropriate for the situation." This prompt allows the system to suggest an improvisational story development, helping to ensure smooth performance on stage.
[0261] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0262] Step 1:
[0263] The user inputs the play's theme, story structure, and character settings into the system. This input is stored in the system as basic data and used for subsequent processing. This sets the fundamental direction for the production.
[0264] Step 2:
[0265] The terminal uses cameras and microphones installed on the stage to acquire real-time information on the performers' movements and voices. The input data includes characteristics such as the performers' gestures, voice tone, and speed. This information is immediately transmitted to the server.
[0266] Step 3:
[0267] The device also acquires audience reactions. Sensors collect laughter, applause, and other reactions from the audience and send them to a server as reaction data. Based on this input, the system evaluates which scenes are effective for the audience.
[0268] Step 4:
[0269] The server receives motion information, audio information, and audience reactions transmitted from the terminals, and analyzes the data using a generative AI model. Based on the input data, it extracts insights into the performers' movements and audience feedback. This provides analytical results that allow for the determination of what kind of staging is necessary.
[0270] Step 5:
[0271] Based on the analysis results, the server generates instructions to adjust the stage equipment, including lighting and sound, using control devices. Specifically, for example, it might instruct the lighting to be dimmed and the music to be played softly during emotional scenes. This allows for the creation of an atmosphere appropriate to the scene.
[0272] Step 6:
[0273] If unexpected actions or improvisations occur, the server uses a generation mechanism to create appropriate corrective instructions. Based on the input data, new action plans are generated to naturally develop the scene while maintaining the continuity of the narrative.
[0274] Step 7:
[0275] Based on the generated results, users can fine-tune the performance as needed. This enables optimal performance tailored to real-time acting, enhancing the audience experience.
[0276] (Application Example 1)
[0277] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0278] The customer experience within a store is difficult to effectively improve with static presentations. There is a need to dynamically adjust the presentation according to the movements and reactions of customers. Also, it is necessary to adapt to the customers' interest in specific products and emphasize the attractiveness of the products.
[0279] The specific processing by the specific processing unit 290 of the data processing apparatus 12 in Application Example 1 is realized by the following respective means.
[0280] In this invention, the server includes an acquisition means for acquiring operation information and voice information, an analysis means for performing analysis based on the information from the acquisition means, a control means for controlling lighting and sound based on the analysis result, a generation means for generating an adaptive presentation based on the movements and voice information of customers, and an adjustment means for emphasizing the specification of products based on the trends within the store. Thereby, it becomes possible to dynamically present the store environment and the attractiveness of products according to the movements and reactions of customers.
[0281] "Operation information" refers to data regarding the movements and actions of customers or audiences.
[0282] "Voice information" refers to data regarding the voices emitted by customers or audiences or the sounds within the store.
[0283] "Acquisition means" refers to equipment or technologies for collecting operation information and voice information.
[0284] "Analysis means" refers to equipment or software for analyzing the acquired operation information and voice information and extracting patterns and insights.
[0285] "Control means" refers to equipment or software for adjusting lighting and sound based on the analysis result.
[0286] "Generation means" refers to equipment or software for creating an adaptive presentation or instruction according to the movements and voice information of customers.
[0287] "Adjustment measures" refer to technologies and devices used to highlight specific products in response to trends within a store or customer interest.
[0288] To implement this invention, a system for dynamically adjusting the in-store environment is required. This system consists of acquisition means, analysis means, control means, generation means, and adjustment means.
[0289] The server uses acquisition methods to collect motion and audio information through cameras and microphones installed in the store. For example, it captures actions such as customers stopping in front of product shelves and audio of conversations.
[0290] Next, the server uses analysis tools to perform a detailed analysis of the acquired behavioral and audio information. This analysis utilizes AI models and machine learning platforms such as TensorFlow. This allows the server to understand customer behavior and reactions and extract necessary insights.
[0291] Based on the analysis results, the lighting and sound are adjusted in real time by the control system. For example, if many customers show interest in a particular product, the lighting will brighten and related music will play to highlight that product.
[0292] Furthermore, it is possible to create visuals that adapt to customer behavior using generation methods. This allows for the emphasis on specific products based on in-store trends, thereby improving the customer experience.
[0293] One concrete example is a store employee wearing smart glasses who walks around the store and provides information about specific products. This allows them to respond immediately to customer questions and increase their desire to purchase.
[0294] An example of a prompt message might be, "Which product is currently attracting the most attention in the store? Please suggest ways to highlight that product."
[0295] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0296] Step 1:
[0297] The terminal acquires motion and audio information through cameras and microphones installed in the store. Inputs include customer location, posture, voice volume, and frequency. Outputs are raw video and audio data. This allows for understanding customer movements within the store.
[0298] Step 2:
[0299] The server receives the acquired raw data and inputs it into the AI model using analysis tools. Data processing includes identifying customer attributes and behavioral patterns through image recognition and recognizing emotions through voice analysis. The output is detailed insights into customer trends and reactions.
[0300] Step 3:
[0301] The server's control system issues instructions to adjust the store's lighting and sound based on the analysis results. The input is insights obtained from the analysis system, and the output is settings for lighting color tone and brightness, and background music selection. For example, an instruction might be issued to brighten the area around a specific product.
[0302] Step 4:
[0303] The generation mechanism generates prompts to provide adaptive presentations and product information in response to changes in behavioral information and customer intentions. The inputs are the customer's current behavior and analysis results. The output is the information displayed to the store employee using smart glasses.
[0304] Step 5:
[0305] The user (store clerk) receives instructions from the server and provides appropriate guidance to customers through smart glasses. For example, based on a prompt sentence such as "Which product in the store is currently attracting the most attention? Please propose a production to highlight that product.", product explanations can be provided to customers in real time.
[0306] Furthermore, an emotion engine for estimating the user's emotions may be combined. That is, the specific processing unit 290 may estimate the user's emotions using the emotion recognition model 59 and perform specific processing using the user's emotions.
[0307] The present invention relates to an "impromptu AI production" system using an emotion engine that recognizes the emotions of performers and audiences in real time. By incorporating an acquisition means, an analysis means, a control means, a generation means, and an emotion engine, this system enables flexible production adjustment in impromptu theater. Hereinafter, the processing of the program of this system will be described in natural language and introduced with specific examples.
[0308] First, the user (producer) uses the terminal to input the theme, storyline, and character settings of the play into the system. This information is stored in the database as basic data required for the progress of the play. Subsequently, the terminal uses the camera and microphone installed on the stage to acquire the actions and lines of the performer in real time. At the same time, the emotion engine analyzes this data and recognizes the emotions from the performer's expressions and voice characteristics.
[0309] This emotion information is received by the server and analyzed through the analysis means to analyze the production guidelines based on the performer's current emotional state. The analysis result is interpreted by the control means to generate instructions on how to effectively utilize the lighting and sound of the stage. As a specific example, when the performer shows anger, an instruction can be issued to increase the sense of tension by making the lighting red and playing a BGM that emphasizes low tones.
[0310] The emotion engine can also capture audience reactions and amplify the performance effects in response to laughter, surprise, and other reactions. For example, when the audience laughs loudly, it can instruct the lighting to change to warmer colors and play positive background music.
[0311] Finally, the user (stage operator) can monitor whether the system's automatic adjustments are being performed accurately and make manual adjustments as needed. In this way, the present invention realizes sophisticated direction that is in line with the emotions of the performers and audience in improvisational theater, thereby improving the quality of the play.
[0312] The following describes the processing flow.
[0313] Step 1:
[0314] The user (director) uses a terminal to input the play's theme, story structure, and character settings into the system. This allows fundamental information for the play to be stored in a database.
[0315] Step 2:
[0316] The device captures the performers' movements and lines in real time through cameras and microphones installed on the stage. Simultaneously, audience responses are collected by voice sensors.
[0317] Step 3:
[0318] The emotion engine analyzes acquired video and audio data to recognize the emotional states of performers and audience members. For example, it infers emotions from changes in performers' facial expressions and tone of voice, and analyzes collective emotional trends from audience reactions.
[0319] Step 4:
[0320] The server thoroughly analyzes the emotional information transmitted from the emotion engine, along with the behavioral and audio information acquired by the terminal, using analysis tools. Based on the analysis results, it constructs a specific performance scenario.
[0321] Step 5:
[0322] The server uses a generation mechanism to create performance instructions based on the analysis results and transmits them to the control mechanism. These instructions include dynamic adjustments to lighting color tone and intensity, sound balance, and background music selection.
[0323] Step 6:
[0324] The terminal controls the equipment on stage and executes the specified effects according to instructions received from the server. This includes switching lighting and adjusting sound effects.
[0325] Step 7:
[0326] The user (stage operator) monitors whether the automated effects provided by the system are working correctly and makes manual adjustments if further fine-tuning is required.
[0327] This allows the system to create high-quality improvisational theater performances based on the emotions of both the performers and the audience.
[0328] (Example 2)
[0329] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0330] In theater, there is a challenge in improving the quality of improvisational plays by instantly grasping the actors' expressions and the audience's reactions, and adjusting the direction in real time. This challenge has been difficult to address with conventional technology, as it was challenging to accurately grasp the emotions of the actors and audience and reflect them in immediate changes to the direction.
[0331] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0332] In this invention, the server includes setting means for inputting the play's theme, storyline, and character settings; acquisition means for acquiring performer's movement information and voice information; emotion analysis means for analyzing emotional information; and generation means for generating control instructions for lighting and sound. This enables real-time adjustment of the performance based on the emotions of the performers and the audience.
[0333] The "setting method" refers to a function that allows performers to input themes, storylines, character settings, and other elements of an improvisational play into the system.
[0334] "Acquisition means" refers to a function for collecting information on performers' movements and voices on stage in real time.
[0335] "Emotional analysis means" refers to an analytical function that estimates the emotional state of a performer based on their actions and voice data.
[0336] The "generation means" refers to a function that generates control instructions for lighting and sound based on analyzed emotional information, and automatically adjusts the stage production.
[0337] The "audience reaction acquisition means" is a function that acquires audience reactions as data and transmits that information to an emotion analysis means for use in adjusting the performance.
[0338] This invention is a system that grasps the emotions of performers and audience in real time and dynamically adjusts the direction of improvisational theater. This system operates effectively by combining a server, terminals, and users.
[0339] First, the user (director or stage operator) inputs the play's theme, storyline, and character settings into the system via the terminal interface. This input information is stored in the server's database as basic data for the production.
[0340] The terminal uses sensors such as cameras and microphones installed on the stage to capture the performers' movements and voices in real time. This terminal has an emotion engine built in that analyzes the acquired raw data. Specifically, it uses image processing technology and voice analysis technology to read emotions from the performers' facial expressions and tone of voice. This makes it possible to determine what emotions the performers are currently expressing.
[0341] The server uses a generative AI model to calculate a performance strategy based on the performer's emotional information and basic data transmitted from the terminal. The server then generates automatic adjustment instructions for stage lighting and sound based on the generated strategy. For example, if a performer expresses anger, the server can issue instructions to turn on red lighting and play background music with emphasized bass.
[0342] The terminal also captures audience reactions using cameras and microphones, analyzing responses such as laughter and surprise. This information is sent to a server and used to fine-tune the performance. For example, if the audience is laughing hysterically, the server can instruct the system to change the lighting to a brighter, warmer tone and play positive background music.
[0343] The user (stage operator) monitors these automated adjustments and manually changes the settings as needed to ensure that the intended performance is being carried out.
[0344] As a concrete example, here is an example of a prompt sentence to be input to a generative AI model: "If the audience is crying during a sad scene for the protagonist, how should the lighting and background music be changed?"
[0345] This system enables timely and flexible direction that responds to the emotions of both performers and audience, thereby improving the quality of improvisational theater.
[0346] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0347] Step 1:
[0348] Users input initial data such as the play's theme, storyline, and character settings into the system via their terminal. This data is sent to the server and stored in a database as foundational data for the production. The input data serves as the basic criteria for future adjustments to the production.
[0349] Step 2:
[0350] The device uses cameras and microphones installed on the stage to capture the performers' movements and voices in real time. The cameras capture the performers' movements and facial expressions as video data, and the microphones collect dialogue and voice tones as audio data. This acquired data is temporarily stored within the device.
[0351] Step 3:
[0352] The device passes the acquired motion data and audio data to the emotion engine. The emotion engine processes the data using image processing algorithms and audio analysis techniques to estimate the performer's emotions. For example, changes in facial muscle movements and tone of voice can be used to classify the performer into a specific emotion such as "joy," "anger," or "sadness." This analysis result is returned to the device as emotion information.
[0353] Step 4:
[0354] The server receives emotional information transmitted from the terminal and uses a generative AI model to formulate a performance plan based on it. This plan includes specific instructions on what lighting and sound effects should be used. The generative AI model combines initial information stored in the database with the results of the emotional analysis to calculate the optimal performance and generate a set of instructions.
[0355] Step 5:
[0356] The server sends the generated performance instructions to the control system, which adjusts the stage lighting and sound in real time. For example, when an actor is expressing "anger," the stage lighting is changed to red, and a bass-heavy background music track is selected and played. This approach makes the actors' emotional expressions more effective for the audience.
[0357] Step 6:
[0358] The terminal further captures audience reactions using a camera and microphone, and performs emotion analysis based on this data. Laughter, surprise, applause, etc., are detected, and the results are sent back to the server. This audience emotion information is then used to fine-tune the stage production.
[0359] Step 7:
[0360] Users (stage operators) can monitor real-time performance adjustment information provided by the server and manually adjust lighting and sound as needed. This allows for further optimization of the performance in response to instantaneous audience reactions and changes in performers.
[0361] Through these steps, the entire system works in coordination, enabling flexible adjustments that respond to the emotions of both performers and audience members while maintaining high quality in the direction of improvisational theater.
[0362] (Application Example 2)
[0363] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0364] In recent years, physical stores have seen a growing demand for flexible environmental adjustments tailored to individual customer emotions in order to enhance the customer experience. However, current technology makes it difficult to accurately recognize customer emotions and automatically adjust store lighting and sound in real time based on those emotions. Therefore, a system that combines emotion recognition and environmental adjustment is necessary to improve the quality of the customer experience within a store.
[0365] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0366] In this invention, the server includes an acquisition means for acquiring information to estimate a person's emotional state, an analysis means for analyzing the emotional state based on the information obtained from the acquisition means, and a control means for controlling a light source and sound based on the emotional state from the analysis means. This enables dynamic environmental adjustments that are in line with the customer's emotions.
[0367] "A person's emotional state" refers to the mental or emotional state an individual is experiencing at a particular moment. This includes specific emotions such as joy, sadness, anger, and surprise.
[0368] "Means of acquisition" refers to means of obtaining information to estimate a person's emotional state. Specifically, this involves using input devices such as cameras and microphones.
[0369] "Analysis means" refers to methods for processing information to analyze a person's emotional state based on the acquired information. Image recognition technology and speech analysis technology may be used for this analysis.
[0370] "Control means" refers to methods for manipulating light sources and sounds based on the results of the analysis. This allows the living environment to be adjusted according to emotions.
[0371] "Generative means" refers to methods for generating specific instructions to adapt the environment to the customer's emotional state. This enables real-time environment configuration.
[0372] "Light sources and sound" refer to elements in physical space that affect vision and hearing, such as lighting equipment and speaker systems.
[0373] "Means of acquiring reactions" refers to means of detecting and acquiring a person's actions and reactions. This includes sensor and camera technologies.
[0374] The system that realizes this invention aims to identify a person's emotional state in real time and dynamically adjust the lighting and sound settings of a physical space, such as a store, accordingly. The system is centered around a server connected to a network.
[0375] The server acquires information such as facial expressions and voice tone of people inside the store through acquisition devices such as cameras and microphones. This acquisition method utilizes OpenCV, a video processing software, and the Google Cloud Speech-to-Text API for speech analysis. By integrating these devices and software, the emotional state of people is estimated, and data based on that estimation is collected.
[0376] The server then processes the acquired data using analysis tools to infer the emotional state. The emotion engine analyzes facial expressions and voice data to recognize specific emotions such as joy, sadness, and anger. The results of the analysis are interpreted by control tools within the server, which then generate environmental adjustment instructions corresponding to the emotional state.
[0377] The server controls the store's lighting and sound systems, adjusting the environmental settings. For example, if a customer smiles, the server instructs the sound system to change the lighting to a warmer color and play rhythmic music.
[0378] In this way, the server can adjust the environment in real time according to emotions. This invention is expected to further enrich the customer experience.
[0379] For example, when customers are in a good mood, the store can be filled with bright lighting and upbeat music to amplify their enjoyment. Another example of a prompt for a generative AI model would be: "Based on the customer's emotions, suggest the optimal lighting and background music settings. For example, if a customer is smiling, consider how to improve the store's atmosphere."
[0380] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0381] Step 1:
[0382] The server captures customers' facial expressions and voices through cameras and microphones placed within the store. The input consists of camera footage and microphone audio, which are collected as digital data. This data is temporarily stored in storage for subsequent analysis.
[0383] Step 2:
[0384] The server processes the collected video data using OpenCV to detect faces and estimate emotions from their expressions. The input is camera video data, and the output is emotional state (e.g., joy, surprise, anger). Emotions are quantified by analyzing facial shape and muscle movements.
[0385] Step 3:
[0386] The server uses the Google Cloud Speech-to-Text API to convert speech data into text and analyzes the tone of the speech. The input is microphone audio data, and the output is an estimate of the emotional state derived from the speech. Emotion is estimated based on volume and intonation.
[0387] Step 4:
[0388] The analyzed emotional information is integrated within the server to evaluate the overall emotional state. The input is the emotional estimation results from steps 2 and 3, and the output is the overall emotional state. This information is used to determine the next environmental adjustment step.
[0389] Step 5:
[0390] The server generates commands to adjust the lighting and sound settings within the store based on the integrated emotional state. The input is the overall emotional state, and the output is a specific control command. For example, if a positive emotion is detected, it recommends brighter lighting and upbeat music.
[0391] Step 6:
[0392] Upon receiving instructions from the server, the lighting and sound systems are automatically adjusted. The server sends control signals to these devices, ensuring that the environmental settings change instantly. Ultimately, a store environment that aligns with the customer's emotional needs is provided.
[0393] This series of steps enables dynamic environmental adjustments that respond to customer emotions.
[0394] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0395] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0396] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0397] [Third Embodiment]
[0398] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0399] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0400] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0401] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0402] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0403] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0404] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0405] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0406] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0407] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0408] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0409] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0410] The "improvisational AI direction" system according to the present invention is for dynamically adjusting the direction on a theatrical stage. This system consists of acquisition means, analysis means, control means, and generation means. The processing of the program in this system is described below in natural language.
[0411] First, the user (director) inputs the play's theme, story structure, and character settings into the system. This prepares the basic information for the play. Next, the terminal captures the actors' movements and conversations in real time through cameras and microphones installed on the stage. In addition, audience reactions are also captured by sensors.
[0412] This data is sent to a server and analyzed in detail by an analysis system. An AI model on the server analyzes the performers' movements and voice information to extract necessary insights regarding the performance. Based on these analysis results, the control system generates instructions to adjust lighting, sound, and other stage equipment in real time.
[0413] Furthermore, if an actor acts spontaneously or makes an unexpected move, the server uses a generation mechanism to instantly create appropriate corrective instructions. This allows for flexible responses that maintain the flow of the story while ensuring the performance is not interrupted.
[0414] For example, if loud laughter erupts from the audience, the server analyzes the reaction and instructs the control system to change the lighting to a warmer color and play upbeat background music at that point. Furthermore, if an actor makes a mistake with their lines, the generation system suggests alternative actions to correct the story in a natural flow. In this way, the present invention improves the quality of direction in improvisational theater.
[0415] The following describes the processing flow.
[0416] Step 1:
[0417] The user (director) uses a terminal to input the play's theme, storyline, character settings, and basic directorial guidelines into the system. The entered information is stored in the system's database and used by the AI for subsequent processing.
[0418] Step 2:
[0419] The devices (cameras and microphones installed on the stage) capture the performers' movements and lines in real time. They also capture audience reactions such as laughter and applause.
[0420] Step 3:
[0421] The server inputs the real-time data stream provided from the terminal into an analysis device, and the performer's actions and voice are analyzed using an AI model. Here, the performer's ad-libs and improvisational actions, the content of the lines, and audience reaction data are identified.
[0422] Step 4:
[0423] Based on the insights gained from the analysis results, the server generates staging instructions that are appropriate for the current stage situation. Specifically, it determines staging elements such as lighting patterns, sound adjustments, and background music selection.
[0424] Step 5:
[0425] The terminal receives performance instructions sent from the server and applies those instructions to the stage's lighting and sound systems. This allows for instantaneous adjustments to lighting color, brightness, and volume, as well as background music switching.
[0426] Step 6:
[0427] The user (stage operator) can monitor whether the automated effects provided by the system are being executed appropriately and make manual adjustments as needed.
[0428] Step 7:
[0429] The server collects feedback data from performers and audience members in response to their reactions, and stores it as points for improvement in the next session, in order to update future production strategies.
[0430] (Example 1)
[0431] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0432] In improvisational theater and live performances, there is a need to respond flexibly in real time to the performers' spontaneous actions and the audience's reactions, thereby improving the quality of the performance. However, conventional systems have the challenge of being unable to respond to changing situations and making appropriate adjustments to maintain the flow of the performance.
[0433] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0434] In this invention, the server includes means for acquiring operational information and audio information, means for performing analysis based on the operational information and audio information from the acquisition means, and means for controlling lighting and sound based on the results of the analysis means. This makes it possible to adjust the performance appropriately in real time even in response to improvisational situational changes.
[0435] "Motion information" refers to data that shows the physical movements of performers and people on stage, such as gestures, movements, and postures.
[0436] "Audio information" refers to data that represents acoustic signals, including sounds from the performers and their surroundings.
[0437] "Means of acquisition" refers to devices and technologies for collecting operational information and audio information.
[0438] "Analytical means" refers to the techniques and processes used to analyze acquired information and derive meaningful results.
[0439] "Control means" refers to devices and technologies used to operate stage equipment such as lighting and sound based on analysis results.
[0440] "Modification instructions" refer to instructions for appropriately changing the stage production in response to the performers' improvisational actions or the audience's reactions.
[0441] "Audience reaction" refers to information that shows the visual, auditory, and other feedback from the audience during a performance.
[0442] "Monitoring methods" refer to devices and technologies used to monitor audience reactions in real time and acquire data.
[0443] "Generative means" refers to the techniques and processes for generating staging and modifications that are appropriate to the improvisational situation.
[0444] The "improvisational AI direction" system of the present invention enables improvisational direction adjustments in theaters and on stage. The system comprises the following components:
[0445] The user inputs basic information into the system, such as the theme, story structure, and character settings of the play to be performed on stage. This information primarily forms the basis for determining the direction of the production. Once the input is complete, the system uses this information to set up the initial configuration.
[0446] The terminal uses cameras and microphones installed on the stage to capture the performers' movements and voices in real time. This data is sent to the server as movement and voice information. In addition, sensors are connected to the terminal to detect audience reactions, thereby capturing feedback information such as laughter and applause.
[0447] The server is responsible for analyzing the acquired data. An advanced generative AI model runs on the server, analyzing the performers' actions and audio information. This analysis provides insights necessary for the performance, generating appropriate direction based on the stage situation. Furthermore, the server can extract insights from audience reactions and incorporate them into the performance in real time.
[0448] For example, in a scene where the audience bursts into laughter, the server analyzes the reaction and generates instructions to change the lighting to a warmer tone and play upbeat music. Furthermore, by using a generative AI model, it is possible to immediately suggest new actions to maintain the flow of the story even if an actor makes a mistake in their lines.
[0449] An example of a prompt to the generative AI model in this system is the instruction, "When the performer makes an unexpected move, please suggest a new story development that is appropriate for the situation." This prompt allows the system to suggest an improvisational story development, helping to ensure smooth performance on stage.
[0450] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0451] Step 1:
[0452] The user inputs the play's theme, story structure, and character settings into the system. This input is stored in the system as basic data and used for subsequent processing. This sets the fundamental direction for the production.
[0453] Step 2:
[0454] The terminal uses cameras and microphones installed on the stage to acquire real-time information on the performers' movements and voices. The input data includes characteristics such as the performers' gestures, voice tone, and speed. This information is immediately transmitted to the server.
[0455] Step 3:
[0456] The device also acquires audience reactions. Sensors collect laughter, applause, and other reactions from the audience and send them to a server as reaction data. Based on this input, the system evaluates which scenes are effective for the audience.
[0457] Step 4:
[0458] The server receives motion information, audio information, and audience reactions transmitted from the terminals, and analyzes the data using a generative AI model. Based on the input data, it extracts insights into the performers' movements and audience feedback. This provides analytical results that allow for the determination of what kind of staging is necessary.
[0459] Step 5:
[0460] Based on the analysis results, the server generates instructions to adjust the stage equipment, including lighting and sound, using control devices. Specifically, for example, it might instruct the lighting to be dimmed and the music to be played softly during emotional scenes. This allows for the creation of an atmosphere appropriate to the scene.
[0461] Step 6:
[0462] If unexpected actions or improvisations occur, the server uses a generation mechanism to create appropriate corrective instructions. Based on the input data, new action plans are generated to naturally develop the scene while maintaining the continuity of the narrative.
[0463] Step 7:
[0464] Based on the generated results, users can fine-tune the performance as needed. This enables optimal performance tailored to real-time acting, enhancing the audience experience.
[0465] (Application Example 1)
[0466] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0467] Improving the customer experience within a store is difficult with static displays alone. Dynamically adjusting the displays in response to customer movements and reactions is essential. Furthermore, it's necessary to adapt to customer interest in specific products and emphasize their appeal.
[0468] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0469] In this invention, the server includes acquisition means for acquiring operational information and audio information, analysis means for performing analysis based on the information from the acquisition means, control means for controlling lighting and sound based on the analysis results, generation means for generating adaptive effects based on customer actions and audio information, and adjustment means for emphasizing product identification based on trends within the store. This makes it possible to dynamically enhance the store environment and the appeal of products in response to customer movements and reactions.
[0470] "Motion information" refers to data about the physical movements and actions of customers and spectators.
[0471] "Audio information" refers to data related to the voices of customers and spectators, as well as sounds within the store.
[0472] "Means of acquisition" refers to equipment and technology for collecting motion information and audio information.
[0473] "Analysis means" refers to equipment and software used to analyze acquired motion information and audio information and extract patterns and insights.
[0474] "Control means" refers to equipment and software used to adjust lighting and sound based on analysis results.
[0475] "Generation means" refers to equipment and software used to create adaptive presentations and instructions in response to customer actions and voice information.
[0476] "Adjustment measures" refer to technologies and devices used to highlight specific products in response to trends within a store or customer interest.
[0477] To implement this invention, a system for dynamically adjusting the in-store environment is required. This system consists of acquisition means, analysis means, control means, generation means, and adjustment means.
[0478] The server uses acquisition methods to collect motion and audio information through cameras and microphones installed in the store. For example, it captures actions such as customers stopping in front of product shelves and audio of conversations.
[0479] Next, the server uses analysis tools to perform a detailed analysis of the acquired behavioral and audio information. This analysis utilizes AI models and machine learning platforms such as TensorFlow. This allows the server to understand customer behavior and reactions and extract necessary insights.
[0480] Based on the analysis results, the lighting and sound are adjusted in real time by the control system. For example, if many customers show interest in a particular product, the lighting will brighten and related music will play to highlight that product.
[0481] Furthermore, it is possible to create visuals that adapt to customer behavior using generation methods. This allows for the emphasis on specific products based on in-store trends, thereby improving the customer experience.
[0482] One concrete example is a store employee wearing smart glasses who walks around the store and provides information about specific products. This allows them to respond immediately to customer questions and increase their desire to purchase.
[0483] An example of a prompt message might be, "Which product is currently attracting the most attention in the store? Please suggest ways to highlight that product."
[0484] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0485] Step 1:
[0486] The terminal acquires motion and audio information through cameras and microphones installed in the store. Inputs include customer location, posture, voice volume, and frequency. Outputs are raw video and audio data. This allows for understanding customer movements within the store.
[0487] Step 2:
[0488] The server receives the acquired raw data and inputs it into the AI model using analysis tools. Data processing includes identifying customer attributes and behavioral patterns through image recognition and recognizing emotions through voice analysis. The output is detailed insights into customer trends and reactions.
[0489] Step 3:
[0490] The server's control system issues instructions to adjust the store's lighting and sound based on the analysis results. The input is insights obtained from the analysis system, and the output is settings for lighting color tone and brightness, and background music selection. For example, an instruction might be issued to brighten the area around a specific product.
[0491] Step 4:
[0492] The generation mechanism generates prompts to provide adaptive presentations and product information in response to changes in behavioral information and customer intentions. The inputs are the customer's current behavior and analysis results. The output is the information displayed to the store employee using smart glasses.
[0493] Step 5:
[0494] The user (store clerk) receives instructions from the server and provides appropriate guidance to customers through smart glasses. For example, based on a prompt such as, "Which product is currently attracting the most attention in the store? Please suggest ways to highlight that product," they can provide product descriptions to customers in real time.
[0495] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0496] This invention relates to an "improvisational AI performance" system that uses an emotion engine to recognize the emotions of performers and audience members in real time. By incorporating acquisition means, analysis means, control means, generation means, and an emotion engine, this system enables flexible adjustment of performances in improvisational theater. The program processing of this system will be explained in natural language below, along with specific examples.
[0497] First, the user (director) uses a terminal to input the play's theme, storyline, and character settings into the system. This information is stored in a database as basic data necessary for the play's progress. Next, the terminal uses cameras and microphones placed on the stage to capture the actors' movements and lines in real time. Simultaneously, an emotion engine analyzes this data and recognizes emotions from the actors' facial expressions and vocal characteristics.
[0498] This emotional information is received by a server and analyzed through an analysis system to determine the performance strategy based on the performer's current emotional state. The analysis results are interpreted by a control system, which generates instructions on how to effectively utilize the stage lighting and sound. For example, if a performer shows anger, the system can instruct the lighting to turn red to increase the sense of tension and play background music with emphasized bass.
[0499] The emotion engine can also capture audience reactions and amplify the performance effects in response to laughter, surprise, and other reactions. For example, when the audience laughs loudly, it can instruct the lighting to change to warmer colors and play positive background music.
[0500] Finally, the user (stage operator) can monitor whether the system's automatic adjustments are being performed accurately and make manual adjustments as needed. In this way, the present invention realizes sophisticated direction that is in line with the emotions of the performers and audience in improvisational theater, thereby improving the quality of the play.
[0501] The following describes the processing flow.
[0502] Step 1:
[0503] The user (director) uses a terminal to input the play's theme, story structure, and character settings into the system. This allows fundamental information for the play to be stored in a database.
[0504] Step 2:
[0505] The device captures the performers' movements and lines in real time through cameras and microphones installed on the stage. Simultaneously, audience responses are collected by voice sensors.
[0506] Step 3:
[0507] The emotion engine analyzes acquired video and audio data to recognize the emotional states of performers and audience members. For example, it infers emotions from changes in performers' facial expressions and tone of voice, and analyzes collective emotional trends from audience reactions.
[0508] Step 4:
[0509] The server thoroughly analyzes the emotional information transmitted from the emotion engine, along with the behavioral and audio information acquired by the terminal, using analysis tools. Based on the analysis results, it constructs a specific performance scenario.
[0510] Step 5:
[0511] The server uses a generation mechanism to create performance instructions based on the analysis results and transmits them to the control mechanism. These instructions include dynamic adjustments to lighting color tone and intensity, sound balance, and background music selection.
[0512] Step 6:
[0513] The terminal controls the equipment on stage and executes the specified effects according to instructions received from the server. This includes switching lighting and adjusting sound effects.
[0514] Step 7:
[0515] The user (stage operator) monitors whether the automated effects provided by the system are working correctly and makes manual adjustments if further fine-tuning is required.
[0516] This allows the system to create high-quality improvisational theater performances based on the emotions of both the performers and the audience.
[0517] (Example 2)
[0518] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0519] In theater, there is a challenge in improving the quality of improvisational plays by instantly grasping the actors' expressions and the audience's reactions, and adjusting the direction in real time. This challenge has been difficult to address with conventional technology, as it was challenging to accurately grasp the emotions of the actors and audience and reflect them in immediate changes to the direction.
[0520] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0521] In this invention, the server includes setting means for inputting the play's theme, storyline, and character settings; acquisition means for acquiring performer's movement information and voice information; emotion analysis means for analyzing emotional information; and generation means for generating control instructions for lighting and sound. This enables real-time adjustment of the performance based on the emotions of the performers and the audience.
[0522] The "setting method" refers to a function that allows performers to input themes, storylines, character settings, and other elements of an improvisational play into the system.
[0523] "Acquisition means" refers to a function for collecting information on performers' movements and voices on stage in real time.
[0524] "Emotional analysis means" refers to an analytical function that estimates the emotional state of a performer based on their actions and voice data.
[0525] The "generation means" refers to a function that generates control instructions for lighting and sound based on analyzed emotional information, and automatically adjusts the stage production.
[0526] The "audience reaction acquisition means" is a function that acquires audience reactions as data and transmits that information to an emotion analysis means for use in adjusting the performance.
[0527] This invention is a system that grasps the emotions of performers and audience in real time and dynamically adjusts the direction of improvisational theater. This system operates effectively by combining a server, terminals, and users.
[0528] First, the user (director or stage operator) inputs the play's theme, storyline, and character settings into the system via a terminal interface. This input information is stored in the server's database as basic data for the production.
[0529] The terminal uses sensors such as cameras and microphones installed on the stage to capture the performers' movements and voices in real time. This terminal has an emotion engine built in that analyzes the acquired raw data. Specifically, it uses image processing technology and voice analysis technology to read emotions from the performers' facial expressions and tone of voice. This makes it possible to determine what emotions the performers are currently expressing.
[0530] The server uses a generative AI model to calculate a performance strategy based on the performer's emotional information and basic data transmitted from the terminal. The server then generates automatic adjustment instructions for stage lighting and sound based on the generated strategy. For example, if a performer expresses anger, the server can issue instructions to turn on red lighting and play background music with emphasized bass.
[0531] The terminal also captures audience reactions using cameras and microphones, analyzing responses such as laughter and surprise. This information is sent to a server and used to fine-tune the performance. For example, if the audience is laughing hysterically, the server can instruct the system to change the lighting to a brighter, warmer tone and play positive background music.
[0532] The user (stage operator) monitors these automated adjustments and manually changes the settings as needed to ensure that the intended performance is being carried out.
[0533] As a concrete example, here is an example of a prompt sentence to be input to a generative AI model: "If the audience is crying during a sad scene for the protagonist, how should the lighting and background music be changed?"
[0534] This system enables timely and flexible direction that responds to the emotions of both performers and audience, thereby improving the quality of improvisational theater.
[0535] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0536] Step 1:
[0537] Users input initial data such as the play's theme, storyline, and character settings into the system via their terminal. This data is sent to the server and stored in a database as foundational data for the production. The input data serves as the basic criteria for future adjustments to the production.
[0538] Step 2:
[0539] The device uses cameras and microphones installed on the stage to capture the performers' movements and voices in real time. The cameras capture the performers' movements and facial expressions as video data, and the microphones collect dialogue and voice tones as audio data. This acquired data is temporarily stored within the device.
[0540] Step 3:
[0541] The device passes the acquired motion data and audio data to the emotion engine. The emotion engine processes the data using image processing algorithms and audio analysis techniques to estimate the performer's emotions. For example, changes in facial muscle movements and tone of voice can be used to classify the performer into a specific emotion such as "joy," "anger," or "sadness." This analysis result is returned to the device as emotion information.
[0542] Step 4:
[0543] The server receives emotional information transmitted from the terminal and uses a generative AI model to formulate a performance plan based on it. This plan includes specific instructions on what lighting and sound effects should be used. The generative AI model combines initial information stored in the database with the results of the emotional analysis to calculate the optimal performance and generate a set of instructions.
[0544] Step 5:
[0545] The server sends the generated performance instructions to the control system, which adjusts the stage lighting and sound in real time. For example, when an actor is expressing "anger," the stage lighting is changed to red, and a bass-heavy background music track is selected and played. This approach makes the actors' emotional expressions more effective for the audience.
[0546] Step 6:
[0547] The terminal further captures audience reactions using its camera and microphone, and performs emotion analysis based on this data. Laughter, surprise, applause, and other reactions are detected, and the results are sent back to the server. This audience emotion information is then used to fine-tune the stage production.
[0548] Step 7:
[0549] Users (stage operators) can monitor real-time performance adjustment information provided by the server and manually adjust lighting and sound as needed. This allows for further optimization of the performance in response to instantaneous audience reactions and changes in performers.
[0550] Through these steps, the entire system works in coordination, enabling flexible adjustments that respond to the emotions of both performers and audience members while maintaining high quality in the direction of improvisational theater.
[0551] (Application Example 2)
[0552] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0553] In recent years, physical stores have seen a growing demand for flexible environmental adjustments tailored to individual customer emotions in order to enhance the customer experience. However, current technology makes it difficult to accurately recognize customer emotions and automatically adjust store lighting and sound in real time based on those emotions. Therefore, a system that combines emotion recognition and environmental adjustment is necessary to improve the quality of the customer experience within a store.
[0554] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0555] In this invention, the server includes an acquisition means for acquiring information to estimate a person's emotional state, an analysis means for analyzing the emotional state based on the information obtained from the acquisition means, and a control means for controlling a light source and sound based on the emotional state from the analysis means. This enables dynamic environmental adjustments that are in line with the customer's emotions.
[0556] "A person's emotional state" refers to the mental or emotional state an individual is experiencing at a particular moment. This includes specific emotions such as joy, sadness, anger, and surprise.
[0557] "Means of acquisition" refers to means of obtaining information to estimate a person's emotional state. Specifically, this involves using input devices such as cameras and microphones.
[0558] "Analysis means" refers to methods for processing information to analyze a person's emotional state based on the acquired information. Image recognition technology and speech analysis technology may be used for this analysis.
[0559] "Control means" refers to methods for manipulating light sources and sounds based on the results of the analysis. This allows the living environment to be adjusted according to emotions.
[0560] "Generative means" refers to methods for generating specific instructions to adapt the environment to the customer's emotional state. This enables real-time environment configuration.
[0561] "Light sources and sound" refer to elements in physical space that affect vision and hearing, such as lighting equipment and speaker systems.
[0562] "Means of acquiring reactions" refers to means of detecting and acquiring a person's actions and reactions. This includes sensor and camera technologies.
[0563] The system that realizes this invention aims to identify a person's emotional state in real time and dynamically adjust the lighting and sound settings of a physical space, such as a store, accordingly. The system is centered around a server connected to a network.
[0564] The server acquires information such as facial expressions and voice tone of people inside the store through acquisition devices such as cameras and microphones. This acquisition method utilizes OpenCV, a video processing software, and the Google Cloud Speech-to-Text API for speech analysis. By integrating these devices and software, the emotional state of people is estimated, and data based on that estimation is collected.
[0565] The server then processes the acquired data using analysis tools to infer the emotional state. The emotion engine analyzes facial expressions and voice data to recognize specific emotions such as joy, sadness, and anger. The results of the analysis are interpreted by control tools within the server, which then generate environmental adjustment instructions corresponding to the emotional state.
[0566] The server controls the store's lighting and sound systems, adjusting the environmental settings. For example, if a customer smiles, the server instructs the sound system to change the lighting to a warmer color and play rhythmic music.
[0567] In this way, the server can adjust the environment in real time according to emotions. This invention is expected to further enrich the customer experience.
[0568] For example, when customers are in a good mood, the store can be filled with bright lighting and upbeat music to amplify their enjoyment. Another example of a prompt for a generative AI model would be: "Based on the customer's emotions, suggest the optimal lighting and background music settings. For example, if a customer is smiling, consider how to improve the store's atmosphere."
[0569] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0570] Step 1:
[0571] The server captures customers' facial expressions and voices through cameras and microphones placed within the store. The input consists of camera footage and microphone audio, which are collected as digital data. This data is temporarily stored in storage for subsequent analysis.
[0572] Step 2:
[0573] The server processes the collected video data using OpenCV to detect faces and estimate emotions from their expressions. The input is camera video data, and the output is emotional state (e.g., joy, surprise, anger). Emotions are quantified by analyzing facial shape and muscle movements.
[0574] Step 3:
[0575] The server uses the Google Cloud Speech-to-Text API to convert speech data into text and analyzes the tone of the speech. The input is microphone audio data, and the output is an estimate of the emotional state derived from the speech. Emotion is estimated based on volume and intonation.
[0576] Step 4:
[0577] The analyzed emotional information is integrated within the server to evaluate the overall emotional state. The input is the emotional estimation results from steps 2 and 3, and the output is the overall emotional state. This information is used to determine the next environmental adjustment step.
[0578] Step 5:
[0579] The server generates commands to adjust the lighting and sound settings within the store based on the integrated emotional state. The input is the overall emotional state, and the output is a specific control command. For example, if a positive emotion is detected, it recommends brighter lighting and upbeat music.
[0580] Step 6:
[0581] Upon receiving instructions from the server, the lighting and sound systems are automatically adjusted. The server sends control signals to these devices, ensuring that the environmental settings change instantly. Ultimately, a store environment that aligns with the customer's emotional needs is provided.
[0582] This series of steps enables dynamic environmental adjustments that respond to customer emotions.
[0583] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0584] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0585] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[0586] [Fourth Embodiment]
[0587] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0588] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0589] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0590] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0591] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0592] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0593] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0594] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0595] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0596] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0597] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0598] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0599] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0600] The "improvisational AI direction" system according to the present invention is for dynamically adjusting the direction on a theatrical stage. This system consists of acquisition means, analysis means, control means, and generation means. The processing of the program in this system is described below in natural language.
[0601] First, the user (director) inputs the play's theme, story structure, and character settings into the system. This prepares the basic information for the play. Next, the terminal captures the actors' movements and conversations in real time through cameras and microphones installed on the stage. In addition, audience reactions are also captured by sensors.
[0602] This data is sent to a server and analyzed in detail by an analysis system. An AI model on the server analyzes the performers' movements and voice information to extract necessary insights regarding the performance. Based on these analysis results, the control system generates instructions to adjust lighting, sound, and other stage equipment in real time.
[0603] Furthermore, if an actor acts spontaneously or makes an unexpected move, the server uses a generation mechanism to instantly create appropriate corrective instructions. This allows for flexible responses that maintain the flow of the story while ensuring the performance is not interrupted.
[0604] For example, if loud laughter erupts from the audience, the server analyzes the reaction and instructs the control system to change the lighting to a warmer color and play upbeat background music at that point. Furthermore, if an actor makes a mistake with their lines, the generation system suggests alternative actions to correct the story in a natural flow. In this way, the present invention improves the quality of direction in improvisational theater.
[0605] The following describes the processing flow.
[0606] Step 1:
[0607] The user (director) uses a terminal to input the play's theme, storyline, character settings, and basic directorial guidelines into the system. The entered information is stored in the system's database and used by the AI for subsequent processing.
[0608] Step 2:
[0609] The devices (cameras and microphones installed on the stage) capture the performers' movements and lines in real time. They also capture audience reactions such as laughter and applause.
[0610] Step 3:
[0611] The server inputs the real-time data stream provided from the terminal into an analysis device, and the performer's actions and voice are analyzed using an AI model. Here, the performer's ad-libs and improvisational actions, the content of the lines, and audience reaction data are identified.
[0612] Step 4:
[0613] Based on the insights gained from the analysis results, the server generates staging instructions that are appropriate for the current stage situation. Specifically, it determines staging elements such as lighting patterns, sound adjustments, and background music selection.
[0614] Step 5:
[0615] The terminal receives performance instructions sent from the server and applies those instructions to the stage's lighting and sound systems. This allows for instantaneous adjustments to lighting color, brightness, and volume, as well as background music switching.
[0616] Step 6:
[0617] The user (stage operator) can monitor whether the automated effects provided by the system are being executed appropriately and make manual adjustments as needed.
[0618] Step 7:
[0619] The server collects feedback data from performers and audience members in response to their reactions, and stores it as points for improvement in the next session, in order to update future production strategies.
[0620] (Example 1)
[0621] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0622] In improvisational theater and live performances, there is a need to respond flexibly in real time to the performers' spontaneous actions and the audience's reactions, thereby improving the quality of the performance. However, conventional systems have the challenge of being unable to respond to changing situations and making appropriate adjustments to maintain the flow of the performance.
[0623] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0624] In this invention, the server includes means for acquiring operational information and audio information, means for performing analysis based on the operational information and audio information from the acquisition means, and means for controlling lighting and sound based on the results of the analysis means. This makes it possible to adjust the performance appropriately in real time even in response to improvisational situational changes.
[0625] "Motion information" refers to data that shows the physical movements of performers and people on stage, such as gestures, movements, and postures.
[0626] "Audio information" refers to data that represents acoustic signals, including sounds from the performers and their surroundings.
[0627] "Means of acquisition" refers to devices and technologies for collecting operational information and audio information.
[0628] "Analytical means" refers to the techniques and processes used to analyze acquired information and derive meaningful results.
[0629] "Control means" refers to devices and technologies used to operate stage equipment such as lighting and sound based on analysis results.
[0630] "Modification instructions" refer to instructions for appropriately changing the stage production in response to the performers' improvisational actions or the audience's reactions.
[0631] "Audience reaction" refers to information that shows the visual, auditory, and other feedback from the audience during a performance.
[0632] "Monitoring methods" refer to devices and technologies used to monitor audience reactions in real time and acquire data.
[0633] "Generative means" refers to the techniques and processes for generating staging and modifications that are appropriate to the improvisational situation.
[0634] The "improvisational AI direction" system of the present invention enables improvisational direction adjustments in theaters and on stage. The system comprises the following components:
[0635] The user inputs basic information into the system, such as the theme, story structure, and character settings of the play to be performed on stage. This information primarily forms the basis for determining the direction of the production. Once the input is complete, the system uses this information to set up the initial configuration.
[0636] The terminal uses cameras and microphones installed on the stage to capture the performers' movements and voices in real time. This data is sent to the server as movement and voice information. In addition, sensors are connected to the terminal to detect audience reactions, thereby capturing feedback information such as laughter and applause.
[0637] The server is responsible for analyzing the acquired data. An advanced generative AI model runs on the server, analyzing the performers' actions and audio information. This analysis provides insights necessary for the performance, generating appropriate direction based on the stage situation. Furthermore, the server can extract insights from audience reactions and incorporate them into the performance in real time.
[0638] For example, in a scene where the audience bursts into laughter, the server analyzes the reaction and generates instructions to change the lighting to a warmer tone and play upbeat music. Furthermore, by using a generative AI model, it is possible to immediately suggest new actions to maintain the flow of the story even if an actor makes a mistake in their lines.
[0639] An example of a prompt to the generative AI model in this system is the instruction, "When the performer makes an unexpected move, please suggest a new story development that is appropriate for the situation." This prompt allows the system to suggest an improvisational story development, helping to ensure smooth performance on stage.
[0640] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0641] Step 1:
[0642] The user inputs the play's theme, story structure, and character settings into the system. This input is stored in the system as basic data and used for subsequent processing. This sets the fundamental direction for the production.
[0643] Step 2:
[0644] The terminal uses cameras and microphones installed on the stage to acquire real-time information on the performers' movements and voices. The input data includes characteristics such as the performers' gestures, voice tone, and speed. This information is immediately transmitted to the server.
[0645] Step 3:
[0646] The device also acquires audience reactions. Sensors collect laughter, applause, and other reactions from the audience and send them to a server as reaction data. Based on this input, the system evaluates which scenes are effective for the audience.
[0647] Step 4:
[0648] The server receives motion information, audio information, and audience reactions transmitted from the terminals, and analyzes the data using a generative AI model. Based on the input data, it extracts insights into the performers' movements and audience feedback. This provides analytical results that allow for the determination of what kind of staging is necessary.
[0649] Step 5:
[0650] Based on the analysis results, the server generates instructions to adjust the stage equipment, including lighting and sound, using control devices. Specifically, for example, it might instruct the lighting to be dimmed and the music to be played softly during emotional scenes. This allows for the creation of an atmosphere appropriate to the scene.
[0651] Step 6:
[0652] If unexpected actions or improvisations occur, the server uses a generation mechanism to create appropriate corrective instructions. Based on the input data, new action plans are generated to naturally develop the scene while maintaining the continuity of the narrative.
[0653] Step 7:
[0654] Based on the generated results, users can fine-tune the performance as needed. This enables optimal performance tailored to real-time acting, enhancing the audience experience.
[0655] (Application Example 1)
[0656] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0657] Improving the customer experience within a store is difficult with static displays alone. Dynamically adjusting the displays in response to customer movements and reactions is essential. Furthermore, it's necessary to adapt to customer interest in specific products and emphasize their appeal.
[0658] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0659] In this invention, the server includes acquisition means for acquiring operational information and audio information, analysis means for performing analysis based on the information from the acquisition means, control means for controlling lighting and sound based on the analysis results, generation means for generating adaptive effects based on customer actions and audio information, and adjustment means for emphasizing product identification based on trends within the store. This makes it possible to dynamically enhance the store environment and the appeal of products in response to customer movements and reactions.
[0660] "Motion information" refers to data about the physical movements and actions of customers and spectators.
[0661] "Audio information" refers to data related to the voices of customers and spectators, as well as sounds within the store.
[0662] "Means of acquisition" refers to equipment and technology for collecting motion information and audio information.
[0663] "Analysis means" refers to equipment and software used to analyze acquired motion information and audio information and extract patterns and insights.
[0664] "Control means" refers to equipment and software used to adjust lighting and sound based on analysis results.
[0665] "Generation means" refers to equipment and software used to create adaptive presentations and instructions in response to customer actions and voice information.
[0666] "Adjustment measures" refer to technologies and devices used to highlight specific products in response to trends within a store or customer interest.
[0667] To implement this invention, a system for dynamically adjusting the in-store environment is required. This system consists of acquisition means, analysis means, control means, generation means, and adjustment means.
[0668] The server uses acquisition methods to collect motion and audio information through cameras and microphones installed in the store. For example, it captures actions such as customers stopping in front of product shelves and audio of conversations.
[0669] Next, the server uses analysis tools to perform a detailed analysis of the acquired behavioral and audio information. This analysis utilizes AI models and machine learning platforms such as TensorFlow. This allows the server to understand customer behavior and reactions and extract necessary insights.
[0670] Based on the analysis results, the lighting and sound are adjusted in real time by the control system. For example, if many customers show interest in a particular product, the lighting will brighten and related music will play to highlight that product.
[0671] Furthermore, it is possible to create visuals that adapt to customer behavior using generation methods. This allows for the emphasis on specific products based on in-store trends, thereby improving the customer experience.
[0672] One concrete example is a store employee wearing smart glasses who walks around the store and provides information about specific products. This allows them to respond immediately to customer questions and increase their desire to purchase.
[0673] An example of a prompt message might be, "Which product is currently attracting the most attention in the store? Please suggest ways to highlight that product."
[0674] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0675] Step 1:
[0676] The terminal acquires motion and audio information through cameras and microphones installed in the store. Inputs include customer location, posture, voice volume, and frequency. Outputs are raw video and audio data. This allows for understanding customer movements within the store.
[0677] Step 2:
[0678] The server receives the acquired raw data and inputs it into the AI model using analysis tools. Data processing includes identifying customer attributes and behavioral patterns through image recognition and recognizing emotions through voice analysis. The output is detailed insights into customer trends and reactions.
[0679] Step 3:
[0680] The server's control system issues instructions to adjust the store's lighting and sound based on the analysis results. The input is insights obtained from the analysis system, and the output is settings for lighting color tone and brightness, and background music selection. For example, an instruction might be issued to brighten the area around a specific product.
[0681] Step 4:
[0682] The generation mechanism generates prompts to provide adaptive presentations and product information in response to changes in behavioral information and customer intentions. The inputs are the customer's current behavior and analysis results. The output is the information displayed to the store employee using smart glasses.
[0683] Step 5:
[0684] The user (store clerk) receives instructions from the server and provides appropriate guidance to customers through smart glasses. For example, based on a prompt such as, "Which product is currently attracting the most attention in the store? Please suggest ways to highlight that product," they can provide product descriptions to customers in real time.
[0685] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0686] This invention relates to an "improvisational AI performance" system that uses an emotion engine to recognize the emotions of performers and audience members in real time. By incorporating acquisition means, analysis means, control means, generation means, and an emotion engine, this system enables flexible adjustment of performances in improvisational theater. The program processing of this system will be explained in natural language below, along with specific examples.
[0687] First, the user (director) uses a terminal to input the play's theme, storyline, and character settings into the system. This information is stored in a database as basic data necessary for the play's progress. Next, the terminal uses cameras and microphones placed on the stage to capture the actors' movements and lines in real time. Simultaneously, an emotion engine analyzes this data and recognizes emotions from the actors' facial expressions and vocal characteristics.
[0688] This emotional information is received by a server and analyzed through an analysis system to determine the performance strategy based on the performer's current emotional state. The analysis results are interpreted by a control system, which generates instructions on how to effectively utilize the stage lighting and sound. For example, if a performer shows anger, the system can instruct the lighting to turn red to increase the sense of tension and play background music with emphasized bass.
[0689] The emotion engine can also capture audience reactions and amplify the performance effects in response to laughter, surprise, and other reactions. For example, when the audience laughs loudly, it can instruct the lighting to change to warmer colors and play positive background music.
[0690] Finally, the user (stage operator) can monitor whether the system's automatic adjustments are being performed accurately and make manual adjustments as needed. In this way, the present invention realizes sophisticated direction that is in line with the emotions of the performers and audience in improvisational theater, thereby improving the quality of the play.
[0691] The following describes the processing flow.
[0692] Step 1:
[0693] The user (director) uses a terminal to input the play's theme, story structure, and character settings into the system. This allows fundamental information for the play to be stored in a database.
[0694] Step 2:
[0695] The device captures the performers' movements and lines in real time through cameras and microphones installed on the stage. Simultaneously, audience responses are collected by voice sensors.
[0696] Step 3:
[0697] The emotion engine analyzes acquired video and audio data to recognize the emotional states of performers and audience members. For example, it infers emotions from changes in performers' facial expressions and tone of voice, and analyzes collective emotional trends from audience reactions.
[0698] Step 4:
[0699] The server thoroughly analyzes the emotional information transmitted from the emotion engine, along with the behavioral and audio information acquired by the terminal, using analysis tools. Based on the analysis results, it constructs a specific performance scenario.
[0700] Step 5:
[0701] The server uses a generation mechanism to create performance instructions based on the analysis results and transmits them to the control mechanism. These instructions include dynamic adjustments to lighting color tone and intensity, sound balance, and background music selection.
[0702] Step 6:
[0703] The terminal controls the equipment on stage and executes the specified effects according to instructions received from the server. This includes switching lighting and adjusting sound effects.
[0704] Step 7:
[0705] The user (stage operator) monitors whether the automated effects provided by the system are working correctly and makes manual adjustments if further fine-tuning is required.
[0706] This allows the system to create high-quality improvisational theater performances based on the emotions of both the performers and the audience.
[0707] (Example 2)
[0708] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0709] In theater, there is a challenge in improving the quality of improvisational plays by instantly grasping the actors' expressions and the audience's reactions, and adjusting the direction in real time. This challenge has been difficult to address with conventional technology, as it was challenging to accurately grasp the emotions of the actors and audience and reflect them in immediate changes to the direction.
[0710] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0711] In this invention, the server includes setting means for inputting the play's theme, storyline, and character settings; acquisition means for acquiring performer's movement information and voice information; emotion analysis means for analyzing emotional information; and generation means for generating control instructions for lighting and sound. This enables real-time adjustment of the performance based on the emotions of the performers and the audience.
[0712] The "setting method" refers to a function that allows performers to input themes, storylines, character settings, and other elements of an improvisational play into the system.
[0713] "Acquisition means" refers to a function for collecting information on performers' movements and voices on stage in real time.
[0714] "Emotional analysis means" refers to an analytical function that estimates the emotional state of a performer based on their actions and voice data.
[0715] The "generation means" refers to a function that generates control instructions for lighting and sound based on analyzed emotional information, and automatically adjusts the stage production.
[0716] The "audience reaction acquisition means" is a function that acquires audience reactions as data and transmits that information to an emotion analysis means for use in adjusting the performance.
[0717] This invention is a system that grasps the emotions of performers and audience in real time and dynamically adjusts the direction of improvisational theater. This system operates effectively by combining a server, terminals, and users.
[0718] First, the user (director or stage operator) inputs the play's theme, storyline, and character settings into the system via a terminal interface. This input information is stored in the server's database as basic data for the production.
[0719] The terminal uses sensors such as cameras and microphones installed on the stage to capture the performers' movements and voices in real time. This terminal has an emotion engine built in that analyzes the acquired raw data. Specifically, it uses image processing technology and voice analysis technology to read emotions from the performers' facial expressions and tone of voice. This makes it possible to determine what emotions the performers are currently expressing.
[0720] The server uses a generative AI model to calculate a performance strategy based on the performer's emotional information and basic data transmitted from the terminal. The server then generates automatic adjustment instructions for stage lighting and sound based on the generated strategy. For example, if a performer expresses anger, the server can issue instructions to turn on red lighting and play background music with emphasized bass.
[0721] The terminal also captures audience reactions using cameras and microphones, analyzing responses such as laughter and surprise. This information is sent to a server and used to fine-tune the performance. For example, if the audience is laughing hysterically, the server can instruct the system to change the lighting to a brighter, warmer tone and play positive background music.
[0722] The user (stage operator) monitors these automated adjustments and manually changes the settings as needed to ensure that the intended performance is being carried out.
[0723] As a concrete example, here is an example of a prompt sentence to be input to a generative AI model: "If the audience is crying during a sad scene for the protagonist, how should the lighting and background music be changed?"
[0724] This system enables timely and flexible direction that responds to the emotions of both performers and audience, thereby improving the quality of improvisational theater.
[0725] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0726] Step 1:
[0727] Users input initial data such as the play's theme, storyline, and character settings into the system via their terminal. This data is sent to the server and stored in a database as foundational data for the production. The input data serves as the basic criteria for future adjustments to the production.
[0728] Step 2:
[0729] The device uses cameras and microphones installed on the stage to capture the performers' movements and voices in real time. The cameras capture the performers' movements and facial expressions as video data, and the microphones collect dialogue and voice tones as audio data. This acquired data is temporarily stored within the device.
[0730] Step 3:
[0731] The device passes the acquired motion data and audio data to the emotion engine. The emotion engine processes the data using image processing algorithms and audio analysis techniques to estimate the performer's emotions. For example, changes in facial muscle movements and tone of voice can be used to classify the performer into a specific emotion such as "joy," "anger," or "sadness." This analysis result is returned to the device as emotion information.
[0732] Step 4:
[0733] The server receives emotional information transmitted from the terminal and uses a generative AI model to formulate a performance plan based on it. This plan includes specific instructions on what lighting and sound effects should be used. The generative AI model combines initial information stored in the database with the results of the emotional analysis to calculate the optimal performance and generate a set of instructions.
[0734] Step 5:
[0735] The server sends the generated performance instructions to the control system, which adjusts the stage lighting and sound in real time. For example, when an actor is expressing "anger," the stage lighting is changed to red, and a bass-heavy background music track is selected and played. This approach makes the actors' emotional expressions more effective for the audience.
[0736] Step 6:
[0737] The terminal further captures audience reactions using its camera and microphone, and performs emotion analysis based on this data. Laughter, surprise, applause, and other reactions are detected, and the results are sent back to the server. This audience emotion information is then used to fine-tune the stage production.
[0738] Step 7:
[0739] Users (stage operators) can monitor real-time performance adjustment information provided by the server and manually adjust lighting and sound as needed. This allows for further optimization of the performance in response to instantaneous audience reactions and changes in performers.
[0740] Through these steps, the entire system works in coordination, enabling flexible adjustments that respond to the emotions of both performers and audience members while maintaining high quality in the direction of improvisational theater.
[0741] (Application Example 2)
[0742] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0743] In recent years, physical stores have seen a growing demand for flexible environmental adjustments tailored to individual customer emotions in order to enhance the customer experience. However, current technology makes it difficult to accurately recognize customer emotions and automatically adjust store lighting and sound in real time based on those emotions. Therefore, a system that combines emotion recognition and environmental adjustment is necessary to improve the quality of the customer experience within a store.
[0744] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0745] In this invention, the server includes an acquisition means for acquiring information to estimate a person's emotional state, an analysis means for analyzing the emotional state based on the information obtained from the acquisition means, and a control means for controlling a light source and sound based on the emotional state from the analysis means. This enables dynamic environmental adjustments that are in line with the customer's emotions.
[0746] "A person's emotional state" refers to the mental or emotional state an individual is experiencing at a particular moment. This includes specific emotions such as joy, sadness, anger, and surprise.
[0747] "Means of acquisition" refers to means of obtaining information to estimate a person's emotional state. Specifically, this involves using input devices such as cameras and microphones.
[0748] "Analysis means" refers to methods for processing information to analyze a person's emotional state based on the acquired information. Image recognition technology and speech analysis technology may be used for this analysis.
[0749] "Control means" refers to methods for manipulating light sources and sounds based on the results of the analysis. This allows the living environment to be adjusted according to emotions.
[0750] "Generative means" refers to methods for generating specific instructions to adapt the environment to the customer's emotional state. This enables real-time environment configuration.
[0751] "Light sources and sound" refer to elements in physical space that affect vision and hearing, such as lighting equipment and speaker systems.
[0752] "Means of acquiring reactions" refers to means of detecting and acquiring a person's actions and reactions. This includes sensor and camera technologies.
[0753] The system that realizes this invention aims to identify a person's emotional state in real time and dynamically adjust the lighting and sound settings of a physical space, such as a store, accordingly. The system is centered around a server connected to a network.
[0754] The server acquires information such as facial expressions and voice tone of people inside the store through acquisition devices such as cameras and microphones. This acquisition method utilizes OpenCV, a video processing software, and the Google Cloud Speech-to-Text API for speech analysis. By integrating these devices and software, the emotional state of people is estimated, and data based on that estimation is collected.
[0755] The server then processes the acquired data using analysis tools to infer the emotional state. The emotion engine analyzes facial expressions and voice data to recognize specific emotions such as joy, sadness, and anger. The results of the analysis are interpreted by control tools within the server, which then generate environmental adjustment instructions corresponding to the emotional state.
[0756] The server controls the store's lighting and sound systems, adjusting the environmental settings. For example, if a customer smiles, the server instructs the sound system to change the lighting to a warmer color and play rhythmic music.
[0757] In this way, the server can adjust the environment in real time according to emotions. This invention is expected to further enrich the customer experience.
[0758] For example, when customers are in a good mood, the store can be filled with bright lighting and upbeat music to amplify their enjoyment. Another example of a prompt for a generative AI model would be: "Based on the customer's emotions, suggest the optimal lighting and background music settings. For example, if a customer is smiling, consider how to improve the store's atmosphere."
[0759] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0760] Step 1:
[0761] The server captures customers' facial expressions and voices through cameras and microphones placed within the store. The input consists of camera footage and microphone audio, which are collected as digital data. This data is temporarily stored in storage for subsequent analysis.
[0762] Step 2:
[0763] The server processes the collected video data using OpenCV to detect faces and estimate emotions from their expressions. The input is camera video data, and the output is emotional state (e.g., joy, surprise, anger). Emotions are quantified by analyzing facial shape and muscle movements.
[0764] Step 3:
[0765] The server uses the Google Cloud Speech-to-Text API to convert speech data into text and analyzes the tone of the speech. The input is microphone audio data, and the output is an estimate of the emotional state derived from the speech. Emotion is estimated based on volume and intonation.
[0766] Step 4:
[0767] The analyzed emotional information is integrated within the server to evaluate the overall emotional state. The input is the emotional estimation results from steps 2 and 3, and the output is the overall emotional state. This information is used to determine the next environmental adjustment step.
[0768] Step 5:
[0769] The server generates commands to adjust the lighting and sound settings within the store based on the integrated emotional state. The input is the overall emotional state, and the output is a specific control command. For example, if a positive emotion is detected, it recommends brighter lighting and upbeat music.
[0770] Step 6:
[0771] Upon receiving instructions from the server, the lighting and sound systems are automatically adjusted. The server sends control signals to these devices, ensuring that the environmental settings change instantly. Ultimately, a store environment that aligns with the customer's emotional needs is provided.
[0772] This series of steps enables dynamic environmental adjustments that respond to customer emotions.
[0773] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0774] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0775] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[0776] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0777] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0778] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0779] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0780] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0781] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0782] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0783] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[0784] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[0785] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0786] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0787] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0788] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0789] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0790] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0791] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0792] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0793] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted as being incorporated by reference.
[0794] The following is further disclosed regarding the embodiments described above.
[0795] (Claim 1)
[0796] A means for acquiring motion information and audio information in a play,
[0797] An analysis means that performs analysis based on the operation information and audio information from the acquisition means,
[0798] A control means for controlling lighting and sound based on the results of the analysis means,
[0799] A generation means for generating corrective instructions to adapt to actions performed spontaneously by the performer,
[0800] A system that includes this.
[0801] (Claim 2)
[0802] The system according to claim 1, comprising an audience reaction acquisition means for acquiring audience reactions and transmitting them to the acquisition means, wherein the analysis means adjusts the performance based on the audience reactions.
[0803] (Claim 3)
[0804] The system according to claim 1, wherein the analysis means transmits instructions to the control means in real time to automatically change the lighting and sound settings.
[0805] "Example 1"
[0806] (Claim 1)
[0807] Means for acquiring operational information and audio information,
[0808] A means for performing analysis based on the operational information and audio information from the aforementioned acquisition means,
[0809] Means for controlling lighting and sound based on the results of the analysis means,
[0810] A means of generating corrective instructions to adapt to actions performed spontaneously by the performer,
[0811] A means of monitoring audience reactions,
[0812] A means for adjusting the stage production using audience reactions obtained through the aforementioned monitoring means,
[0813] A system that includes this.
[0814] (Claim 2)
[0815] The system according to claim 1, wherein the analysis means transmits instructions to the control means in real time to automatically change the lighting and sound settings.
[0816] (Claim 3)
[0817] The system according to claim 1, wherein the aforementioned modification instructions include new performance suggestions that correspond to the performer's improvisational actions.
[0818] "Application Example 1"
[0819] (Claim 1)
[0820] An acquisition means for acquiring operational information and audio information,
[0821] An analysis means that performs analysis based on the operation information and audio information from the acquisition means,
[0822] A control means for controlling lighting and sound based on the results of the analysis means,
[0823] A generation means for generating adaptive performances based on customer actions and voice information,
[0824] A means of adjustment that emphasizes product identification based on trends within the store,
[0825] A system that includes this.
[0826] (Claim 2)
[0827] The system according to claim 1, comprising a reaction acquisition means for acquiring audience reactions and transmitting them to the acquisition means, wherein the analysis means adjusts the performance based on the reactions.
[0828] (Claim 3)
[0829] The system according to claim 1, wherein the analysis means transmits instructions to the control means in real time to automatically change the lighting and sound settings, and adjusts the display content to highlight a specific product.
[0830] "Example 2 of combining an emotion engine"
[0831] (Claim 1)
[0832] A setting method for inputting the play's theme, storyline, and character settings,
[0833] A means for acquiring performer's movement information and audio information,
[0834] An emotion analysis means analyzes emotion information based on the operation information and audio information from the aforementioned acquisition means,
[0835] A generation means that generates control instructions for lighting and sound based on the analysis results by the emotion analysis means,
[0836] A system that includes this.
[0837] (Claim 2)
[0838] The system according to claim 1, comprising an audience reaction acquisition means for acquiring audience reactions and transmitting them to the emotion analysis means, wherein the generation means adjusts the performance based on the audience reactions.
[0839] (Claim 3)
[0840] The system according to claim 1, wherein the generation means generates instructions in real time to automatically change the lighting and sound settings.
[0841] "Application example 2 when combining with an emotional engine"
[0842] (Claim 1)
[0843] A means for obtaining information to estimate a person's emotional state,
[0844] Based on the information obtained from the acquisition means, an analysis means for analyzing the emotional state,
[0845] Based on the emotional state obtained from the analysis means, a control means controls the light source and sound,
[0846] A generation means for generating instructions to adapt the environment according to the customer's emotional state,
[0847] A system that includes this.
[0848] (Claim 2)
[0849] The system according to claim 1, comprising a reaction acquisition means for acquiring a person's reaction and transmitting it to the acquisition means, wherein the analysis means adjusts the environmental settings based on the person's reaction.
[0850] (Claim 3)
[0851] The system according to claim 1, wherein the analysis means transmits instructions to the control means in real time to automatically change the settings of the light source and sound. [Explanation of Symbols]
[0852] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. A means for acquiring motion information and audio information in a play, An analysis means that performs analysis based on the operation information and audio information from the acquisition means, A control means for controlling lighting and sound based on the results of the analysis means, A generation means for generating corrective instructions to adapt to actions performed spontaneously by the performer, A system that includes this.
2. The system according to claim 1, comprising an audience reaction acquisition means for acquiring audience reactions and transmitting them to the acquisition means, wherein the analysis means adjusts the performance based on the audience reactions.
3. The system according to claim 1, wherein the analysis means transmits instructions to the control means in real time to automatically change the lighting and sound settings.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A