System and method for providing feedback information from audience response information of musical performance through artificial intelligence

By using cameras and sensing devices to collect audience information during music performances and utilizing neural network analysis feedback systems, the problem of reflecting audience reactions in real time in existing technologies has been solved, thereby improving the immersion and quality of the performance.

CN121997092APending Publication Date: 2026-05-08BUTTERFLY TO THE ISLAND OF THE CORP +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BUTTERFLY TO THE ISLAND OF THE CORP
Filing Date
2024-11-18
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

Existing music performance systems struggle to reflect audience reactions in real time, limiting immersion and performance quality, especially in small musicals with limited budgets and resources where high-quality feedback is difficult to obtain.

Method used

Video and biometric information is collected by camera devices installed on the stage and sensing devices worn by the audience. The neural network is used to analyze the audience’s movements and reactions, and provides real-time feedback information to the director’s terminal.

Benefits of technology

It enables real-time analysis of audience emotions and reactions, allowing directors to adjust performances based on feedback, thereby enhancing audience immersion and performance quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121997092A_ABST
    Figure CN121997092A_ABST
Patent Text Reader

Abstract

Embodiments propose a system and method for providing feedback information based on response information to a viewer of a musical performance through artificial intelligence. A method according to one embodiment receives image information about a musical performance viewer from a camera device mounted on a musical performance stage, and receives biometric information about the musical performance viewer from a sensing device worn on the musical performance viewer. And the biological characteristic information of the music performance includes heart rate data, body temperature data, respiratory rate data, and skin conductance data, and obtaining motion information about the audience of the music performance by using a motion analysis model based on a first neural network. Based on the motion information and the biometric information, response information of the audience of the musical performance is determined by a response analysis model using a second neural network, and response information of the audience of the musical performance is determined based on the response. The method may include determining feedback information about music performance, and sending the feedback information about the music performance to a command terminal. The feedback information on the musical performance may include feedback information on lines and songs included in each of a plurality of scenes constituting the musical performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to techniques for providing feedback information about musical performances, and to systems and methods for providing audiences of musical performances with feedback information based on their reactions using artificial intelligence. Background Technology

[0002] In recent years, various directing techniques and technologies have been introduced into the performing arts, particularly in music performances, to enhance audience engagement and immersion. However, existing performances still primarily rely on fixed productions and storylines, limiting their ability to reflect audience reactions in real time. Introducing a system capable of analyzing audience responses and adjusting the performance direction accordingly could significantly increase immersion and maximize the audience experience.

[0003] Furthermore, existing performance feedback systems primarily rely on post-performance evaluations such as reviews and surveys to achieve performance improvement, but these systems suffer from difficulties in real-time reflection, timely feedback, and reflection. Audience reactions are also a concern. Moreover, it is difficult to expect high-quality feedback from audience feedback, and the possibility of deliberately giving positive or negative evaluations limits the availability of accurate and reliable feedback.

[0004] Furthermore, small-scale musicals face numerous limitations in terms of budget and infrastructure compared to large-scale productions, making it difficult to set budgets that incorporate audience feedback. In other words, operating with limited personnel and resources, small-scale musicals rely heavily on traditional feedback methods, making it challenging to demonstrate improvements and potentially limiting the enhancement of performance quality.

[0005] Therefore, there is a need for a system and method that uses artificial intelligence to provide feedback based on audience responses to musical performances. Summary of the Invention

[0006] The problem that the invention aims to solve

[0007] Embodiments of this disclosure may provide a system and method for providing audiences of musical performances with feedback information based on reaction information through artificial intelligence.

[0008] The technical challenges to be addressed in the embodiments are not limited to those described above, and those skilled in the art may consider other technical challenges not mentioned in the various embodiments described below.

[0009] Methods for solving problems

[0010] A method according to an embodiment for a server to provide feedback information about a musical performance based on audience reaction information includes: receiving video information about the audience of the musical performance from a camera device mounted on the stage of the musical performance; and receiving biometric information of the audience members from sensing devices worn by the audience members based on the image information, the biometric information including heart rate data, body temperature data, respiratory rate data, and skin conductance data; determining the motion information of the audience members using a motion analysis model of a first neural network; determining the audience members' reaction information based on the motion information and the biometric information using a response analysis model of a second neural network; and determining the response may include determining feedback information for the musical performance based on the information and preset expected response information, and sending the feedback information for the musical performance to a command terminal. The feedback information about the musical performance may include feedback information about the dialogue and songs included in each of a plurality of scenes constituting the musical performance.

[0011] Invention Effects

[0012] As described above, the present invention has the following effects:

[0013] According to an embodiment, the server can collect audience reaction information during a performance by analyzing their facial expressions, movements, and biometric information in real time, and generate feedback information based on this information. This allows the director to understand the audience's emotional state and reaction patterns in real time and adjust the performance production as needed to maximize audience immersion.

[0014] According to an embodiment, the server uses an artificial intelligence model to predict the audience's external emotional state (e.g., joy, sadness, surprise, etc.) and internal emotional state (e.g., immersion, tension, fatigue, etc.) in real time, and can provide performance feedback by comparing the information with the director's intentions. In this way, the director can check whether the audience's reactions match the expected reactions and improve the quality of the performance accordingly.

[0015] The effects that can be obtained from the embodiments are not limited to those described above, and other effects not mentioned can be clearly derived and understood by those skilled in the art based on the following detailed description. Attached Figure Description

[0016] Figure 1 A system is shown, according to an embodiment, for providing audiences of a musical performance with feedback information based on their reaction information using artificial intelligence.

[0017] Figure 2 This is a flowchart of a method according to an embodiment, wherein the server provides feedback information about a musical performance based on reaction information from the audience of the musical performance.

[0018] Figure 3 This is a block diagram illustrating the configuration of a server according to an embodiment. Detailed Implementation

[0019] The following embodiments combine elements and features of the embodiments in a predetermined manner. Unless otherwise explicitly stated, each component or feature can be considered optional. Each component or feature can be implemented without being combined with other components or features. Furthermore, various embodiments can be configured by combining some components and / or features. The order of operations described in the various embodiments can be changed. Some features or characteristics of one embodiment can be included in other embodiments, or can be replaced by corresponding features or characteristics of other embodiments.

[0020] In the description of the accompanying drawings, processes or steps that may obscure the essence of the various embodiments are not described, nor are processes or steps that can be understood by those skilled in the art.

[0021] Throughout this specification, when a section is referred to as “comprising or including” an element, it means that it does not exclude other elements, but may include other elements, unless explicitly stated otherwise. Additionally, terms such as “…unit,” “…module,” and “module” as used herein refer to a unit that performs at least one function or operation, meaning hardware, software, or a combination of hardware and software that can be implemented as follows: Furthermore, the terms “a or an,” “an,” “the,” and similar related terms are used herein in the context of describing various embodiments (particularly in the context of the appended claims) unless otherwise stated or obviously contradictory. In this context, it can be used in both singular and plural terms.

[0022] In the following, embodiments according to various embodiments will be described in detail with reference to the accompanying drawings. The detailed description below, taken in conjunction with the drawings, is intended to describe exemplary embodiments of various embodiments and is not intended to represent the only embodiments.

[0023] In addition, specific terminology used in the various embodiments is provided to aid in understanding the embodiments, and the use of these specific terminology may be changed to other forms without departing from the technical spirit of the various embodiments.

[0024] In a network environment, an electronic device communicates with another electronic device via a first network (e.g., a short-range wireless communication network) or with another electronic device or a server via a second network (e.g., a long-range wireless communication network). It can communicate with at least one of these. According to one embodiment, the electronic device can communicate with another electronic device via a server. According to one embodiment, the electronic device includes a processor, memory, an input module, an audio output module, a display module, an audio module, a sensor module, an interface, a connection terminal, a haptic module, a camera module, a power management module, a battery, and a communication module, which may include an identification module or an antenna module. In some embodiments, at least one of these components (e.g., a connection terminal) may be omitted, or one or more other components may be added to the electronic device. In some embodiments, some of these components (e.g., a sensor module, a camera module, or an antenna module) may be integrated into a single component (e.g., a display module). The electronic device may also be referred to as a client, a terminal, or a peer.

[0025] For example, a processor can execute software (e.g., a program) to control at least one other component (e.g., a hardware or software component) of an electronic device connected to the processor and perform various data processing or operations. According to one embodiment, as at least part of data processing or computation, the processor stores commands or data received from another component (e.g., a sensor module or a communication module) in volatile memory, and stores commands or data stored in volatile memory in non-volatile memory. It can be processed, and the resulting data can be stored in non-volatile memory. According to one embodiment, the processor is a main processor (e.g., a central processing unit or application processor) or an auxiliary processor (e.g., a graphics processing unit, neural processing unit (NPU), image signal processor, sensor, hub processor, or communication processor) that can operate independently or together with the main processor. For example, when an electronic device includes a main processor and an auxiliary processor, the auxiliary processor can be configured to use less power than the main processor or be dedicated to a specific function. A coprocessor can be implemented separately from the main processor or as part of the main processor.

[0026] A coprocessor can act on behalf of the main processor, for example, when the main processor is inactive (e.g., in a sleep state), or act in conjunction with the main processor when the main processor is active (e.g., in an application execution state). The electronic device can control at least some functions or states associated with at least one component (e.g., a display module, a sensor module, or a communication module). According to one embodiment, the coprocessor (e.g., an image signal processor or a communication processor) can be implemented as part of another functionally related component (e.g., a camera module or a communication module). According to one embodiment, an auxiliary processor (e.g., a neural network processing unit) can include hardware structures specifically designed for processing artificial intelligence models.

[0027] Artificial intelligence models can be created through machine learning. For example, such learning can be performed within the electronic device itself that executes the AI ​​model, or it can be performed via a separate server. Learning algorithms can include, for example, but are not limited to, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning. AI models can include multiple layers of artificial neural networks. Artificial neural networks include deep neural networks (DNNs), convolutional neural networks (CNNs), recurrent neural networks (RNNs), restricted Boltzmann machines (RBMs), belief deep networks (DBNs), bidirectional recurrent deep neural networks (BRDNNs), etc. It can be one of deep Q-networks or a combination of two or more of the above networks, but is not limited to the examples above. In addition to hardware architecture, AI models may additionally or alternatively include software architecture.

[0028] Memory can store various data used by at least one component of an electronic device (e.g., a processor or sensor module). Data may include, for example, input or output data of software (e.g., a program) and associated instructions. Memory may be volatile or non-volatile.

[0029] Programs can be stored in memory as software and can include, for example, operating systems, middleware, or applications.

[0030] An input module can receive commands or data from outside the electronic device (e.g., a user) to be used in a component of the electronic device (e.g., a processor). An input module may include, for example, a microphone, mouse, keyboard, keys (e.g., buttons), or a digital pen (e.g., a stylus).

[0031] A sound output module can output sound signals to an external part of an electronic device. The sound output module may include, for example, a speaker or a receiver. The speaker can be used for general purposes, such as multimedia playback or recording playback. The receiver can be used to receive incoming calls. According to one embodiment, the receiver can be implemented separately from the speaker or as part of it.

[0032] A display module can visually provide information to the outside of an electronic device (e.g., to a user). The display module may include, for example, a display, a holographic device, or a projector, and control circuitry for controlling the device. According to one embodiment, the display module may include a touch sensor configured to detect a touch, or a pressure sensor configured to measure the intensity of the force generated by a touch.

[0033] An audio module can convert sound into electrical signals, or vice versa. According to one embodiment, the audio module can acquire sound through an input module, a sound output module, or an external electronic device (e.g., a speaker or headphones) directly or wirelessly connected to the electronic device.

[0034] The sensor module can detect the operating status of an electronic device (e.g., power or temperature) or the status of the external environment (e.g., user status) and generate electrical signals or data values ​​corresponding to the detected status. According to one embodiment, the sensor module includes, for example, motion sensors, gyroscope sensors, barometric pressure sensors, magnetic sensors, accelerometers, grip force sensors, proximity sensors, color sensors, infrared sensors, biometric sensors, and acoustic sensors; it may also include temperature sensors, humidity sensors, or light sensors.

[0035] The interface may support one or more specified protocols that can be used for direct or wireless connection between an electronic device and an external electronic device (e.g., another electronic device). According to one embodiment, the interface may include, for example, a High Definition Multimedia Interface (HDMI), a Universal Serial Bus (USB) interface, an SD card interface, or an audio interface.

[0036] The connection terminal may include a connector through which an electronic device can be physically connected to an external electronic device (e.g., an electronic device). According to one embodiment, the connection terminal may include, for example, an HDMI connector, a USB connector, an SD card connector, or an audio connector (e.g., a headphone connector).

[0037] A haptic module can convert electrical signals into mechanical stimuli (e.g., vibration or motion) or electrical stimuli that a user can perceive through touch or kinesthesia. According to one embodiment, the haptic module may include, for example, a motor, a piezoelectric element, or an electrical stimulation device.

[0038] The camera module can capture still images and videos. According to one embodiment, the camera module may include one or more lenses, an image sensor, an image signal processor, or a flash.

[0039] A power management module can manage the power supplied to an electronic device. According to one embodiment, the power management module can be implemented as at least part of, for example, a power management integrated circuit (PMIC).

[0040] A battery can power at least one component of an electronic device. According to one embodiment, the battery may include, for example, a non-rechargeable primary battery, a rechargeable secondary battery, or a fuel cell.

[0041] A communication module can support the establishment of direct (e.g., wired) or wireless communication channels between an electronic device and an external electronic device (e.g., an electronic device or a server), and perform communication through the established communication channels. The communication module operates independently of a processor (e.g., an application processor) and may include one or more communication processors that support direct (e.g., wired) or wireless communication. According to one embodiment, the communication module is a wireless communication module (e.g., a cellular communication module, a short-range wireless communication module, or a Global Navigation Satellite System (GNSS) communication module) or a wired communication module (e.g., a local area network (LAN) communication module or a power line communication module). Among these communication modules, the corresponding communication module is a first network (e.g., a short-range communication network such as Bluetooth, Wi-Fi Direct, or Infrared Data Association (IrDA)) or a second network (e.g., a traditional network). Cellular networks, 5G networks, and telecommunications networks such as next-generation communication networks, the Internet, or computer networks (e.g., LANs or WANs) can communicate with external electronic devices. These various types of communication modules can be integrated into a single component (e.g., a single chip) or implemented as multiple separate components (e.g., multiple chips). The wireless communication module can use subscriber information (e.g., International Mobile Subscriber Identifier (IMSI)) stored in the subscriber identification module to identify or authenticate electronic devices within a communication network (e.g., a first network or a second network).

[0042] The wireless communication module can support 5G networks after 4G networks and next-generation communication technologies, such as NR access technology (New Radio Access Technology). NR access technology provides high-speed transmission of large amounts of data (eMBB (enhanced Mobile Broadband)), minimized terminal power consumption and multi-terminal access (mMTC (massive Machine-Type Communication)), or high reliability and low latency (URLLC can support ultra-reliable and low-latency). The wireless communication module 192 can support, for example, high-frequency bands (e.g., millimeter-wave bands) to achieve high data rates. The wireless communication module uses various technologies to ensure high-frequency band performance, such as beamforming, massive MIMO (Multiple-Input Multiple-Output), and full-dimensional multiple-input / output (FD) technologies. Multi-dimensional MIMO, array antennas, analog beamforming, or massive MIMO are also used. The wireless communication module can support various requirements specified in electronic devices, external electronic devices, or network systems (e.g., auxiliary networks). According to one embodiment, the wireless communication module has a peak data rate (e.g., 20 Gbps or higher) for implementing eMBB, a loss coverage range (e.g., 164 dB or lower) for implementing mMTC, or a U-plane delay (e.g., downtime) for implementing URLLC. It can support link (DL) and uplink (UL) latency of 0.5 ms or less each, or round-trip latency of 1 ms or less.

[0043] Antenna modules can transmit or receive signals or power to or from external sources (e.g., external electronic devices). According to one embodiment, an antenna module may include an antenna comprising a radiator made of a conductor or conductive pattern formed on a substrate (e.g., a PCB). According to one embodiment, an antenna module may include multiple antennas (e.g., an array antenna). In this case, for example, a communication module can select at least one antenna suitable for a communication method used in a communication network such as a first network or a second network. Signals or power can be transmitted or received between the communication module and the external electronic device through the selected at least one antenna. According to some embodiments, in addition to the radiator, other components (e.g., a radio frequency integrated circuit (RFIC)) may be additionally incorporated into the antenna module.

[0044] According to various embodiments, the antenna module can form a millimeter-wave antenna module. According to one embodiment, the millimeter-wave antenna module includes a printed circuit board, an RFIC disposed on or adjacent to a first side (e.g., bottom side) of the printed circuit board and capable of supporting a specified high-frequency band (e.g., millimeter-wave band), and multiple antennas (e.g., array antennas) may be disposed on or near a second side (e.g., top or side) of the printed circuit board and capable of transmitting or receiving signals in the specified high-frequency band.

[0045] At least some components are connected to each other via communication methods between peripheral devices (e.g., buses, general purpose input and output (GPIO), serial peripheral interfaces (SPI), or mobile industry processor interfaces (MIPI)) and signals. (e.g., commands or data) can be exchanged with each other.

[0046] According to one embodiment, commands or data can be sent or received between an electronic device and external electronic devices via a server connected to a second network. Each external electronic device may be of the same or different type as the electronic device. According to one embodiment, all or part of the operations performed in the electronic device may be performed in one or more external electronic devices. For example, when the electronic device is required to perform a function or service automatically or in response to a request from a user or another device, the electronic device may perform one or more functions or services in lieu of performing that function or service, or in addition to performing that function or service. The service itself may be requested to perform at least part of the function or service. One or more external electronic devices receiving the request may perform at least a portion of the requested function or service, or additional functions or services related to the request, and send the execution result to the electronic device. The electronic device may process the result as is or additionally and provide it as at least part of the response to the request. For this purpose, cloud computing, distributed computing, mobile edge computing (MEC), or client-server computing technologies may be used, for example. The electronic device may use distributed computing or mobile edge computing, etc., to provide ultra-low latency services. In another embodiment, the external electronic device may include an Internet of Things (IoT) device. The server may be an intelligent server using machine learning and / or neural networks. According to one embodiment, an external electronic device or server may be included in the second network. The electronic device can be used for smart services (such as smart homes, smart cities, smart cars, or healthcare) based on 5G communication technology and Internet of Things (IoT) related technologies.

[0047] A server connects to electronic devices and can provide services to those devices. Additionally, the server can handle membership registration, store and manage various information about registered users, and provide various purchase and payment functions related to the service. Furthermore, the server can share execution data of service applications running on multiple electronic devices in real time, enabling service sharing among users. The server can have the same hardware configuration as a typical web server or service server. However, in terms of software, it can be implemented in any language such as C, C++, Java, Python, Golang, and Kotlin, and can include program modules that perform various functions. Furthermore, a server is generally a computer system that connects to an unspecified number of clients and / or other servers via an open computer network such as the Internet, receives work performance requests from clients or other servers, and exports and provides work results. It refers to the computer software (server program) installed for this purpose. In addition to the server program mentioned above, a server also includes a series of applications running on the server, and in some cases, various built-in or external databases (DBs, hereinafter referred to as "DBs"), which should be understood as a broad concept depending on the situation. Therefore, the server categorizes member registration information and various game-related information and data, stores them in a database (DB), and manages that DB, which can be implemented internally or externally. Furthermore, the server can be implemented using general-purpose server hardware and server programs, which are provided in various ways depending on the operating system (e.g., Windows, Linux, UNIX, and Macintosh). Representative examples include web services implemented using IIS (Internet Information Services). CERN, NCSA, APPACH, and TOMCAT are used in Windows environments and in Unix environments. Additionally, the server can link to authentication and payment systems for user authentication or payment for service-related purchases.

[0048] The first and second networks refer to connection structures that allow each node (e.g., terminal and server) to exchange information, or networks connecting servers and electronic devices. The first and second networks include the Internet, LAN (Local Area Network), wireless LAN (Wireless Local Area Network), WAN (Wide Area Network), PAN (Personal Area Network), 3G, 4G, LTE, 5G, Wi-Fi, etc., but are not limited to these. The first and second networks can be closed networks such as LANs or WANs, but are preferably open networks such as the Internet. The Internet includes protocols such as TCP / IP, TCP and UDP (User Datagram Protocol), and various upper-layer services such as HTTP (Hypertext Transfer Protocol), Telnet, FTP (File Transfer Protocol), DNS (Domain Name System), SMTP (Simple Mail Transfer Protocol), SNMP (Simple Network Management Protocol), NFS (Network File Service), and NIS (Network Information Service).

[0049] A database can have a general data structure implemented in the storage space (hard disk or memory) of a computer system using a database management system (DBMS). A database can have a data storage format that allows free searching (retrieval), deletion, editing, and adding of data. A database can be a relational database management system (RDBMS) such as Oracle, Informix, Sybase, and DB2, or an object-oriented database management system such as Gemston, Orion, and O2. Using an object-oriented database management system (OODBMS) and an XML-native database (such as Excelon, Tamino, and Sekaiju) with its own functionality, you can have appropriate fields or elements.

[0050] According to one embodiment, the program may include an operating system, middleware, or an application executable on an operating system for controlling one or more resources of an electronic device. The operating system may include, for example, Android™, iOS™, Windows™, Symbian™, Tizen™, or Bada™. At least some of the program may be pre-loaded into the electronic device, for example, during manufacturing, or may be downloaded or updated from an external electronic device (e.g., an electronic device or a server) when used by a user. All or part of the program may include a neural network.

[0051] An operating system can control the management (e.g., allocation or reclamation) of one or more system resources (e.g., processes, memory, or power) of an electronic device. Additionally or alternatively, the operating system may include other hardware devices of the electronic device, such as input modules, audio output modules, display modules, audio modules, sensor modules, interfaces, haptic modules, camera modules, power management modules, and batteries. It may include one or more drivers for driving communication modules, subscriber identification modules, or antenna modules.

[0052] Middleware can provide various functionalities to applications, enabling applications to use functions or information provided from one or more resources of an electronic device. According to one embodiment, middleware can dynamically remove existing components or add new ones. According to one embodiment, at least a portion of the middleware can be included as part of an operating system or can be implemented as separate software different from the operating system.

[0053] The application may include, for example, various applications. According to one embodiment, the application may also include an information exchange application capable of supporting information exchange between an electronic device and an external electronic device. The information exchange application may include, for example, a notification relay application configured to transmit specified information (e.g., a call, message, or alarm) to an external electronic device, or a device management application configured to manage an external electronic device. For example, a notification relay application may deliver notification information corresponding to a specified event (e.g., email reception) generated in another application on the electronic device (e.g., email application 269) to an external electronic device. Alternatively or concurrently, the notification relay application may receive notification information from an external electronic device and provide it to the user of the electronic device.

[0054] Device management applications can, for example, control the power (e.g., turn on or off) or functionality of external electronic devices or certain components thereof (e.g., the display module or camera module of the external electronic device) that communicate with external electronic devices. They can control electronic devices (e.g., brightness, resolution, or focus). Device management applications may additionally or alternatively support the installation, removal, or updating of applications running on external electronic devices.

[0055] Throughout this specification, neural networks, neural network functions, and network functions can be used with the same meaning. A neural network may consist of a set of interconnected computational units, commonly referred to as "nodes." These "nodes" may also be called "neurons." A neural network consists of at least two or more nodes. The nodes (or neurons) that make up a neural network can be interconnected through one or more "links."

[0056] In neural networks, two or more nodes connected by links can form relative relationships of input and output nodes. The concepts of input and output nodes are relative; any node that is an output node to one node can potentially be an input node to another, and vice versa. As described above, input-to-output relationships can be created around links. One or more output nodes can be connected to an input node via links, and vice versa.

[0057] In a relationship between input and output nodes connected by a link, the value of the output node can be determined based on the data input to the input nodes. Here, the nodes connecting the input and output nodes can have weights. These weights can be variable and can be changed by the user or algorithm so that the neural network performs the desired function. The edges or links interconnecting the input and output nodes have weights that can be variably applied by the user or algorithm to perform the function required by the neural network. For example, when one or more input nodes are connected to an output node via their respective links, the output node is set as the output of the input nodes connected to the output node and the corresponding links for each input node. Node values ​​can be determined based on the weights.

[0058] As mentioned above, in a neural network, two or more nodes are interconnected by one or more links to form the input and output node relationships within the network. The characteristics of a neural network can be determined based on the number of nodes and links, the correlation between nodes and links, and the weights assigned to each link. For example, if two neural networks have the same number of nodes and links but different weights between the links, these two neural networks might be identified as different from each other.

[0059] Figure 1 A system is shown, according to an embodiment, for providing audiences of a musical performance with feedback information based on their reaction information using artificial intelligence. Figure 1 The embodiments described herein can be combined with various embodiments of this disclosure.

[0060] Reference Figure 1 A system 10 (hereinafter referred to as the feedback providing system) that provides feedback information based on reaction information to the audience of a music performance through artificial intelligence may include a camera device 110, a sensing device 120, a server 130, and a supervisor terminal 140.

[0061] The feedback system 10 provides video information of the audience watching the musical performance via a camera device 110 installed on the stage where the performance takes place, and acquires information about the audience from sensing devices 120 installed on the audience members. The server 130 determines the audience's reaction information using artificial intelligence based on the video and biometric information, and uses this reaction information to provide feedback on the musical performance. The server 130 may be a system provided to the director's terminal 140.

[0062] According to one embodiment, the feedback providing system 10 can provide feedback information about the musical performance to the director terminal 140 after the musical performance ends.

[0063] According to one embodiment, as soon as the scene ends during a musical performance, the feedback providing system 10 can send feedback information about the scene to an in-ear device installed on the musical performer.

[0064] Camera device 110 can be mounted on a music performance stage to film the audience of the music performance. For example, camera device 110 can be mounted on a music performance stage to film the entire audience position of the music performance. For example, camera device 110 can capture still images and videos of the audience of the music performance. For example, camera device 110 may include one or more lenses, image sensors, image signal processors or flashes, as well as memory and communication units.

[0065] Sensing device 120 is mounted on an audience member of a musical performance and can determine biometric information about the audience member. For example, sensing device 120 may be a wristband device. For example, sensing device 120 may be mounted on the wrist of an audience member of a musical performance. For example, sensing device 120 can detect the state of the audience member and generate an electrical signal or data value corresponding to the detected state. For example, sensing device 120 may include a processor, memory, multiple sensors, and a communication unit. For example, multiple sensors may include a photoplethysmography (PPG) sensor that measures heart rate data, an infrared sensor that measures body temperature data, an accelerometer that measures respiratory rate data, and a skin conductance sensor that measures skin conductance data. For example, biostatistical information may include heart rate data, body temperature data, respiratory rate data, and skin conductance data. For example, a PPG sensor can measure heart rate by projecting light onto the skin using an LED and detecting the amount of light absorbed and reflected based on blood flow. For example, an infrared sensor can measure body temperature by detecting infrared radiation emitted from the skin surface. For example, an accelerometer can measure respiratory rate by detecting wrist movement. For example, a skin conductivity sensor can measure skin conductivity by using two electrodes. Alternatively, it can measure skin conductivity as a current value by applying a current to the two electrodes. The unit of skin conductivity can be the reciprocal of resistance.

[0066] Server 130 uses artificial intelligence to determine audience reaction information regarding the music performance using video and biometric information of the audience, and director terminal 140 can provide feedback information about the music performance based on this feedback. For example, server 130 can provide feedback information about the music performance through a graphical user interface, graphical user experience, and / or a webpage. Alternatively, server 130 can provide feedback information about the music performance through an application pre-installed on director terminal 140. For example, server 130 may include... Figure 1 Server 108.

[0067] Additionally, according to one embodiment, whenever a scene during a musical performance ends, server 130 uses video information and biometric information of the audience for that scene's musical performance to update information about the musical performance in terms of reaction information. The audience can be identified, and based on the audience reaction information for that scene's musical performance, feedback information about the musical performance can be provided via an in-ear device installed on the performer. The in-ear device can be a listening device inserted into the ear to accurately hear human voices. For example, the in-ear device installed on the musical performer can convert feedback information about the musical performance into vibration signals. In this case, the vibration signals can be converted according to indication values ​​included in the feedback information.

[0068] Director terminal 140 can refer to the terminal of a director who wants to provide feedback on the musical performance through server 130. Supervisor terminal 140 can include desktop computers, laptops, notebooks, smartphones, tablet PCs, mobile phones, smartwatches, and smart glasses with communication capabilities. For example, supervisor terminal 140 may include… Figure 1 Electronic device 102.

[0069] exist Figure 3 In this illustration, server 130 is shown as a single device; however, according to embodiments, server 130 may include multiple devices operatively connected to server 130.

[0070] Figure 2 This is a flowchart of a method according to an embodiment, wherein the server provides feedback information about a musical performance based on reaction information from the audience of the musical performance. Figure 2 The embodiments described herein can be combined with various embodiments of this disclosure.

[0071] Reference Figure 2 In step S210, the server can receive image information about the audience of the music performance from the camera equipment installed on the music performance stage.

[0072] For example, camera equipment installed on a musical performance stage can film the audience from the moment the performance begins. Similarly, a server can receive real-time video information about the audience from the camera equipment installed on the stage, starting from the beginning of the performance.

[0073] In step S220, the server can receive biometric information about the audience of the music performance from the sensing devices worn by the audience.

[0074] For example, sensing devices can be installed on audience members when they enter to watch a musical performance. For instance, the sensing device could be a wristband-type device. For example, the sensing device can be worn on the audience member's wrist while watching the musical performance.

[0075] For example, sensing devices can measure biometric information about the audience of a music performance from the moment the performance begins. Biometric information could include, for example, heart rate data, body temperature data, respiratory rate data, and skin conductance data. Alternatively, a server could receive real-time biometric information about the audience of a music performance from the moment the performance begins, via sensing devices worn by the audience.

[0076] For example, sensing devices may include a PPG sensor that measures heart rate data, an infrared sensor that measures body temperature data, an accelerometer that measures respiratory rate data, and a skin conductance sensor that measures skin conductance data.

[0077] In step S230, the server can determine motion information about the audience of the music performance based on the video information by using a motion analysis model of a first neural network.

[0078] Motion analysis models can be GRU (Gated Recurrent Unit) models that include CNN (Convolutional Neural Network). For example, a GRU model that includes CNN can be called a ConvGRU model.

[0079] CNNs (Convolutional Neural Networks) are deep learning neural network models used for image recognition. CNNs extract image features through convolution operations, compress data through pooling to extract important information, and classify images based on these important features. GRU models are likely modifications of RNNs (Recurrent Neural Networks), which rely on past observations and therefore suffer from vanishing or exploding gradients. To address this issue, LSTM (Long Short-Term Memory) networks were developed. By replacing nodes within an LSTM with memory units, information can be accumulated or some past information can be deleted. This compensates for the gradient loss or large gradient values ​​in RNNs. GRU is a model that improves processing speed by simply modifying the LSTM structure.

[0080] The ConvGRU model can be used to identify patterns in image data with time series data. The ConvGRU model extracts spatial features such as edges and shapes from image frames using a CNN, and inputs the CNN output of each image frame into the GRU. By learning the continuity and change patterns between image frames, motion information can be determined.

[0081] For example, image vectors containing pixel values ​​of images across multiple time periods can be generated through image information preprocessing. For instance, a server can generate image vectors containing pixel values ​​of images across multiple time periods through image information preprocessing.

[0082] Additionally, images can be classified based on multiple time periods of image information. For example, image contrast correction, tilt correction, and noise removal can be performed to specify viewer regions for multiple images within each time period. Image contrast correction is an operation that sharpens boundaries to separate the background from the image, and can be performed, for example, using CLAHE (Contrast-Limited Adaptive Histogram Equalization). CLAHE is a technique that divides an image into blocks of a certain size and performs histogram equalization on each block to equalize the entire image. Tilt correction adjusts the angle of the tilted image to align it horizontally and vertically, improving image recognition accuracy. Denoising can further remove noise that was not removed during image binarization and image contrast correction using various filters such as Gaussian blur, center blur, and bidirectional filtering. This allows for improved processing speed of motion analysis models by reducing the image size.

[0083] For example, motion information about the audience at a musical performance can be determined based on image vectors input into a motion analysis model. For instance, a server can determine motion information about the audience at a musical performance by inputting image vectors into a motion analysis model.

[0084] For example, motion information about the audience at a musical performance could include values ​​related to the motion of each of multiple audience members over multiple time periods. For example, motion-related values ​​could include values ​​related to facial orientation, eye shape, mouth shape, upper body movement, and hand movement. Values ​​related to facial orientation could include values ​​for left-right face rotation, up-down face rotation, and face tilt. Values ​​related to eye shape might include blink frequency and eye opening / closing ratio. Values ​​related to mouth shape could include mouth opening ratio and changes in mouth corner position. For example, changes in mouth corner position could include changes in the x-axis and y-axis of the mouth corner based on the xy coordinates of the face's center point. That is, assuming the mouth corners are symmetrical, changes in the x-axis and y-axis of one corner of the mouth could be included. In this case, the x-axis could have a change based on a ratio of half the face width, and the y-axis could have a change based on a ratio of half the face length. You can. For example, if the value of the change in the corner of the mouth is (0.3, 0.2), it means that the corner of the mouth has moved to the right by 30% of half the width of the face and moved upwards by 20%, exceeding half the length of the face. Values ​​related to upper body movement can include the tilt angle of the upper body and the distance the upper body moves. Values ​​related to hand movement can be the distance the hand moves and the frequency of hand position changes. For example, the frequency of hand position changes can be the number of times the hand position changes more than a preset value.

[0085] For example, a motion analysis model can be learned based on multiple image vectors, multiple reference image vectors, and multiple correct motion information.

[0086] A motion analysis model may include a first input layer, two or more first hidden layers, and a first output layer. The two or more first hidden layers may include one or more CNN layers and one or more GRU layers.

[0087] Learning data, consisting of multiple image vectors and multiple correct motion information, is input into a first input layer, passes through two or more first hidden layers and a first output layer, and is output as a first output vector. This first output vector is the first output vector 1, which is the input to the first loss function layer connected to the output layer. The first loss function layer uses a first loss function to compare the first output vector with the correct answer to output a first loss value. For each training data vector, parameters of the motion analysis model can be learned in the direction that reduces the first loss value.

[0088] One or more CNN layers can include one or more convolutional layers and one or more pooling layers. For example, multiple image vectors can be filtered in convolutional layers, and feature maps can be formed through convolutional layers. For example, feature maps can be used to represent multiple landmarks of a viewer's face, upper body, and hands. For example, in pooling layers, fixed vectors associated with the features are selected for dimensionality reduction based on the formed feature maps, and the formed feature maps are subsampled to generate multiple correct motion information from the vectorized time series. Data can be extracted. For example, pooling layers can be max pooling layers that extract the maximum value. For example, pooling layers can be average pooling layers that extract the average value. For example, in this case, parameters can include parameters associated with convolutional and pooling layers (feature map size, filter size, depth, stride, zero padding).

[0089] One or more GRU layers comprise one or more GRU blocks, and a GRU block may include reset gates and update gates. Here, reset gates and update gates may comprise sigmoid layers. For example, a sigmoid layer may be a layer whose activation function is the sigmoid function(). For example, the hidden state may be controlled by reset gates and update gates, and each gate and input may have weights.

[0090] An image vector can consist of a reference image vector, correct motion information, and a training dataset. For example, multiple training datasets can be pre-stored on a server.

[0091] Reference image vectors can be vectors used as a standard for comparison with image vectors. For example, multiple reference image vectors can be pre-determined from each of multiple images having values ​​associated with multiple operations. That is, multiple reference image vectors are pre-generated based on images of various motions, and thus can be used as learning data for motion analysis models. Videos of various movements can be collected based on videos of multiple viewers watching a musical.

[0092] Correct motion information can include values ​​related to the motion of the image that generates a reference image vector. For example, each of several correct motion information can include values ​​related to different motions.

[0093] For example, a motion analysis model identifies the reference image vector among multiple reference image vectors that has the highest similarity to the input image vector, and uses the correct motion information of the reference image vector with the highest similarity as the motion information of the relevant image vector. Viewers can learn to do this.

[0094] For example, the similarity between an image vector and a reference image vector can be determined using Equation 1 below.

[0095] [Formula 1]

[0096]

[0097] In Equation 1, Sim is the similarity between the image vector and the reference image vector, T is the total number of time periods, L is the total number of landmarks, Fj,t is the j-th landmark in the j-th location vector image vector, and F*j,t can be the location vector of the j-th landmark in the reference image vector at time interval t.

[0098] For example, "X"_u could be a function that calculates the Euclidean distance to X. Landmarks could include points that identify facial features such as eyebrows, eyes, nose, and lips; points that identify hand features; and points that identify upper body features. The locations of the landmarks could be preset.

[0099] For example, the similarity between an image vector and a reference image vector can be determined as a value greater than 0 and less than or equal to 1.

[0100] Therefore, the motion analysis model identifies motion by analyzing changes in facial features, upper body features, and hand gestures within the image information of the audience during a musical performance, providing information about the motion to determine its nature. Thus, the server analyzes audience motion information by using a roundabout method to determine groups of motion information from highly similar reference videos, rather than directly analyzing the motion of each individual audience member during the musical performance. This reduces the time required.

[0101] In step S240, the server can determine the audience's reaction information regarding the musical performance based on motion information and biometric information using a reaction analysis model of a second neural network.

[0102] For example, a GRU model using an attention algorithm can be used as a response analysis model. Here, the attention algorithm can be an algorithm that allows the neural network model to focus on specific parts of the input data.

[0103] For example, motion vectors for each of multiple viewers can be generated through data preprocessing of motion information. For instance, a server can generate motion vectors for each of multiple viewers through data preprocessing of motion information. The motion vectors can include values ​​associated with motion at each of multiple time intervals.

[0104] For example, biometric vectors for each of multiple viewers can be generated through data preprocessing of biometric information. For instance, a server can generate biometric vectors for each of multiple viewers through data preprocessing of biometric information. Biometric vectors can include values ​​of heart rate, body temperature, respiratory rate, and skin conductance over multiple time intervals.

[0105] For example, audience reaction information for a musical performance can be determined based on the motion vectors and biometric vectors of each of multiple audience members input into the reaction analysis model. For instance, a server can determine audience reaction information for a musical performance by inputting the motion vectors and biometric vectors of each of multiple audience members into the reaction analysis model.

[0106] Information about the audience's response to a musical performance may include scores for each of multiple external emotional states for each of multiple time periods and scores for each of multiple internal emotional states. For example, multiple external emotional states may include happiness, sadness, surprise, fear, anger, and neutrality. Multiple internal emotional states may include engagement, tension, fatigue, and anxiety. Scores may be determined as values ​​greater than 0 and less than or equal to 10.

[0107] For example, based on the motion vectors and biometric vectors of each of multiple audience members input into the response analysis model, the scores of each of multiple external emotional states and the scores of each of multiple internal emotional states can determine audience reaction information for each of multiple time intervals regarding the included musical performance. For instance, the server inputs the motion vectors and biometric vectors of each of multiple audience members into the response analysis model to calculate the scores of each of multiple external emotional states and the scores of each of multiple internal emotional states for each of multiple time intervals, thus obtaining information about the audience's reaction to the included musical performance.

[0108] For example, a response analysis model can be learned based on multiple motion vectors, multiple biometric vectors, and multiple correct response information.

[0109] A motion vector and a biometric vector can consist of a correct response message and a training dataset. For example, multiple training datasets can be pre-stored on a server. A correct response message can include scores for each of multiple external emotional states for each of multiple time periods and scores for each of multiple internal emotional states for each of multiple time periods. For example, each of the multiple correct response messages can have different combinations of scores for each of the multiple external emotional states and scores for each of the multiple internal emotional states.

[0110] For example, a GRU model used in response analysis can include a forward GRU layer, a backward GRU layer, an attention layer, and an output layer. That is, the GRU model can be a bidirectional GRU model connected to the attention layer. In this way, the GRU model can address the long-term limitations of existing GRU models by selectively learning important information through the attention layer and improving predictive performance by learning more information through the interactive GRU layer. The forward GRU layer processes the input sequence from front to back and can use past information to compute the hidden state. The backward GRU layer processes the input sequence from back to front and can use future information to compute the hidden state. The attention layer summarizes the final information by selecting important information from the hidden states of the forward and backward GRU layers and assigning weights to them. The output layer computes the final predicted value based on the output value of the attention layer.

[0111] For example, the motion vector and biometric vector of each of multiple viewers can be fed into forward GRU layers and backward GRU layers in multiple input sequences. For example, the multiple input sequences can include input sequences for each of the multiple viewers. An input sequence can include motion vectors and biometric vectors.

[0112] For example, multiple input sequences can be forward-processed values ​​generated by a forward GRU layer. Alternatively, multiple input sequences can be processed in the opposite direction by a backward GRU layer to generate values. For instance, a combined sequence of forward and backward processed values ​​can be input to the attention layer as the hidden state value for each time interval.

[0113] For example, combined sequences can generate weighted result vectors through attention layers. The resulting vector can be a vector that selectively highlights important information in each time segment of the combined sequence. That is, the resulting vector can be a vector that assigns weights to each time segment in the combined sequence, calculates a weighted sum, and finally aggregates the vectors into a single vector.

[0114] For example, the resulting vector can be output as a score for each of multiple external emotional states and a score for each of multiple internal emotional states through an output layer. In this case, the output layer can include multiple dense layers. A dense layer is a fully connected (FC) layer that generates output values ​​by multiplying by weights, adding biases, and applying activation functions.

[0115] In step S250, the server can determine feedback information about the musical performance based on the reaction information and the preset expected reaction information.

[0116] For example, the pre-defined expected response information may include expected scores for each of multiple external emotional states and multiple internal emotional states for each of multiple time intervals. For example, the expected score may be determined as a value greater than 0 and less than or equal to 10. For example, the command terminal may send expected response information for a musical performance to the server. For example, the command terminal may pre-define expected response information by sending expected response information for a musical performance to the server before the musical performance begins.

[0117] For example, feedback on a musical performance can include feedback on the lines and songs included in each of the multiple scenes that make up the musical performance.

[0118] For example, feedback on a musical performance can be determined based on a first difference between the score of each of multiple external emotional states and the expected score, and a second difference between the score of each external emotional state and the expected score. You can, for example, have a server provide feedback on a musical performance based on a first difference between the score of each of multiple external emotional states and the expected score, and a second difference between the score of each external emotional state and the expected score. Multiple internal emotional states can be determined.

[0119] Additionally, for example, the server can output a first difference between the score and the expected score for each of multiple external emotional states, and a second difference between the score and the expected score for each of multiple internal emotional states. Too much cannot be determined over multiple time intervals.

[0120] For example, a server can divide multiple lines of dialogue and each song in a musical performance into multiple time segments. Alternatively, a server can group lines and songs divided into multiple time segments into multiple scenes that constitute a musical performance.

[0121] For example, the server can provide a first difference between the score and the expected score for each of multiple external emotional states, and a second difference between the score and the expected score for each of multiple internal emotional states. Each of the multiple scenes constituting a musical performance can be matched with each of the multiple lines and songs included in the scene. For example, the server can provide preset feedback of a combination of the first difference between the score and the expected score for each of the multiple external emotional states, and the second difference between the score and the expected score for each of the multiple external emotional states. You can match the lines and songs included in the scene.

[0122] For example, feedback for each line of dialogue and song can be preset for each combination of the first difference of each of multiple external emotional states and the second difference of multiple internal emotional states. The feedback might be something like, "The audience's external reaction is decreasing, tension is increasing. Reduce tension by softening the production," or "This is the part where audience interest is decreasing." It might include specific phrases such as, "We need to strengthen our delivery of lines or stage effects." Similarly, for each combination of the first difference of each of multiple external emotional states and the second difference of multiple internal emotional states, feedback for multiple lines can be preset on the server. Likewise, for each combination of the first difference of each of multiple external emotional states and the second difference of multiple internal emotional states, feedback for multiple songs can be preset on the server.

[0123] In other words, the server provides pre-defined feedback to a combination of a first difference between the score of each of the multiple external emotional states and the expected score, and a second difference between the score of each external emotional state and the expected score. By matching the dialogue and songs of each scene, feedback information regarding the musical performance can be determined.

[0124] In step S260, the server can send feedback information about the musical performance to the director terminal.

[0125] For example, the server can send feedback information about the musical performance to the director's terminal via a pre-installed application.

[0126] Additionally, according to one embodiment, when a scene included in a musical performance is in progress, the server performs motion analysis using the aforementioned first neural network based on video information of the audience of the musical performance in that scene, thereby determining the behavioral information of the audience in the scene.

[0127] For example, when a scene included in a musical performance is in progress, the server uses the aforementioned second neural network to run a response analysis model based on motion information and biometric information about the audience in the scene. This allows you to determine the audience's reaction information within the scene.

[0128] For example, after a scene in a musical performance ends, the server generates a first response for each of the multiple external emotional states for that scene based on the difference between the audience's reaction information and the preset expected reaction information for that scene. A second difference for multiple internal emotional states can then be determined.

[0129] For example, the server provides preset feedback to the dialogue and songs included in the scene, consisting of a first difference for each of multiple external emotional states and a second difference for multiple internal emotional states. By matching these, the scene's feedback information can be determined. For instance, the server can send the scene's feedback information to the director's terminal.

[0130] For example, a server can determine an audience's reaction score based on a first difference among multiple external emotional states and a second difference among multiple internal emotional states in a scene. For example, the audience's reaction score can be determined based on: N being the number of external emotional states, M being the number of internal emotional states, Di being the first difference of the i-th external emotional state, D_j^' being the second difference of the j-th internal emotional state, wi being the weight of the i-th external emotional state, and vj being the weight of the j-th internal emotional state. For example, each of the weights of the multiple external emotional states and the multiple internal emotional states can be set to a value greater than 0 and less than 1. For example, the weights of the multiple external emotional states and the multiple internal emotional states can be set differently for each of the multiple scenes.

[0131] For example, a server can determine one of multiple branch points based on audience reaction scores. For example, the server can send an indication value corresponding to the determined branch point to an in-ear device installed on the musician. For example, the server can pre-set a scoring range matching each of the multiple branch points. For example, when the in-ear device receives an instruction value, it can generate a vibration pattern matching the instruction value. For example, the in-ear device can include a processor, memory, communication unit, audio output module, and haptic module. For example, information about the multiple vibration patterns and the indication value matching each of the multiple vibration patterns can be preset in the in-ear device. The sound output module can output sound signals to the outside of the in-ear device. The sound output module can include, for example, a speaker or a receiver. The receiver can be implemented separately from the speaker or as part of the speaker. The haptic module can convert electrical signals into mechanical stimuli (e.g., vibration or motion) that the user can perceive through touch or kinesthesia. For example, the haptic module can include, for example, a motor, piezoelectric element, or electrical stimulation device. For example, when the in-ear device receives an instruction value, it can generate a vibration pattern matching the instruction value through the haptic module.

[0132] The GRU model used in response analysis models can include a forward GRU layer, a backward GRU layer, an attention layer, and an output layer.

[0133] For example, a forward GRU layer and a backward GRU layer can each consist of 60 units, and the combination of two GRU layers can produce a layer with 120 units.

[0134] For example, the forward GRU layer and the reverse GRU layer may include one or more second hidden layers. Specifically, the one or more second hidden layers may include one or more GRU blocks, and a GRU block may include a reset gate and an update gate.

[0135] The reset gate resets past information and can determine the weights r(t) derived from previous hidden layers. For example, when the motion vectors and biometric vectors of each of multiple viewers are input into the reset gate as multiple input sequences, and the input value (xt) at the current time point is generated based on the multiple input sequences, the current time point is multiplied by the weight Wr, the hidden state (h(t-1)) generated from the multiple input sequences at the previous time point is multiplied by the weight Ur at the previous time point, and finally the two values ​​are added to the sigmoid function, resulting in 0 and 1. It can output a value between 0 and 1. Through these values ​​between 0 and 1, it is possible to determine how much of the hidden state value from the previous point will be utilized.

[0136] For example, multiple input sequences are fed into a reset gate. When the current input value (xt) generated based on these sequences is input into the reset gate, it is multiplied by the current weight Wr. The previous hidden state (h(t-1)) generated by the reset gate based on the sequence is then multiplied by the previous weight Ur. Finally, the two values ​​are added together and input into the sigmoid function. The result is between 0 and 1 and can be output as a value. These values ​​between 0 and 1 determine how much of the previous hidden state value will be utilized.

[0137] The update gate determines the update rate of past and present information, and z(t) can be defined as the amount of information at the current time. For example, when the input value (xt) at the current time is input, a dot product is performed with the weight Wz at the current time point, and a dot product is performed with the hidden state (h(t-1)) at the previous time point and the weight Uz at the previous time point. Finally, the two values ​​are added together and input into the sigmoid function, and the result can be output as a value between 0 and 1. 1-z(t) can be multiplied by the information from the previous hidden layer (h(t-1)).

[0138] In this way, z(t) can reflect how much current information will be used, and 1-z(t) can reflect how much past information will be used.

[0139] The candidate information group for the current time t can be determined by multiplying the result by the reset gate. For example, when the input value (xt) at the current time is input, the dot product is the weight Wh at the current time, and the hidden state (h(t-1)) at the previous time is the weight Uh of the previous dot product of Wh and t. The value multiplied by r(t) can be added and input into the tanh function. For example, tanh represents a nonlinear activation function (hyperbolic tangent function).

[0140] By combining the results of the update gate and the candidate group, the weights of the hidden layer at the current time can be determined. For example, the weights of the hidden layer at the current time can be determined by summing the output value z(t) of the update gate multiplied by the hidden state (h(t)) at the current time and the value 1-z(t) discarded from the update gate multiplied by the hidden state (h(t-1)) at the current time.

[0141] The attention layer utilizes the values ​​output by the forward and backward GRU layers to focus on the more important parts of the prediction, i.e., predicting the score of each of the multiple external emotional states and the operation of scoring for each of the multiple internal emotional states, which can be weighted. For example, the attention layer can determine eij based on ht (the hidden state output by the bidirectional GRU layer), where eij represents the importance of each time interval. For example, the combined sequence can be ht, which is the hidden state output by the forward and backward GRU layers. For example, eij is an intermediate value used to determine the attention score and can be a value calculated by computing the similarity between the hidden state at time t and a specific context vector. For example, the similarity can be calculated by a dot product operation. For example, the attention layer can determine the attention score aij based on eij. For example, the attention score can be determined as follows. In other words, by applying a softmax function to eij, the attention layer can convert the similarity of all time intervals into probability values ​​and assign higher weights to the most important time intervals. For example, the attention layer can generate a context vector ci by multiplying the hidden state ht by the attention score aij and using the attention score as the weight. For example, ci can be determined as follows.

[0142] For example, in the attention layer, an early stopping method can be applied to the number of epochs to prevent overfitting. For instance, the learning rate can be set to 0.0001, and the batch size to 8192. In this case, the decay rate can be set to 0.96 to gradually reduce the learning rate as the epoch progresses. For example, when performing one epoch, the learning rate can be multiplied by the decay rate to obtain the new learning rate.

[0143] The output layer can consist of four fully connected (FC) layers (1st FC to 4th FC). In this case, the activation function for the first to third FC layers can be the ReLU function. The activation function for the fourth FC layer, the last FC layer, can be the softmax activation function. Multiple external emotions can be used to obtain the score of each of the multiple external emotional states, and each of the multiple internal emotional states can have an output node corresponding to the total number of states and the multiple internal emotional states.

[0144] Alternatively, for example, the score of the emotional state included in the correct response information can be determined by Equation 2 below.

[0145] [Formula 2]

[0146]

[0147] In equation 2, where It is the score of the i-th emotional state within the time interval t. It is the change in the j-th element of the motion vector within the time interval t. Let t be the change in the k-th element t of the biometric vector within the time interval, where α_(i,j) is the weight of the j-th element of the motion vector for the i-th emotional state, and β_(i,k) is the weight of the k-th element of the biometric vector for the i-th emotional state. Here, the motion vector and biometric vector can be vectors composed of correct response information and a training dataset.

[0148] For example, α_(i,j) and β_(i,k) can be set to values ​​greater than 0 and less than 1. For example, α_(i,j) and β_(i,k) can be set differently for multiple external emotional states and multiple internal emotional states.

[0149] For example, the change of the j-th element of the motion vector within the time interval t is the absolute value of the difference between the value of the j-th element of the motion vector and the reference value of the j-th element of the motion vector. It can be the value divided by the reference value of the element.

[0150] For example, the change of the kth element of a biological vector within a time interval t is the absolute value of the difference between the value of the kth element of the biological vector and the reference value of the kth element of the biological vector. It can be the value divided by the reference value of the element.

[0151] For example, the scores of emotional states included in the correct response information can be converted to a range of 0 to 10 by normalization and scaling.

[0152] Figure 3 This is a block diagram illustrating the configuration of a server according to an embodiment. Figure 3 One embodiment can be combined with various embodiments of this disclosure.

[0153] like Figure 3 As shown, server 300 may include processor 310, communication unit 320, and memory 330. However, not all of them are... Figure 3 All components shown are necessary components of server 300. Server 300 can be used in more ways than... Figure 3 The components shown can be implemented with more components, or server 300 can be used with more components. Figure 3 The components shown are implemented with fewer components. For example, in addition to the processor 310, communication unit 320 and memory 330, the server 300 according to some embodiments may also include a user input interface (not shown), an output unit (not shown), etc.

[0154] Processor 310 typically controls the overall operation of server 300. Processor 310 may include one or more processors and control other components included in server 300. For example, processor 310 can typically control communication unit 320 and memory 330 by executing programs stored in memory 330. Additionally, processor 310 can perform other tasks by executing programs stored in memory 330. Figure 1 and Figure 2 The functions of server 300 shown are illustrated.

[0155] The communication unit 320 may include one or more components that allow the server 300 to communicate with other devices (not shown) and servers (not shown). Other devices (not shown) may be computing devices or sensing devices such as the server 300, but are not limited thereto. The communication unit 320 may receive user input from another electronic device or receive data stored in an external device via a network.

[0156] The memory 330 can store programs for processing and control by the processor 310. For example, the memory 330 can store information input to the server or information received from another device via a network. Additionally, the memory 330 can store data generated by the processor 310. The memory 330 can also store information input to or output from the server 300.

[0157] The memory 330 is of the flash memory type, hard disk type, multimedia card micro type, card type memory (e.g., SD or XD memory), and RAM (RAM, random access memory), SRAM (static RAM). It can also include random access memory, ROM (read-only memory), EEPROM (electrically erasable programmable read-only memory), PROM (programmable read-only memory), magnetic storage, disk, and may include at least one of these types of storage media, optical disc.

[0158] For example, camera devices, sensing devices, and in-ear devices may include the aforementioned processor, memory, and communication unit.

[0159] The above embodiments can be implemented using hardware components, software components, and / or combinations of hardware and software components. For example, the apparatus, methods, and components described in the embodiments may include, for example, processors, controllers, arithmetic logic units (ALUs), digital signal processors, microcomputers, and field-programmable gate arrays (FPGAs). It can be implemented using one or more general-purpose or special-purpose computers, such as arrays, programmable logic units (PLUs), microprocessors, or any other device capable of executing and responding to instructions. The processing device may execute an operating system (OS) and one or more software applications running on the operating system. Additionally, the processing device may access, store, manipulate, process, and generate data in response to the execution of software. For ease of understanding, it may be described as using a single processing device; however, those skilled in the art will understand that a processing device may include multiple processing elements and / or various types of processing elements. For example, a processing apparatus may include multiple processors or one processor and one controller. Furthermore, other processing configurations, such as parallel processors, are also possible.

[0160] Software can include computer programs, code, instructions, or combinations thereof, which can configure processing units to operate as needed, or can independently or jointly command devices. Software and / or data can be used on any type of machine, component, physical device, virtual device, computer storage medium, or device to be interpreted by or to provide instructions or data to processing devices, or can be permanently or temporarily embodied in transmitted signal waves. Software can be distributed across networked computer systems and thus stored or executed in a distributed manner. Software and data can be stored on one or more computer-readable recording media.

[0161] The method according to the embodiments can be implemented in the form of program instructions executable by various computer devices and recorded on a computer-readable medium. The computer-readable medium may include program instructions, data files, data structures, etc., individually or in combination. The program instructions recorded on the medium may be specifically designed and configured for the embodiments, or may be known and available to those skilled in the art of computer software. Examples of computer-readable recording media include magnetic media such as hard disks, floppy disks, and magnetic tapes; optical media such as CD-ROMs and DVDs; and magnetic media such as floppy disks—including optical media (magneto-optical media). Hardware devices specifically designed for storing and executing program instructions, such as ROM, RAM, flash memory, etc. Examples of program instructions include machine language code, such as code generated by a compiler, and high-level language code that can be executed by a computer using an interpreter. The aforementioned hardware devices may be configured to operate as one or more software modules to perform the operations of the embodiments, and vice versa.

[0162] Although embodiments have been described above using limited accompanying drawings, those skilled in the art can make various modifications and variations based on the above application of techniques. For example, the described techniques may be performed in a different order than the described methods, and / or components of the described systems, structures, devices, circuits, etc. may be combined or combined in a different form than the described methods, or other components may be substituted or replaced with equivalents, and appropriate results may be obtained.

[0163] Therefore, other implementations, other embodiments, and equivalents of the claims also fall within the scope of the claims described below.

Claims

1. A method for a server to provide feedback information about a musical performance based on audience reaction information, comprising: Receive video information about the audience of the music performance from camera equipment installed on the music performance stage; The system receives biometric information from the audience members wearing sensor devices. This biometric information includes heart rate data, body temperature data, respiratory rate data, and skin conductance data. Based on video information, a first neural network is used to determine motion information of the audience regarding a musical performance through a motion analysis model; Based on motion information and biometric information, a response analysis model using a second neural network is used to determine audience response information regarding a musical performance. The feedback information for the musical performance is determined based on the response information and the pre-set expected response information; and This includes transmitting feedback information from the musical performance to the director's terminal. Feedback on the musical performance includes feedback on the dialogue and songs included in each of the multiple scenes that make up the musical performance.