system

The system addresses the inflexibility of robot dance systems by using music analysis and generative AI to create real-time dance sequences based on user input, ensuring immediate and adaptable performances.

JP2026062287APending Publication Date: 2026-04-09SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-09-30
Publication Date
2026-04-09

AI Technical Summary

Technical Problem

Existing robot dance systems require detailed programming in advance and lack flexibility to generate and execute real-time dances in response to music and user-specified conditions, resulting in a lack of immediacy and adaptability for events and promotions.

Method used

A system that acquires music data from the environment, analyzes it using Fast Fourier Transform, and uses generative artificial intelligence to generate dance sequences based on user-specified conditions, encoding these sequences as robot motion commands for real-time execution.

Benefits of technology

Enables instant generation and execution of dance performances synchronized with music and user instructions, providing high flexibility and immediacy in events and promotional settings.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026062287000001_ABST
    Figure 2026062287000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] A means for acquiring music data from the surrounding environment, A means for analyzing acquired music data and extracting characteristic data of the music, A means of receiving dance conditions from users, A means for generating a dance sequence using generative artificial intelligence based on extracted music characteristic data and received dance conditions, A means for encoding the generated dance sequence as robot motion command data and transmitting it to the robot, A means by which a robot performs dance movements based on motion command data it receives, A system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] The current robot dance system has the problem that it requires detailed programming in advance and it is difficult for users to generate and execute flexible dance performances immediately. Furthermore, there are few systems that can generate real-time dances in accordance with music, and it is also difficult for them to perform operations according to conditions specified by users. As a result, it lacks immediacy and flexibility in events and promotions.

Means for Solving the Problems

[0005] To solve this problem, the present invention proposes the following configuration. That is,

[0006] A system including means for acquiring music data from the surrounding environment, means for analyzing the acquired music data and extracting musical characteristic data, means for receiving dance conditions from the user, means for generating a dance sequence using generative artificial intelligence based on the extracted musical characteristic data and received dance conditions, means for encoding the generated dance sequence as robot motion command data and transmitting it to the robot, and means for the robot to execute dance movements based on the received motion command data, makes it possible to instantly generate and execute dance performances based on user instructions. Furthermore, by analyzing music data using Fast Fourier Transform and using a generative artificial intelligence model that considers tempo and movement patterns, real-time dance generation synchronized with music can be achieved.

[0007] "Music data" refers to data that includes audio signals acquired from the surrounding environment.

[0008] "Music characteristic data" refers to data that includes information such as volume, tempo, time signature, and pitch, extracted from music data.

[0009] "Dance conditions" are parameters that indicate the style or atmosphere of the dance specified by the user (e.g., "fun feeling," "calm feeling").

[0010] "Generative artificial intelligence" is an artificial intelligence system that uses a pre-trained model to generate appropriate output from input data.

[0011] A "dance sequence" is data that includes a series of action commands generated based on dance conditions and music characteristic data.

[0012] A "robot" is a machine that performs physical actions based on the motion command data it receives.

[0013] "Motion command data" refers to data that contains specific instructions for controlling the robot's movements.

[0014] "Means for acquiring music data" refers to devices and methods for capturing and collecting music data from the surrounding environment.

[0015] "Means for analyzing music data and extracting characteristic data of music" refers to devices and methods for processing music data and extracting characteristic data of music.

[0016] "Means for receiving dance conditions" refers to devices or methods for receiving dance conditions from users as input.

[0017] "Means for generating dance sequences" refers to devices and methods for generating dance sequences using artificial intelligence based on music characteristic data and dance conditions.

[0018] "Means for encoding a dance sequence as robot motion command data and transmitting it to a robot" refers to a device or method for converting a generated dance sequence into a format understandable to a robot and transmitting it to the robot.

[0019] "Means for a robot to perform dance movements based on motion command data it has received" refers to a device or method for a robot to actually perform dance movements in accordance with the transmitted motion command data.

[0020] The Fast Fourier Transform is a mathematical method for converting time-domain signals into frequency-domain signals.

[0021] "Tempo and movement patterns" are parameters that refer to the speed and rhythm of music, as well as the types and sequences of movements. [Brief explanation of the drawing]

[0022] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of the data processing device and smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which a plurality of emotions are mapped. [Figure 10] It shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when an emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when an emotion engine is combined.

Embodiments for Carrying Out the Invention

[0023] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0024] First, the language used in the following description will be explained.

[0025] In the following embodiments, the signed processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Furthermore, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), and APU (Accelerated Processing Unit).

[0026] In the following embodiments, signed RAM (Random Access Memory) is a memory that temporarily stores information and is used as work memory by the processor.

[0027] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.

[0028] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0029] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0030] [First Embodiment]

[0031] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0032] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0033] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0034] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0035] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0036] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0037] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0038] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0039] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0040] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0041] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0042] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0043] This invention relates to a system in which a robot generates and performs dances in real time based on instructions from a user. The system analyzes the characteristics of music, generates an appropriate dance sequence based on dance conditions specified by the user, and transmits motion commands to the robot, enabling the robot to perform the dance as instructed.

[0044] System Overview

[0045] This system consists of the following main components:

[0046] 1. Means for acquiring music data

[0047] 2. Music Analysis Methods

[0048] 3. Means of receiving dance conditions

[0049] 4. Dance generation means

[0050] 5. Means for transmitting operation command data

[0051] 6. Robot motion execution means

[0052] Natural language explanation of program processing

[0053] 1. Acquisition of music data

[0054] The server acquires surrounding music data in real time via a microphone. This music data includes audio signals, which are treated as input data.

[0055] 2. Music Analysis

[0056] The server analyzes the acquired music data using algorithms such as FFT (Fast Fourier Transform). As a result of the analysis, characteristic data of the music, such as volume, tempo, and pitch, is extracted, and further processing is performed based on this data.

[0057] 3. Receiving dance conditions

[0058] The user inputs dance conditions such as "fun feeling" or "calm feeling" through the interface. This input data is sent to the server.

[0059] 4. Dance Generation

[0060] The server uses a generative artificial intelligence (generative AI) model to generate dance sequences, taking music characteristic data and dance conditions from the user as input.

[0061] The generation AI is pre-trained with various dance patterns and generates movements in real time that are appropriate to the tempo and rhythm of the music.

[0062] 5. Generation and transmission of operation command data

[0063] The server encodes the generated dance sequence into motion command data so that the robot can understand it. This data is transmitted to the terminal (robot) via the network.

[0064] 6. Execution of actions by the robot

[0065] The terminal (robot) analyzes the received motion command data and controls each motor and actuator. Based on the control signals, the robot performs the predetermined movements and performs a dance.

[0066] Specific example

[0067] Situation: Robot dancing at an event venue

[0068] 1. Pop music is playing at the event venue.

[0069] 2. The user uses a smartphone or dedicated terminal to instruct the robot to "do a fun dance."

[0070] 3. The server picks up the music from the venue via the microphone and acquires the music data.

[0071] 4. The server analyzes the music data using FFT and extracts characteristic data such as volume, tempo, and rhythm.

[0072] 5. The server provides the analysis results and the user-specified condition of "fun feeling" to the generating AI, and generates a dance sequence based on that.

[0073] 6. The server encodes the generated dance sequence as motion command data and transmits it to the terminal (robot) via the network.

[0074] 7. The terminal (robot) operates its motors and actuators based on the received motion command data to perform a "fun-feeling" dance.

[0075] Thus, this system can instantly generate and execute appropriate dance performances in response to user instructions, providing high flexibility and immediacy in events and promotions.

[0076] The following describes the processing flow.

[0077] Step 1:

[0078] Users use their smartphones or dedicated terminals to instruct the robot to "do a fun dance."

[0079] Step 2:

[0080] The server acquires music data in real time from the surroundings via a microphone sensor. This music data includes audio signals.

[0081] Step 3:

[0082] The server analyzes the acquired music data using FFT (Fast Fourier Transform). This analysis extracts characteristic data of the music, such as volume, tempo, and pitch.

[0083] Step 4:

[0084] The server temporarily stores music characteristic data in memory and receives dance conditions from the user (e.g., "happy feeling," "calm feeling").

[0085] Step 5:

[0086] The server generates a dance sequence using generative artificial intelligence (generative AI) based on the received dance conditions and music characteristics data. The generative AI selects appropriate movement patterns from the input data and outputs them as a dance sequence.

[0087] Step 6:

[0088] The server encodes the generated dance sequence as motion command data for robot operation. This motion command data is then converted into a format that the robot can understand and execute.

[0089] Step 7:

[0090] The server transmits operation command data to the terminal (robot) via the network.

[0091] Step 8:

[0092] The terminal analyzes the received operation command data and generates signals to control each motor and actuator based on that data.

[0093] Step 9:

[0094] The terminal then operates the motors and actuators appropriately according to the generated control signals. This allows the robot to perform the specified dance movements (e.g., waving arms, spinning).

[0095] Step 10:

[0096] The terminal (robot) performs a "fun-sounding" dance based on user instructions and the characteristics of the music, livening up the atmosphere at event venues and other locations.

[0097] As described above, it is possible for the robot to generate and execute dances in real time that match the characteristics of the music, based on user instructions.

[0098] (Example 1)

[0099] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0100] Conventional robot dance systems only reproduce pre-programmed movements and cannot flexibly respond to changes in music or environment. Furthermore, they struggle to reflect user emotions and intentions in real time, lacking the immediacy and flexibility needed for events and promotions. Therefore, there is a need to develop a system that generates dance sequences in real time based on music characteristics and user-specified conditions, and has the robot execute them.

[0101] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0102] In this invention, the server includes means for acquiring music data from the surrounding environment, means for analyzing the acquired music data and extracting music characteristic data, means for receiving dance conditions from the user, means for generating a dance sequence using generative artificial intelligence based on the extracted music characteristic data and the received dance conditions, means for encoding the generated dance sequence as robot motion command data and transmitting it to the robot, means for the robot to execute dance movements based on the received motion command data, means for analyzing the acquired music data using the Fast Fourier Transform, means for the generative artificial intelligence model to consider tempo and movement patterns in the generation of the dance sequence, and means for transmitting the generated motion command data to the robot using wireless communication. This enables rapid response to changes in music and environment, and allows for dance performances that reflect the user's emotions and intentions in real time.

[0103] "Surrounding environment" refers to the location where the system is installed and the physical space surrounding it.

[0104] "Music data" refers to audio signals and music information such as background music acquired from the surrounding environment.

[0105] "Music characteristic data" refers to attribute information of music, such as volume, tempo, rhythm, and pitch, obtained by analyzing music data.

[0106] "Dance conditions" refer to the user's requirements regarding the style and atmosphere of the dance, such as "fun" or "relaxed."

[0107] "Generative artificial intelligence" refers to a technology that uses machine learning algorithms to generate new data and information based on input data.

[0108] A "dance sequence" refers to a combination of motion commands that a robot executes.

[0109] "Motion command data" refers to the control signals and commands necessary for a robot to perform a specific action.

[0110] A "robot" refers to an autonomous or semi-autonomous device that performs specified actions through a mechanical structure and computer control.

[0111] The "Fast Fourier Transform" refers to an efficient algorithm for converting time-domain data into the frequency domain.

[0112] "Wireless communication" refers to the technology of sending and receiving data using radio waves or light waves.

[0113] "Tempo" refers to the speed of music, specifically the number of beats contained within a given time.

[0114] An "action pattern" refers to a sequence of actions performed under specific conditions.

[0115] This invention relates to a system in which a robot generates and performs dances in real time based on instructions from a user. The system analyzes the characteristics of music, generates an appropriate dance sequence based on dance conditions specified by the user, and transmits motion commands to the robot, enabling the robot to perform the dance as instructed.

[0116] System Configuration

[0117] This system consists of the following main components:

[0118] 1. Means for acquiring music data

[0119] 2. Music Analysis Methods

[0120] 3. Means of receiving dance conditions

[0121] 4. Dance generation means

[0122] 5. Means for transmitting operation command data

[0123] 6. Robot motion execution means

[0124] Music data acquisition method

[0125] The server acquires ambient music data in real time through a high-sensitivity microphone. Specifically, it uses a high-quality audio capture device (e.g., a high-sensitivity microphone) to capture ambient sound as a digital signal. This acquired music data is temporarily stored in the server's buffer memory.

[0126] Music analysis methods

[0127] The server uses the Python libraries SciPy and NumPy to perform a Fast Fourier Transform (FFT) to analyze the acquired music data. This analysis extracts frequency components, and then the LibROSA library is used to calculate characteristic data of the music, such as volume, tempo, and pitch. This characteristic data is used in the next step.

[0128] Dance condition receiving method

[0129] Users specify dance conditions through an application on their smartphone or a dedicated tablet device. The user interface is developed with React Native, allowing for easy input of conditions such as "fun" or "calm." This input data is sent to the server in real time.

[0130] Dance generation means

[0131] The server uses a generative artificial intelligence (generative AI) model to generate dance sequences, taking music characteristic data and dance conditions from the user as input. The generative AI is built using TENSORFLOW® and is pre-trained on a variety of dance styles. For example, if the music tempo is fast and the dance condition is "fun," the generative AI will generate an energetic and rhythmic dance sequence.

[0132] Operation command data transmission means

[0133] The server encodes the generated dance sequence into motion command data that the robot can understand. Specifically, it uses ROBO-API to convert the dance sequence into control signals for the servo motors. The generated motion command data is then transmitted to the terminal (robot) via WiFi or Bluetooth.

[0134] Robot motion execution means

[0135] The terminal (robot) analyzes the received motion command data. Based on the analyzed data, it controls each motor and actuator (e.g., Dynamixel servo motor) to execute the specified action in real time. For example, if a "fun" dance command is given, the robot will take light steps and dance cheerfully while swinging both arms.

[0136] Specific example

[0137] Situation: Robot dancing at an event venue

[0138] 1. Pop music is playing at the event venue.

[0139] 2. The user uses a smartphone or dedicated device to instruct the robot to "do a fun dance." The user selects "fun" from the app's dropdown menu and clicks the send button.

[0140] 3. The server captures the music from the venue via the microphone and acquires the music data. The moment the microphone picks up sound, the server starts processing.

[0141] 4. The server analyzes the music data using FFT and LibROSA, and extracts characteristic data such as volume, tempo, and rhythm. After analysis, the characteristic data is stored in memory.

[0142] 5. The server provides the generation AI with the analysis results and the user-specified condition of "fun feeling," and generates a dance sequence based on this. The generation AI generates the sequence in a few seconds, and the server checks the results.

[0143] 6. The server encodes the generated dance sequence as motion command data and sends it to the terminal (robot) via the network. As soon as the transmission is complete, the terminal sends a confirmation of receipt to the server.

[0144] 7. The terminal (robot) operates its motors and actuators based on the received motion command data to perform a "fun-feeling" dance. For example, the robot takes light steps and dances cheerfully, raising and swinging both arms.

[0145] Examples of prompt statements

[0146] The following are some possible prompts to input into the generative AI model:

[0147] "Please generate a robot dance sequence based on the following conditions: Music tempo is 120 BPM, and the user-specified emotion is 'happy'. The output format should be motion command data."

[0148] In this way, each step of the entire system is intricately coordinated, enabling robot dances to be performed in real time in response to user instructions.

[0149] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0150] Step 1:

[0151] The server acquires music data from its surroundings. Specifically, it uses a high-sensitivity microphone to capture audio signals in real time. It receives analog audio signals from the microphone as input, converts them to digital signals, and stores them in buffer memory. The output is digital music data.

[0152] Step 2:

[0153] The server analyzes the acquired music data. Specifically, it performs an FFT (Fast Fourier Transform) using the Python libraries SciPy and NumPy. It receives digital music data stored in buffer memory as input and extracts frequency components. It also calculates characteristic data such as volume, tempo, rhythm, and pitch using the LibROSA library. The output is characteristic data of the music.

[0154] Step 3:

[0155] Users specify dance conditions through an application on their smartphone or a dedicated tablet device. Users input dance conditions such as "fun feeling" or "calm feeling" through the interface. The input consists of the user's specified conditions and is transmitted from the application to the server in real time. The output is the user's dance condition data.

[0156] Step 4:

[0157] The server uses a generative artificial intelligence (generative AI) model to generate dance sequences, taking music characteristic data and dance condition data from the user as input. The generative AI is pre-trained on a variety of dance styles and generates the optimal dance sequence according to the input data. The input is music characteristic data and dance condition data, and the output is the generated dance sequence.

[0158] Step 5:

[0159] The server encodes the generated dance sequence as robot motion command data. Specifically, it uses ROBO-API to convert the dance sequence into servo motor control signals. The input is the generated dance sequence, and the output is the motion command data. This motion command data is transmitted to the terminal (robot) via WiFi or Bluetooth.

[0160] Step 6:

[0161] The terminal (robot) analyzes the received motion command data. Based on the analyzed data, it controls each motor and actuator to execute the specified action in real time. The input is motion command data, and the output is the actual robot's movement. For example, if a "fun" dance command is given, the robot will dance with light steps and swing its arms.

[0162] Through these processing steps, the system can perform robotic dances in real time in response to user instructions.

[0163] (Application Example 1)

[0164] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0165] Conventional robot dance systems only execute pre-programmed movements, making it difficult for them to respond to changes in the surrounding environment or music in real time. Furthermore, they lack a mechanism for users to spontaneously instruct the robots to dance, resulting in a lack of flexibility in entertainment and promotional settings. Additionally, they are unable to properly consider tempo and movement patterns when generating dance sequences, making it difficult to provide an engaging performance for audiences.

[0166] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0167] In this invention, the server includes means for acquiring music data from the surrounding environment, means for analyzing the acquired music data and extracting music characteristic data, means for receiving dance conditions from the user, means for generating a dance sequence using artificial intelligence based on the extracted music characteristic data and the received dance conditions, means for encoding the generated dance sequence as robot motion command data and transmitting it to the robot, means for the robot to execute dance movements based on the received motion command data, and means for the user to give improvisational dance instructions to the robot via a smartphone. As a result, the robot can perform dances in real time according to the characteristics of the surrounding music and respond to improvisational instructions from the user, providing high flexibility and immediacy for events and promotions.

[0168] "Means for acquiring music data" refers to devices or methods for acquiring music data from the surrounding environment.

[0169] "Music analysis means" refers to devices or methods that analyze acquired music data and extract characteristic data of the music.

[0170] A "dance condition receiving method" refers to a device or method for receiving dance conditions from users.

[0171] A "dance generation means" is a device or method that generates a dance sequence using artificial intelligence based on extracted music characteristic data and received dance conditions.

[0172] "Motion command data transmission means" refers to a device or method that encodes the generated dance sequence as motion command data for a robot and transmits it to the robot.

[0173] A "robot motion execution means" is a device or method that allows a robot to perform dance movements based on motion command data it has received.

[0174] A "smartphone" is a portable information terminal used by users to give improvised dance instructions to a robot.

[0175] "Generative artificial intelligence" is an artificial intelligence technology that generates dance sequences based on music characteristic data and user conditions.

[0176] This invention relates to a system in which a robot generates and performs dances in real time based on instructions from a user. The system analyzes the characteristics of music, generates an appropriate dance sequence based on dance conditions specified by the user, and transmits motion commands to the robot, enabling the robot to perform the dance as instructed.

[0177] System Overview

[0178] This system consists of the following main components:

[0179] 1. Means for acquiring music data

[0180] The server acquires music data from its surrounding environment. Specifically, it acquires music playing in environments such as event venues and stores in real time via microphones.

[0181] 2. Music Analysis Methods

[0182] The server analyzes the acquired music data using algorithms such as FFT (Fast Fourier Transform). As a result of the analysis, characteristic data of the music, such as volume, tempo, and pitch, is extracted.

[0183] 3. Means of receiving dance conditions

[0184] Users input dance conditions for the robot via their smartphones, such as "fun" or "elegant."

[0185] 4. Dance generation means

[0186] The server generates dance sequences using a generative AI model based on the extracted music characteristic data and received dance conditions. The generative AI model is pre-trained with various dance patterns and generates movements in real time that are appropriate to the tempo and rhythm of the music.

[0187] 5. Means for transmitting operation command data

[0188] The server encodes the generated dance sequence as robot motion command data and transmits it to the robot via the network. Bluetooth and Wi-Fi are used as communication methods.

[0189] 6. Robot motion execution means

[0190] The robot controls its motors and actuators based on the motion command data it receives, and performs the specified dance movements.

[0191] Specific example

[0192] Situation: Robot dancing at an event venue

[0193] 1. Pop music is playing at the event venue.

[0194] 2. The user uses a smartphone application to instruct the robot to "do a fun dance."

[0195] 3. The server captures the music from the venue via microphones and acquires the music data.

[0196] 4. The server analyzes the music data using FFT and extracts characteristic data such as volume, tempo, and rhythm.

[0197] 5. The server provides the analysis results and the user-specified condition of "fun feeling" to the generating AI, and generates a dance sequence based on that.

[0198] 6. The server encodes the generated dance sequence as motion command data and transmits it to the robot via the network.

[0199] 7. The robot operates its motors and actuators based on the received motion command data to perform a "fun-feeling" dance.

[0200] Thus, this system can instantly generate and execute appropriate dance performances in response to user instructions, providing high flexibility and immediacy in events and promotions.

[0201] Examples of prompt statements used

[0202] Based on the user's specified "fun" dance conditions, analyze the tempo and volume of the acquired music data and generate a suitable dance sequence.

[0203] The above describes specific embodiments for carrying out the present invention.

[0204] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0205] Step 1:

[0206] The server acquires music data from the surrounding environment. Specifically, it collects music data in real time through a microphone. The input for this step is the audio signal of the music, and the output is the acquired raw music data.

[0207] Step 2:

[0208] The server analyzes the acquired music data. It extracts characteristic music data using algorithms such as FFT (Fast Fourier Transform). The input to this step is the acquired music data, and the output is music characteristic data such as volume, tempo, and rhythm.

[0209] Step 3:

[0210] The user inputs dance conditions for the robot via their smartphone. For example, conditions such as "fun" or "elegant" can be specified. The input in this step is the user's dance conditions, and the output is data for those dance conditions.

[0211] Step 4:

[0212] The server generates a dance sequence using a generative AI model based on the music characteristic data and the received dance conditions. The input for this step is the music characteristic data and the user's dance conditions, and the output is the generated dance sequence.

[0213] Step 5:

[0214] The server encodes the generated dance sequence as robot motion command data and transmits it to the robot over the network. The input to this step is the generated dance sequence, and the output is the motion command data sent to the robot.

[0215] Step 6:

[0216] The robot controls its motors and actuators based on the received motion command data to perform the specified dance movements. The input for this step is the motion command data, and the output is the actual dance movements of the robot.

[0217] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0218] This invention relates to a robot dance system that incorporates an emotion engine that recognizes user emotions. This system analyzes the characteristics of music, generates an appropriate dance sequence according to the user's emotions and specified dance conditions, and transmits motion commands to the robot, enabling the robot to perform the dance in real time.

[0219] System Overview

[0220] This system consists of the following main components:

[0221] 1. Means for acquiring music data

[0222] 2. Music Analysis Methods

[0223] 3. Emotional Engine

[0224] 4. Means of receiving dance conditions

[0225] 5. Dance generation means

[0226] 6. Operation command data transmission means

[0227] 7. Robot motion execution means

[0228] Natural language explanation of program processing

[0229] 1. Acquisition of music data

[0230] The server acquires surrounding music data in real time via a microphone. This music data includes audio signals, which are treated as input data.

[0231] 2. Music Analysis

[0232] The server analyzes the acquired music data using algorithms such as FFT (Fast Fourier Transform). As a result of the analysis, characteristic data of the music, such as volume, tempo, and pitch, is extracted, and further processing is performed based on this data.

[0233] 3. Receiving dance conditions

[0234] The user inputs dance conditions such as "fun feeling" or "calm feeling" through the interface. This input data is sent to the server.

[0235] 4. Use of the Emotion Engine

[0236] The emotion engine integrated into the server recognizes emotions through the user's facial expressions, voice, and gestures. This uses cameras and microphones to collect data in real time.

[0237] 5. Emotion analysis

[0238] The server uses an emotion engine to analyze the collected data and determine the user's emotions (e.g., joy, sadness, surprise). This emotion data is also incorporated into the dance conditions.

[0239] 6. Dance Generation

[0240] The server uses a generative artificial intelligence (generative AI) model to generate dance sequences, taking music characteristic data, dance conditions from the user, and emotional data as input.

[0241] The generating AI is pre-trained with various dance patterns and generates movements in real time that are appropriate to the music's tempo and rhythm, as well as the user's emotions.

[0242] 7. Generation and transmission of operation command data

[0243] The server encodes the generated dance sequence as motion command data for robot operation. This data is transmitted to the terminal (robot) via the network.

[0244] 8. Execution of actions by the robot

[0245] The terminal (robot) analyzes the received motion command data and generates signals to control each motor and actuator.

[0246] The terminal then operates the motors and actuators appropriately according to the generated control signals. This allows the robot to perform predetermined movements and dance.

[0247] Specific example

[0248] Situation: Robot dancing at an event venue

[0249] 1. Pop music is playing at the event venue.

[0250] 2. The user uses a smartphone or dedicated terminal to instruct the robot to "do a fun dance."

[0251] 3. The server picks up the music in the venue via a microphone sensor and acquires the music data.

[0252] 4. The server analyzes the music data using FFT and extracts characteristic data such as volume, tempo, and rhythm.

[0253] 5. The server uses cameras and microphones to collect the user's facial expressions, voice, and gestures in real time, and uses an emotion engine to analyze the user's emotions.

[0254] 6. The server provides the analysis results and the user-specified condition of "fun feeling" to the generating AI, and generates a dance sequence based on that.

[0255] 7. The server encodes the generated dance sequence as motion command data and transmits it to the terminal (robot) via the network.

[0256] 8. The terminal (robot) operates its motors and actuators based on the received motion command data to perform a "fun-feeling" dance.

[0257] Thus, this system can instantly generate and execute appropriate dance performances in response to user instructions and emotions, providing high flexibility and immediacy in events and promotions.

[0258] The following describes the processing flow.

[0259] Step 1:

[0260] Users use their smartphones or dedicated terminals to instruct the robot to "do a fun dance."

[0261] Step 2:

[0262] The server acquires music data in real time from the surroundings via a microphone sensor. This music data includes audio signals.

[0263] Step 3:

[0264] The server analyzes the acquired music data using FFT (Fast Fourier Transform). This analysis extracts characteristic data of the music, such as volume, tempo, and pitch.

[0265] Step 4:

[0266] The server temporarily stores music characteristic data in memory.

[0267] Step 5:

[0268] The server uses cameras and microphones to collect the user's facial expressions, voice, and gestures in real time, and analyzes them using an emotion engine. The analysis recognizes the user's emotions (e.g., joy, surprise).

[0269] Step 6:

[0270] The server receives the analyzed user emotion data along with dance conditions from the user (e.g., "feeling happy").

[0271] Step 7:

[0272] The server generates dance sequences using generative artificial intelligence (generative AI) based on music characteristic data, dance conditions from the user, and emotional data. The generative AI selects appropriate dance patterns from the input data and outputs them as dance sequences.

[0273] Step 8:

[0274] The server encodes the generated dance sequence as operation command data for robot operation. This operation command data is converted into a format that the robot can understand and execute.

[0275] Step 9:

[0276] The server transmits the operation command data to the terminal (robot) via the network.

[0277] Step 10:

[0278] The terminal analyzes the received operation command data and generates signals for controlling each motor and actuator based on it.

[0279] Step 11:

[0280] The terminal operates the motors and actuators appropriately according to the generated control signals. Thereby, the robot executes a predetermined operation and performs a dance.

[0281] Step 12:

[0282] The terminal (robot) performs a "fun" dance based on the user's instructions and the characteristics of the music, enhancing the atmosphere of an event venue or the like.

[0283] In this way, it is possible for the robot to generate and execute a dance in accordance with the characteristics of the music in real time based on the user's instructions and emotion recognition.

[0284] (Example 2)

[0285] Next, Example 2 will be described. In the following description, the data processing device 12 is referred to as the "server", and the smart device 14 is referred to as the "terminal".

[0286] In conventional robot dance systems, it has been mainstream to simply execute movements based on the tempo and rhythm of music, and it has been difficult to reflect the user's emotions and specific dance conditions in real time. As a result, the user experience has been limited, and the expressiveness and flexibility of robot dance have been insufficient. The present invention aims to generate a dance sequence according to the user's emotions and specified dance conditions, enabling the robot to perform a more dynamic and adaptable dance in real time.

[0287] The specific processing by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for acquiring music data from the surrounding environment, means for analyzing the acquired music data and extracting characteristic data of the music, means for receiving dance conditions from the user, means for recognizing emotions through the user's expressions, voices, and gestures, means for analyzing the user's emotion data, means for generating a dance sequence using a generative artificial intelligence based on the extracted music characteristic data, received dance conditions, and recognized emotion data, means for encoding the generated dance sequence as operation command data for the robot and transmitting it to the robot, and means for executing a dance operation based on the operation command data received by the robot. Thereby, it becomes possible to execute a high-level and real-time robot dance in accordance with the user's emotions and dance conditions.

[0288] The "means for acquiring music data from the surrounding environment" is a function of capturing ambient audio signals in real time using sensors such as microphones built into the robot or server.

[0289] The "means for analyzing the acquired music data and extracting characteristic data of the music" is a function of applying algorithms such as FFT (Fast Fourier Transform) to the music data and calculating and extracting feature data such as volume, tempo, rhythm, and musical scale.

[0290] "Means of receiving dance conditions from users" refers to a function that inputs and receives dance conditions specified by the user (e.g., "a fun feeling" or "a calm feeling") via an interface such as a smartphone app or dedicated terminal.

[0291] "Means of recognizing emotions through the user's facial expressions, voice, and gestures" refers to a function that uses cameras and microphones to capture the user's facial expressions, voice tone, and gestures in real time, and recognizes emotions through an emotion engine.

[0292] "Methods for analyzing user emotion data" refers to a function that analyzes facial expressions, voice, and gesture data acquired in real time and converts the user's emotions into labels such as joy, sadness, and surprise.

[0293] "A means of generating a dance sequence using generative artificial intelligence based on extracted music characteristic data, received dance conditions, and recognized emotion data" refers to a function that uses a generative artificial intelligence model (generative AI model) to generate a suitable dance sequence in real time, taking music characteristic data, dance conditions, and emotion data as input.

[0294] "Means for encoding the generated dance sequence as robot motion command data and transmitting it to the robot" refers to a function that converts the generated dance sequence into specific control signals for robot operation and transmits them to the robot via a network.

[0295] "Means for executing dance movements based on motion command data received by the robot" refers to a function that, based on the control signals received by the robot, appropriately operates each motor and actuator to execute a predetermined dance movement.

[0296] The present invention relates to a robot dance system combined with an emotion engine that recognizes the emotions of users. This system analyzes the characteristics of music, generates an appropriate dance sequence according to the emotions of the user and the specified dance conditions, transmits an operation command to the robot, and enables the robot to perform a dance in real time.

[0297] Configuration of the system

[0298] This system is composed of the following main components:

[0299] 1. Music data acquisition means

[0300] The server uses sensors such as microphones built into the robot to acquire ambient music in real time, thereby collecting audio signals.

[0301] 2. Music analysis means

[0302] The server analyzes the acquired music data using algorithms such as FFT (Fast Fourier Transform) and extracts characteristic data such as volume, tempo, rhythm, and pitch.

[0303] 3. Dance condition receiving means

[0304] The user inputs dance conditions such as "happy feeling" or "calm feeling" via a smartphone app or a dedicated terminal. The data is transmitted to the server through the Internet.

[0305] 4. Emotion engine (emotion recognition means)

[0306] The server uses a camera and additional microphones to collect the user's expressions, voices, and gestures in real time, and recognizes the user's emotions using an emotion engine.

[0307] 5. Emotion analysis means

[0308] The server analyzes data collected through the emotion engine to determine the user's emotions (joy, sadness, surprise). This emotion data is also incorporated into the dance conditions.

[0309] 6. Dance generation means

[0310] The server generates dance sequences using a generative AI model based on music characteristic data, dance conditions from the user, and recognized emotion data.

[0311] 7. Operation command data transmission means

[0312] The server encodes the generated dance sequence as motion command data for robot operation and transmits it to the robot (terminal) via the network.

[0313] 8. Robot motion execution means

[0314] The terminal (robot) analyzes the received motion command data, generates signals to control each motor and actuator, and executes the specified dance movements.

[0315] Specific example

[0316] Situation: Robot dancing at an event venue

[0317] 1. Pop music is playing at the event venue.

[0318] 2. The user uses a smartphone or dedicated terminal to instruct the robot to "do a fun dance."

[0319] 3. The server picks up the music in the venue via a microphone sensor and acquires the music data.

[0320] 4. The server analyzes the music data using FFT and extracts characteristic data such as volume, tempo, and rhythm.

[0321] 5. The server uses cameras and microphones to collect the user's facial expressions, voice, and gestures in real time, and uses an emotion engine to analyze the user's emotions.

[0322] 6. The server provides the analysis results and the user-specified condition of "fun feeling" as prompts to the generating AI, and generates a dance sequence based on them.

[0323] 7. The server encodes the generated dance sequence as motion command data and transmits it to the terminal (robot) via the network.

[0324] 8. The terminal (robot) operates its motors and actuators based on the received motion command data to perform a "fun-feeling" dance.

[0325] Example of a prompt

[0326] "Generate a dance sequence that fits a fun, upbeat pop song. The user is expressing joy. The tempo is 120 BPM, and the volume is medium."

[0327] By inputting this prompt into the AI ​​model, it is possible to generate a dance sequence that is suitable for the user's emotions and musical characteristics.

[0328] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0329] Step 1: Obtain music data

[0330] The server uses sensors, such as microphones built into the robot, to acquire ambient music in real time. The music data acquired at this stage (input) is stored in a buffer as an audio signal (output).

[0331] Specific operation: The server captures audio data at regular time intervals and monitors it continuously.

[0332] Step 2: Music Analysis

[0333] The server analyzes the acquired music data (input) using FFT (Fast Fourier Transform) and extracts characteristic data (output) such as frequency components, volume, tempo, rhythm, and pitch.

[0334] Specific operation: The server performs an FFT on each time window to extract time-frequency domain features. A beat detection algorithm is then applied to calculate tempo data.

[0335] Step 3: Receiving the dance conditions

[0336] Users access the interface via a smartphone app or dedicated device and input dance conditions (inputs) such as "fun feeling" or "calm feeling." This input is sent to a server via the internet and stored as dance condition data (output).

[0337] Specific operation: The user selects the dance conditions on the device's UI and presses the send button.

[0338] Step 4: Using the Emotion Engine

[0339] The server uses cameras and additional microphones to collect the user's facial expressions, voice, and gestures (input) in real time, and recognizes emotional data (output) through an emotion engine.

[0340] Specific operation: Video data acquired by the camera is processed by a face recognition algorithm, and facial expression data is extracted. Similarly, a voice analysis algorithm is also applied.

[0341] Step 5: Emotion Analysis

[0342] The server analyzes real-time collected facial, voice, and gesture data (input) and uses an emotion engine to determine the user's emotions (e.g., joy, sadness, surprise). This emotion data (output) is then incorporated into the dance conditions.

[0343] Specific operation: The emotion engine uses a facial expression recognition model and a voice tone analysis model to generate emotion labels.

[0344] Step 6: Dance Generation

[0345] The server sends prompts to the generative AI model based on music characteristic data, dance conditions from the user, and recognized emotion data (input), and generates a dance sequence (output).

[0346] Specific operation: The server inputs a prompt message such as "Generate a dance sequence that fits a fun pop song" into the AI ​​model and retrieves the generated dance sequence data.

[0347] Step 7: Generate and transmit operation command data

[0348] The server encodes the generated dance sequence (input) as motion command data for robot operation and transmits it to the robot (terminal) via the network. This motion command data is generated as output.

[0349] Specific operation: The server converts the generated dance sequence into control data for each joint and motor, and transmits it at the appropriate timing.

[0350] Step 8: Robot execution of actions

[0351] The terminal (robot) analyzes the received motion command data (input), generates signals to control each motor and actuator, and executes the specified dance movement (output).

[0352] Specific operation: The terminal's control board analyzes the received data and sends the appropriate current and voltage to each motor to execute the specified operation in real time.

[0353] (Application Example 2)

[0354] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0355] Conventional robot dance systems automatically generate dances based on user emotions and music characteristics. However, to improve work efficiency on a factory production line, it is necessary to consider the emotions and work conditions of the workers. This can improve work efficiency and motivation. However, existing systems lack the technology to appropriately acquire emotional data and generate response sequences. Therefore, there is a need for a robot system that can perform appropriate responses according to the emotions and work conditions of workers on a factory production line.

[0356] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0357] In this invention, the server includes means for acquiring music data, means for analyzing music data, means for receiving dance conditions from a user, means for generating a dance sequence based on extracted music characteristic data and received dance conditions, means for encoding the generated dance sequence as robot motion command data, means for transmitting it to a robot, means for the robot to execute dance movements based on the received motion command data, means for acquiring worker emotion data from the work environment, means for generating a response sequence for improving work efficiency based on the acquired emotion data, means for encoding the generated response sequence as robot motion command data, means for transmitting it, and means for the robot to execute response movements based on the received motion command data. This makes it possible to provide appropriate responses in real time according to the worker's emotions and work situation in the work environment.

[0358] "Music data" refers to the audio signals and waveform data that make up music.

[0359] "Characteristic data" refers to distinctive information such as tempo, volume, and rhythm extracted from music data.

[0360] "Dance conditions" refer to the requirements specified by the user, such as the type of dance and the atmosphere.

[0361] "Generative artificial intelligence" refers to artificial intelligence technology used to generate new data or sequences based on existing information.

[0362] A "dance sequence" refers to a continuous flow or pattern of movements. It specifically refers to a series of actions in a dance performance.

[0363] "Action command data" refers to a set of instructions that cause a robot to perform a specific action.

[0364] A "robot" refers to a mechanical device that performs actions automatically.

[0365] "Emotional data" refers to emotional information analyzed from the workers' facial expressions, voices, gestures, etc.

[0366] A "response sequence" refers to a sequence of actions and reactions generated to respond appropriately to emotions and situations.

[0367] "Work environment" refers to the location where work is performed, such as a factory or production line, and the surrounding conditions.

[0368] "Workers" refers to people who perform tasks on production lines or similar environments.

[0369] This invention is a system that controls the movements of a robot based on music data and emotion data, thereby improving work efficiency and motivation in the work environment. The following describes specific embodiments of this system.

[0370] System configuration and operation

[0371] 1. Acquisition of music data

[0372] The server acquires music data from the surrounding environment in real time using a microphone. This music data includes audio signals.

[0373] 2. Analysis of music data

[0374] The server analyzes the acquired music data using algorithms such as FFT (Fast Fourier Transform) to extract characteristic data such as volume, tempo, and rhythm. This process utilizes, for example, high-performance computing servers or cloud computing resources.

[0375] 3. Receiving dance conditions

[0376] The server receives dance conditions such as "feeling happy" or "feeling calm" from the user via smartphone or interface. This information is transmitted to the server over the network.

[0377] 4. Acquisition of emotional data

[0378] The server collects the workers' facial expressions, voices, and gestures in real time using cameras and microphones, and analyzes them using an emotion engine. This emotion data is also stored on the server.

[0379] 5. Generation of the dance sequence

[0380] The server uses a generative AI model (e.g., GPT-4®) to generate dance sequences and response sequences, taking music characteristic data, user dance conditions, and emotional data as input. The generated sequences are then coded as command data for robot movements.

[0381] 6. Transmission of command data

[0382] The server transmits the generated motion command data to the robot via the network. The robot's terminal analyzes this data and sends control signals to each motor and actuator.

[0383] 7. Execution of actions by the robot

[0384] The robot controls its motors and actuators appropriately based on motion command data received from the server to perform specified dances and work response actions.

[0385] Specific examples of hardware and software

[0386] Camera: High-resolution, night vision camera (e.g., Logitech C920s Pro)

[0387] Microphone: High-sensitivity microphone (e.g., Blue Yeti USB Microphone)

[0388] Sensors: Belt conveyor speed sensors and quality control sensors (e.g., Honeywell FF-SY series)

[0389] Server: High-performance server (e.g., NVIDIA DGX Station) or cloud computing resources

[0390] Generative AI models: Generative models based on GPT-4 and Transformers

[0391] Network: Stable WiFi or Ethernet connection

[0392] Specific example

[0393] Situation: Improving production efficiency within a factory

[0394] 1. The server acquires data in real time from conveyor belt speed sensors and quality control sensors within the factory.

[0395] 2. The server uses a camera and microphone to collect the worker's facial expressions, voice, and gestures, and analyzes them using an emotion engine. Based on this analysis, it recognizes that the worker is tired.

[0396] 3. The server provides the generation AI model with prompts like the following to generate a response sequence:

[0397] "If the workers are tired, please generate a refreshing dance to increase production efficiency."

[0398] 4. The generated response sequence is sent to the robot, which then performs nimble steps and simple exercises to increase worker motivation and improve efficiency.

[0399] This system makes it possible to generate and execute appropriate robot movements in real time based on music data and emotion data, contributing to improved productivity and employee motivation in the work environment.

[0400] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0401] Step 1:

[0402] The server acquires music data from the surrounding environment. Specifically, it uses a microphone to collect audio signals in real time and saves them as audio data. The input for this step is ambient music, and the output is the acquired audio data.

[0403] Step 2:

[0404] The server analyzes the acquired music data. Specifically, it uses the FFT (Fast Fourier Transform) algorithm to extract characteristic data of the music (volume, tempo, rhythm, etc.). The input for this step is audio data, and the output is characteristic data.

[0405] Step 3:

[0406] The server receives dance conditions from the user. Specifically, it receives conditions such as "feeling happy" or "feeling calm" that the user sends via smartphone or interface over the network. The input for this step is the user's dance conditions, and the output is the received dance conditions.

[0407] Step 4:

[0408] The server acquires emotional data from the work environment. Specifically, it uses cameras and microphones to collect the workers' facial expressions, voices, and gestures in real time, and analyzes them using an emotion engine. The input for this step is the workers' facial expressions, voices, and gestures, and the output is the analyzed emotional data.

[0409] Step 5:

[0410] The server uses a generative AI model (e.g., GPT-4) to generate dance sequences and response sequences, taking music characteristic data, dance conditions from the user, and emotion data as input. Specifically, it provides the generative AI model with prompts like the following to generate appropriate sequences. The input for this step is characteristic data, dance conditions, and emotion data, and the output is the generated sequence.

[0411] Example prompt: "If the workers are tired, generate a refresh dance sequence to increase production efficiency."

[0412] Step 6:

[0413] The server encodes the generated sequence as robot motion command data and transmits it to the robot via the network. Specifically, it converts the generated sequence into control signals and sends them to the robot. The input to this step is the generated sequence, and the output is the robot motion command data.

[0414] Step 7:

[0415] The terminal (robot) controls its motors and actuators based on motion command data received from the server, and performs the specified actions. Specifically, it performs dances and response actions according to the actions commanded by the server. The input for this step is motion command data, and the output is the performed action.

[0416] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0417] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0418] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0419] [Second Embodiment]

[0420] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0421] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0422] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0423] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0424] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0425] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0426] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0427] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0428] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0429] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0430] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0431] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0432] This invention relates to a system in which a robot generates and performs dances in real time based on instructions from a user. The system analyzes the characteristics of music, generates an appropriate dance sequence based on dance conditions specified by the user, and transmits motion commands to the robot, enabling the robot to perform the dance as instructed.

[0433] System Overview

[0434] This system consists of the following main components:

[0435] 1. Means for acquiring music data

[0436] 2. Music Analysis Methods

[0437] 3. Means of receiving dance conditions

[0438] 4. Dance generation means

[0439] 5. Means for transmitting operation command data

[0440] 6. Robot motion execution means

[0441] Natural language explanation of program processing

[0442] 1. Acquisition of music data

[0443] The server acquires surrounding music data in real time via a microphone. This music data includes audio signals, which are treated as input data.

[0444] 2. Music Analysis

[0445] The server analyzes the acquired music data using algorithms such as FFT (Fast Fourier Transform). As a result of the analysis, characteristic data of the music, such as volume, tempo, and pitch, is extracted, and further processing is performed based on this data.

[0446] 3. Receiving dance conditions

[0447] The user inputs dance conditions such as "fun feeling" or "calm feeling" through the interface. This input data is sent to the server.

[0448] 4. Dance Generation

[0449] The server uses a generative artificial intelligence (generative AI) model to generate dance sequences, taking music characteristic data and dance conditions from the user as input.

[0450] The generation AI is pre-trained with various dance patterns and generates movements in real time that are appropriate to the tempo and rhythm of the music.

[0451] 5. Generation and transmission of operation command data

[0452] The server encodes the generated dance sequence into motion command data so that the robot can understand it. This data is transmitted to the terminal (robot) via the network.

[0453] 6. Execution of actions by the robot

[0454] The terminal (robot) analyzes the received motion command data and controls each motor and actuator. Based on the control signals, the robot performs the predetermined movements and performs a dance.

[0455] Specific example

[0456] Situation: Robot dancing at an event venue

[0457] 1. Pop music is playing at the event venue.

[0458] 2. The user uses a smartphone or dedicated terminal to instruct the robot to "do a fun dance."

[0459] 3. The server picks up the music from the venue via the microphone and acquires the music data.

[0460] 4. The server analyzes the music data using FFT and extracts characteristic data such as volume, tempo, and rhythm.

[0461] 5. The server provides the analysis results and the user-specified condition of "fun feeling" to the generating AI, and generates a dance sequence based on that.

[0462] 6. The server encodes the generated dance sequence as motion command data and transmits it to the terminal (robot) via the network.

[0463] 7. The terminal (robot) operates its motors and actuators based on the received motion command data to perform a "fun-feeling" dance.

[0464] Thus, this system can instantly generate and execute appropriate dance performances in response to user instructions, providing high flexibility and immediacy in events and promotions.

[0465] The following describes the processing flow.

[0466] Step 1:

[0467] Users use their smartphones or dedicated terminals to instruct the robot to "do a fun dance."

[0468] Step 2:

[0469] The server acquires music data in real time from the surroundings via a microphone sensor. This music data includes audio signals.

[0470] Step 3:

[0471] The server analyzes the acquired music data using FFT (Fast Fourier Transform). This analysis extracts characteristic data of the music, such as volume, tempo, and pitch.

[0472] Step 4:

[0473] The server temporarily stores music characteristic data in memory and receives dance conditions from the user (e.g., "happy feeling," "calm feeling").

[0474] Step 5:

[0475] The server generates a dance sequence using generative artificial intelligence (generative AI) based on the received dance conditions and music characteristics data. The generative AI selects appropriate movement patterns from the input data and outputs them as a dance sequence.

[0476] Step 6:

[0477] The server encodes the generated dance sequence as motion command data for robot operation. This motion command data is then converted into a format that the robot can understand and execute.

[0478] Step 7:

[0479] The server transmits operation command data to the terminal (robot) via the network.

[0480] Step 8:

[0481] The terminal analyzes the received operation command data and generates signals to control each motor and actuator based on that data.

[0482] Step 9:

[0483] The terminal then operates the motors and actuators appropriately according to the generated control signals. This allows the robot to perform the specified dance movements (e.g., waving arms, spinning).

[0484] Step 10:

[0485] The terminal (robot) performs a "fun-sounding" dance based on user instructions and the characteristics of the music, livening up the atmosphere at event venues and other locations.

[0486] As described above, it is possible for the robot to generate and execute dances in real time that match the characteristics of the music, based on user instructions.

[0487] (Example 1)

[0488] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0489] Conventional robot dance systems only reproduce pre-programmed movements and cannot flexibly respond to changes in music or environment. Furthermore, they struggle to reflect user emotions and intentions in real time, lacking the immediacy and flexibility needed for events and promotions. Therefore, there is a need to develop a system that generates dance sequences in real time based on music characteristics and user-specified conditions, and has the robot execute them.

[0490] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0491] In this invention, the server includes means for acquiring music data from the surrounding environment, means for analyzing the acquired music data and extracting music characteristic data, means for receiving dance conditions from the user, means for generating a dance sequence using generative artificial intelligence based on the extracted music characteristic data and the received dance conditions, means for encoding the generated dance sequence as robot motion command data and transmitting it to the robot, means for the robot to execute dance movements based on the received motion command data, means for analyzing the acquired music data using the Fast Fourier Transform, means for the generative artificial intelligence model to consider tempo and movement patterns in the generation of the dance sequence, and means for transmitting the generated motion command data to the robot using wireless communication. This enables rapid response to changes in music and environment, and allows for dance performances that reflect the user's emotions and intentions in real time.

[0492] "Surrounding environment" refers to the location where the system is installed and the physical space surrounding it.

[0493] "Music data" refers to audio signals and music information such as background music acquired from the surrounding environment.

[0494] "Music characteristic data" refers to attribute information of music, such as volume, tempo, rhythm, and pitch, obtained by analyzing music data.

[0495] "Dance conditions" refer to the user's requirements regarding the style and atmosphere of the dance, such as "fun" or "relaxed."

[0496] "Generative artificial intelligence" refers to a technology that uses machine learning algorithms to generate new data and information based on input data.

[0497] A "dance sequence" refers to a combination of motion commands that a robot executes.

[0498] "Motion command data" refers to the control signals and commands necessary for a robot to perform a specific action.

[0499] A "robot" refers to an autonomous or semi-autonomous device that performs specified actions through a mechanical structure and computer control.

[0500] The "Fast Fourier Transform" refers to an efficient algorithm for converting time-domain data into the frequency domain.

[0501] "Wireless communication" refers to the technology of sending and receiving data using radio waves or light waves.

[0502] "Tempo" refers to the speed of music, specifically the number of beats contained within a given time.

[0503] An "action pattern" refers to a sequence of actions performed under specific conditions.

[0504] This invention relates to a system in which a robot generates and performs dances in real time based on instructions from a user. The system analyzes the characteristics of music, generates an appropriate dance sequence based on dance conditions specified by the user, and transmits motion commands to the robot, enabling the robot to perform the dance as instructed.

[0505] System Configuration

[0506] This system consists of the following main components:

[0507] 1. Means for acquiring music data

[0508] 2. Music Analysis Methods

[0509] 3. Means of receiving dance conditions

[0510] 4. Dance generation means

[0511] 5. Means for transmitting operation command data

[0512] 6. Robot motion execution means

[0513] Music data acquisition method

[0514] The server acquires ambient music data in real time through a high-sensitivity microphone. Specifically, it uses a high-quality audio capture device (e.g., a high-sensitivity microphone) to capture ambient sound as a digital signal. This acquired music data is temporarily stored in the server's buffer memory.

[0515] Music analysis methods

[0516] The server uses the Python libraries SciPy and NumPy to perform a Fast Fourier Transform (FFT) to analyze the acquired music data. This analysis extracts frequency components, and then the LibROSA library is used to calculate characteristic data of the music, such as volume, tempo, and pitch. This characteristic data is used in the next step.

[0517] Dance condition receiving method

[0518] Users specify dance conditions through an application on their smartphone or a dedicated tablet device. The user interface is developed with React Native, allowing for easy input of conditions such as "fun" or "calm." This input data is sent to the server in real time.

[0519] Dance generation means

[0520] The server uses a generative artificial intelligence (generative AI) model to generate dance sequences, taking music characteristic data and dance conditions from the user as input. The generative AI is built using TensorFlow and has been pre-trained on a variety of dance styles. For example, if the music tempo is fast and the dance condition is "fun," the generative AI will generate an energetic and rhythmic dance sequence.

[0521] Operation command data transmission means

[0522] The server encodes the generated dance sequence into motion command data that the robot can understand. Specifically, it uses ROBO-API to convert the dance sequence into control signals for the servo motors. The generated motion command data is then transmitted to the terminal (robot) via WiFi or Bluetooth.

[0523] Robot motion execution means

[0524] The terminal (robot) analyzes the received motion command data. Based on the analyzed data, it controls each motor and actuator (e.g., Dynamixel servo motor) to execute the specified action in real time. For example, if a "fun" dance command is given, the robot will take light steps and dance cheerfully while swinging both arms.

[0525] Specific example

[0526] Situation: Robot dancing at an event venue

[0527] 1. Pop music is playing at the event venue.

[0528] 2. The user uses a smartphone or dedicated device to instruct the robot to "do a fun dance." The user selects "fun" from the app's dropdown menu and clicks the send button.

[0529] 3. The server captures the music from the venue via the microphone and acquires the music data. The server processing begins the moment the microphone picks up sound.

[0530] 4. The server analyzes the music data using FFT and LibROSA, and extracts characteristic data such as volume, tempo, and rhythm. After analysis, the characteristic data is stored in memory.

[0531] 5. The server provides the generation AI with the analysis results and the user-specified condition of "fun feeling," and generates a dance sequence based on this. The generation AI generates the sequence in a few seconds, and the server checks the results.

[0532] 6. The server encodes the generated dance sequence as motion command data and sends it to the terminal (robot) via the network. As soon as the transmission is complete, the terminal sends a receipt confirmation to the server.

[0533] 7. The terminal (robot) operates its motors and actuators based on the received motion command data to perform a "fun-feeling" dance. For example, the robot takes light steps and dances cheerfully, raising and swinging both arms.

[0534] Examples of prompt statements

[0535] The following are some possible prompts to input into the generative AI model:

[0536] "Please generate a robot dance sequence based on the following conditions: Music tempo is 120 BPM, and the user-specified emotion is 'happy'. The output format should be motion command data."

[0537] In this way, each step of the entire system is intricately coordinated, enabling robot dances to be performed in real time in response to user instructions.

[0538] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0539] Step 1:

[0540] The server acquires music data from its surroundings. Specifically, it uses a high-sensitivity microphone to capture audio signals in real time. It receives analog audio signals from the microphone as input, converts them to digital signals, and stores them in buffer memory. The output is digital music data.

[0541] Step 2:

[0542] The server analyzes the acquired music data. Specifically, it performs an FFT (Fast Fourier Transform) using the Python libraries SciPy and NumPy. It receives digital music data stored in buffer memory as input and extracts frequency components. It also calculates characteristic data such as volume, tempo, rhythm, and pitch using the LibROSA library. The output is characteristic data of the music.

[0543] Step 3:

[0544] Users specify dance conditions through an application on their smartphone or a dedicated tablet device. Users input dance conditions such as "fun feeling" or "calm feeling" through the interface. The input consists of the user's specified conditions and is transmitted from the application to the server in real time. The output is the user's dance condition data.

[0545] Step 4:

[0546] The server uses a generative artificial intelligence (generative AI) model to generate dance sequences, taking music characteristic data and dance condition data from the user as input. The generative AI is pre-trained on a variety of dance styles and generates the optimal dance sequence according to the input data. The input is music characteristic data and dance condition data, and the output is the generated dance sequence.

[0547] Step 5:

[0548] The server encodes the generated dance sequence as robot motion command data. Specifically, it uses ROBO-API to convert the dance sequence into servo motor control signals. The input is the generated dance sequence, and the output is the motion command data. This motion command data is transmitted to the terminal (robot) via WiFi or Bluetooth.

[0549] Step 6:

[0550] The terminal (robot) analyzes the received motion command data. Based on the analyzed data, it controls each motor and actuator to execute the specified action in real time. The input is motion command data, and the output is the actual robot's movement. For example, if a "fun" dance command is given, the robot will dance with light steps and swing its arms.

[0551] Through these processing steps, the system can perform robotic dances in real time in response to user instructions.

[0552] (Application Example 1)

[0553] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0554] Conventional robot dance systems only execute pre-programmed movements, making it difficult for them to respond to changes in the surrounding environment or music in real time. Furthermore, they lack a mechanism for users to spontaneously instruct the robots to dance, resulting in a lack of flexibility in entertainment and promotional settings. Additionally, they are unable to properly consider tempo and movement patterns when generating dance sequences, making it difficult to provide an engaging performance for audiences.

[0555] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0556] In this invention, the server includes means for acquiring music data from the surrounding environment, means for analyzing the acquired music data and extracting music characteristic data, means for receiving dance conditions from the user, means for generating a dance sequence using artificial intelligence based on the extracted music characteristic data and the received dance conditions, means for encoding the generated dance sequence as robot motion command data and transmitting it to the robot, means for the robot to execute dance movements based on the received motion command data, and means for the user to give improvisational dance instructions to the robot via a smartphone. As a result, the robot can perform dances in real time according to the characteristics of the surrounding music and respond to improvisational instructions from the user, providing high flexibility and immediacy for events and promotions.

[0557] "Means for acquiring music data" refers to devices or methods for acquiring music data from the surrounding environment.

[0558] "Music analysis means" refers to devices or methods that analyze acquired music data and extract characteristic data of the music.

[0559] A "dance condition receiving method" refers to a device or method for receiving dance conditions from users.

[0560] A "dance generation means" is a device or method that generates a dance sequence using artificial intelligence based on extracted music characteristic data and received dance conditions.

[0561] "Motion command data transmission means" refers to a device or method that encodes the generated dance sequence as motion command data for a robot and transmits it to the robot.

[0562] A "robot motion execution means" is a device or method that allows a robot to perform dance movements based on motion command data it has received.

[0563] A "smartphone" is a portable information terminal used by users to give improvised dance instructions to a robot.

[0564] "Generative artificial intelligence" is an artificial intelligence technology that generates dance sequences based on music characteristic data and user conditions.

[0565] This invention relates to a system in which a robot generates and performs dances in real time based on instructions from a user. The system analyzes the characteristics of music, generates an appropriate dance sequence based on dance conditions specified by the user, and transmits motion commands to the robot, enabling the robot to perform the dance as instructed.

[0566] System Overview

[0567] This system consists of the following main components:

[0568] 1. Means for acquiring music data

[0569] The server acquires music data from its surrounding environment. Specifically, it acquires music playing in environments such as event venues and stores in real time via microphones.

[0570] 2. Music Analysis Methods

[0571] The server analyzes the acquired music data using algorithms such as FFT (Fast Fourier Transform). As a result of the analysis, characteristic data of the music, such as volume, tempo, and pitch, is extracted.

[0572] 3. Means of receiving dance conditions

[0573] Users input dance conditions for the robot via their smartphones, such as "fun" or "elegant."

[0574] 4. Dance generation means

[0575] The server generates dance sequences using a generative AI model based on the extracted music characteristic data and received dance conditions. The generative AI model is pre-trained with various dance patterns and generates movements in real time that are appropriate to the tempo and rhythm of the music.

[0576] 5. Means for transmitting operation command data

[0577] The server encodes the generated dance sequence as robot motion command data and transmits it to the robot via the network. Bluetooth and Wi-Fi are used as communication methods.

[0578] 6. Robot motion execution means

[0579] The robot controls its motors and actuators based on the motion command data it receives, and performs the specified dance movements.

[0580] Specific example

[0581] Situation: Robot dancing at an event venue

[0582] 1. Pop music is playing at the event venue.

[0583] 2. The user uses a smartphone application to instruct the robot to "do a fun dance."

[0584] 3. The server captures the music from the venue via microphones and acquires the music data.

[0585] 4. The server analyzes the music data using FFT and extracts characteristic data such as volume, tempo, and rhythm.

[0586] 5. The server provides the analysis results and the user-specified condition of "fun feeling" to the generating AI, and generates a dance sequence based on that.

[0587] 6. The server encodes the generated dance sequence as motion command data and transmits it to the robot via the network.

[0588] 7. The robot operates its motors and actuators based on the received motion command data to perform a "fun-feeling" dance.

[0589] Thus, this system can instantly generate and execute appropriate dance performances in response to user instructions, providing high flexibility and immediacy in events and promotions.

[0590] Examples of prompt statements used

[0591] Based on the user's specified "fun" dance conditions, analyze the tempo and volume of the acquired music data and generate a suitable dance sequence.

[0592] The above describes specific embodiments for carrying out the present invention.

[0593] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0594] Step 1:

[0595] The server acquires music data from the surrounding environment. Specifically, it collects music data in real time through a microphone. The input for this step is the audio signal of the music, and the output is the acquired raw music data.

[0596] Step 2:

[0597] The server analyzes the acquired music data. It extracts characteristic music data using algorithms such as FFT (Fast Fourier Transform). The input to this step is the acquired music data, and the output is music characteristic data such as volume, tempo, and rhythm.

[0598] Step 3:

[0599] The user inputs dance conditions for the robot via their smartphone. For example, conditions such as "fun" or "elegant" can be specified. The input in this step is the user's dance conditions, and the output is data for those dance conditions.

[0600] Step 4:

[0601] The server generates a dance sequence using a generative AI model based on the music characteristic data and the received dance conditions. The input for this step is the music characteristic data and the user's dance conditions, and the output is the generated dance sequence.

[0602] Step 5:

[0603] The server encodes the generated dance sequence as robot motion command data and transmits it to the robot over the network. The input to this step is the generated dance sequence, and the output is the motion command data sent to the robot.

[0604] Step 6:

[0605] The robot controls its motors and actuators based on the received motion command data to perform the specified dance movements. The input for this step is the motion command data, and the output is the actual dance movements of the robot.

[0606] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0607] This invention relates to a robot dance system that incorporates an emotion engine that recognizes user emotions. This system analyzes the characteristics of music, generates an appropriate dance sequence according to the user's emotions and specified dance conditions, and transmits motion commands to the robot, enabling the robot to perform the dance in real time.

[0608] System Overview

[0609] This system consists of the following main components:

[0610] 1. Means for acquiring music data

[0611] 2. Music Analysis Methods

[0612] 3. Emotional Engine

[0613] 4. Means of receiving dance conditions

[0614] 5. Dance generation means

[0615] 6. Operation command data transmission means

[0616] 7. Robot motion execution means

[0617] Natural language explanation of program processing

[0618] 1. Acquisition of music data

[0619] The server acquires surrounding music data in real time via a microphone. This music data includes audio signals, which are treated as input data.

[0620] 2. Music Analysis

[0621] The server analyzes the acquired music data using algorithms such as FFT (Fast Fourier Transform). As a result of the analysis, characteristic data of the music, such as volume, tempo, and pitch, is extracted, and further processing is performed based on this data.

[0622] 3. Receiving dance conditions

[0623] The user inputs dance conditions such as "fun feeling" or "calm feeling" through the interface. This input data is sent to the server.

[0624] 4. Use of the Emotion Engine

[0625] The emotion engine integrated into the server recognizes emotions through the user's facial expressions, voice, and gestures. This uses cameras and microphones to collect data in real time.

[0626] 5. Emotion analysis

[0627] The server uses an emotion engine to analyze the collected data and determine the user's emotions (e.g., joy, sadness, surprise). This emotion data is also incorporated into the dance conditions.

[0628] 6. Dance Generation

[0629] The server uses a generative artificial intelligence (generative AI) model to generate dance sequences, taking music characteristic data, dance conditions from the user, and emotional data as input.

[0630] The generating AI is pre-trained with various dance patterns and generates movements in real time that are appropriate to the music's tempo and rhythm, as well as the user's emotions.

[0631] 7. Generation and transmission of operation command data

[0632] The server encodes the generated dance sequence as motion command data for robot operation. This data is transmitted to the terminal (robot) via the network.

[0633] 8. Execution of actions by the robot

[0634] The terminal (robot) analyzes the received motion command data and generates signals to control each motor and actuator.

[0635] The terminal then operates the motors and actuators appropriately according to the generated control signals. This allows the robot to perform predetermined movements and dance.

[0636] Specific example

[0637] Situation: Robot dancing at an event venue

[0638] 1. Pop music is playing at the event venue.

[0639] 2. The user uses a smartphone or dedicated terminal to instruct the robot to "do a fun dance."

[0640] 3. The server picks up the music in the venue via a microphone sensor and acquires the music data.

[0641] 4. The server analyzes the music data using FFT and extracts characteristic data such as volume, tempo, and rhythm.

[0642] 5. The server uses cameras and microphones to collect the user's facial expressions, voice, and gestures in real time, and uses an emotion engine to analyze the user's emotions.

[0643] 6. The server provides the analysis results and the user-specified condition of "fun feeling" to the generating AI, and generates a dance sequence based on that.

[0644] 7. The server encodes the generated dance sequence as motion command data and transmits it to the terminal (robot) via the network.

[0645] 8. The terminal (robot) operates its motors and actuators based on the received motion command data to perform a "fun-feeling" dance.

[0646] Thus, this system can instantly generate and execute appropriate dance performances in response to user instructions and emotions, providing high flexibility and immediacy in events and promotions.

[0647] The following describes the processing flow.

[0648] Step 1:

[0649] Users use their smartphones or dedicated terminals to instruct the robot to "do a fun dance."

[0650] Step 2:

[0651] The server acquires music data in real time from the surroundings via a microphone sensor. This music data includes audio signals.

[0652] Step 3:

[0653] The server analyzes the acquired music data using FFT (Fast Fourier Transform). This analysis extracts characteristic data of the music, such as volume, tempo, and pitch.

[0654] Step 4:

[0655] The server temporarily stores music characteristic data in memory.

[0656] Step 5:

[0657] The server uses cameras and microphones to collect the user's facial expressions, voice, and gestures in real time, and analyzes them using an emotion engine. The analysis recognizes the user's emotions (e.g., joy, surprise).

[0658] Step 6:

[0659] The server receives the analyzed user emotion data along with dance conditions from the user (e.g., "feeling happy").

[0660] Step 7:

[0661] The server generates dance sequences using generative artificial intelligence (generative AI) based on music characteristic data, dance conditions from the user, and emotional data. The generative AI selects appropriate dance patterns from the input data and outputs them as dance sequences.

[0662] Step 8:

[0663] The server encodes the generated dance sequence as motion command data for robot operation. This motion command data is then converted into a format that the robot can understand and execute.

[0664] Step 9:

[0665] The server transmits operation command data to the terminal (robot) via the network.

[0666] Step 10:

[0667] The terminal analyzes the received operation command data and generates signals to control each motor and actuator based on that data.

[0668] Step 11:

[0669] The terminal then operates the motors and actuators appropriately according to the generated control signals. This allows the robot to perform predetermined movements and dance.

[0670] Step 12:

[0671] The terminal (robot) performs a "fun-sounding" dance based on user instructions and the characteristics of the music, livening up the atmosphere at event venues and other locations.

[0672] In this way, the robot can generate and perform dances in real time that match the characteristics of the music, based on user instructions and emotion recognition.

[0673] (Example 2)

[0674] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0675] Conventional robot dance systems primarily perform movements based solely on the tempo and rhythm of music, making it difficult to reflect user emotions or specific dance conditions in real time. This limits the user experience and reduces the expressiveness and flexibility of robot dance. This invention aims to generate dance sequences in response to user emotions and specified dance conditions, enabling the robot to perform more dynamic and adaptive dances in real time.

[0676] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for acquiring music data from the surrounding environment, means for analyzing the acquired music data and extracting music characteristic data, means for receiving dance conditions from the user, means for recognizing emotions through the user's facial expressions, voice, and gestures, means for analyzing the user's emotion data, means for generating a dance sequence using generative artificial intelligence based on the extracted music characteristic data, the received dance conditions, and the recognized emotion data, means for encoding the generated dance sequence as robot motion command data and transmitting it to the robot, and means for the robot to execute dance movements based on the received motion command data. This makes it possible to perform advanced, real-time robot dances that are in line with the user's emotions and dance conditions.

[0677] "Means of acquiring music data from the surrounding environment" refers to a function that captures ambient audio signals in real time using sensors such as microphones built into robots or servers.

[0678] "Methods for analyzing acquired music data and extracting characteristic music data" refers to a function that applies algorithms such as FFT (Fast Fourier Transform) to music data to calculate and extract characteristic data such as volume, tempo, rhythm, and pitch.

[0679] "Means of receiving dance conditions from users" refers to a function that inputs and receives dance conditions specified by the user (e.g., "a fun feeling" or "a calm feeling") via an interface such as a smartphone app or dedicated terminal.

[0680] "Means of recognizing emotions through the user's facial expressions, voice, and gestures" refers to a function that uses cameras and microphones to capture the user's facial expressions, voice tone, and gestures in real time, and recognizes emotions through an emotion engine.

[0681] "Methods for analyzing user emotion data" refers to a function that analyzes facial expressions, voice, and gesture data acquired in real time and converts the user's emotions into labels such as joy, sadness, and surprise.

[0682] "A means of generating a dance sequence using generative artificial intelligence based on extracted music characteristic data, received dance conditions, and recognized emotion data" refers to a function that uses a generative artificial intelligence model (generative AI model) to generate a suitable dance sequence in real time, taking music characteristic data, dance conditions, and emotion data as input.

[0683] "Means for encoding the generated dance sequence as robot motion command data and transmitting it to the robot" refers to a function that converts the generated dance sequence into specific control signals for robot operation and transmits them to the robot via a network.

[0684] "Means for executing dance movements based on motion command data received by the robot" refers to a function that, based on the control signals received by the robot, appropriately operates each motor and actuator to execute a predetermined dance movement.

[0685] This invention relates to a robot dance system that incorporates an emotion engine that recognizes user emotions. This system analyzes the characteristics of music, generates an appropriate dance sequence according to the user's emotions and specified dance conditions, and transmits motion commands to the robot, enabling the robot to perform the dance in real time.

[0686] System Configuration

[0687] This system consists of the following main components:

[0688] 1. Means for acquiring music data

[0689] The server uses sensors, such as microphones built into the robot, to acquire ambient music in real time. This allows it to collect audio signals.

[0690] 2. Music Analysis Methods

[0691] The server analyzes the acquired music data using algorithms such as FFT (Fast Fourier Transform) to extract characteristic data such as volume, tempo, rhythm, and pitch.

[0692] 3. Means of receiving dance conditions

[0693] Users input dance conditions such as "fun feeling" or "calm feeling" via a smartphone app or dedicated device. The data is then transmitted to a server via the internet.

[0694] 4. Emotion Engine (Means of Emotion Recognition)

[0695] The server uses cameras and additional microphones to collect the user's facial expressions, voice, and gestures in real time, and uses an emotion engine to recognize the user's emotions.

[0696] 5. Emotion analysis method

[0697] The server analyzes data collected through the emotion engine to determine the user's emotions (joy, sadness, surprise). This emotion data is also incorporated into the dance conditions.

[0698] 6. Dance generation means

[0699] The server generates dance sequences using a generative AI model based on music characteristic data, dance conditions from the user, and recognized emotion data.

[0700] 7. Operation command data transmission means

[0701] The server encodes the generated dance sequence as motion command data for robot operation and transmits it to the robot (terminal) via the network.

[0702] 8. Robot motion execution means

[0703] The terminal (robot) analyzes the received motion command data, generates signals to control each motor and actuator, and executes the specified dance movements.

[0704] Specific example

[0705] Situation: Robot dancing at an event venue

[0706] 1. Pop music is playing at the event venue.

[0707] 2. The user uses a smartphone or dedicated terminal to instruct the robot to "do a fun dance."

[0708] 3. The server picks up the music in the venue via a microphone sensor and acquires the music data.

[0709] 4. The server analyzes the music data using FFT and extracts characteristic data such as volume, tempo, and rhythm.

[0710] 5. The server uses cameras and microphones to collect the user's facial expressions, voice, and gestures in real time, and uses an emotion engine to analyze the user's emotions.

[0711] 6. The server provides the analysis results and the user-specified condition of "fun feeling" as prompts to the generating AI, and generates a dance sequence based on them.

[0712] 7. The server encodes the generated dance sequence as motion command data and transmits it to the terminal (robot) via the network.

[0713] 8. The terminal (robot) operates its motors and actuators based on the received motion command data to perform a "fun-feeling" dance.

[0714] Example of a prompt

[0715] "Generate a dance sequence that fits a fun, upbeat pop song. The user is expressing joy. The tempo is 120 BPM, and the volume is medium."

[0716] By inputting this prompt into the AI ​​model, it is possible to generate a dance sequence that is suitable for the user's emotions and musical characteristics.

[0717] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0718] Step 1: Obtain music data

[0719] The server uses sensors, such as microphones built into the robot, to acquire ambient music in real time. The music data acquired at this stage (input) is stored in a buffer as an audio signal (output).

[0720] Specific operation: The server captures audio data at regular time intervals and monitors it continuously.

[0721] Step 2: Music Analysis

[0722] The server analyzes the acquired music data (input) using FFT (Fast Fourier Transform) and extracts characteristic data (output) such as frequency components, volume, tempo, rhythm, and pitch.

[0723] Specific operation: The server performs an FFT on each time window to extract time-frequency domain features. A beat detection algorithm is then applied to calculate tempo data.

[0724] Step 3: Receiving the dance conditions

[0725] Users access the interface via a smartphone app or dedicated device and input dance conditions (inputs) such as "fun feeling" or "calm feeling." This input is sent to a server via the internet and stored as dance condition data (output).

[0726] Specific operation: The user selects the dance conditions on the device's UI and presses the send button.

[0727] Step 4: Using the Emotion Engine

[0728] The server uses cameras and additional microphones to collect the user's facial expressions, voice, and gestures (input) in real time, and recognizes emotional data (output) through an emotion engine.

[0729] Specific operation: Video data acquired by the camera is processed by a face recognition algorithm, and facial expression data is extracted. Similarly, a voice analysis algorithm is also applied.

[0730] Step 5: Emotion Analysis

[0731] The server analyzes real-time collected facial, voice, and gesture data (input) and uses an emotion engine to determine the user's emotions (e.g., joy, sadness, surprise). This emotion data (output) is then incorporated into the dance conditions.

[0732] Specific operation: The emotion engine uses a facial expression recognition model and a voice tone analysis model to generate emotion labels.

[0733] Step 6: Dance Generation

[0734] The server sends prompts to the generative AI model based on music characteristic data, dance conditions from the user, and recognized emotion data (input), and generates a dance sequence (output).

[0735] Specific operation: The server inputs a prompt message such as "Generate a dance sequence that fits a fun pop song" into the AI ​​model and retrieves the generated dance sequence data.

[0736] Step 7: Generate and transmit operation command data

[0737] The server encodes the generated dance sequence (input) as motion command data for robot operation and transmits it to the robot (terminal) via the network. This motion command data is generated as output.

[0738] Specific operation: The server converts the generated dance sequence into control data for each joint and motor, and transmits it at the appropriate timing.

[0739] Step 8: Robot execution of actions

[0740] The terminal (robot) analyzes the received motion command data (input), generates signals to control each motor and actuator, and executes the specified dance movement (output).

[0741] Specific operation: The terminal's control board analyzes the received data and sends the appropriate current and voltage to each motor to execute the specified operation in real time.

[0742] (Application Example 2)

[0743] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0744] Conventional robot dance systems automatically generate dances based on user emotions and music characteristics. However, to improve work efficiency on a factory production line, it is necessary to consider the emotions and work conditions of the workers. This can improve work efficiency and motivation. However, existing systems lack the technology to appropriately acquire emotional data and generate response sequences. Therefore, there is a need for a robot system that can perform appropriate responses according to the emotions and work conditions of workers on a factory production line.

[0745] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0746] In this invention, the server includes means for acquiring music data, means for analyzing music data, means for receiving dance conditions from a user, means for generating a dance sequence based on extracted music characteristic data and received dance conditions, means for encoding the generated dance sequence as robot motion command data, means for transmitting it to a robot, means for the robot to execute dance movements based on the received motion command data, means for acquiring worker emotion data from the work environment, means for generating a response sequence for improving work efficiency based on the acquired emotion data, means for encoding the generated response sequence as robot motion command data, means for transmitting it, and means for the robot to execute response movements based on the received motion command data. This makes it possible to provide appropriate responses in real time according to the worker's emotions and work situation in the work environment.

[0747] "Music data" refers to the audio signals and waveform data that make up music.

[0748] "Characteristic data" refers to distinctive information such as tempo, volume, and rhythm extracted from music data.

[0749] "Dance conditions" refer to the requirements specified by the user, such as the type of dance and the atmosphere.

[0750] "Generative artificial intelligence" refers to artificial intelligence technology used to generate new data or sequences based on existing information.

[0751] A "dance sequence" refers to a continuous flow or pattern of movements. It specifically refers to a series of actions in a dance performance.

[0752] "Action command data" refers to a set of instructions that cause a robot to perform a specific action.

[0753] A "robot" refers to a mechanical device that performs actions automatically.

[0754] "Emotional data" refers to emotional information analyzed from the workers' facial expressions, voices, gestures, etc.

[0755] A "response sequence" refers to a sequence of actions and reactions generated to respond appropriately to emotions and situations.

[0756] "Work environment" refers to the location where work is performed, such as a factory or production line, and the surrounding conditions.

[0757] "Workers" refers to people who perform tasks on production lines or similar environments.

[0758] This invention is a system that controls the movements of a robot based on music data and emotion data, thereby improving work efficiency and motivation in the work environment. The following describes specific embodiments of this system.

[0759] System configuration and operation

[0760] 1. Acquisition of music data

[0761] The server acquires music data from the surrounding environment in real time using a microphone. This music data includes audio signals.

[0762] 2. Analysis of music data

[0763] The server analyzes the acquired music data using algorithms such as FFT (Fast Fourier Transform) to extract characteristic data such as volume, tempo, and rhythm. This process utilizes, for example, high-performance computing servers or cloud computing resources.

[0764] 3. Receiving dance conditions

[0765] The server receives dance conditions such as "feeling happy" or "feeling calm" from the user via smartphone or interface. This information is transmitted to the server over the network.

[0766] 4. Acquisition of emotional data

[0767] The server collects the workers' facial expressions, voices, and gestures in real time using cameras and microphones, and analyzes them using an emotion engine. This emotion data is also stored on the server.

[0768] 5. Generation of the dance sequence

[0769] The server uses a generative AI model (e.g., GPT-4) to generate dance sequences and response sequences, taking music characteristic data, dance conditions from the user, and emotional data as input. The generated sequences are then coded as command data for robot movements.

[0770] 6. Transmission of command data

[0771] The server transmits the generated motion command data to the robot via the network. The robot's terminal analyzes this data and sends control signals to each motor and actuator.

[0772] 7. Execution of actions by the robot

[0773] The robot controls its motors and actuators appropriately based on motion command data received from the server to perform specified dances and work response actions.

[0774] Specific examples of hardware and software

[0775] Camera: High-resolution, night vision camera (e.g., Logitech C920s Pro)

[0776] Microphone: High-sensitivity microphone (e.g., Blue Yeti USB Microphone)

[0777] Sensors: Belt conveyor speed sensors and quality control sensors (e.g., Honeywell FF-SY series)

[0778] Server: High-performance server (e.g., NVIDIA DGX Station) or cloud computing resources

[0779] Generative AI models: Generative models based on GPT-4 and Transformers

[0780] Network: Stable WiFi or Ethernet connection

[0781] Specific example

[0782] Situation: Improving production efficiency within a factory

[0783] 1. The server acquires data in real time from conveyor belt speed sensors and quality control sensors within the factory.

[0784] 2. The server uses a camera and microphone to collect the worker's facial expressions, voice, and gestures, and analyzes them using an emotion engine. Based on this analysis, it recognizes that the worker is tired.

[0785] 3. The server provides the generation AI model with prompts like the following to generate a response sequence:

[0786] "If the workers are tired, please generate a refreshing dance to increase production efficiency."

[0787] 4. The generated response sequence is sent to the robot, which then performs nimble steps and simple exercises to increase worker motivation and improve efficiency.

[0788] This system makes it possible to generate and execute appropriate robot movements in real time based on music data and emotion data, contributing to improved productivity and employee motivation in the work environment.

[0789] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0790] Step 1:

[0791] The server acquires music data from the surrounding environment. Specifically, it uses a microphone to collect audio signals in real time and saves them as audio data. The input for this step is ambient music, and the output is the acquired audio data.

[0792] Step 2:

[0793] The server analyzes the acquired music data. Specifically, it uses the FFT (Fast Fourier Transform) algorithm to extract characteristic data of the music (volume, tempo, rhythm, etc.). The input for this step is audio data, and the output is characteristic data.

[0794] Step 3:

[0795] The server receives dance conditions from the user. Specifically, it receives conditions such as "feeling happy" or "feeling calm" that the user sends via smartphone or interface over the network. The input for this step is the user's dance conditions, and the output is the received dance conditions.

[0796] Step 4:

[0797] The server acquires emotional data from the work environment. Specifically, it uses cameras and microphones to collect the workers' facial expressions, voices, and gestures in real time, and analyzes them using an emotion engine. The input for this step is the workers' facial expressions, voices, and gestures, and the output is the analyzed emotional data.

[0798] Step 5:

[0799] The server uses a generative AI model (e.g., GPT-4) to generate dance sequences and response sequences, taking music characteristic data, dance conditions from the user, and emotion data as input. Specifically, it provides the generative AI model with prompts like the following to generate appropriate sequences. The input for this step is characteristic data, dance conditions, and emotion data, and the output is the generated sequence.

[0800] Example prompt: "If the workers are tired, generate a refresh dance sequence to increase production efficiency."

[0801] Step 6:

[0802] The server encodes the generated sequence as robot motion command data and transmits it to the robot via the network. Specifically, it converts the generated sequence into control signals and sends them to the robot. The input to this step is the generated sequence, and the output is the robot motion command data.

[0803] Step 7:

[0804] The terminal (robot) controls its motors and actuators based on motion command data received from the server, and performs the specified actions. Specifically, it performs dances and response actions according to the actions commanded by the server. The input for this step is motion command data, and the output is the performed action.

[0805] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0806] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0807] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0808] [Third Embodiment]

[0809] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0810] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0811] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0812] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0813] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0814] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0815] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0816] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0817] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0818] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0819] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0820] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0821] This invention relates to a system in which a robot generates and performs dances in real time based on instructions from a user. The system analyzes the characteristics of music, generates an appropriate dance sequence based on dance conditions specified by the user, and transmits motion commands to the robot, enabling the robot to perform the dance as instructed.

[0822] System Overview

[0823] This system consists of the following main components:

[0824] 1. Means for acquiring music data

[0825] 2. Music Analysis Methods

[0826] 3. Means of receiving dance conditions

[0827] 4. Dance generation means

[0828] 5. Means for transmitting operation command data

[0829] 6. Robot motion execution means

[0830] Natural language explanation of program processing

[0831] 1. Acquisition of music data

[0832] The server acquires surrounding music data in real time via a microphone. This music data includes audio signals, which are treated as input data.

[0833] 2. Music Analysis

[0834] The server analyzes the acquired music data using algorithms such as FFT (Fast Fourier Transform). As a result of the analysis, characteristic data of the music, such as volume, tempo, and pitch, is extracted, and further processing is performed based on this data.

[0835] 3. Receiving dance conditions

[0836] The user inputs dance conditions such as "fun feeling" or "calm feeling" through the interface. This input data is sent to the server.

[0837] 4. Dance Generation

[0838] The server uses a generative artificial intelligence (generative AI) model to generate dance sequences, taking music characteristic data and dance conditions from the user as input.

[0839] The generation AI is pre-trained with various dance patterns and generates movements in real time that are appropriate to the tempo and rhythm of the music.

[0840] 5. Generation and transmission of operation command data

[0841] The server encodes the generated dance sequence into motion command data so that the robot can understand it. This data is transmitted to the terminal (robot) via the network.

[0842] 6. Execution of actions by the robot

[0843] The terminal (robot) analyzes the received motion command data and controls each motor and actuator. Based on the control signals, the robot performs the predetermined movements and performs a dance.

[0844] Specific example

[0845] Situation: Robot dancing at an event venue

[0846] 1. Pop music is playing at the event venue.

[0847] 2. The user uses a smartphone or dedicated terminal to instruct the robot to "do a fun dance."

[0848] 3. The server picks up the music from the venue via the microphone and acquires the music data.

[0849] 4. The server analyzes the music data using FFT and extracts characteristic data such as volume, tempo, and rhythm.

[0850] 5. The server provides the analysis results and the user-specified condition of "fun feeling" to the generating AI, and generates a dance sequence based on that.

[0851] 6. The server encodes the generated dance sequence as motion command data and transmits it to the terminal (robot) via the network.

[0852] 7. The terminal (robot) operates its motors and actuators based on the received motion command data to perform a "fun-feeling" dance.

[0853] Thus, this system can instantly generate and execute appropriate dance performances in response to user instructions, providing high flexibility and immediacy in events and promotions.

[0854] The following describes the processing flow.

[0855] Step 1:

[0856] Users use their smartphones or dedicated terminals to instruct the robot to "do a fun dance."

[0857] Step 2:

[0858] The server acquires music data in real time from the surroundings via a microphone sensor. This music data includes audio signals.

[0859] Step 3:

[0860] The server analyzes the acquired music data using FFT (Fast Fourier Transform). This analysis extracts characteristic data of the music, such as volume, tempo, and pitch.

[0861] Step 4:

[0862] The server temporarily stores music characteristic data in memory and receives dance conditions from the user (e.g., "happy feeling," "calm feeling").

[0863] Step 5:

[0864] The server generates a dance sequence using generative artificial intelligence (generative AI) based on the received dance conditions and music characteristics data. The generative AI selects appropriate movement patterns from the input data and outputs them as a dance sequence.

[0865] Step 6:

[0866] The server encodes the generated dance sequence as motion command data for robot operation. This motion command data is then converted into a format that the robot can understand and execute.

[0867] Step 7:

[0868] The server transmits operation command data to the terminal (robot) via the network.

[0869] Step 8:

[0870] The terminal analyzes the received operation command data and generates signals to control each motor and actuator based on that data.

[0871] Step 9:

[0872] The terminal then operates the motors and actuators appropriately according to the generated control signals. This allows the robot to perform the specified dance movements (e.g., waving arms, spinning).

[0873] Step 10:

[0874] The terminal (robot) performs a "fun-sounding" dance based on user instructions and the characteristics of the music, livening up the atmosphere at event venues and other locations.

[0875] As described above, it is possible for the robot to generate and execute dances in real time that match the characteristics of the music, based on user instructions.

[0876] (Example 1)

[0877] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0878] Conventional robot dance systems only reproduce pre-programmed movements and cannot flexibly respond to changes in music or environment. Furthermore, they struggle to reflect user emotions and intentions in real time, lacking the immediacy and flexibility needed for events and promotions. Therefore, there is a need to develop a system that generates dance sequences in real time based on music characteristics and user-specified conditions, and has the robot execute them.

[0879] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0880] In this invention, the server includes means for acquiring music data from the surrounding environment, means for analyzing the acquired music data and extracting music characteristic data, means for receiving dance conditions from the user, means for generating a dance sequence using generative artificial intelligence based on the extracted music characteristic data and the received dance conditions, means for encoding the generated dance sequence as robot motion command data and transmitting it to the robot, means for the robot to execute dance movements based on the received motion command data, means for analyzing the acquired music data using the Fast Fourier Transform, means for the generative artificial intelligence model to consider tempo and movement patterns in the generation of the dance sequence, and means for transmitting the generated motion command data to the robot using wireless communication. This enables rapid response to changes in music and environment, and allows for dance performances that reflect the user's emotions and intentions in real time.

[0881] "Surrounding environment" refers to the location where the system is installed and the physical space surrounding it.

[0882] "Music data" refers to audio signals and music information such as background music acquired from the surrounding environment.

[0883] "Music characteristic data" refers to attribute information of music, such as volume, tempo, rhythm, and pitch, obtained by analyzing music data.

[0884] "Dance conditions" refer to the user's requirements regarding the style and atmosphere of the dance, such as "fun" or "relaxed."

[0885] "Generative artificial intelligence" refers to a technology that uses machine learning algorithms to generate new data and information based on input data.

[0886] A "dance sequence" refers to a combination of motion commands that a robot executes.

[0887] "Motion command data" refers to the control signals and commands necessary for a robot to perform a specific action.

[0888] A "robot" refers to an autonomous or semi-autonomous device that performs specified actions through a mechanical structure and computer control.

[0889] The "Fast Fourier Transform" refers to an efficient algorithm for converting time-domain data into the frequency domain.

[0890] "Wireless communication" refers to the technology of sending and receiving data using radio waves or light waves.

[0891] "Tempo" refers to the speed of music, specifically the number of beats contained within a given time.

[0892] An "action pattern" refers to a sequence of actions performed under specific conditions.

[0893] This invention relates to a system in which a robot generates and performs dances in real time based on instructions from a user. The system analyzes the characteristics of music, generates an appropriate dance sequence based on dance conditions specified by the user, and transmits motion commands to the robot, enabling the robot to perform the dance as instructed.

[0894] System Configuration

[0895] This system consists of the following main components:

[0896] 1. Means for acquiring music data

[0897] 2. Music Analysis Methods

[0898] 3. Means of receiving dance conditions

[0899] 4. Dance generation means

[0900] 5. Means for transmitting operation command data

[0901] 6. Robot motion execution means

[0902] Music data acquisition method

[0903] The server acquires ambient music data in real time through a high-sensitivity microphone. Specifically, it uses a high-quality audio capture device (e.g., a high-sensitivity microphone) to capture ambient sound as a digital signal. This acquired music data is temporarily stored in the server's buffer memory.

[0904] Music analysis methods

[0905] The server uses the Python libraries SciPy and NumPy to perform a Fast Fourier Transform (FFT) to analyze the acquired music data. This analysis extracts frequency components, and then the LibROSA library is used to calculate characteristic data of the music, such as volume, tempo, and pitch. This characteristic data is used in the next step.

[0906] Dance condition receiving method

[0907] Users specify dance conditions through an application on their smartphone or a dedicated tablet device. The user interface is developed with React Native, allowing for easy input of conditions such as "fun" or "calm." This input data is sent to the server in real time.

[0908] Dance generation means

[0909] The server uses a generative artificial intelligence (generative AI) model to generate dance sequences, taking music characteristic data and dance conditions from the user as input. The generative AI is built using TensorFlow and has been pre-trained on a variety of dance styles. For example, if the music tempo is fast and the dance condition is "fun," the generative AI will generate an energetic and rhythmic dance sequence.

[0910] Operation command data transmission means

[0911] The server encodes the generated dance sequence into motion command data that the robot can understand. Specifically, it uses ROBO-API to convert the dance sequence into control signals for the servo motors. The generated motion command data is then transmitted to the terminal (robot) via WiFi or Bluetooth.

[0912] Robot motion execution means

[0913] The terminal (robot) analyzes the received motion command data. Based on the analyzed data, it controls each motor and actuator (e.g., Dynamixel servo motor) to execute the specified action in real time. For example, if a "fun" dance command is given, the robot will take light steps and dance cheerfully while swinging both arms.

[0914] Specific example

[0915] Situation: Robot dancing at an event venue

[0916] 1. Pop music is playing at the event venue.

[0917] 2. The user uses a smartphone or dedicated device to instruct the robot to "do a fun dance." The user selects "fun" from the app's dropdown menu and clicks the send button.

[0918] 3. The server captures the music from the venue via the microphone and acquires the music data. The server processing begins the moment the microphone picks up sound.

[0919] 4. The server analyzes the music data using FFT and LibROSA, and extracts characteristic data such as volume, tempo, and rhythm. After analysis, the characteristic data is stored in memory.

[0920] 5. The server provides the generation AI with the analysis results and the user-specified condition of "fun feeling," and generates a dance sequence based on this. The generation AI generates the sequence in a few seconds, and the server checks the results.

[0921] 6. The server encodes the generated dance sequence as motion command data and sends it to the terminal (robot) via the network. As soon as the transmission is complete, the terminal sends a receipt confirmation to the server.

[0922] 7. The terminal (robot) operates its motors and actuators based on the received motion command data to perform a "fun-feeling" dance. For example, the robot takes light steps and dances cheerfully, raising and swinging both arms.

[0923] Examples of prompt statements

[0924] The following are some possible prompts to input into the generative AI model:

[0925] "Please generate a robot dance sequence based on the following conditions: Music tempo is 120 BPM, and the user-specified emotion is 'happy'. The output format should be motion command data."

[0926] In this way, each step of the entire system is intricately coordinated, enabling robot dances to be performed in real time in response to user instructions.

[0927] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0928] Step 1:

[0929] The server acquires music data from its surroundings. Specifically, it uses a high-sensitivity microphone to capture audio signals in real time. It receives analog audio signals from the microphone as input, converts them to digital signals, and stores them in buffer memory. The output is digital music data.

[0930] Step 2:

[0931] The server analyzes the acquired music data. Specifically, it performs an FFT (Fast Fourier Transform) using the Python libraries SciPy and NumPy. It receives digital music data stored in buffer memory as input and extracts frequency components. It also calculates characteristic data such as volume, tempo, rhythm, and pitch using the LibROSA library. The output is characteristic data of the music.

[0932] Step 3:

[0933] Users specify dance conditions through an application on their smartphone or a dedicated tablet device. Users input dance conditions such as "fun feeling" or "calm feeling" through the interface. The input consists of the user's specified conditions and is transmitted from the application to the server in real time. The output is the user's dance condition data.

[0934] Step 4:

[0935] The server uses a generative artificial intelligence (generative AI) model to generate dance sequences, taking music characteristic data and dance condition data from the user as input. The generative AI is pre-trained on a variety of dance styles and generates the optimal dance sequence according to the input data. The input is music characteristic data and dance condition data, and the output is the generated dance sequence.

[0936] Step 5:

[0937] The server encodes the generated dance sequence as robot motion command data. Specifically, it uses ROBO-API to convert the dance sequence into servo motor control signals. The input is the generated dance sequence, and the output is the motion command data. This motion command data is transmitted to the terminal (robot) via WiFi or Bluetooth.

[0938] Step 6:

[0939] The terminal (robot) analyzes the received motion command data. Based on the analyzed data, it controls each motor and actuator to execute the specified action in real time. The input is motion command data, and the output is the actual robot's movement. For example, if a "fun" dance command is given, the robot will dance with light steps and swing its arms.

[0940] Through these processing steps, the system can perform robotic dances in real time in response to user instructions.

[0941] (Application Example 1)

[0942] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0943] Conventional robot dance systems only execute pre-programmed movements, making it difficult for them to respond to changes in the surrounding environment or music in real time. Furthermore, they lack a mechanism for users to spontaneously instruct the robots to dance, resulting in a lack of flexibility in entertainment and promotional settings. Additionally, they are unable to properly consider tempo and movement patterns when generating dance sequences, making it difficult to provide an engaging performance for audiences.

[0944] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0945] In this invention, the server includes means for acquiring music data from the surrounding environment, means for analyzing the acquired music data and extracting music characteristic data, means for receiving dance conditions from the user, means for generating a dance sequence using artificial intelligence based on the extracted music characteristic data and the received dance conditions, means for encoding the generated dance sequence as robot motion command data and transmitting it to the robot, means for the robot to execute dance movements based on the received motion command data, and means for the user to give improvisational dance instructions to the robot via a smartphone. As a result, the robot can perform dances in real time according to the characteristics of the surrounding music and respond to improvisational instructions from the user, providing high flexibility and immediacy for events and promotions.

[0946] "Means for acquiring music data" refers to devices or methods for acquiring music data from the surrounding environment.

[0947] "Music analysis means" refers to devices or methods that analyze acquired music data and extract characteristic data of the music.

[0948] A "dance condition receiving method" refers to a device or method for receiving dance conditions from users.

[0949] A "dance generation means" is a device or method that generates a dance sequence using artificial intelligence based on extracted music characteristic data and received dance conditions.

[0950] "Motion command data transmission means" refers to a device or method that encodes the generated dance sequence as motion command data for a robot and transmits it to the robot.

[0951] A "robot motion execution means" is a device or method that allows a robot to perform dance movements based on motion command data it has received.

[0952] A "smartphone" is a portable information terminal used by users to give improvised dance instructions to a robot.

[0953] "Generative artificial intelligence" is an artificial intelligence technology that generates dance sequences based on music characteristic data and user conditions.

[0954] This invention relates to a system in which a robot generates and performs dances in real time based on instructions from a user. The system analyzes the characteristics of music, generates an appropriate dance sequence based on dance conditions specified by the user, and transmits motion commands to the robot, enabling the robot to perform the dance as instructed.

[0955] System Overview

[0956] This system consists of the following main components:

[0957] 1. Means for acquiring music data

[0958] The server acquires music data from its surrounding environment. Specifically, it acquires music playing in environments such as event venues and stores in real time via microphones.

[0959] 2. Music Analysis Methods

[0960] The server analyzes the acquired music data using algorithms such as FFT (Fast Fourier Transform). As a result of the analysis, characteristic data of the music, such as volume, tempo, and pitch, is extracted.

[0961] 3. Means of receiving dance conditions

[0962] Users input dance conditions for the robot via their smartphones, such as "fun" or "elegant."

[0963] 4. Dance generation means

[0964] The server generates dance sequences using a generative AI model based on the extracted music characteristic data and received dance conditions. The generative AI model is pre-trained with various dance patterns and generates movements in real time that are appropriate to the tempo and rhythm of the music.

[0965] 5. Means for transmitting operation command data

[0966] The server encodes the generated dance sequence as robot motion command data and transmits it to the robot via the network. Bluetooth and Wi-Fi are used as communication methods.

[0967] 6. Robot motion execution means

[0968] The robot controls its motors and actuators based on the motion command data it receives, and performs the specified dance movements.

[0969] Specific example

[0970] Situation: Robot dancing at an event venue

[0971] 1. Pop music is playing at the event venue.

[0972] 2. The user uses a smartphone application to instruct the robot to "do a fun dance."

[0973] 3. The server captures the music from the venue via microphones and acquires the music data.

[0974] 4. The server analyzes the music data using FFT and extracts characteristic data such as volume, tempo, and rhythm.

[0975] 5. The server provides the analysis results and the user-specified condition of "fun feeling" to the generating AI, and generates a dance sequence based on that.

[0976] 6. The server encodes the generated dance sequence as motion command data and transmits it to the robot via the network.

[0977] 7. The robot operates its motors and actuators based on the received motion command data to perform a "fun-feeling" dance.

[0978] Thus, this system can instantly generate and execute appropriate dance performances in response to user instructions, providing high flexibility and immediacy in events and promotions.

[0979] Examples of prompt statements used

[0980] Based on the user's specified "fun" dance conditions, analyze the tempo and volume of the acquired music data and generate a suitable dance sequence.

[0981] The above describes specific embodiments for carrying out the present invention.

[0982] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0983] Step 1:

[0984] The server acquires music data from the surrounding environment. Specifically, it collects music data in real time through a microphone. The input for this step is the audio signal of the music, and the output is the acquired raw music data.

[0985] Step 2:

[0986] The server analyzes the acquired music data. It extracts characteristic music data using algorithms such as FFT (Fast Fourier Transform). The input to this step is the acquired music data, and the output is music characteristic data such as volume, tempo, and rhythm.

[0987] Step 3:

[0988] The user inputs dance conditions for the robot via their smartphone. For example, conditions such as "fun" or "elegant" can be specified. The input in this step is the user's dance conditions, and the output is data for those dance conditions.

[0989] Step 4:

[0990] The server generates a dance sequence using a generative AI model based on the music characteristic data and the received dance conditions. The input for this step is the music characteristic data and the user's dance conditions, and the output is the generated dance sequence.

[0991] Step 5:

[0992] The server encodes the generated dance sequence as robot motion command data and transmits it to the robot over the network. The input to this step is the generated dance sequence, and the output is the motion command data sent to the robot.

[0993] Step 6:

[0994] The robot controls its motors and actuators based on the received motion command data to perform the specified dance movements. The input for this step is the motion command data, and the output is the actual dance movements of the robot.

[0995] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0996] This invention relates to a robot dance system that incorporates an emotion engine that recognizes user emotions. This system analyzes the characteristics of music, generates an appropriate dance sequence according to the user's emotions and specified dance conditions, and transmits motion commands to the robot, enabling the robot to perform the dance in real time.

[0997] System Overview

[0998] This system consists of the following main components:

[0999] 1. Means for acquiring music data

[1000] 2. Music Analysis Methods

[1001] 3. Emotional Engine

[1002] 4. Means of receiving dance conditions

[1003] 5. Dance generation means

[1004] 6. Operation command data transmission means

[1005] 7. Robot motion execution means

[1006] Natural language explanation of program processing

[1007] 1. Acquisition of music data

[1008] The server acquires surrounding music data in real time via a microphone. This music data includes audio signals, which are treated as input data.

[1009] 2. Music Analysis

[1010] The server analyzes the acquired music data using algorithms such as FFT (Fast Fourier Transform). As a result of the analysis, characteristic data of the music, such as volume, tempo, and pitch, is extracted, and further processing is performed based on this data.

[1011] 3. Receiving dance conditions

[1012] The user inputs dance conditions such as "fun feeling" or "calm feeling" through the interface. This input data is sent to the server.

[1013] 4. Use of the Emotion Engine

[1014] The emotion engine integrated into the server recognizes emotions through the user's facial expressions, voice, and gestures. This uses cameras and microphones to collect data in real time.

[1015] 5. Emotion analysis

[1016] The server uses an emotion engine to analyze the collected data and determine the user's emotions (e.g., joy, sadness, surprise). This emotion data is also incorporated into the dance conditions.

[1017] 6. Dance Generation

[1018] The server uses a generative artificial intelligence (generative AI) model to generate dance sequences, taking music characteristic data, dance conditions from the user, and emotional data as input.

[1019] The generating AI is pre-trained with various dance patterns and generates movements in real time that are appropriate to the music's tempo and rhythm, as well as the user's emotions.

[1020] 7. Generation and transmission of operation command data

[1021] The server encodes the generated dance sequence as motion command data for robot operation. This data is transmitted to the terminal (robot) via the network.

[1022] 8. Execution of actions by the robot

[1023] The terminal (robot) analyzes the received motion command data and generates signals to control each motor and actuator.

[1024] The terminal then operates the motors and actuators appropriately according to the generated control signals. This allows the robot to perform predetermined movements and dance.

[1025] Specific example

[1026] Situation: Robot dancing at an event venue

[1027] 1. Pop music is playing at the event venue.

[1028] 2. The user uses a smartphone or dedicated terminal to instruct the robot to "do a fun dance."

[1029] 3. The server picks up the music in the venue via a microphone sensor and acquires the music data.

[1030] 4. The server analyzes the music data using FFT and extracts characteristic data such as volume, tempo, and rhythm.

[1031] 5. The server uses cameras and microphones to collect the user's facial expressions, voice, and gestures in real time, and uses an emotion engine to analyze the user's emotions.

[1032] 6. The server provides the analysis results and the user-specified condition of "fun feeling" to the generating AI, and generates a dance sequence based on that.

[1033] 7. The server encodes the generated dance sequence as motion command data and transmits it to the terminal (robot) via the network.

[1034] 8. The terminal (robot) operates its motors and actuators based on the received motion command data to perform a "fun-feeling" dance.

[1035] Thus, this system can instantly generate and execute appropriate dance performances in response to user instructions and emotions, providing high flexibility and immediacy in events and promotions.

[1036] The following describes the processing flow.

[1037] Step 1:

[1038] Users use their smartphones or dedicated terminals to instruct the robot to "do a fun dance."

[1039] Step 2:

[1040] The server acquires music data in real time from the surroundings via a microphone sensor. This music data includes audio signals.

[1041] Step 3:

[1042] The server analyzes the acquired music data using FFT (Fast Fourier Transform). This analysis extracts characteristic data of the music, such as volume, tempo, and pitch.

[1043] Step 4:

[1044] The server temporarily stores music characteristic data in memory.

[1045] Step 5:

[1046] The server uses cameras and microphones to collect the user's facial expressions, voice, and gestures in real time, and analyzes them using an emotion engine. The analysis recognizes the user's emotions (e.g., joy, surprise).

[1047] Step 6:

[1048] The server receives the analyzed user emotion data along with dance conditions from the user (e.g., "feeling happy").

[1049] Step 7:

[1050] The server generates dance sequences using generative artificial intelligence (generative AI) based on music characteristic data, dance conditions from the user, and emotional data. The generative AI selects appropriate dance patterns from the input data and outputs them as dance sequences.

[1051] Step 8:

[1052] The server encodes the generated dance sequence as motion command data for robot operation. This motion command data is then converted into a format that the robot can understand and execute.

[1053] Step 9:

[1054] The server transmits operation command data to the terminal (robot) via the network.

[1055] Step 10:

[1056] The terminal analyzes the received operation command data and generates signals to control each motor and actuator based on that data.

[1057] Step 11:

[1058] The terminal then operates the motors and actuators appropriately according to the generated control signals. This allows the robot to perform predetermined movements and dance.

[1059] Step 12:

[1060] The terminal (robot) performs a "fun-sounding" dance based on user instructions and the characteristics of the music, livening up the atmosphere at event venues and other locations.

[1061] In this way, the robot can generate and perform dances in real time that match the characteristics of the music, based on user instructions and emotion recognition.

[1062] (Example 2)

[1063] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[1064] Conventional robot dance systems primarily perform movements based solely on the tempo and rhythm of music, making it difficult to reflect user emotions or specific dance conditions in real time. This limits the user experience and reduces the expressiveness and flexibility of robot dance. This invention aims to generate dance sequences in response to user emotions and specified dance conditions, enabling the robot to perform more dynamic and adaptive dances in real time.

[1065] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for acquiring music data from the surrounding environment, means for analyzing the acquired music data and extracting music characteristic data, means for receiving dance conditions from the user, means for recognizing emotions through the user's facial expressions, voice, and gestures, means for analyzing the user's emotion data, means for generating a dance sequence using generative artificial intelligence based on the extracted music characteristic data, the received dance conditions, and the recognized emotion data, means for encoding the generated dance sequence as robot motion command data and transmitting it to the robot, and means for the robot to execute dance movements based on the received motion command data. This makes it possible to perform advanced, real-time robot dances that are in line with the user's emotions and dance conditions.

[1066] "Means of acquiring music data from the surrounding environment" refers to a function that captures ambient audio signals in real time using sensors such as microphones built into robots or servers.

[1067] "Methods for analyzing acquired music data and extracting characteristic music data" refers to a function that applies algorithms such as FFT (Fast Fourier Transform) to music data to calculate and extract characteristic data such as volume, tempo, rhythm, and pitch.

[1068] "Means of receiving dance conditions from users" refers to a function that inputs and receives dance conditions specified by the user (e.g., "a fun feeling" or "a calm feeling") via an interface such as a smartphone app or dedicated terminal.

[1069] "Means of recognizing emotions through the user's facial expressions, voice, and gestures" refers to a function that uses cameras and microphones to capture the user's facial expressions, voice tone, and gestures in real time, and recognizes emotions through an emotion engine.

[1070] "Methods for analyzing user emotion data" refers to a function that analyzes facial expressions, voice, and gesture data acquired in real time and converts the user's emotions into labels such as joy, sadness, and surprise.

[1071] "A means of generating a dance sequence using generative artificial intelligence based on extracted music characteristic data, received dance conditions, and recognized emotion data" refers to a function that uses a generative artificial intelligence model (generative AI model) to generate a suitable dance sequence in real time, taking music characteristic data, dance conditions, and emotion data as input.

[1072] "Means for encoding the generated dance sequence as robot motion command data and transmitting it to the robot" refers to a function that converts the generated dance sequence into specific control signals for robot operation and transmits them to the robot via a network.

[1073] "Means for executing dance movements based on motion command data received by the robot" refers to a function that, based on the control signals received by the robot, appropriately operates each motor and actuator to execute a predetermined dance movement.

[1074] This invention relates to a robot dance system that incorporates an emotion engine that recognizes user emotions. This system analyzes the characteristics of music, generates an appropriate dance sequence according to the user's emotions and specified dance conditions, and transmits motion commands to the robot, enabling the robot to perform the dance in real time.

[1075] System Configuration

[1076] This system consists of the following main components:

[1077] 1. Means for acquiring music data

[1078] The server uses sensors, such as microphones built into the robot, to acquire ambient music in real time. This allows it to collect audio signals.

[1079] 2. Music Analysis Methods

[1080] The server analyzes the acquired music data using algorithms such as FFT (Fast Fourier Transform) to extract characteristic data such as volume, tempo, rhythm, and pitch.

[1081] 3. Means of receiving dance conditions

[1082] Users input dance conditions such as "fun feeling" or "calm feeling" via a smartphone app or dedicated device. The data is then transmitted to a server via the internet.

[1083] 4. Emotion Engine (Means of Emotion Recognition)

[1084] The server uses cameras and additional microphones to collect the user's facial expressions, voice, and gestures in real time, and uses an emotion engine to recognize the user's emotions.

[1085] 5. Emotion analysis method

[1086] The server analyzes data collected through the emotion engine to determine the user's emotions (joy, sadness, surprise). This emotion data is also incorporated into the dance conditions.

[1087] 6. Dance generation means

[1088] The server generates dance sequences using a generative AI model based on music characteristic data, dance conditions from the user, and recognized emotion data.

[1089] 7. Operation command data transmission means

[1090] The server encodes the generated dance sequence as motion command data for robot operation and transmits it to the robot (terminal) via the network.

[1091] 8. Robot motion execution means

[1092] The terminal (robot) analyzes the received motion command data, generates signals to control each motor and actuator, and executes the specified dance movements.

[1093] Specific example

[1094] Situation: Robot dancing at an event venue

[1095] 1. Pop music is playing at the event venue.

[1096] 2. The user uses a smartphone or dedicated terminal to instruct the robot to "do a fun dance."

[1097] 3. The server picks up the music in the venue via a microphone sensor and acquires the music data.

[1098] 4. The server analyzes the music data using FFT and extracts characteristic data such as volume, tempo, and rhythm.

[1099] 5. The server uses cameras and microphones to collect the user's facial expressions, voice, and gestures in real time, and uses an emotion engine to analyze the user's emotions.

[1100] 6. The server provides the analysis results and the user-specified condition of "fun feeling" as prompts to the generating AI, and generates a dance sequence based on them.

[1101] 7. The server encodes the generated dance sequence as motion command data and transmits it to the terminal (robot) via the network.

[1102] 8. The terminal (robot) operates its motors and actuators based on the received motion command data to perform a "fun-feeling" dance.

[1103] Example of a prompt

[1104] "Generate a dance sequence that fits a fun, upbeat pop song. The user is expressing joy. The tempo is 120 BPM, and the volume is medium."

[1105] By inputting this prompt into the AI ​​model, it is possible to generate a dance sequence that is suitable for the user's emotions and musical characteristics.

[1106] The flow of the specific processing in Example 2 will be explained using Figure 13.

[1107] Step 1: Obtain music data

[1108] The server uses sensors, such as microphones built into the robot, to acquire ambient music in real time. The music data acquired at this stage (input) is stored in a buffer as an audio signal (output).

[1109] Specific operation: The server captures audio data at regular time intervals and monitors it continuously.

[1110] Step 2: Music Analysis

[1111] The server analyzes the acquired music data (input) using FFT (Fast Fourier Transform) and extracts characteristic data (output) such as frequency components, volume, tempo, rhythm, and pitch.

[1112] Specific operation: The server performs an FFT on each time window to extract time-frequency domain features. A beat detection algorithm is then applied to calculate tempo data.

[1113] Step 3: Receiving the dance conditions

[1114] Users access the interface via a smartphone app or dedicated device and input dance conditions (inputs) such as "fun feeling" or "calm feeling." This input is sent to a server via the internet and stored as dance condition data (output).

[1115] Specific operation: The user selects the dance conditions on the device's UI and presses the send button.

[1116] Step 4: Using the Emotion Engine

[1117] The server uses cameras and additional microphones to collect the user's facial expressions, voice, and gestures (input) in real time, and recognizes emotional data (output) through an emotion engine.

[1118] Specific operation: Video data acquired by the camera is processed by a face recognition algorithm, and facial expression data is extracted. Similarly, a voice analysis algorithm is also applied.

[1119] Step 5: Emotion Analysis

[1120] The server analyzes real-time collected facial, voice, and gesture data (input) and uses an emotion engine to determine the user's emotions (e.g., joy, sadness, surprise). This emotion data (output) is then incorporated into the dance conditions.

[1121] Specific operation: The emotion engine uses a facial expression recognition model and a voice tone analysis model to generate emotion labels.

[1122] Step 6: Dance Generation

[1123] The server sends prompts to the generative AI model based on music characteristic data, dance conditions from the user, and recognized emotion data (input), and generates a dance sequence (output).

[1124] Specific operation: The server inputs a prompt message such as "Generate a dance sequence that fits a fun pop song" into the AI ​​model and retrieves the generated dance sequence data.

[1125] Step 7: Generate and transmit operation command data

[1126] The server encodes the generated dance sequence (input) as motion command data for robot operation and transmits it to the robot (terminal) via the network. This motion command data is generated as output.

[1127] Specific operation: The server converts the generated dance sequence into control data for each joint and motor, and transmits it at the appropriate timing.

[1128] Step 8: Robot execution of actions

[1129] The terminal (robot) analyzes the received motion command data (input), generates signals to control each motor and actuator, and executes the specified dance movement (output).

[1130] Specific operation: The terminal's control board analyzes the received data and sends the appropriate current and voltage to each motor to execute the specified operation in real time.

[1131] (Application Example 2)

[1132] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[1133] Conventional robot dance systems automatically generate dances based on user emotions and music characteristics. However, to improve work efficiency on a factory production line, it is necessary to consider the emotions and work conditions of the workers. This can improve work efficiency and motivation. However, existing systems lack the technology to appropriately acquire emotional data and generate response sequences. Therefore, there is a need for a robot system that can perform appropriate responses according to the emotions and work conditions of workers on a factory production line.

[1134] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[1135] In this invention, the server includes means for acquiring music data, means for analyzing music data, means for receiving dance conditions from a user, means for generating a dance sequence based on extracted music characteristic data and received dance conditions, means for encoding the generated dance sequence as robot motion command data, means for transmitting it to a robot, means for the robot to execute dance movements based on the received motion command data, means for acquiring worker emotion data from the work environment, means for generating a response sequence for improving work efficiency based on the acquired emotion data, means for encoding the generated response sequence as robot motion command data, means for transmitting it, and means for the robot to execute response movements based on the received motion command data. This makes it possible to provide appropriate responses in real time according to the worker's emotions and work situation in the work environment.

[1136] "Music data" refers to the audio signals and waveform data that make up music.

[1137] "Characteristic data" refers to distinctive information such as tempo, volume, and rhythm extracted from music data.

[1138] "Dance conditions" refer to the requirements specified by the user, such as the type of dance and the atmosphere.

[1139] "Generative artificial intelligence" refers to artificial intelligence technology used to generate new data or sequences based on existing information.

[1140] A "dance sequence" refers to a continuous flow or pattern of movements. It specifically refers to a series of actions in a dance performance.

[1141] "Action command data" refers to a set of instructions that cause a robot to perform a specific action.

[1142] A "robot" refers to a mechanical device that performs actions automatically.

[1143] "Emotional data" refers to emotional information analyzed from the workers' facial expressions, voices, gestures, etc.

[1144] A "response sequence" refers to a sequence of actions and reactions generated to respond appropriately to emotions and situations.

[1145] "Work environment" refers to the location where work is performed, such as a factory or production line, and the surrounding conditions.

[1146] "Workers" refers to people who perform tasks on production lines or similar environments.

[1147] This invention is a system that controls the movements of a robot based on music data and emotion data, thereby improving work efficiency and motivation in the work environment. The following describes specific embodiments of this system.

[1148] System configuration and operation

[1149] 1. Acquisition of music data

[1150] The server acquires music data from the surrounding environment in real time using a microphone. This music data includes audio signals.

[1151] 2. Analysis of music data

[1152] The server analyzes the acquired music data using algorithms such as FFT (Fast Fourier Transform) to extract characteristic data such as volume, tempo, and rhythm. This process utilizes, for example, high-performance computing servers or cloud computing resources.

[1153] 3. Receiving dance conditions

[1154] The server receives dance conditions such as "feeling happy" or "feeling calm" from the user via smartphone or interface. This information is transmitted to the server over the network.

[1155] 4. Acquisition of emotional data

[1156] The server collects the workers' facial expressions, voices, and gestures in real time using cameras and microphones, and analyzes them using an emotion engine. This emotion data is also stored on the server.

[1157] 5. Generation of the dance sequence

[1158] The server uses a generative AI model (e.g., GPT-4) to generate dance sequences and response sequences, taking music characteristic data, dance conditions from the user, and emotional data as input. The generated sequences are then coded as command data for robot movements.

[1159] 6. Transmission of command data

[1160] The server transmits the generated motion command data to the robot via the network. The robot's terminal analyzes this data and sends control signals to each motor and actuator.

[1161] 7. Execution of actions by the robot

[1162] The robot controls its motors and actuators appropriately based on motion command data received from the server to perform specified dances and work response actions.

[1163] Specific examples of hardware and software

[1164] Camera: High-resolution, night vision camera (e.g., Logitech C920s Pro)

[1165] Microphone: High-sensitivity microphone (e.g., Blue Yeti USB Microphone)

[1166] Sensors: Belt conveyor speed sensors and quality control sensors (e.g., Honeywell FF-SY series)

[1167] Server: High-performance server (e.g., NVIDIA DGX Station) or cloud computing resources

[1168] Generative AI models: Generative models based on GPT-4 and Transformers

[1169] Network: Stable WiFi or Ethernet connection

[1170] Specific example

[1171] Situation: Improving production efficiency within a factory

[1172] 1. The server acquires data in real time from conveyor belt speed sensors and quality control sensors within the factory.

[1173] 2. The server uses a camera and microphone to collect the worker's facial expressions, voice, and gestures, and analyzes them using an emotion engine. Based on this analysis, it recognizes that the worker is tired.

[1174] 3. The server provides the generation AI model with prompts like the following to generate a response sequence:

[1175] "If the workers are tired, please generate a refreshing dance to increase production efficiency."

[1176] 4. The generated response sequence is sent to the robot, which then performs nimble steps and simple exercises to increase worker motivation and improve efficiency.

[1177] This system makes it possible to generate and execute appropriate robot movements in real time based on music data and emotion data, contributing to improved productivity and employee motivation in the work environment.

[1178] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[1179] Step 1:

[1180] The server acquires music data from the surrounding environment. Specifically, it uses a microphone to collect audio signals in real time and saves them as audio data. The input for this step is ambient music, and the output is the acquired audio data.

[1181] Step 2:

[1182] The server analyzes the acquired music data. Specifically, it uses the FFT (Fast Fourier Transform) algorithm to extract characteristic data of the music (volume, tempo, rhythm, etc.). The input for this step is audio data, and the output is characteristic data.

[1183] Step 3:

[1184] The server receives dance conditions from the user. Specifically, it receives conditions such as "feeling happy" or "feeling calm" that the user sends via smartphone or interface over the network. The input for this step is the user's dance conditions, and the output is the received dance conditions.

[1185] Step 4:

[1186] The server acquires emotional data from the work environment. Specifically, it uses cameras and microphones to collect the workers' facial expressions, voices, and gestures in real time, and analyzes them using an emotion engine. The input for this step is the workers' facial expressions, voices, and gestures, and the output is the analyzed emotional data.

[1187] Step 5:

[1188] The server uses a generative AI model (e.g., GPT-4) to generate dance sequences and response sequences, taking music characteristic data, dance conditions from the user, and emotion data as input. Specifically, it provides the generative AI model with prompts like the following to generate appropriate sequences. The input for this step is characteristic data, dance conditions, and emotion data, and the output is the generated sequence.

[1189] Example prompt: "If the workers are tired, generate a refresh dance sequence to increase production efficiency."

[1190] Step 6:

[1191] The server encodes the generated sequence as robot motion command data and transmits it to the robot via the network. Specifically, it converts the generated sequence into control signals and sends them to the robot. The input to this step is the generated sequence, and the output is the robot motion command data.

[1192] Step 7:

[1193] The terminal (robot) controls its motors and actuators based on motion command data received from the server, and performs the specified actions. Specifically, it performs dances and response actions according to the actions commanded by the server. The input for this step is motion command data, and the output is the performed action.

[1194] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[1195] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1196] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[1197] [Fourth Embodiment]

[1198] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[1199] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1200] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1201] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[1202] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[1203] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[1204] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[1205] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[1206] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[1207] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1208] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1209] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[1210] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1211] This invention relates to a system in which a robot generates and performs dances in real time based on instructions from a user. The system analyzes the characteristics of music, generates an appropriate dance sequence based on dance conditions specified by the user, and transmits motion commands to the robot, enabling the robot to perform the dance as instructed.

[1212] System Overview

[1213] This system consists of the following main components:

[1214] 1. Means for acquiring music data

[1215] 2. Music Analysis Methods

[1216] 3. Means of receiving dance conditions

[1217] 4. Dance generation means

[1218] 5. Means for transmitting operation command data

[1219] 6. Robot motion execution means

[1220] Natural language explanation of program processing

[1221] 1. Acquisition of music data

[1222] The server acquires surrounding music data in real time via a microphone. This music data includes audio signals, which are treated as input data.

[1223] 2. Music Analysis

[1224] The server analyzes the acquired music data using algorithms such as FFT (Fast Fourier Transform). As a result of the analysis, characteristic data of the music, such as volume, tempo, and pitch, is extracted, and further processing is performed based on this data.

[1225] 3. Receiving dance conditions

[1226] The user inputs dance conditions such as "fun feeling" or "calm feeling" through the interface. This input data is sent to the server.

[1227] 4. Dance Generation

[1228] The server uses a generative artificial intelligence (generative AI) model to generate dance sequences, taking music characteristic data and dance conditions from the user as input.

[1229] The generation AI is pre-trained with various dance patterns and generates movements in real time that are appropriate to the tempo and rhythm of the music.

[1230] 5. Generation and transmission of operation command data

[1231] The server encodes the generated dance sequence into motion command data so that the robot can understand it. This data is transmitted to the terminal (robot) via the network.

[1232] 6. Execution of actions by the robot

[1233] The terminal (robot) analyzes the received motion command data and controls each motor and actuator. Based on the control signals, the robot performs the predetermined movements and performs a dance.

[1234] Specific example

[1235] Situation: Robot dancing at an event venue

[1236] 1. Pop music is playing at the event venue.

[1237] 2. The user uses a smartphone or dedicated terminal to instruct the robot to "do a fun dance."

[1238] 3. The server picks up the music from the venue via the microphone and acquires the music data.

[1239] 4. The server analyzes the music data using FFT and extracts characteristic data such as volume, tempo, and rhythm.

[1240] 5. The server provides the analysis results and the user-specified condition of "fun feeling" to the generating AI, and generates a dance sequence based on that.

[1241] 6. The server encodes the generated dance sequence as motion command data and transmits it to the terminal (robot) via the network.

[1242] 7. The terminal (robot) operates its motors and actuators based on the received motion command data to perform a "fun-feeling" dance.

[1243] Thus, this system can instantly generate and execute appropriate dance performances in response to user instructions, providing high flexibility and immediacy in events and promotions.

[1244] The following describes the processing flow.

[1245] Step 1:

[1246] Users use their smartphones or dedicated terminals to instruct the robot to "do a fun dance."

[1247] Step 2:

[1248] The server acquires music data in real time from the surroundings via a microphone sensor. This music data includes audio signals.

[1249] Step 3:

[1250] The server analyzes the acquired music data using FFT (Fast Fourier Transform). This analysis extracts characteristic data of the music, such as volume, tempo, and pitch.

[1251] Step 4:

[1252] The server temporarily stores music characteristic data in memory and receives dance conditions from the user (e.g., "happy feeling," "calm feeling").

[1253] Step 5:

[1254] The server generates a dance sequence using generative artificial intelligence (generative AI) based on the received dance conditions and music characteristics data. The generative AI selects appropriate movement patterns from the input data and outputs them as a dance sequence.

[1255] Step 6:

[1256] The server encodes the generated dance sequence as motion command data for robot operation. This motion command data is then converted into a format that the robot can understand and execute.

[1257] Step 7:

[1258] The server transmits operation command data to the terminal (robot) via the network.

[1259] Step 8:

[1260] The terminal analyzes the received operation command data and generates signals to control each motor and actuator based on that data.

[1261] Step 9:

[1262] The terminal then operates the motors and actuators appropriately according to the generated control signals. This allows the robot to perform the specified dance movements (e.g., waving arms, spinning).

[1263] Step 10:

[1264] The terminal (robot) performs a "fun-sounding" dance based on user instructions and the characteristics of the music, livening up the atmosphere at event venues and other locations.

[1265] As described above, it is possible for the robot to generate and execute dances in real time that match the characteristics of the music, based on user instructions.

[1266] (Example 1)

[1267] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1268] Conventional robot dance systems only reproduce pre-programmed movements and cannot flexibly respond to changes in music or environment. Furthermore, they struggle to reflect user emotions and intentions in real time, lacking the immediacy and flexibility needed for events and promotions. Therefore, there is a need to develop a system that generates dance sequences in real time based on music characteristics and user-specified conditions, and has the robot execute them.

[1269] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[1270] In this invention, the server includes means for acquiring music data from the surrounding environment, means for analyzing the acquired music data and extracting music characteristic data, means for receiving dance conditions from the user, means for generating a dance sequence using generative artificial intelligence based on the extracted music characteristic data and the received dance conditions, means for encoding the generated dance sequence as robot motion command data and transmitting it to the robot, means for the robot to execute dance movements based on the received motion command data, means for analyzing the acquired music data using the Fast Fourier Transform, means for the generative artificial intelligence model to consider tempo and movement patterns in the generation of the dance sequence, and means for transmitting the generated motion command data to the robot using wireless communication. This enables rapid response to changes in music and environment, and allows for dance performances that reflect the user's emotions and intentions in real time.

[1271] "Surrounding environment" refers to the location where the system is installed and the physical space surrounding it.

[1272] "Music data" refers to audio signals and music information such as background music acquired from the surrounding environment.

[1273] "Music characteristic data" refers to attribute information of music, such as volume, tempo, rhythm, and pitch, obtained by analyzing music data.

[1274] "Dance conditions" refer to the user's requirements regarding the style and atmosphere of the dance, such as "fun" or "relaxed."

[1275] "Generative artificial intelligence" refers to a technology that uses machine learning algorithms to generate new data and information based on input data.

[1276] A "dance sequence" refers to a combination of motion commands that a robot executes.

[1277] "Motion command data" refers to the control signals and commands necessary for a robot to perform a specific action.

[1278] A "robot" refers to an autonomous or semi-autonomous device that performs specified actions through a mechanical structure and computer control.

[1279] The "Fast Fourier Transform" refers to an efficient algorithm for converting time-domain data into the frequency domain.

[1280] "Wireless communication" refers to the technology of sending and receiving data using radio waves or light waves.

[1281] "Tempo" refers to the speed of music, specifically the number of beats contained within a given time.

[1282] An "action pattern" refers to a sequence of actions performed under specific conditions.

[1283] This invention relates to a system in which a robot generates and performs dances in real time based on instructions from a user. The system analyzes the characteristics of music, generates an appropriate dance sequence based on dance conditions specified by the user, and transmits motion commands to the robot, enabling the robot to perform the dance as instructed.

[1284] System Configuration

[1285] This system consists of the following main components:

[1286] 1. Means for acquiring music data

[1287] 2. Music Analysis Methods

[1288] 3. Means of receiving dance conditions

[1289] 4. Dance generation means

[1290] 5. Means for transmitting operation command data

[1291] 6. Robot motion execution means

[1292] Music data acquisition method

[1293] The server acquires ambient music data in real time through a high-sensitivity microphone. Specifically, it uses a high-quality audio capture device (e.g., a high-sensitivity microphone) to capture ambient sound as a digital signal. This acquired music data is temporarily stored in the server's buffer memory.

[1294] Music analysis methods

[1295] The server uses the Python libraries SciPy and NumPy to perform a Fast Fourier Transform (FFT) to analyze the acquired music data. This analysis extracts frequency components, and then the LibROSA library is used to calculate characteristic data of the music, such as volume, tempo, and pitch. This characteristic data is used in the next step.

[1296] Dance condition receiving method

[1297] Users specify dance conditions through an application on their smartphone or a dedicated tablet device. The user interface is developed with React Native, allowing for easy input of conditions such as "fun" or "calm." This input data is sent to the server in real time.

[1298] Dance generation means

[1299] The server uses a generative artificial intelligence (generative AI) model to generate dance sequences, taking music characteristic data and dance conditions from the user as input. The generative AI is built using TensorFlow and has been pre-trained on a variety of dance styles. For example, if the music tempo is fast and the dance condition is "fun," the generative AI will generate an energetic and rhythmic dance sequence.

[1300] Operation command data transmission means

[1301] The server encodes the generated dance sequence into motion command data that the robot can understand. Specifically, it uses ROBO-API to convert the dance sequence into control signals for the servo motors. The generated motion command data is then transmitted to the terminal (robot) via WiFi or Bluetooth.

[1302] Robot motion execution means

[1303] The terminal (robot) analyzes the received motion command data. Based on the analyzed data, it controls each motor and actuator (e.g., Dynamixel servo motor) to execute the specified action in real time. For example, if a "fun" dance command is given, the robot will take light steps and dance cheerfully while swinging both arms.

[1304] Specific example

[1305] Situation: Robot dancing at an event venue

[1306] 1. Pop music is playing at the event venue.

[1307] 2. The user uses a smartphone or dedicated device to instruct the robot to "do a fun dance." The user selects "fun" from the app's dropdown menu and clicks the send button.

[1308] 3. The server captures the music from the venue via the microphone and acquires the music data. The server processing begins the moment the microphone picks up sound.

[1309] 4. The server analyzes the music data using FFT and LibROSA, and extracts characteristic data such as volume, tempo, and rhythm. After analysis, the characteristic data is stored in memory.

[1310] 5. The server provides the generation AI with the analysis results and the user-specified condition of "fun feeling," and generates a dance sequence based on this. The generation AI generates the sequence in a few seconds, and the server checks the results.

[1311] 6. The server encodes the generated dance sequence as motion command data and sends it to the terminal (robot) via the network. As soon as the transmission is complete, the terminal sends a receipt confirmation to the server.

[1312] 7. The terminal (robot) operates its motors and actuators based on the received motion command data to perform a "fun-feeling" dance. For example, the robot takes light steps and dances cheerfully, raising and swinging both arms.

[1313] Examples of prompt statements

[1314] The following are some possible prompts to input into the generative AI model:

[1315] "Please generate a robot dance sequence based on the following conditions: Music tempo is 120 BPM, and the user-specified emotion is 'happy'. The output format should be motion command data."

[1316] In this way, each step of the entire system is intricately coordinated, enabling robot dances to be performed in real time in response to user instructions.

[1317] The flow of the specific processing in Example 1 will be explained using Figure 11.

[1318] Step 1:

[1319] The server acquires music data from its surroundings. Specifically, it uses a high-sensitivity microphone to capture audio signals in real time. It receives analog audio signals from the microphone as input, converts them to digital signals, and stores them in buffer memory. The output is digital music data.

[1320] Step 2:

[1321] The server analyzes the acquired music data. Specifically, it performs an FFT (Fast Fourier Transform) using the Python libraries SciPy and NumPy. It receives digital music data stored in buffer memory as input and extracts frequency components. It also calculates characteristic data such as volume, tempo, rhythm, and pitch using the LibROSA library. The output is characteristic data of the music.

[1322] Step 3:

[1323] Users specify dance conditions through an application on their smartphone or a dedicated tablet device. Users input dance conditions such as "fun feeling" or "calm feeling" through the interface. The input consists of the user's specified conditions and is transmitted from the application to the server in real time. The output is the user's dance condition data.

[1324] Step 4:

[1325] The server uses a generative artificial intelligence (generative AI) model to generate dance sequences, taking music characteristic data and dance condition data from the user as input. The generative AI is pre-trained on a variety of dance styles and generates the optimal dance sequence according to the input data. The input is music characteristic data and dance condition data, and the output is the generated dance sequence.

[1326] Step 5:

[1327] The server encodes the generated dance sequence as robot motion command data. Specifically, it uses ROBO-API to convert the dance sequence into servo motor control signals. The input is the generated dance sequence, and the output is the motion command data. This motion command data is transmitted to the terminal (robot) via WiFi or Bluetooth.

[1328] Step 6:

[1329] The terminal (robot) analyzes the received motion command data. Based on the analyzed data, it controls each motor and actuator to execute the specified action in real time. The input is motion command data, and the output is the actual robot's movement. For example, if a "fun" dance command is given, the robot will dance with light steps and swing its arms.

[1330] Through these processing steps, the system can perform robotic dances in real time in response to user instructions.

[1331] (Application Example 1)

[1332] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1333] Conventional robot dance systems only execute pre-programmed movements, making it difficult for them to respond to changes in the surrounding environment or music in real time. Furthermore, they lack a mechanism for users to spontaneously instruct the robots to dance, resulting in a lack of flexibility in entertainment and promotional settings. Additionally, they are unable to properly consider tempo and movement patterns when generating dance sequences, making it difficult to provide an engaging performance for audiences.

[1334] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[1335] In this invention, the server includes means for acquiring music data from the surrounding environment, means for analyzing the acquired music data and extracting music characteristic data, means for receiving dance conditions from the user, means for generating a dance sequence using artificial intelligence based on the extracted music characteristic data and the received dance conditions, means for encoding the generated dance sequence as robot motion command data and transmitting it to the robot, means for the robot to execute dance movements based on the received motion command data, and means for the user to give improvisational dance instructions to the robot via a smartphone. As a result, the robot can perform dances in real time according to the characteristics of the surrounding music and respond to improvisational instructions from the user, providing high flexibility and immediacy for events and promotions.

[1336] "Means for acquiring music data" refers to devices or methods for acquiring music data from the surrounding environment.

[1337] "Music analysis means" refers to devices or methods that analyze acquired music data and extract characteristic data of the music.

[1338] A "dance condition receiving method" refers to a device or method for receiving dance conditions from users.

[1339] A "dance generation means" is a device or method that generates a dance sequence using artificial intelligence based on extracted music characteristic data and received dance conditions.

[1340] "Motion command data transmission means" refers to a device or method that encodes the generated dance sequence as motion command data for a robot and transmits it to the robot.

[1341] A "robot motion execution means" is a device or method that allows a robot to perform dance movements based on motion command data it has received.

[1342] A "smartphone" is a portable information terminal used by users to give improvised dance instructions to a robot.

[1343] "Generative artificial intelligence" is an artificial intelligence technology that generates dance sequences based on music characteristic data and user conditions.

[1344] This invention relates to a system in which a robot generates and performs dances in real time based on instructions from a user. The system analyzes the characteristics of music, generates an appropriate dance sequence based on dance conditions specified by the user, and transmits motion commands to the robot, enabling the robot to perform the dance as instructed.

[1345] System Overview

[1346] This system consists of the following main components:

[1347] 1. Means for acquiring music data

[1348] The server acquires music data from its surrounding environment. Specifically, it acquires music playing in environments such as event venues and stores in real time via microphones.

[1349] 2. Music Analysis Methods

[1350] The server analyzes the acquired music data using algorithms such as FFT (Fast Fourier Transform). As a result of the analysis, characteristic data of the music, such as volume, tempo, and pitch, is extracted.

[1351] 3. Means of receiving dance conditions

[1352] Users input dance conditions for the robot via their smartphones, such as "fun" or "elegant."

[1353] 4. Dance generation means

[1354] The server generates dance sequences using a generative AI model based on the extracted music characteristic data and received dance conditions. The generative AI model is pre-trained with various dance patterns and generates movements in real time that are appropriate to the tempo and rhythm of the music.

[1355] 5. Means for transmitting operation command data

[1356] The server encodes the generated dance sequence as robot motion command data and transmits it to the robot via the network. Bluetooth and Wi-Fi are used as communication methods.

[1357] 6. Robot motion execution means

[1358] The robot controls its motors and actuators based on the motion command data it receives, and performs the specified dance movements.

[1359] Specific example

[1360] Situation: Robot dancing at an event venue

[1361] 1. Pop music is playing at the event venue.

[1362] 2. The user uses a smartphone application to instruct the robot to "do a fun dance."

[1363] 3. The server captures the music from the venue via microphones and acquires the music data.

[1364] 4. The server analyzes the music data using FFT and extracts characteristic data such as volume, tempo, and rhythm.

[1365] 5. The server provides the analysis results and the user-specified condition of "fun feeling" to the generating AI, and generates a dance sequence based on that.

[1366] 6. The server encodes the generated dance sequence as motion command data and transmits it to the robot via the network.

[1367] 7. The robot operates its motors and actuators based on the received motion command data to perform a "fun-feeling" dance.

[1368] Thus, this system can instantly generate and execute appropriate dance performances in response to user instructions, providing high flexibility and immediacy in events and promotions.

[1369] Examples of prompt statements used

[1370] Based on the user's specified "fun" dance conditions, analyze the tempo and volume of the acquired music data and generate a suitable dance sequence.

[1371] The above describes specific embodiments for carrying out the present invention.

[1372] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[1373] Step 1:

[1374] The server acquires music data from the surrounding environment. Specifically, it collects music data in real time through a microphone. The input for this step is the audio signal of the music, and the output is the acquired raw music data.

[1375] Step 2:

[1376] The server analyzes the acquired music data. It extracts characteristic music data using algorithms such as FFT (Fast Fourier Transform). The input to this step is the acquired music data, and the output is music characteristic data such as volume, tempo, and rhythm.

[1377] Step 3:

[1378] The user inputs dance conditions for the robot via their smartphone. For example, conditions such as "fun" or "elegant" can be specified. The input in this step is the user's dance conditions, and the output is data for those dance conditions.

[1379] Step 4:

[1380] The server generates a dance sequence using a generative AI model based on the music characteristic data and the received dance conditions. The input for this step is the music characteristic data and the user's dance conditions, and the output is the generated dance sequence.

[1381] Step 5:

[1382] The server encodes the generated dance sequence as robot motion command data and transmits it to the robot over the network. The input to this step is the generated dance sequence, and the output is the motion command data sent to the robot.

[1383] Step 6:

[1384] The robot controls its motors and actuators based on the received motion command data to perform the specified dance movements. The input for this step is the motion command data, and the output is the actual dance movements of the robot.

[1385] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[1386] This invention relates to a robot dance system that incorporates an emotion engine that recognizes user emotions. This system analyzes the characteristics of music, generates an appropriate dance sequence according to the user's emotions and specified dance conditions, and transmits motion commands to the robot, enabling the robot to perform the dance in real time.

[1387] System Overview

[1388] This system consists of the following main components:

[1389] 1. Means for acquiring music data

[1390] 2. Music Analysis Methods

[1391] 3. Emotional Engine

[1392] 4. Means of receiving dance conditions

[1393] 5. Dance generation means

[1394] 6. Operation command data transmission means

[1395] 7. Robot motion execution means

[1396] Natural language explanation of program processing

[1397] 1. Acquisition of music data

[1398] The server acquires surrounding music data in real time via a microphone. This music data includes audio signals, which are treated as input data.

[1399] 2. Music Analysis

[1400] The server analyzes the acquired music data using algorithms such as FFT (Fast Fourier Transform). As a result of the analysis, characteristic data of the music, such as volume, tempo, and pitch, is extracted, and further processing is performed based on this data.

[1401] 3. Receiving dance conditions

[1402] The user inputs dance conditions such as "fun feeling" or "calm feeling" through the interface. This input data is sent to the server.

[1403] 4. Use of the Emotion Engine

[1404] The emotion engine integrated into the server recognizes emotions through the user's facial expressions, voice, and gestures. This uses cameras and microphones to collect data in real time.

[1405] 5. Emotion analysis

[1406] The server uses an emotion engine to analyze the collected data and determine the user's emotions (e.g., joy, sadness, surprise). This emotion data is also incorporated into the dance conditions.

[1407] 6. Dance Generation

[1408] The server uses a generative artificial intelligence (generative AI) model to generate dance sequences, taking music characteristic data, dance conditions from the user, and emotional data as input.

[1409] The generating AI is pre-trained with various dance patterns and generates movements in real time that are appropriate to the music's tempo and rhythm, as well as the user's emotions.

[1410] 7. Generation and transmission of operation command data

[1411] The server encodes the generated dance sequence as motion command data for robot operation. This data is transmitted to the terminal (robot) via the network.

[1412] 8. Execution of actions by the robot

[1413] The terminal (robot) analyzes the received motion command data and generates signals to control each motor and actuator.

[1414] The terminal then operates the motors and actuators appropriately according to the generated control signals. This allows the robot to perform predetermined movements and dance.

[1415] Specific example

[1416] Situation: Robot dancing at an event venue

[1417] 1. Pop music is playing at the event venue.

[1418] 2. The user uses a smartphone or dedicated terminal to instruct the robot to "do a fun dance."

[1419] 3. The server picks up the music in the venue via a microphone sensor and acquires the music data.

[1420] 4. The server analyzes the music data using FFT and extracts characteristic data such as volume, tempo, and rhythm.

[1421] 5. The server uses cameras and microphones to collect the user's facial expressions, voice, and gestures in real time, and uses an emotion engine to analyze the user's emotions.

[1422] 6. The server provides the analysis results and the user-specified condition of "fun feeling" to the generating AI, and generates a dance sequence based on that.

[1423] 7. The server encodes the generated dance sequence as motion command data and transmits it to the terminal (robot) via the network.

[1424] 8. The terminal (robot) operates its motors and actuators based on the received motion command data to perform a "fun-feeling" dance.

[1425] Thus, this system can instantly generate and execute appropriate dance performances in response to user instructions and emotions, providing high flexibility and immediacy in events and promotions.

[1426] The following describes the processing flow.

[1427] Step 1:

[1428] Users use their smartphones or dedicated terminals to instruct the robot to "do a fun dance."

[1429] Step 2:

[1430] The server acquires music data in real time from the surroundings via a microphone sensor. This music data includes audio signals.

[1431] Step 3:

[1432] The server analyzes the acquired music data using FFT (Fast Fourier Transform). This analysis extracts characteristic data of the music, such as volume, tempo, and pitch.

[1433] Step 4:

[1434] The server temporarily stores music characteristic data in memory.

[1435] Step 5:

[1436] The server uses cameras and microphones to collect the user's facial expressions, voice, and gestures in real time, and analyzes them using an emotion engine. The analysis recognizes the user's emotions (e.g., joy, surprise).

[1437] Step 6:

[1438] The server receives the analyzed user emotion data along with dance conditions from the user (e.g., "feeling happy").

[1439] Step 7:

[1440] The server generates dance sequences using generative artificial intelligence (generative AI) based on music characteristic data, dance conditions from the user, and emotional data. The generative AI selects appropriate dance patterns from the input data and outputs them as dance sequences.

[1441] Step 8:

[1442] The server encodes the generated dance sequence as motion command data for robot operation. This motion command data is then converted into a format that the robot can understand and execute.

[1443] Step 9:

[1444] The server transmits operation command data to the terminal (robot) via the network.

[1445] Step 10:

[1446] The terminal analyzes the received operation command data and generates signals to control each motor and actuator based on that data.

[1447] Step 11:

[1448] The terminal then operates the motors and actuators appropriately according to the generated control signals. This allows the robot to perform predetermined movements and dance.

[1449] Step 12:

[1450] The terminal (robot) performs a "fun-sounding" dance based on user instructions and the characteristics of the music, livening up the atmosphere at event venues and other locations.

[1451] In this way, the robot can generate and perform dances in real time that match the characteristics of the music, based on user instructions and emotion recognition.

[1452] (Example 2)

[1453] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1454] Conventional robot dance systems primarily perform movements based solely on the tempo and rhythm of music, making it difficult to reflect user emotions or specific dance conditions in real time. This limits the user experience and reduces the expressiveness and flexibility of robot dance. This invention aims to generate dance sequences in response to user emotions and specified dance conditions, enabling the robot to perform more dynamic and adaptive dances in real time.

[1455] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for acquiring music data from the surrounding environment, means for analyzing the acquired music data and extracting music characteristic data, means for receiving dance conditions from the user, means for recognizing emotions through the user's facial expressions, voice, and gestures, means for analyzing the user's emotion data, means for generating a dance sequence using generative artificial intelligence based on the extracted music characteristic data, the received dance conditions, and the recognized emotion data, means for encoding the generated dance sequence as robot motion command data and transmitting it to the robot, and means for the robot to execute dance movements based on the received motion command data. This makes it possible to perform advanced, real-time robot dances that are in line with the user's emotions and dance conditions.

[1456] "Means of acquiring music data from the surrounding environment" refers to a function that captures ambient audio signals in real time using sensors such as microphones built into robots or servers.

[1457] "Methods for analyzing acquired music data and extracting characteristic music data" refers to a function that applies algorithms such as FFT (Fast Fourier Transform) to music data to calculate and extract characteristic data such as volume, tempo, rhythm, and pitch.

[1458] "Means of receiving dance conditions from users" refers to a function that inputs and receives dance conditions specified by the user (e.g., "a fun feeling" or "a calm feeling") via an interface such as a smartphone app or dedicated terminal.

[1459] "Means of recognizing emotions through the user's facial expressions, voice, and gestures" refers to a function that uses cameras and microphones to capture the user's facial expressions, voice tone, and gestures in real time, and recognizes emotions through an emotion engine.

[1460] "Methods for analyzing user emotion data" refers to a function that analyzes facial expressions, voice, and gesture data acquired in real time and converts the user's emotions into labels such as joy, sadness, and surprise.

[1461] "A means of generating a dance sequence using generative artificial intelligence based on extracted music characteristic data, received dance conditions, and recognized emotion data" refers to a function that uses a generative artificial intelligence model (generative AI model) to generate a suitable dance sequence in real time, taking music characteristic data, dance conditions, and emotion data as input.

[1462] "Means for encoding the generated dance sequence as robot motion command data and transmitting it to the robot" refers to a function that converts the generated dance sequence into specific control signals for robot operation and transmits them to the robot via a network.

[1463] "Means for executing dance movements based on motion command data received by the robot" refers to a function that, based on the control signals received by the robot, appropriately operates each motor and actuator to execute a predetermined dance movement.

[1464] This invention relates to a robot dance system that incorporates an emotion engine that recognizes user emotions. This system analyzes the characteristics of music, generates an appropriate dance sequence according to the user's emotions and specified dance conditions, and transmits motion commands to the robot, enabling the robot to perform the dance in real time.

[1465] System Configuration

[1466] This system consists of the following main components:

[1467] 1. Means for acquiring music data

[1468] The server uses sensors, such as microphones built into the robot, to acquire ambient music in real time. This allows it to collect audio signals.

[1469] 2. Music Analysis Methods

[1470] The server analyzes the acquired music data using algorithms such as FFT (Fast Fourier Transform) to extract characteristic data such as volume, tempo, rhythm, and pitch.

[1471] 3. Means of receiving dance conditions

[1472] Users input dance conditions such as "fun feeling" or "calm feeling" via a smartphone app or dedicated device. The data is then transmitted to a server via the internet.

[1473] 4. Emotion Engine (Means of Emotion Recognition)

[1474] The server uses cameras and additional microphones to collect the user's facial expressions, voice, and gestures in real time, and uses an emotion engine to recognize the user's emotions.

[1475] 5. Emotion analysis method

[1476] The server analyzes data collected through the emotion engine to determine the user's emotions (joy, sadness, surprise). This emotion data is also incorporated into the dance conditions.

[1477] 6. Dance generation means

[1478] The server generates dance sequences using a generative AI model based on music characteristic data, dance conditions from the user, and recognized emotion data.

[1479] 7. Operation command data transmission means

[1480] The server encodes the generated dance sequence as motion command data for robot operation and transmits it to the robot (terminal) via the network.

[1481] 8. Robot motion execution means

[1482] The terminal (robot) analyzes the received motion command data, generates signals to control each motor and actuator, and executes the specified dance movements.

[1483] Specific example

[1484] Situation: Robot dancing at an event venue

[1485] 1. Pop music is playing at the event venue.

[1486] 2. The user uses a smartphone or dedicated terminal to instruct the robot to "do a fun dance."

[1487] 3. The server picks up the music in the venue via a microphone sensor and acquires the music data.

[1488] 4. The server analyzes the music data using FFT and extracts characteristic data such as volume, tempo, and rhythm.

[1489] 5. The server uses cameras and microphones to collect the user's facial expressions, voice, and gestures in real time, and uses an emotion engine to analyze the user's emotions.

[1490] 6. The server provides the analysis results and the user-specified condition of "fun feeling" as prompts to the generating AI, and generates a dance sequence based on them.

[1491] 7. The server encodes the generated dance sequence as motion command data and transmits it to the terminal (robot) via the network.

[1492] 8. The terminal (robot) operates its motors and actuators based on the received motion command data to perform a "fun-feeling" dance.

[1493] Example of a prompt

[1494] "Generate a dance sequence that fits a fun, upbeat pop song. The user is expressing joy. The tempo is 120 BPM, and the volume is medium."

[1495] By inputting this prompt into the AI ​​model, it is possible to generate a dance sequence that is suitable for the user's emotions and musical characteristics.

[1496] The flow of the specific processing in Example 2 will be explained using Figure 13.

[1497] Step 1: Obtain music data

[1498] The server uses sensors, such as microphones built into the robot, to acquire ambient music in real time. The music data acquired at this stage (input) is stored in a buffer as an audio signal (output).

[1499] Specific operation: The server captures audio data at regular time intervals and monitors it continuously.

[1500] Step 2: Music Analysis

[1501] The server analyzes the acquired music data (input) using FFT (Fast Fourier Transform) and extracts characteristic data (output) such as frequency components, volume, tempo, rhythm, and pitch.

[1502] Specific operation: The server performs an FFT on each time window to extract time-frequency domain features. A beat detection algorithm is then applied to calculate tempo data.

[1503] Step 3: Receiving the dance conditions

[1504] Users access the interface via a smartphone app or dedicated device and input dance conditions (inputs) such as "fun feeling" or "calm feeling." This input is sent to a server via the internet and stored as dance condition data (output).

[1505] Specific operation: The user selects the dance conditions on the device's UI and presses the send button.

[1506] Step 4: Using the Emotion Engine

[1507] The server uses cameras and additional microphones to collect the user's facial expressions, voice, and gestures (input) in real time, and recognizes emotional data (output) through an emotion engine.

[1508] Specific operation: Video data acquired by the camera is processed by a face recognition algorithm, and facial expression data is extracted. Similarly, a voice analysis algorithm is also applied.

[1509] Step 5: Emotion Analysis

[1510] The server analyzes real-time collected facial, voice, and gesture data (input) and uses an emotion engine to determine the user's emotions (e.g., joy, sadness, surprise). This emotion data (output) is then incorporated into the dance conditions.

[1511] Specific operation: The emotion engine uses a facial expression recognition model and a voice tone analysis model to generate emotion labels.

[1512] Step 6: Dance Generation

[1513] The server sends prompts to the generative AI model based on music characteristic data, dance conditions from the user, and recognized emotion data (input), and generates a dance sequence (output).

[1514] Specific operation: The server inputs a prompt message such as "Generate a dance sequence that fits a fun pop song" into the AI ​​model and retrieves the generated dance sequence data.

[1515] Step 7: Generate and transmit operation command data

[1516] The server encodes the generated dance sequence (input) as motion command data for robot operation and transmits it to the robot (terminal) via the network. This motion command data is generated as output.

[1517] Specific operation: The server converts the generated dance sequence into control data for each joint and motor, and transmits it at the appropriate timing.

[1518] Step 8: Robot execution of actions

[1519] The terminal (robot) analyzes the received motion command data (input), generates signals to control each motor and actuator, and executes the specified dance movement (output).

[1520] Specific operation: The terminal's control board analyzes the received data and sends the appropriate current and voltage to each motor to execute the specified operation in real time.

[1521] (Application Example 2)

[1522] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1523] Conventional robot dance systems automatically generate dances based on user emotions and music characteristics. However, to improve work efficiency on a factory production line, it is necessary to consider the emotions and work conditions of the workers. This can improve work efficiency and motivation. However, existing systems lack the technology to appropriately acquire emotional data and generate response sequences. Therefore, there is a need for a robot system that can perform appropriate responses according to the emotions and work conditions of workers on a factory production line.

[1524] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[1525] In this invention, the server includes means for acquiring music data, means for analyzing music data, means for receiving dance conditions from a user, means for generating a dance sequence based on extracted music characteristic data and received dance conditions, means for encoding the generated dance sequence as robot motion command data, means for transmitting it to a robot, means for the robot to execute dance movements based on the received motion command data, means for acquiring worker emotion data from the work environment, means for generating a response sequence for improving work efficiency based on the acquired emotion data, means for encoding the generated response sequence as robot motion command data, means for transmitting it, and means for the robot to execute response movements based on the received motion command data. This makes it possible to provide appropriate responses in real time according to the worker's emotions and work situation in the work environment.

[1526] "Music data" refers to the audio signals and waveform data that make up music.

[1527] "Characteristic data" refers to distinctive information such as tempo, volume, and rhythm extracted from music data.

[1528] "Dance conditions" refer to the requirements specified by the user, such as the type of dance and the atmosphere.

[1529] "Generative artificial intelligence" refers to artificial intelligence technology used to generate new data or sequences based on existing information.

[1530] A "dance sequence" refers to a continuous flow or pattern of movements. It specifically refers to a series of actions in a dance performance.

[1531] "Action command data" refers to a set of instructions that cause a robot to perform a specific action.

[1532] A "robot" refers to a mechanical device that performs actions automatically.

[1533] "Emotional data" refers to emotional information analyzed from the workers' facial expressions, voices, gestures, etc.

[1534] A "response sequence" refers to a sequence of actions and reactions generated to respond appropriately to emotions and situations.

[1535] "Work environment" refers to the location where work is performed, such as a factory or production line, and the surrounding conditions.

[1536] "Workers" refers to people who perform tasks on production lines or similar environments.

[1537] This invention is a system that controls the movements of a robot based on music data and emotion data, thereby improving work efficiency and motivation in the work environment. The following describes specific embodiments of this system.

[1538] System configuration and operation

[1539] 1. Acquisition of music data

[1540] The server acquires music data from the surrounding environment in real time using a microphone. This music data includes audio signals.

[1541] 2. Analysis of music data

[1542] The server analyzes the acquired music data using algorithms such as FFT (Fast Fourier Transform) to extract characteristic data such as volume, tempo, and rhythm. This process utilizes, for example, high-performance computing servers or cloud computing resources.

[1543] 3. Receiving dance conditions

[1544] The server receives dance conditions such as "feeling happy" or "feeling calm" from the user via smartphone or interface. This information is transmitted to the server over the network.

[1545] 4. Acquisition of emotional data

[1546] The server collects the workers' facial expressions, voices, and gestures in real time using cameras and microphones, and analyzes them using an emotion engine. This emotion data is also stored on the server.

[1547] 5. Generation of the dance sequence

[1548] The server uses a generative AI model (e.g., GPT-4) to generate dance sequences and response sequences, taking music characteristic data, dance conditions from the user, and emotional data as input. The generated sequences are then coded as command data for robot movements.

[1549] 6. Transmission of command data

[1550] The server transmits the generated motion command data to the robot via the network. The robot's terminal analyzes this data and sends control signals to each motor and actuator.

[1551] 7. Execution of actions by the robot

[1552] The robot controls its motors and actuators appropriately based on motion command data received from the server to perform specified dances and work response actions.

[1553] Specific examples of hardware and software

[1554] Camera: High-resolution, night vision camera (e.g., Logitech C920s Pro)

[1555] Microphone: High-sensitivity microphone (e.g., Blue Yeti USB Microphone)

[1556] Sensors: Belt conveyor speed sensors and quality control sensors (e.g., Honeywell FF-SY series)

[1557] Server: High-performance server (e.g., NVIDIA DGX Station) or cloud computing resources

[1558] Generative AI models: Generative models based on GPT-4 and Transformers

[1559] Network: Stable WiFi or Ethernet connection

[1560] Specific example

[1561] Situation: Improving production efficiency within a factory

[1562] 1. The server acquires data in real time from conveyor belt speed sensors and quality control sensors within the factory.

[1563] 2. The server uses a camera and microphone to collect the worker's facial expressions, voice, and gestures, and analyzes them using an emotion engine. Based on this analysis, it recognizes that the worker is tired.

[1564] 3. The server provides the generation AI model with prompts like the following to generate a response sequence:

[1565] "If the workers are tired, please generate a refreshing dance to increase production efficiency."

[1566] 4. The generated response sequence is sent to the robot, which then performs nimble steps and simple exercises to increase worker motivation and improve efficiency.

[1567] This system makes it possible to generate and execute appropriate robot movements in real time based on music data and emotion data, contributing to improved productivity and employee motivation in the work environment.

[1568] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[1569] Step 1:

[1570] The server acquires music data from the surrounding environment. Specifically, it uses a microphone to collect audio signals in real time and saves them as audio data. The input for this step is ambient music, and the output is the acquired audio data.

[1571] Step 2:

[1572] The server analyzes the acquired music data. Specifically, it uses the FFT (Fast Fourier Transform) algorithm to extract characteristic data of the music (volume, tempo, rhythm, etc.). The input for this step is audio data, and the output is characteristic data.

[1573] Step 3:

[1574] The server receives dance conditions from the user. Specifically, it receives conditions such as "feeling happy" or "feeling calm" that the user sends via smartphone or interface over the network. The input for this step is the user's dance conditions, and the output is the received dance conditions.

[1575] Step 4:

[1576] The server acquires emotional data from the work environment. Specifically, it uses cameras and microphones to collect the workers' facial expressions, voices, and gestures in real time, and analyzes them using an emotion engine. The input for this step is the workers' facial expressions, voices, and gestures, and the output is the analyzed emotional data.

[1577] Step 5:

[1578] The server uses a generative AI model (e.g., GPT-4) to generate dance sequences and response sequences, taking music characteristic data, dance conditions from the user, and emotion data as input. Specifically, it provides the generative AI model with prompts like the following to generate appropriate sequences. The input for this step is characteristic data, dance conditions, and emotion data, and the output is the generated sequence.

[1579] Example prompt: "If the workers are tired, generate a refresh dance sequence to increase production efficiency."

[1580] Step 6:

[1581] The server encodes the generated sequence as robot motion command data and transmits it to the robot via the network. Specifically, it converts the generated sequence into control signals and sends them to the robot. The input to this step is the generated sequence, and the output is the robot motion command data.

[1582] Step 7:

[1583] The terminal (robot) controls its motors and actuators based on motion command data received from the server, and performs the specified actions. Specifically, it performs dances and response actions according to the actions commanded by the server. The input for this step is motion command data, and the output is the performed action.

[1584] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[1585] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1586] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[1587] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1588] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[1589] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[1590] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[1591] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[1592] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[1593] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[1594] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[1595] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[1596] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[1597] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1598] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[1599] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[1600] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[1601] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[1602] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[1603] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[1604] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[1605] The following is further disclosed regarding the embodiments described above.

[1606] (Claim 1)

[1607] A means of acquiring music data from the surrounding environment,

[1608] A means for analyzing acquired music data and extracting characteristic data of the music,

[1609] A means of receiving dance conditions from users,

[1610] A means for generating a dance sequence using generative artificial intelligence based on extracted music characteristic data and received dance conditions,

[1611] A means for encoding the generated dance sequence as robot motion command data and transmitting it to the robot,

[1612] A means by which a robot performs dance movements based on motion command data it receives,

[1613] A system that includes this.

[1614] (Claim 2)

[1615] The system according to claim 1 for analyzing music data using the Fast Fourier Transform.

[1616] (Claim 3)

[1617] The system according to claim 1, which uses a generative artificial intelligence model that takes into account tempo and movement patterns in generating dance sequences.

[1618] "Example 1"

[1619] (Claim 1)

[1620] A means of acquiring music data from the surrounding environment,

[1621] A means for analyzing acquired music data and extracting characteristic data of the music,

[1622] A means of receiving dance conditions from users,

[1623] A means for generating a dance sequence using generative artificial intelligence based on extracted music characteristic data and received dance conditions,

[1624] A means for encoding the generated dance sequence as robot motion command data and transmitting it to the robot,

[1625] A means by which a robot performs dance movements based on motion command data it receives,

[1626] A method for analyzing acquired music data using the Fast Fourier Transform,

[1627] A means by which a generative artificial intelligence model considers tempo and movement patterns in generating dance sequences,

[1628] A means for transmitting the generated motion command data to the robot using wireless communication,

[1629] A system that includes this.

[1630] (Claim 2)

[1631] The system according to claim 1, which analyzes the acquired music data in real time.

[1632] (Claim 3)

[1633] The system according to claim 1, wherein dance conditions are input via a user interface.

[1634] "Application Example 1"

[1635] (Claim 1)

[1636] A means of acquiring music data from the surrounding environment,

[1637] A means for analyzing acquired music data and extracting characteristic data of the music,

[1638] A means of receiving dance conditions from users,

[1639] A means for generating a dance sequence using generative artificial intelligence based on extracted music characteristic data and received dance conditions,

[1640] A means for encoding the generated dance sequence as robot motion command data and transmitting it to the robot,

[1641] A means by which a robot performs dance movements based on motion command data it receives,

[1642] A means by which users can give impromptu dance instructions to a robot via a smartphone,

[1643] A system that includes this.

[1644] (Claim 2)

[1645] The system according to claim 1 for analyzing music data using the Fast Fourier Transform.

[1646] (Claim 3)

[1647] The system according to claim 1, which uses a generative artificial intelligence model that takes into account tempo and movement patterns in generating dance sequences.

[1648] "Example 2 of combining an emotion engine"

[1649] (Claim 1)

[1650] A means of acquiring music data from the surrounding environment,

[1651] A means for analyzing acquired music data and extracting characteristic data of the music,

[1652] A means of receiving dance conditions from users,

[1653] A means of recognizing emotions through the user's facial expressions, voice, and gestures,

[1654] A means of analyzing user sentiment data,

[1655] A means for generating a dance sequence using generative artificial intelligence based on extracted music characteristic data, received dance conditions, and recognized emotion data,

[1656] A means for encoding the generated dance sequence as robot motion command data and transmitting it to the robot,

[1657] A means by which a robot performs dance movements based on motion command data it receives,

[1658] A system that includes this.

[1659] (Claim 2)

[1660] The system according to claim 1 for analyzing music data using the Fast Fourier Transform.

[1661] (Claim 3)

[1662] The system according to claim 1, which uses a generative artificial intelligence model that takes into account tempo and movement patterns in generating dance sequences.

[1663] "Application example 2 of combining emotional engines"

[1664] (Claim 1)

[1665] A means of acquiring music data from the surrounding environment,

[1666] A means for analyzing acquired music data and extracting characteristic data of the music,

[1667] A means of receiving dance conditions from users,

[1668] A means for generating a dance sequence using generative artificial intelligence based on extracted music characteristic data and received dance conditions,

[1669] A means for encoding the generated dance sequence as robot motion command data and transmitting it to the robot,

[1670] A means by which a robot performs dance movements based on motion command data it receives,

[1671] A means of obtaining worker emotional data from the work environment,

[1672] A means for generating response sequences to improve work efficiency based on acquired emotional data,

[1673] A means for encoding the generated response sequence as robot motion command data and transmitting it to the robot,

[1674] A means for the robot to perform a response action based on the motion command data it has received,

[1675] A system that includes this.

[1676] (Claim 2)

[1677] The system according to claim 1 for analyzing music data using the Fast Fourier Transform.

[1678] (Claim 3)

[1679] The system according to claim 1, which uses a generative artificial intelligence model that takes into account tempo and movement patterns in generating dance sequences. [Explanation of symbols]

[1680] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. A means of acquiring music data from the surrounding environment, A means for analyzing acquired music data and extracting characteristic data of the music, A means of receiving dance conditions from users, A means for generating a dance sequence using generative artificial intelligence based on extracted music characteristic data and received dance conditions, A means for encoding the generated dance sequence as robot motion command data and transmitting it to the robot, A means by which a robot performs dance movements based on motion command data it receives, A system that includes this.

2. The system according to claim 1 for analyzing music data using the Fast Fourier Transform.

3. The system according to claim 1, which uses a generative artificial intelligence model that takes into account tempo and movement patterns in generating dance sequences.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A