System
The system uses a generative model to recreate cheers and excitement for remote spectators by generating on-site sounds from their devices and transmitting venue audio, addressing the lack of realism and unity in remote viewing.
Patent Information
- Application Number
- JP2024129403
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-08-05
- Publication Date
- 2026-02-18
AI Technical Summary
Remote spectators experience a poor viewing experience due to the lack of realism and unity in cheering, as conventional systems struggle to generate and provide real-time audio feedback from remote locations.
A system utilizing a generative model that analyzes input signals from remote spectators' communication devices to generate cheering sounds, which are played on-site, while also recording and transmitting audio from the venue to the spectators' devices, creating an immersive experience.
This system allows remote spectators to feel as if they are present at the event by recreating the cheers and excitement, enhancing the viewing experience with real-time interaction and audio feedback.
Smart Images

Figure 2026026982000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] In recent years, the ways of watching sports have diversified, with an increasing number of people watching remotely. However, remote spectators are unable to fully experience the excitement and realism of being there, resulting in a poor viewing experience. Furthermore, cheering from remote locations does not reach the venue, resulting in a lack of unity. Conventional systems have difficulty generating audio and providing feedback on remote cheering in real time, creating a need for a solution to these issues. [Means for solving the problem]
[0005] The present invention provides a system that uses a generative model to reproduce the cheers and excitement of a spectator's location. This system transmits input signals from the remote spectators' communication devices to a server, analyzes the input signals, generates cheering sounds, and outputs them from a sound device. Furthermore, the system records the audio from the spectator's location, transmits it to the communication device, and plays it back on that device, providing the remote spectators with a sense of realism.
[0006] Specifically, the system analyzes audio feeds from the viewing location, and when a specific event is detected, generates audio based on that event using a generative model. Vibration and motion data from the remote cheering device is sent to a server, which analyzes the data to generate cheering sounds, which are then played on the on-site sound system. This technology allows remote spectators to feel as if they are actually there, providing a truly immersive viewing experience.
[0007] A "generative model" is an algorithm that generates output data such as audio or images based on input data.
[0008] "Spectator venue" refers to the location where a sporting event or sporting event is held.
[0009] "Cheers" refers to the loud noises made by spectators at sporting events, concerts, etc.
[0010] "Excitement" refers to the excited and passionate atmosphere of spectators and participants at an event or sporting event.
[0011] "Communications terminal" means a device capable of sending and receiving data over the Internet or other communications network.
[0012] "Input signal" refers to data or information transmitted from a communication terminal.
[0013] A "server" refers to a computer system that provides services to multiple client devices over a network.
[0014] "Analysis" refers to the process of analyzing received data in detail and extracting useful information from it.
[0015] "Cheering sounds" refers to sounds that express support, such as cheers and applause from the audience.
[0016] "Audio equipment" refers to equipment or systems for reproducing sound.
[0017] "Recording" refers to the process of recording sound.
[0018] "Streaming" refers to a technology for playing data as it is transmitted and received in real time.
[0019] A "remote support device" refers to a device used to convey the intention to support from a remote location.
[0020] "Vibration Data" refers to data relating to the movement or vibration of a device.
[0021] "Motion Data" refers to data relating to the movement or positional changes of a device.
[0022] An "event" refers to a series of actions or occurrences, such as a sporting event or a concert. [Brief explanation of the drawings]
[0023] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5]FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0024] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0025] First, the terms used in the following description will be explained.
[0026] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0027] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0028] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0029] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0030] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0031] [First embodiment]
[0032] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0033] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0034] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0035] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0036] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0037] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0038] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0039] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0040] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0041] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0042] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0043] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0044] This system uses a generative model to recreate the cheers and excitement of the spectator's location, providing a realistic experience for remote spectators. The system is primarily composed of a server, communication terminals, remote cheering devices, and audio equipment.
[0045] Program processing and explanation
[0046] Initializing a generative AI model
[0047] 1. The server initializes the generative AI model
[0048] On startup, the server loads the generative AI model and sets the necessary parameters, so the model is ready to generate cheers and other acoustic patterns for the spectator area.
[0049] Launching the app and connecting
[0050] 2. The user launches the app and selects a match.
[0051] Users launch the app on their smartphone or other communication device and select the game they want to watch. Once the selection is complete, the device sends a connection request to the server.
[0052] 3. The server establishes the connection
[0053] The server receives a connection request from the user's device, establishes the connection through an authentication process, notifies the device that the connection is successful, and the user can begin watching the game.
[0054] Real-time analysis of match data
[0055] 4. The server receives and analyzes the live match feed
[0056] The server receives live video and audio feeds of the match in real time, allowing it to detect events that occur during the match (e.g., goals scored, falls) and input them into the generative model.
[0057] AI-powered voice generation
[0058] 5. The server generates real-time audio using the generative model
[0059] The server uses a generative model to generate cheers and other supportive sounds based on detected events and local audio feeds.
[0060] 6. The server streams the generated audio
[0061] The server encodes the generated audio data and streams it to the communication terminal, which plays the audio in real time.
[0062] Remote cheering bat data transmission
[0063] 7. User uses a remote rooting device
[0064] The user starts cheering by shaking the remote cheering device, which then transmits vibration and motion data to the communication terminal.
[0065] 8. The device sends the support data to the server
[0066] The communication terminal sends the received data from the cheering device to the server, which analyzes the data and uses it to generate cheering sounds.
[0067] Cheering sound generation and output
[0068] 9. The server analyzes the cheering data and generates voice
[0069] The server analyzes the cheering data and generates cheering sounds using a generative model, which are then played from the sound equipment in the stadium.
[0070] 10. Record audio from the viewing area and send it to a communication device
[0071] Microphones installed in the viewing area capture audio and transmit it to a server, which then streams it to the communication device and plays it back on the device.
[0072] Specific examples
[0073] For example, consider the case where User A is watching a soccer match remotely. When User A launches the app and selects a match, the app connects to the server and streams the match video and AI-generated cheering sounds to the device.
[0074] When User A shakes the remote cheering device during a match, the vibration and motion data from the device is sent to the server via the terminal. The server analyzes the data and generates cheering sounds using a generation AI. The generated cheering sounds are played over the stadium's sound system, allowing on-site spectators to hear the cheers.
[0075] Furthermore, microphones installed in the viewing area capture the sounds from the venue in real time and transmit them via a server to User A's device, allowing User A to experience the cheers and excitement of the crowd in real time through earphones.
[0076] The processing flow will be explained below.
[0077] Step 1:
[0078] The server initializes the generated AI model.
[0079] At startup, the server loads the AI model and configures the model's parameters, which includes loading the model's training data.
[0080] The server caches past match data and audience reaction data into the AI model, preparing it for real-time processing.
[0081] Step 2:
[0082] The user launches the app and selects a match
[0083] The user launches the app on their smartphone or tablet.
[0084] Select the game you want to watch from the app's main screen and click the "Start Watching" button.
[0085] Step 3:
[0086] The device sends a connection request to the server
[0087] The user's terminal sends a request to start watching to the server.
[0088] The device connects to the server's API endpoint and sends authentication information.
[0089] Step 4:
[0090] The server establishes the connection and performs authentication
[0091] The server receives the connection request and authenticates the user.
[0092] Once the authentication is complete, the server establishes a session and notifies the terminal of the successful connection.
[0093] Step 5:
[0094] The server receives a live feed of the match
[0095] The server receives live video and audio feeds of the match in real time.
[0096] The server analyzes this data and prepares to detect important events (e.g. goals, falls).
[0097] Step 6:
[0098] The server detects an important event
[0099] The server analyzes the video and audio of the match in real time to detect specific events.
[0100] When a specific event such as a goal or fall is detected, that information is passed on to the generating AI.
[0101] Step 7:
[0102] The server generates the voice using AI
[0103] The server uses a generative AI model to generate cheers and sounds based on the detected events.
[0104] The generated audio data is encoded in real time.
[0105] Step 8:
[0106] The server streams the generated audio to the device
[0107] The server streams the encoded audio data to the user terminal.
[0108] The user terminal plays back the received audio in real time.
[0109] Step 9:
[0110] User uses remote rooting device
[0111] Users cheer by waving the remote cheering device.
[0112] Vibration and motion data from the device is sent to the user's terminal.
[0113] Step 10:
[0114] The device sends the support data to the server.
[0115] The user's terminal transmits the data received from the remote support device to the server.
[0116] The data sent to the server includes vibration data and motion data.
[0117] Step 11:
[0118] The server analyzes the support data
[0119] The server analyzes the received support data in real time.
[0120] Based on the analysis results, the server determines what kind of cheering sound to generate.
[0121] Step 12:
[0122] The server generates cheering sounds using AI
[0123] The server uses a generative AI model to generate cheering sounds, including applause and cheers.
[0124] The generated cheering sound is encoded and transmitted to an audio device.
[0125] Step 13:
[0126] Play cheering sounds from the sound system
[0127] The cheering sound data transmitted from the server is received by the sound device.
[0128] The sound system will then play the received cheering sounds in real time, allowing on-site spectators to feel the cheering from the remote area.
[0129] Step 14:
[0130] Record audio from the viewing location and send it to your device
[0131] The server captures audio from microphones installed at the viewing area and transmits it to the server in real time.
[0132] The server encodes the captured audio for streaming to the user terminal.
[0133] Step 15:
[0134] The device plays audio from the viewing area
[0135] The user terminal receives the streamed audio data and plays it back in real time.
[0136] Users can experience the cheers and atmosphere of the venue in real time.
[0137] Example 1
[0138] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0139] The aim is to solve the problem that it is difficult for remote spectators to experience the sense of presence and cheers of being at the venue in real time, and that cheers from remote cheering devices are not easily conveyed to on-site spectators.
[0140] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0141] In this invention, the server includes a means for recreating the cheers and excitement of the spectator location using a generative model, a means for receiving and analyzing a live feed of the game and detecting events, and a means for encoding audio data and streaming it to a communication terminal, allowing remote spectators to experience realistic cheers in real time and for cheering via cheering devices to be reflected at the venue.
[0142] A "generative model" is an artificial intelligence technology that generates new data and information based on learned algorithms and data.
[0143] "Spectator venue" refers to a local venue or stadium where a sporting event, concert, etc. is held.
[0144] "Cheering and excitement" refers to the vocal and emotional excitement of spectators expressing their excitement and support for a game or event.
[0145] A "communication terminal" is an electronic device that can send and receive data over a network, such as a smartphone, tablet PC, or personal computer.
[0146] A "server" is a high performance computing device for storing, processing, and distributing data over a network.
[0147] An "input signal" is an electrical signal that contains instructions or data from a user or device.
[0148] "Cheering sounds" are sounds and sound effects made by spectators when cheering on a game or event.
[0149] "Sound equipment" means equipment such as speakers and amplifiers for reproducing sound.
[0150] "Recording" is the act of recording sound in digital or analog form.
[0151] "Streaming" refers to the technology of continuously transmitting and playing data in real time over the Internet.
[0152] A "remote support device" is a device that enables support from a remote location, and has the function of transmitting vibration data and motion data mainly to a communication terminal.
[0153] This system uses a generative model to recreate the cheers and excitement of the spectator's location, providing a realistic experience for remote spectators. The system mainly consists of a server, communication terminals, remote cheering devices, and audio equipment.
[0154] The server first initializes the generative AI model by loading a pre-trained model file into memory using a machine learning library such as TensorFlow or PyTorch. It then sets the necessary parameters and prepares the model for operation, ready to generate cheers and other acoustic patterns for the spectators.
[0155] Users launch a dedicated app on their communication device, such as a smartphone or tablet, and select the game they want to watch. Once a game is selected, the device sends the selection information to the server and requests a connection to the server. The server receives the connection request from the user device and performs an authentication process based on the authentication information. If authentication is successful, the session begins and the user can begin remote viewing.
[0156] The server receives live video and audio feeds of the game in real time. It then analyzes the data to detect goals and important events. The results of this analysis are converted into a format that can be used by the generative AI model and fed into it. The resulting audio data is then encoded and streamed to the user's device.
[0157] Users begin cheering by shaking the remote cheering device. This device transmits vibration and motion data to a communication terminal. The communication terminal then transmits this data to a server, which then generates cheering sounds based on that data. The generated cheering sounds are played from the sound equipment in the stadium, allowing on-site spectators to hear the cheers.
[0158] Furthermore, microphones installed in the viewing area capture the sound from the venue in real time and transmit it to a server, which then streams the sound to communication devices, allowing users to experience the cheers and excitement of the venue in real time.
[0159] As a concrete example, consider the case where User A is watching a soccer match remotely. When User A launches the app and selects a match, the app connects to the server, and game footage and cheering sounds generated by AI are streamed to the device. When User A shakes the remote cheering device during the match, the vibration and motion data from the device are sent to the server, and cheering sounds are generated and played from the on-site sound equipment. In addition, audio captured at the venue is sent to the device in real time, allowing User A to enjoy the immersive experience through earphones.
[0160] Examples of prompts for a generative AI model include:
[0161] "In order to generate audio that reproduces the cheers and excitement of the viewing area, we provide the following information:
[0162] 1. The crowd's reaction when a goal is scored
[0163] 2. The crowd booing when the fall happened
[0164] 3. The crowd roaring during halftime
[0165] Based on this, generate realistic cheer sounds in real time.
[0166] This system will enable remote spectators to experience the excitement and cheers of being at the venue in real time, improving the viewing experience.
[0167] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0168] Step 1: Initializing the generative AI model
[0169] The server initializes the generated AI model.
[0170] Input: trained model file, required parameters
[0171] Output: Initialized generative AI model
[0172] What it does: When the server starts up, it uses machine learning libraries like TensorFlow and PyTorch to load a trained generative AI model into memory, reads the model's parameters (e.g., settings for generating voices), and prepares the model for operation, ready to generate cheers and other acoustic patterns for the spectators.
[0173] Step 2: Launch the app and connect
[0174] The user launches the app and selects a match
[0175] Input: Select the match you want to watch
[0176] Output: A connection request is sent
[0177] Specific operation: The user launches the dedicated app on their smartphone or tablet and selects the game they want to watch from a list on the main screen. Once a game is selected, the app sends a connection request to the server with the selected information.
[0178] The server establishes the connection
[0179] Input: Connection request, user credentials
[0180] Output: Connection established notification
[0181] Specific operation: The server receives a connection request from a user terminal and performs an authentication process based on authentication information (e.g., user ID and password). If authentication is successful, the server starts a session and notifies the communication terminal that the connection has been established.
[0182] Step 3: Real-time analysis of match data
[0183] The server receives and analyzes the live match feed
[0184] Input: Live video and audio feed of the match
[0185] Output: Event data (scoring scenes, falls, etc.)
[0186] How it works: The server receives live video and audio feeds from live game broadcasts in real time. It analyzes the input video and audio data to detect goals and important events. The analysis results are then converted into a format that can be used by the generative AI model.
[0187] Step 4: AI-powered voice generation
[0188] The server generates real-time audio using the generative model
[0189] Input: Event data, local audio feed
[0190] Output: Generated audio data
[0191] How it works: The server uses a generative AI model to generate realistic cheering and cheering sounds based on the detected event data and local audio feeds. The generated audio data is then encoded and prepared for transmission to the communication device.
[0192] The server streams the generated audio
[0193] Input: Generated audio data
[0194] Output: Audio stream
[0195] Specific operation: The server streams the encoded audio data to the user's communication device in real time, allowing the user to play back the realistic audio generated on the device.
[0196] Step 5: Send data to your remote rooting device
[0197] User uses remote rooting device
[0198] Input: Vibration data and motion data of the remote rooting device
[0199] Output: Send data to a communication terminal
[0200] Specific operation: A user expresses his / her support by shaking a remote support device (e.g., a Bluetooth-connected support bat). This device transmits vibration data and motion data to a communication terminal.
[0201] The device sends the support data to the server.
[0202] Input: Vibration data and motion data of the remote rooting device
[0203] Output: Send data to the server
[0204] Specific operation: The communication device sends the received vibration and motion data to the server, which receives this data, analyzes it, and uses it to generate cheering sounds.
[0205] Step 6: Generate and output cheer sounds
[0206] The server analyzes the cheering data and generates voice
[0207] Input: Remote rooting device data
[0208] Output: Generated cheer sound
[0209] Specific operation: The server analyzes the received data and generates sounds according to the intensity and frequency of the cheering. The generated cheering sounds are sent to the sound equipment in the stadium and played back.
[0210] Record audio from the viewing location and send it to a communication device
[0211] Input: Audio data captured at the viewing location
[0212] Output: Audio stream to communication device
[0213] How it works: Microphones installed in the viewing area capture audio in real time. The captured audio data is sent to a server and then encoded and transmitted to a communication device. Users can play this audio on their device, allowing them to experience the realism of being at the venue.
[0214] (Application example 1)
[0215] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0216] Conventional remote viewing systems make it difficult to experience the cheers and excitement of the fans in real time, and they lack the interactivity of cheering. Furthermore, there is a lack of a way to efficiently analyze data from remote cheering devices and generate realistic cheering sounds. There is a need to solve these issues and provide a more realistic viewing experience.
[0217] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0218] In this invention, the server includes means for reproducing the cheers and excitement of the spectator location using a generative model, means for transmitting input signals from spectators to the server using a communication device, means for analyzing the input signals to generate cheering sounds and outputting them from a sound device, means for recording audio from the spectator location and transmitting it to the communication device, means for playing the audio on the communication device, means for performing real-time event analysis and generating cheers using a generative model, and means for acquiring vibration data and motion data from the cheering device and transmitting it to the server. This provides remote spectators with the same sense of realism as if they were watching the game in person, allowing users to send their cheers interactively through their devices.
[0219] A "generative model" is a model that uses artificial intelligence technology to generate specific patterns, sounds, images, etc.
[0220] A "spectator location" is a local location where entertainment or competition, such as a sporting event or concert, takes place.
[0221] "Cheers" are the cheers and excited voices that spectators make at the viewing venue in response to a game or performance.
[0222] "Excitement" refers to the atmosphere of excitement and enthusiasm created by the audience.
[0223] A "communication device" is a device for sending and receiving information, including smartphones, tablets, and computers.
[0224] A "spectator" is someone who watches events such as sports or concerts remotely or in person.
[0225] An "input signal" is data or instructions sent by a spectator through a communication device.
[0226] "Cheering sounds" are sounds generated in response to the cheering actions of spectators.
[0227] "Audio device" means a device for reproducing sound, including speakers and headphones.
[0228] "Real-time event analysis" is a process that instantly analyzes the status and behavior of an event currently in progress.
[0229] "Vibration data" is data recorded of the movement of the support device when it vibrates.
[0230] "Motion data" is data that records the movement of the support device.
[0231] A "server" is a computer that manages and provides information over a network.
[0232] The system embodying this invention uses a generative AI model to recreate the cheers and excitement of the spectator's location, providing a sense of realism to remote spectators. The system is primarily composed of a server, communication equipment, a remote cheering device, and an audio device.
[0233] 1. System Program
[0234] The system's program is built around a generative AI model and mainly performs the following processes:
[0235] 1. Initializing the generative AI model
[0236] The server loads the generative AI model and sets the necessary parameters, so the model is ready to generate cheers and other sound patterns for the spectator area.
[0237] 2. Connecting communication devices
[0238] The user starts up the communication device and selects the event they want to watch. The communication device sends a connection request to the server, and the server establishes the connection. If the connection is successful, the user can start watching.
[0239] 3. Event Data Analysis
[0240] The server receives live video and audio feeds of the event in real time, detects and analyzes specific events, and feeds the results into a generative AI model.
[0241] 4. Real-time speech generation
[0242] The server uses a generative AI model to generate cheering and support sounds in real time, and this generated audio data is sent to a communication device and played back in real time.
[0243] 5. Use of remote support devices
[0244] When a user shakes the remote cheering device, the device transmits vibration and motion data to the communication device, which then analyzes the data and generates cheering sounds.
[0245] 6. Data transmission and reception between the server and communication device
[0246] The server plays the generated cheering sounds from the on-site sound equipment and records the audio from the viewing location and transmits it to the communication device, which then plays back the audio, allowing the user to experience the cheers and realism of the venue.
[0247] 2. Hardware and Software Details
[0248] Hardware
[0249] Server: Run the system on a high-performance server (e.g., AWS EC2).
[0250] Communication devices: Common devices such as smartphones and tablets.
[0251] Remote rooting device: An IoT device for capturing user actions.
[0252] Sound device: An audio playback device such as a speaker or headphones.
[0253] software
[0254] Generative AI model: For example, we use OpenAI's GPT-3 model for speech generation.
[0255] Communication and Data Processing: Build WebSocket server and client functions using Python.
[0256] 3. Specific Examples
[0257] For example, if a user wants to watch a live concert remotely, the process would be as follows:
[0258] 1. A user turns on a communication device and selects a live concert. The communication device connects to a server and receives a real-time live feed.
[0259] 2. The server analyzes the live feed and detects specific events (e.g., highlights or excitement).
[0260] 3. The server uses a generative AI model to generate cheers and cheering sounds based on these events.
[0261] 4. When the user shakes the remote cheering device, the motion data is sent to the server via the communication device, and a cheering sound is generated.
[0262] 5. The generated audio is played on the communication device and also played on-site.
[0263] Prompt Sentence Examples
[0264] 1. Initializing the generative AI model
[0265] "Load a generative model and generate cheers and cheers in real time."
[0266] 2. User's live concert viewing
[0267] "Users simply launch the app, select a specific live concert, and watch it. The live feed, including the sounds of cheering and cheering from the venue, is streamed to their smartphone."
[0268] 3. Use a remote support device
[0269] "When a user remotely shakes a cheering device, the vibration and motion data from the device is sent to the server, which generates cheering sounds that can be played by other spectators watching in real time."
[0270] In this way, this invention allows users to enjoy a realistic viewing experience even from a remote location.
[0271] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0272] Step 1: Initializing the generative AI model
[0273] The server loads the generative AI model and sets the necessary parameters, so it is ready to generate cheers and other sound patterns for the spectator area. It takes the model's reference path as input and the generative AI model initialization as output.
[0274] Step 2: Connecting communication devices
[0275] A user starts up a communication device and selects the event they want to watch. The communication device sends a connection request to the server, and the server establishes the connection. The input is the event information selected by the user, and the output is the established connection.
[0276] Step 3: Receive live feeds and parse event data
[0277] The server receives live video and audio feeds of the event in real time. It detects and analyzes specific events (e.g., goals scored). This process has the live feed data as input and the analyzed event data as output.
[0278] Step 4: Real-time speech generation
[0279] The server uses a generative AI model to generate cheers and cheering sounds in real time. This generated audio data is sent to a communication device. The input is event data, and the output is generated audio data.
[0280] Step 5: Encode and transmit the audio data
[0281] The server encodes the generated audio data and streams it to the communication device, with the generated audio data as input and the encoded audio data as output.
[0282] Step 6: Real-time audio playback
[0283] The communication device plays back the received audio data in real time, with the encoded audio data as input and the audio being played back as output.
[0284] Step 7: Send data to your remote rooting device
[0285] The user cheers by shaking the remote cheering device. The device collects vibration and motion data and sends it to the communication device. The motion data is input and sent to the communication device as output.
[0286] Step 8: Sending remote rooting data to the server
[0287] The communication device receives data from the remote support device and transmits it to the server, which has the remote support data as input and transmits the data to the server as output.
[0288] Step 9: Analyzing cheering data and generating cheering sounds
[0289] The server analyzes the remote cheering data and generates cheering sounds using a generative model. The remote cheering data is input, and the generated cheering sounds are output.
[0290] Step 10: Playing the cheering sounds and sending the audio feed
[0291] The server plays the generated cheering sounds from an audio device and records the audio from the viewing location and sends it to a communication device.The generated cheering sounds and on-site audio are input, and the audio is played back and an audio feed is sent as output.
[0292] Step 11: Playback of local audio via communication device
[0293] The communication device plays back the received local audio. It has received audio data as input and plays back the audio as output.
[0294] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0295] This invention is a system that uses a generative model to reproduce the cheers and excitement of the spectator's location, providing a realistic experience for remote spectators. This system is realized by combining a server, communication terminals, remote cheering devices, audio equipment, and an emotion engine that recognizes the user's emotions.
[0296] Program processing and explanation
[0297] Initializing a generative AI model
[0298] 1. The server initializes the generative AI model
[0299] On startup, the server loads the generative AI model and sets the necessary parameters, which includes loading the model's training data.
[0300] The server caches past match data and audience reaction data into the AI model, preparing it for real-time processing.
[0301] Launching the app and connecting
[0302] 2. The user launches the app and selects a match.
[0303] Users launch the app on their smartphone or tablet and select the game they want to watch. Once the selection is complete, the device sends a connection request to the server.
[0304] 3. The server establishes the connection
[0305] The server receives a connection request from the user's device, establishes the connection through an authentication process, notifies the device that the connection is successful, and the user can begin watching the game.
[0306] Real-time analysis of match data
[0307] 4. The server receives and analyzes the live match feed
[0308] The server receives live video and audio feeds of the match in real time, allowing it to detect events that occur during the match (e.g., goals, falls) and input them into the generative model.
[0309] AI-powered voice generation
[0310] 5. The server generates real-time audio using the generative model
[0311] The server uses a generative model to generate cheers and other supportive sounds based on detected events and local audio feeds.
[0312] 6. The server streams the generated audio to the device
[0313] The server encodes the generated audio data and streams it to the communication terminal, which plays the audio in real time.
[0314] Remote cheering bat data transmission
[0315] 7. User uses a remote rooting device
[0316] Users cheer by shaking the remote cheering device, and vibration and motion data from the device is sent to the user's device.
[0317] 8. The device sends the support data to the server
[0318] The user's device sends the data received from the remote support device to the server, including vibration data and motion data.
[0319] Cheering sound generation and output
[0320] 9. The server analyzes the cheering data and generates voice
[0321] The server analyzes the cheering data and generates cheering sounds using a generative model, which are then played from the sound equipment in the stadium.
[0322] 10. Record audio from the viewing location and send it to your device
[0323] Microphones installed in the viewing area capture audio and transmit it to a server, which then streams it to the communication device and plays it back on the device.
[0324] Using Emotion Data with an Emotion Engine
[0325] 11. The emotion engine recognizes emotions on the user device
[0326] The emotion engine recognizes the user's emotions by analyzing their facial expressions and tone of voice.
[0327] The recognized emotion data is sent to the server in real time.
[0328] 12. Analyze emotional data and reflect it in speech generation
[0329] The server analyzes the received emotion data and dynamically changes the intensity and content of the generated voice based on the user's emotion. For example, if the user is excited, it generates louder cheers.
[0330] Specific examples
[0331] For example, consider the case where User A is watching a soccer match remotely. When User A launches the app and selects a match, the app connects to the server and streams the match video and AI-generated cheering sounds to the device.
[0332] When User A shakes the remote cheering device during a match, the vibration and motion data from the device is sent to the server via the terminal. The server analyzes the data and generates cheering sounds using a generation AI. The generated cheering sounds are played over the stadium's sound system, allowing on-site spectators to hear the cheers.
[0333] Furthermore, the emotion engine recognizes User A's emotions, and if User A is in an excited state, for example, louder cheers and cheering sounds are generated. This allows User A to feel an even stronger sense of presence and unity with the local area.
[0334] A microphone installed at the viewing location captures the sounds from the venue in real time and sends them to User A's device via a server. This allows User A to experience the cheers and excitement of the venue in real time through earphones. This system allows remote spectators to feel as if they are actually there, without feeling any physical distance.
[0335] The processing flow will be explained below.
[0336] Step 1:
[0337] The server initializes the generated AI model.
[0338] On startup, the server loads the generative AI model, sets the necessary parameters, and optimizes it using training data, ready to generate game cheers and sounds.
[0339] Step 2:
[0340] The user launches the app and selects a match
[0341] Users launch the app on their smartphone or tablet and select the game they want to watch. When the user clicks the "Start Watching" button, a connection request is sent from the device to the server.
[0342] Step 3:
[0343] The device sends a connection request to the server
[0344] The user's device sends a connection request to the server's API endpoint, which includes authentication information.
[0345] Step 4:
[0346] The server establishes the connection and performs authentication
[0347] The server receives the connection request and verifies the user's authentication information. If authentication is successful, the server establishes a session and notifies the terminal that the connection is successful.
[0348] Step 5:
[0349] The server receives a live feed of the match
[0350] The server receives live video and audio feeds of the match in real time and prepares to analyze the progress of the match.
[0351] Step 6:
[0352] The server analyzes the match data
[0353] The server analyzes the game footage in real time to detect specific events (e.g., goals, falls), and sends the detected event information to the generative AI model.
[0354] Step 7:
[0355] The server generates the voice using the generative AI model
[0356] The server inputs the detected event information into a generative AI model to generate cheering and cheering sounds in real time, and the generated audio data is encoded.
[0357] Step 8:
[0358] The server streams the generated audio to the device
[0359] The server transmits the encoded audio data in streaming format to the user's device, which plays the audio in real time.
[0360] Step 9:
[0361] User uses remote rooting device
[0362] Users can show their support by shaking the remote cheering device, and vibration and motion data from the device are sent to the user's device.
[0363] Step 10:
[0364] The device sends the support data to the server.
[0365] The user's device then sends the received device data, including vibration and motion data, to the server.
[0366] Step 11:
[0367] The server analyzes the support data
[0368] The server analyzes the received cheering device data and issues instructions to the generation AI model based on the results, which then generates the cheering sound.
[0369] Step 12:
[0370] The server generates cheering sounds using AI
[0371] The server generates cheering sounds based on the analysis results, and the generated audio data is encoded and sent to the stadium's sound system.
[0372] Step 13:
[0373] Play cheering sounds from the sound system
[0374] The sound equipment receives the cheering sound data sent from the server and plays it back in real time, allowing on-site spectators to feel the remote cheering.
[0375] Step 14:
[0376] Record audio from the viewing location and send it to your device
[0377] Microphones installed in the viewing area capture audio in real time and transmit it to a server, which encodes the captured audio for streaming to user devices.
[0378] Step 15:
[0379] The device plays audio from the viewing area
[0380] The user's device receives the streamed audio data and plays it back in real time, allowing the user to experience the cheers and atmosphere of the venue in real time.
[0381] Step 16:
[0382] An emotion engine recognizes emotions on the user's device
[0383] The user's device analyzes the user's facial expressions and tone of voice through a camera and microphone, and recognizes emotions using an emotion engine.
[0384] Step 17:
[0385] The device sends emotion data to the server.
[0386] The user's device transmits the recognized emotion data to the server in real time.
[0387] Step 18:
[0388] The server analyzes the emotional data and reflects it in the speech generation.
[0389] The server analyzes the emotion data and dynamically changes the intensity and content of the generated voices based on the user's emotions, for example, generating louder cheers if the user is excited.
[0390] Specific examples
[0391] For example, consider the case where User A is watching a soccer match remotely. When User A launches the app and selects a match, the app connects to the server and streams the match video and AI-generated cheering sounds to the device.
[0392] When User A shakes the remote cheering device during a match, the vibration and motion data from the device is sent to the server via the terminal. The server analyzes the data and generates cheering sounds using a generation AI. The generated cheering sounds are played over the stadium's sound system, allowing on-site spectators to hear the cheers.
[0393] Furthermore, the emotion engine recognizes User A's emotions, and if User A is in an excited state, for example, louder cheers and cheering sounds are generated. This allows User A to feel an even stronger sense of presence and unity with the local area.
[0394] A microphone installed at the viewing location captures the sounds from the venue in real time and sends them to User A's device via a server. This allows User A to experience the cheers and excitement of the venue in real time through earphones. This system allows remote spectators to feel as if they are actually there, without feeling any physical distance.
[0395] Example 2
[0396] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0397] Current remote viewing systems have difficulty fully reproducing the cheers and excitement of the stadium, meaning that remote spectators cannot fully experience the sense of presence and unity of being at the stadium. Furthermore, there is a lack of technology to effectively generate cheering sounds that utilize spectators' emotions and input signals from remote cheering devices.
[0398] The specific processing by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes a means for initializing the generative model and setting necessary parameters, a means for receiving live video and audio feeds of the game and detecting specific events, a means for transmitting vibration and motion data from the remote cheering device to the server via the communication terminal, a means for using an emotion engine that recognizes the user's emotions, and a means for transmitting the user's emotion data from the communication terminal to the server and reflecting it in sound generation. This allows the cheers and excitement of the viewing location to be reproduced in real time, allowing remote spectators to feel the presence of the venue. Furthermore, by utilizing the spectator's input signals and emotion data, more personalized cheering sounds can be generated.
[0399] A "generative model" is an algorithm that uses machine learning and artificial intelligence techniques to generate new data or patterns based on specific input data.
[0400] "Communication terminal" refers to a device such as a smartphone, tablet, or PC that can send and receive data via the Internet or other networks.
[0401] "Input signals" refer to data and commands sent from spectators or devices to the server, including vibration data, motion data from remote cheering devices, and user emotional data.
[0402] A "server" is a computer system that provides specific services or functions over a network, including data processing, storage, and management.
[0403] "Remote cheering device" refers to a device used by spectators to cheer from a remote location. Specifically, it includes stick-shaped or handheld devices that transmit vibration and motion data to a server.
[0404] An "emotion engine" refers to software or algorithms that analyze a user's facial expressions, tone of voice, physical movements, etc. to recognize emotions, and then generate appropriate responses in real time based on that data.
[0405] "Sound equipment" refers to speakers and sound systems that output the generated cheering sounds and cheers and reproduce them audibly.
[0406] "Live Game Video and Audio Feed" means a data stream that transmits real-time video and audio from the location where the Game is being played.
[0407] "User Emotion Data" refers to data that indicates the user's emotional state analyzed by the emotion engine. This data is used in the speech generation process using the generative model.
[0408] The present invention is a system that uses a generative model to reproduce the cheers and excitement of the spectator's location, providing a realistic experience for remote spectators. This system is realized by a server, communication terminals, remote cheering devices, audio equipment, and an emotion engine.
[0409] Hardware and software used
[0410] The system includes the following hardware and software:
[0411] 1. Server:
[0412] Computer system for executing generative AI models
[0413] Storage for caching past match data and audience reaction data
[0414] Network equipment for receiving and analyzing live video and audio feeds
[0415] 2. Communication terminal:
[0416] Smartphones, tablets, and computers used by spectators
[0417] Connectivity for receiving and sending data from remote rooted devices to the server
[0418] Speakers and earphones for playing audio data from the server
[0419] 3. Remote rooting devices:
[0420] A handheld device that spectators wave to cheer on the game.
[0421] A function that generates vibration and motion data and sends it to a communication device
[0422] 4. Sound equipment:
[0423] A speaker system that outputs cheering sounds and cheers within the stadium
[0424] 5. Emotion Engine:
[0425] Software that recognizes emotions by analyzing the user's facial expressions and tone of voice
[0426] A function to send the recognized emotion data to the server
[0427] Data processing and calculation
[0428] The main processes of this system are:
[0429] Initialize the generative AI model:
[0430] The server loads the generative AI model and sets the necessary parameters, and caches past match data and audience reaction data to prepare for real-time processing.
[0431] Receive live video and audio feeds:
[0432] The server receives live video and audio feeds of the match in real time and detects specific events.
[0433] Voice generation:
[0434] The server uses a generative model to generate cheers and sounds based on the detected events and streams them to the communication device.
[0435] Processing input data from a remote rooting device:
[0436] Vibration and motion data from the remote cheering device is transmitted to a server via a communication terminal, and cheering sounds are generated.
[0437] Use of emotion data:
[0438] The system analyzes the user's facial expressions and tone of voice, and sends the emotional data recognized by the emotion engine to the server, which then dynamically changes the intensity and content of the generated voice based on this data.
[0439] Specific examples
[0440] For example, consider the case where User A is watching a soccer match remotely. When User A launches the app and selects a match, the app connects to the server and streams the match video and AI-generated cheering sounds to the device.
[0441] When User A shakes the remote cheering device during a match, the vibration and motion data from the device is sent to the server via the terminal. The server analyzes the data and generates cheering sounds using a generation AI. The generated cheering sounds are played through the sound system at the stadium, so that on-site spectators can also hear the cheers.
[0442] Furthermore, the emotion engine recognizes User A's emotions, and if User A is in an excited state, for example, louder cheers and cheering sounds are generated. This allows User A to feel an even stronger sense of presence and unity with the local area.
[0443] Prompt Sentence Examples
[0444] An example of a prompt to be input to a generative AI model is written as follows:
[0445] "When a goal is scored during a game, how do we generate audio that synchronizes with the cheers of the crowd?"
[0446] "How to change the intensity of the cheering sounds generated when the user is recognized as excited"
[0447] "What kind of cheering sound should be generated based on the vibration data received from the remote cheering device?"
[0448] This allows the entire system to work together, providing spectators with a realistic experience in real time.
[0449] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0450] The flow of this system's program processing
[0451] Step 1:
[0452] Initializing a generative AI model
[0453] Server loads the model: When the server starts up, it loads the generative AI model from disk and loads it into memory, making the model immediately available for use.
[0454] Server sets parameters: The server sets the hyperparameters required for the model to operate, including the learning rate and batch size. The set parameters are important because they directly affect the model's performance.
[0455] Cache past data: The server caches past match data and audience reaction data, which speeds up real-time processing. Cached data improves response time when an event occurs.
[0456] Input: Generative AI model and configuration parameters from disk
[0457] Output: A generative AI model that can be run in memory
[0458] Step 2:
[0459] Launch the app and select a match
[0460] User launches app: The spectator launches the dedicated app on their smartphone or tablet. The app displays the home screen and shows a list of available matches.
[0461] User selects a game: The user selects the game they want to watch from the list. This selection information is sent from the communication device to the server, which then prepares the corresponding live feed.
[0462] Server accepts connection request: The server accepts the user's connection request and performs authentication. If authentication is successful, the server sends a connection establishment notification to the user's terminal.
[0463] Input: User's match selection information
[0464] Output: Connection establishment and match information on the server side
[0465] Step 3:
[0466] Receiving and analyzing live feeds
[0467] Server receives live video and audio: The server receives live video and audio feeds from the stadium in real time. These feeds are transmitted via a communications network.
[0468] Server detects specific events: The server analyzes and detects specific events that occur during the match (e.g., goals, falls). This data is fed into a generative AI model and used as the basis for generating cheering sounds and cheers.
[0469] Input: Live video and audio feed
[0470] Output: Detected event data
[0471] Step 4:
[0472] Real-time voice generation
[0473] Server generates sounds: Cheering and cheering sounds are generated by a generative AI model based on detected event data. The generated sounds change dynamically depending on the situation of the match.
[0474] The server encodes the generated audio: the encoded audio data is prepared for streaming, providing high-quality audio to the user in real time.
[0475] Input: Detected event data
[0476] Output: Generated cheering and cheering sound data
[0477] Step 5:
[0478] Streaming generated audio
[0479] The server streams audio data to the device: The server sends encoded audio data in real time to the user's communication device, which receives the data and plays it through its built-in speaker or earphones.
[0480] Input: Generated audio data
[0481] Output: Streamed audio data
[0482] Step 6:
[0483] Sending data from a remote rooting device
[0484] The user operates the device: The user performs cheering actions using the remote cheering device. The device generates vibration and motion data and transmits it to the communication terminal.
[0485] The terminal transmits the data to the server: The communication terminal transmits the data received from the device to the server, which analyzes the data and generates additional cheering sounds.
[0486] Input: Vibration and motion data from the remote rooting device
[0487] Output: Data sent to the server
[0488] Step 7:
[0489] Cheering sound generation
[0490] The server analyzes the data: vibration and motion data is analyzed, and a generative AI model is used to generate cheering sounds, which are then played over the on-site sound system.
[0491] Input: Vibration and motion data
[0492] Output: Generated cheer sound
[0493] Step 8:
[0494] Audio recording from the viewing area
[0495] The server captures the audio: Microphones installed in the viewing area capture the audio from the venue and send it to the server.
[0496] The server streams the audio to the device: The captured audio is sent to the communication device and played back in real time. This process is important for conveying the cheers and realism of the actual event to the user.
[0497] Input: Audio data from microphones installed in the viewing area
[0498] Output: Audio data streamed to the user's device
[0499] Step 9:
[0500] Analysis and use of emotional data
[0501] User device collects emotional data: The emotion engine analyzes the user's facial expressions and tone of voice to collect emotional data.
[0502] The device sends emotional data to the server: The analyzed emotional data is sent to the server in real time. The server analyzes this data and dynamically changes the intensity and content of the generated voice.
[0503] Input: User facial expressions and tone of voice
[0504] Output: Emotion data sent to the server.
[0505] (Application example 2)
[0506] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0507] Current remote viewing systems have the problem of being unable to fully reproduce the sense of presence and unity that users feel at a real viewing location. In particular, it is difficult to experience the cheers and sounds of the fans at the venue in real time, and they are unable to generate cheering sounds that reflect the emotions of the remote spectators, limiting the viewing experience. Furthermore, they lack the functionality to generate cheering sounds using data from remote cheering devices. Technology that solves these issues is needed.
[0508] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0509] In this invention, the server includes means for reproducing the cheers and excitement of the spectator location using a generative model, means for transmitting input signals from spectators using a communication terminal to the server, means for analyzing the input signals, generating cheering sounds, and outputting them from a sound device, means for recording audio from the spectator location and transmitting the audio to the communication terminal, means for playing the audio on the communication terminal, means for recognizing the emotions of spectators using an emotion engine and transmitting the emotion data to the server in real time, and means for dynamically changing the intensity and content of the cheering sounds based on the emotion data. This allows remote spectators to experience a sense of realism similar to that of being at the venue, and cheering sounds that correspond to their emotions can be generated.
[0510] A "generative model" is a machine learning model that uses artificial intelligence to generate new data.
[0511] A "spectator venue" is the location where a sport or event actually takes place and where spectators physically gather.
[0512] "Cheers and enthusiasm" refers to the cheers and support emitted by spectators at the viewing venue, and are the sounds and atmosphere that indicate the excitement of the event.
[0513] A "communication terminal" is an electronic device used to send and receive data over the Internet, including smartphones and tablets.
[0514] "Input signals from spectators" refer to data that spectators send to the server via their communication terminals, and are signals that reflect the actions and emotions of the spectators.
[0515] A "server" is a computer system that processes and manages data over a network.
[0516] "Cheering sounds" are sounds generated based on input signals from spectators and generative models, and include sounds of cheering and cheering.
[0517] "Audio device" refers to equipment for outputting sound, including speakers and earphones.
[0518] An "emotion engine" is software that analyzes and recognizes the user's emotions from their facial expressions and movements.
[0519] "Real time" is a time concept that refers to instantaneous reaction and processing without delay.
[0520] A "remote cheering device" is a device that spectators use to cheer from home or another location, and has the ability to transmit vibration and motion data to a server.
[0521] The system for implementing this invention includes a generative AI model, an emotion engine, a communication terminal, a server, and an audio device. The operation of the entire system is as follows.
[0522] First, the user launches the remote viewing application on their communication device (smartphone or tablet). Within the application, the user selects the game or event they want to watch. The communication device then sends the selection information to the server, which then receives it.
[0523] The server receives live game feeds (video and audio) and inputs them into a generative AI model to generate cheering and support sounds for the spectators in real time. The generated audio is encoded, streamed to a communication device, and played back to the user in real time. The generative AI model then detects specific events (e.g., a goal or a foul) and dynamically generates audio based on those events.
[0524] Furthermore, the communication terminal is equipped with an emotion engine that recognizes the user's emotions. The emotion engine analyzes the user's facial expressions and tone of voice and transmits the recognized emotion data to the server in real time. The server analyzes this emotion data and dynamically changes the intensity and content of the cheering sounds according to the user's emotions. For example, if the user is excited, a louder cheer will be generated.
[0525] Users can also use a remote cheering device. Vibration and motion data from the remote cheering device is sent to a communication terminal, which then transmits the data to a server. The server analyzes the data and generates cheering sounds using a generative AI model, which are then played on the on-site sound equipment.
[0526] As a concrete example, consider the case where a user is watching a soccer match remotely. When the user launches the application and selects a match, live video and cheering sounds generated by a generative AI are streamed to the communication terminal. When the user shakes the remote cheering device during the match, vibration and motion data from the device are sent to the server via the communication terminal. The server analyzes the data and generates cheering sounds using a generative AI model. These cheering sounds are then played over the sound equipment at the viewing location.
[0527] An example of this prompt would be:
[0528] "A user is watching a soccer match remotely. The emotion engine on the user's smartphone recognizes the user's facial expressions and tone of voice, and the generation AI generates the cheering sounds of the local crowd in real time accordingly. As the user's excitement increases, the generation AI generates louder cheers and cheering sounds, providing a realistic spectator experience."
[0529] As described above, using this system allows remote spectators to experience the same sense of realism as if they were in person, and cheering sounds can be generated to suit their emotions, greatly improving the sense of realism and unity of remote spectators.
[0530] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0531] Step 1:
[0532] The server initializes the generated AI model.
[0533] Input: Past match data and audience reaction data
[0534] Output: Initialized generative AI model
[0535] At startup, the server loads the generative AI model, sets the necessary parameters, reads the model's training data, and prepares it for real-time processing.
[0536] Step 2:
[0537] The user starts the application on the communication terminal and selects the game they want to watch.
[0538] Input: User's match selection information
[0539] Output: Match selection information sent to server
[0540] The user launches the application and selects the game they want to watch. The selection information is sent to the server, which receives it.
[0541] Step 3:
[0542] The server receives a live feed of the game and uses a generative AI model to generate the cheers and support sounds of the spectators.
[0543] Input: Live video and audio feed
[0544] Output: Generated cheers and cheering sounds
[0545] The server receives live video and audio feeds of the game, feeds them into a generative AI model, and generates cheers and cheering sounds in real time, which are then encoded and streamed to communication devices.
[0546] Step 4:
[0547] The communication terminal reproduces the generated cheers and cheering sounds in real time.
[0548] Input: Audio data streamed from the server
[0549] Output: Real-time cheers and cheering sounds played
[0550] The communication terminal decodes the received audio data and plays it back in real time, allowing the user to experience realistic audio.
[0551] Step 5:
[0552] The emotion engine recognizes the user's emotions and transmits the emotion data to the server in real time.
[0553] Input: User facial expressions and tone of voice
[0554] Output: Emotion data sent to the server
[0555] The emotion engine analyzes the user's facial expressions and tone of voice on the communication device to recognize their emotions, and the recognized emotion data is sent to the server.
[0556] Step 6:
[0557] The server analyzes the emotional data and uses a generative AI model to dynamically change the intensity and content of the cheering sounds.
[0558] Input: User emotion data
[0559] Output: Cheering sounds with dynamically changing intensity and content
[0560] The server analyzes the received emotional data and dynamically changes the intensity and content of the cheering sounds based on the user's emotions using a generative AI model. For example, if the user is excited, a louder cheer will be generated.
[0561] Step 7:
[0562] The user uses the remote support device and transmits the vibration data and motion data to the communication terminal.
[0563] Input: Vibration and motion data from the remote rooting device
[0564] Output: Data sent to the communication device
[0565] Users cheer by shaking the remote cheering device, and vibration and motion data from the device are sent to the communication terminal.
[0566] Step 8:
[0567] The communication terminal transmits the data received from the remote support device to the server.
[0568] Input: Vibration and motion data from the remote rooting device
[0569] Output: Device data sent to the server
[0570] The communication terminal transmits the received vibration data and motion data to the server.
[0571] Step 9:
[0572] The server analyzes the data from the remote cheering device and generates cheering sounds using a generative AI model.
[0573] Input: Vibration and motion data
[0574] Output: Generated cheer sound
[0575] The server analyzes the data from the remote cheering device and generates cheering sounds using a generative AI model. The generated cheering sounds are then played over the sound equipment at the viewing area.
[0576] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0577] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0578] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0579] [Second embodiment]
[0580] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0581] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0582] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0583] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0584] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0585] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0586] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0587] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0588] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0589] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0590] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0591] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0592] This system uses a generative model to recreate the cheers and excitement of the spectator's location, providing a realistic experience for remote spectators. The system is primarily composed of a server, communication terminals, remote cheering devices, and audio equipment.
[0593] Program processing and explanation
[0594] Initializing a generative AI model
[0595] 1. The server initializes the generative AI model
[0596] On startup, the server loads the generative AI model and sets the necessary parameters, so the model is ready to generate cheers and other acoustic patterns for the spectator area.
[0597] Launching the app and connecting
[0598] 2. The user launches the app and selects a match.
[0599] Users launch the app on their smartphone or other communication device and select the game they want to watch. Once the selection is complete, the device sends a connection request to the server.
[0600] 3. The server establishes the connection
[0601] The server receives a connection request from the user's device, establishes the connection through an authentication process, notifies the device that the connection is successful, and the user can begin watching the game.
[0602] Real-time analysis of match data
[0603] 4. The server receives and analyzes the live match feed
[0604] The server receives live video and audio feeds of the match in real time, allowing it to detect events that occur during the match (e.g., goals scored, falls) and input them into the generative model.
[0605] AI-powered voice generation
[0606] 5. The server generates real-time audio using the generative model
[0607] The server uses a generative model to generate cheers and other supportive sounds based on detected events and local audio feeds.
[0608] 6. The server streams the generated audio
[0609] The server encodes the generated audio data and streams it to the communication terminal, which plays the audio in real time.
[0610] Remote cheering bat data transmission
[0611] 7. User uses a remote rooting device
[0612] The user starts cheering by shaking the remote cheering device, which then transmits vibration and motion data to the communication terminal.
[0613] 8. The device sends the support data to the server
[0614] The communication terminal sends the received data from the cheering device to the server, which analyzes the data and uses it to generate cheering sounds.
[0615] Cheering sound generation and output
[0616] 9. The server analyzes the cheering data and generates voice
[0617] The server analyzes the cheering data and generates cheering sounds using a generative model, which are then played from the sound equipment in the stadium.
[0618] 10. Record audio from the viewing area and send it to a communication device
[0619] Microphones installed in the viewing area capture audio and transmit it to a server, which then streams it to the communication device and plays it back on the device.
[0620] Specific examples
[0621] For example, consider the case where User A is watching a soccer match remotely. When User A launches the app and selects a match, the app connects to the server and streams the match video and AI-generated cheering sounds to the device.
[0622] When User A shakes the remote cheering device during a match, the vibration and motion data from the device is sent to the server via the terminal. The server analyzes the data and generates cheering sounds using a generation AI. The generated cheering sounds are played over the stadium's sound system, allowing on-site spectators to hear the cheers.
[0623] Furthermore, microphones installed in the viewing area capture the sounds from the venue in real time and transmit them via a server to User A's device, allowing User A to experience the cheers and excitement of the crowd in real time through earphones.
[0624] The processing flow will be explained below.
[0625] Step 1:
[0626] The server initializes the generated AI model.
[0627] At startup, the server loads the AI model and configures the model's parameters, which includes loading the model's training data.
[0628] The server caches past match data and audience reaction data into the AI model, preparing it for real-time processing.
[0629] Step 2:
[0630] The user launches the app and selects a match
[0631] The user launches the app on their smartphone or tablet.
[0632] Select the game you want to watch from the app's main screen and click the "Start Watching" button.
[0633] Step 3:
[0634] The device sends a connection request to the server
[0635] The user's terminal sends a request to start watching to the server.
[0636] The device connects to the server's API endpoint and sends authentication information.
[0637] Step 4:
[0638] The server establishes the connection and performs authentication
[0639] The server receives the connection request and authenticates the user.
[0640] Once the authentication is complete, the server establishes a session and notifies the terminal of the successful connection.
[0641] Step 5:
[0642] The server receives a live feed of the match
[0643] The server receives live video and audio feeds of the match in real time.
[0644] The server analyzes this data and prepares to detect important events (e.g. goals, falls).
[0645] Step 6:
[0646] The server detects an important event
[0647] The server analyzes the video and audio of the match in real time to detect specific events.
[0648] When a specific event such as a goal or fall is detected, that information is passed on to the generating AI.
[0649] Step 7:
[0650] The server generates the voice using AI
[0651] The server uses a generative AI model to generate cheers and sounds based on the detected events.
[0652] The generated audio data is encoded in real time.
[0653] Step 8:
[0654] The server streams the generated audio to the device
[0655] The server streams the encoded audio data to the user terminal.
[0656] The user terminal plays back the received audio in real time.
[0657] Step 9:
[0658] User uses remote rooting device
[0659] Users cheer by waving the remote cheering device.
[0660] Vibration and motion data from the device is sent to the user's terminal.
[0661] Step 10:
[0662] The device sends the support data to the server.
[0663] The user's terminal transmits the data received from the remote support device to the server.
[0664] The data sent to the server includes vibration data and motion data.
[0665] Step 11:
[0666] The server analyzes the support data
[0667] The server analyzes the received support data in real time.
[0668] Based on the analysis results, the server determines what kind of cheering sound to generate.
[0669] Step 12:
[0670] The server generates cheering sounds using AI
[0671] The server uses a generative AI model to generate cheering sounds, including applause and cheers.
[0672] The generated cheering sound is encoded and transmitted to an audio device.
[0673] Step 13:
[0674] Play cheering sounds from the sound system
[0675] The cheering sound data transmitted from the server is received by the sound device.
[0676] The sound system will then play the received cheering sounds in real time, allowing on-site spectators to feel the cheering from the remote area.
[0677] Step 14:
[0678] Record audio from the viewing location and send it to your device
[0679] The server captures audio from microphones installed at the viewing area and transmits it to the server in real time.
[0680] The server encodes the captured audio for streaming to the user terminal.
[0681] Step 15:
[0682] The device plays audio from the viewing area
[0683] The user terminal receives the streamed audio data and plays it back in real time.
[0684] Users can experience the cheers and atmosphere of the venue in real time.
[0685] Example 1
[0686] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0687] The aim is to solve the problem that it is difficult for remote spectators to experience the sense of presence and cheers of being at the venue in real time, and that cheers from remote cheering devices are not easily conveyed to on-site spectators.
[0688] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0689] In this invention, the server includes a means for recreating the cheers and excitement of the spectator location using a generative model, a means for receiving and analyzing a live feed of the game and detecting events, and a means for encoding audio data and streaming it to a communication terminal, allowing remote spectators to experience realistic cheers in real time and for cheering via cheering devices to be reflected at the venue.
[0690] A "generative model" is an artificial intelligence technology that generates new data and information based on learned algorithms and data.
[0691] "Spectator venue" refers to a local venue or stadium where a sporting event, concert, etc. is held.
[0692] "Cheering and excitement" refers to the vocal and emotional excitement of spectators expressing their excitement and support for a game or event.
[0693] A "communication terminal" is an electronic device that can send and receive data over a network, such as a smartphone, tablet PC, or personal computer.
[0694] A "server" is a high performance computing device for storing, processing, and distributing data over a network.
[0695] An "input signal" is an electrical signal that contains instructions or data from a user or device.
[0696] "Cheering sounds" are sounds and sound effects made by spectators when cheering on a game or event.
[0697] "Sound equipment" means equipment such as speakers and amplifiers for reproducing sound.
[0698] "Recording" is the act of recording sound in digital or analog form.
[0699] "Streaming" refers to the technology of continuously transmitting and playing data in real time over the Internet.
[0700] A "remote support device" is a device that enables support from a remote location, and has the function of transmitting vibration data and motion data mainly to a communication terminal.
[0701] This system uses a generative model to recreate the cheers and excitement of the spectator's location, providing a realistic experience for remote spectators. The system mainly consists of a server, communication terminals, remote cheering devices, and audio equipment.
[0702] The server first initializes the generative AI model by loading a pre-trained model file into memory using a machine learning library such as TensorFlow or PyTorch. It then sets the necessary parameters and prepares the model for operation, ready to generate cheers and other acoustic patterns for the spectators.
[0703] Users launch a dedicated app on their communication device, such as a smartphone or tablet, and select the game they want to watch. Once a game is selected, the device sends the selection information to the server and requests a connection to the server. The server receives the connection request from the user device and performs an authentication process based on the authentication information. If authentication is successful, the session begins and the user can begin remote viewing.
[0704] The server receives live video and audio feeds of the game in real time. It then analyzes the data to detect goals and important events. The results of this analysis are converted into a format that can be used by the generative AI model and fed into it. The resulting audio data is then encoded and streamed to the user's device.
[0705] Users begin cheering by shaking the remote cheering device. This device transmits vibration and motion data to a communication terminal. The communication terminal then transmits this data to a server, which then generates cheering sounds based on that data. The generated cheering sounds are played from the sound equipment in the stadium, allowing on-site spectators to hear the cheers.
[0706] Furthermore, microphones installed in the viewing area capture the sound from the venue in real time and transmit it to a server, which then streams the sound to communication devices, allowing users to experience the cheers and excitement of the venue in real time.
[0707] As a concrete example, consider the case where User A is watching a soccer match remotely. When User A launches the app and selects a match, the app connects to the server, and game footage and cheering sounds generated by AI are streamed to the device. When User A shakes the remote cheering device during the match, the vibration and motion data from the device are sent to the server, and cheering sounds are generated and played from the on-site sound equipment. In addition, audio captured at the venue is sent to the device in real time, allowing User A to enjoy the immersive experience through earphones.
[0708] Examples of prompts for a generative AI model include:
[0709] "In order to generate audio that reproduces the cheers and excitement of the viewing area, we provide the following information:
[0710] 1. The crowd's reaction when a goal is scored
[0711] 2. The crowd booing when the fall happened
[0712] 3. The crowd roaring during halftime
[0713] Based on this, generate realistic cheer sounds in real time.
[0714] This system will enable remote spectators to experience the excitement and cheers of being at the venue in real time, improving the viewing experience.
[0715] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0716] Step 1: Initializing the generative AI model
[0717] The server initializes the generated AI model.
[0718] Input: trained model file, required parameters
[0719] Output: Initialized generative AI model
[0720] What it does: When the server starts up, it uses machine learning libraries like TensorFlow and PyTorch to load a trained generative AI model into memory, reads the model's parameters (e.g., settings for generating voices), and prepares the model for operation, ready to generate cheers and other acoustic patterns for the spectators.
[0721] Step 2: Launch the app and connect
[0722] The user launches the app and selects a match
[0723] Input: Select the match you want to watch
[0724] Output: A connection request is sent
[0725] Specific operation: The user launches the dedicated app on their smartphone or tablet and selects the game they want to watch from a list on the main screen. Once a game is selected, the app sends a connection request to the server with the selected information.
[0726] The server establishes the connection
[0727] Input: Connection request, user credentials
[0728] Output: Connection established notification
[0729] Specific operation: The server receives a connection request from a user terminal and performs an authentication process based on authentication information (e.g., user ID and password). If authentication is successful, the server starts a session and notifies the communication terminal that the connection has been established.
[0730] Step 3: Real-time analysis of match data
[0731] The server receives and analyzes the live match feed
[0732] Input: Live video and audio feed of the match
[0733] Output: Event data (scoring scenes, falls, etc.)
[0734] How it works: The server receives live video and audio feeds from live game broadcasts in real time. It analyzes the input video and audio data to detect goals and important events. The analysis results are then converted into a format that can be used by the generative AI model.
[0735] Step 4: AI-powered voice generation
[0736] The server generates real-time audio using the generative model
[0737] Input: Event data, local audio feed
[0738] Output: Generated audio data
[0739] How it works: The server uses a generative AI model to generate realistic cheering and cheering sounds based on the detected event data and local audio feeds. The generated audio data is then encoded and prepared for transmission to the communication device.
[0740] The server streams the generated audio
[0741] Input: Generated audio data
[0742] Output: Audio stream
[0743] Specific operation: The server streams the encoded audio data to the user's communication device in real time, allowing the user to play back the realistic audio generated on the device.
[0744] Step 5: Send data to your remote rooting device
[0745] User uses remote rooting device
[0746] Input: Vibration data and motion data of the remote rooting device
[0747] Output: Send data to a communication terminal
[0748] Specific operation: A user expresses his / her support by shaking a remote support device (e.g., a Bluetooth-connected support bat). This device transmits vibration data and motion data to a communication terminal.
[0749] The device sends the support data to the server.
[0750] Input: Vibration data and motion data of the remote rooting device
[0751] Output: Send data to the server
[0752] Specific operation: The communication device sends the received vibration and motion data to the server, which receives this data, analyzes it, and uses it to generate cheering sounds.
[0753] Step 6: Generate and output cheer sounds
[0754] The server analyzes the cheering data and generates voice
[0755] Input: Remote rooting device data
[0756] Output: Generated cheer sound
[0757] Specific operation: The server analyzes the received data and generates sounds according to the intensity and frequency of the cheering. The generated cheering sounds are sent to the sound equipment in the stadium and played back.
[0758] Record audio from the viewing location and send it to a communication device
[0759] Input: Audio data captured at the viewing location
[0760] Output: Audio stream to communication device
[0761] How it works: Microphones installed in the viewing area capture audio in real time. The captured audio data is sent to a server and then encoded and transmitted to a communication device. Users can play this audio on their device, allowing them to experience the realism of being at the venue.
[0762] (Application example 1)
[0763] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0764] Conventional remote viewing systems make it difficult to experience the cheers and excitement of the fans in real time, and they lack the interactivity of cheering. Furthermore, there is a lack of a way to efficiently analyze data from remote cheering devices and generate realistic cheering sounds. There is a need to solve these issues and provide a more realistic viewing experience.
[0765] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0766] In this invention, the server includes means for reproducing the cheers and excitement of the spectator location using a generative model, means for transmitting input signals from spectators to the server using a communication device, means for analyzing the input signals to generate cheering sounds and outputting them from a sound device, means for recording audio from the spectator location and transmitting it to the communication device, means for playing the audio on the communication device, means for performing real-time event analysis and generating cheers using a generative model, and means for acquiring vibration data and motion data from the cheering device and transmitting it to the server. This provides remote spectators with the same sense of realism as if they were watching the game in person, allowing users to send their cheers interactively through their devices.
[0767] A "generative model" is a model that uses artificial intelligence technology to generate specific patterns, sounds, images, etc.
[0768] A "spectator location" is a local location where entertainment or competition, such as a sporting event or concert, takes place.
[0769] "Cheers" are the cheers and excited voices that spectators make at the viewing venue in response to a game or performance.
[0770] "Excitement" refers to the atmosphere of excitement and enthusiasm created by the audience.
[0771] A "communication device" is a device for sending and receiving information, including smartphones, tablets, and computers.
[0772] A "spectator" is someone who watches events such as sports or concerts remotely or in person.
[0773] An "input signal" is data or instructions sent by a spectator through a communication device.
[0774] "Cheering sounds" are sounds generated in response to the cheering actions of spectators.
[0775] "Audio device" means a device for reproducing sound, including speakers and headphones.
[0776] "Real-time event analysis" is a process that instantly analyzes the status and behavior of an event currently in progress.
[0777] "Vibration data" is data recorded of the movement of the support device when it vibrates.
[0778] "Motion data" is data that records the movement of the support device.
[0779] A "server" is a computer that manages and provides information over a network.
[0780] The system embodying this invention uses a generative AI model to recreate the cheers and excitement of the spectator's location, providing a sense of realism to remote spectators. The system is primarily composed of a server, communication equipment, a remote cheering device, and an audio device.
[0781] 1. System Program
[0782] The system's program is built around a generative AI model and mainly performs the following processes:
[0783] 1. Initializing the generative AI model
[0784] The server loads the generative AI model and sets the necessary parameters, so the model is ready to generate cheers and other sound patterns for the spectator area.
[0785] 2. Connecting communication devices
[0786] The user starts up the communication device and selects the event they want to watch. The communication device sends a connection request to the server, and the server establishes the connection. If the connection is successful, the user can start watching.
[0787] 3. Event Data Analysis
[0788] The server receives live video and audio feeds of the event in real time, detects and analyzes specific events, and feeds the results into a generative AI model.
[0789] 4. Real-time speech generation
[0790] The server uses a generative AI model to generate cheers and cheering sounds in real time, and this generated audio data is sent to a communication device and played back in real time.
[0791] 5. Use of remote support devices
[0792] When a user shakes the remote cheering device, the device transmits vibration and motion data to the communication device, which then analyzes the data and generates cheering sounds.
[0793] 6. Data transmission and reception between the server and communication device
[0794] The server plays the generated cheering sounds from the on-site sound equipment and records the audio from the viewing location and transmits it to the communication device, which then plays back the audio, allowing the user to experience the cheers and realism of the venue.
[0795] 2. Hardware and Software Details
[0796] Hardware
[0797] Server: Run the system on a high-performance server (e.g., AWS EC2).
[0798] Communication devices: Common devices such as smartphones and tablets.
[0799] Remote rooting device: An IoT device for capturing user actions.
[0800] Sound device: An audio playback device such as a speaker or headphones.
[0801] software
[0802] Generative AI model: For example, we use OpenAI's GPT-3 model for speech generation.
[0803] Communication and Data Processing: Build WebSocket server and client functions using Python.
[0804] 3. Specific Examples
[0805] For example, if a user wants to watch a live concert remotely, the process would be as follows:
[0806] 1. A user turns on a communication device and selects a live concert. The communication device connects to a server and receives a real-time live feed.
[0807] 2. The server analyzes the live feed and detects specific events (e.g., highlights or excitement).
[0808] 3. The server uses a generative AI model to generate cheers and cheering sounds based on these events.
[0809] 4. When the user shakes the remote cheering device, the motion data is sent to the server via the communication device, and a cheering sound is generated.
[0810] 5. The generated audio is played on the communication device and also played on-site.
[0811] Prompt Sentence Examples
[0812] 1. Initializing the generative AI model
[0813] "Load a generative model and generate cheers and cheers in real time."
[0814] 2. User's live concert viewing
[0815] "Users simply launch the app, select a specific live concert, and watch it. The live feed, including the sounds of cheering and cheering from the venue, is streamed to their smartphone."
[0816] 3. Use a remote support device
[0817] "When a user remotely shakes a cheering device, the vibration and motion data from the device is sent to the server, which generates cheering sounds that can be played by other spectators watching in real time."
[0818] In this way, this invention allows users to enjoy a realistic viewing experience even from a remote location.
[0819] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0820] Step 1: Initializing the generative AI model
[0821] The server loads the generative AI model and sets the necessary parameters, so it is ready to generate cheers and other sound patterns for the spectator area. It takes the model's reference path as input and the generative AI model initialization as output.
[0822] Step 2: Connecting communication devices
[0823] A user starts up a communication device and selects the event they want to watch. The communication device sends a connection request to the server, and the server establishes the connection. The input is the event information selected by the user, and the output is the established connection.
[0824] Step 3: Receive live feeds and parse event data
[0825] The server receives live video and audio feeds of the event in real time. It detects and analyzes specific events (e.g., goals scored). This process has the live feed data as input and the analyzed event data as output.
[0826] Step 4: Real-time speech generation
[0827] The server uses a generative AI model to generate cheers and cheering sounds in real time. This generated audio data is sent to a communication device. The input is event data, and the output is generated audio data.
[0828] Step 5: Encode and transmit the audio data
[0829] The server encodes the generated audio data and streams it to the communication device, with the generated audio data as input and the encoded audio data as output.
[0830] Step 6: Real-time audio playback
[0831] The communication device plays back the received audio data in real time, with the encoded audio data as input and the audio being played back as output.
[0832] Step 7: Send data to your remote rooting device
[0833] The user cheers by shaking the remote cheering device. The device collects vibration and motion data and sends it to the communication device. The motion data is input and sent to the communication device as output.
[0834] Step 8: Sending remote rooting data to the server
[0835] The communication device receives data from the remote support device and transmits it to the server, which has the remote support data as input and transmits the data to the server as output.
[0836] Step 9: Analyzing cheering data and generating cheering sounds
[0837] The server analyzes the remote cheering data and generates cheering sounds using a generative model. The remote cheering data is input, and the generated cheering sounds are output.
[0838] Step 10: Playing the cheering sounds and sending the audio feed
[0839] The server plays the generated cheering sounds from an audio device and records the audio from the viewing location and sends it to a communication device.The generated cheering sounds and on-site audio are input, and the audio is played back and an audio feed is sent as output.
[0840] Step 11: Playback of local audio via communication device
[0841] The communication device plays back the received local audio. It has received audio data as input and plays back the audio as output.
[0842] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0843] This invention is a system that uses a generative model to reproduce the cheers and excitement of the spectator's location, providing a realistic experience for remote spectators. This system is realized by combining a server, communication terminals, remote cheering devices, audio equipment, and an emotion engine that recognizes the user's emotions.
[0844] Program processing and explanation
[0845] Initializing a generative AI model
[0846] 1. The server initializes the generative AI model
[0847] On startup, the server loads the generative AI model and sets the necessary parameters, which includes loading the model's training data.
[0848] The server caches past match data and audience reaction data into the AI model, preparing it for real-time processing.
[0849] Launching the app and connecting
[0850] 2. The user launches the app and selects a match.
[0851] Users launch the app on their smartphone or tablet and select the game they want to watch. Once the selection is complete, the device sends a connection request to the server.
[0852] 3. The server establishes the connection
[0853] The server receives a connection request from the user's device, establishes the connection through an authentication process, notifies the device that the connection is successful, and the user can begin watching the game.
[0854] Real-time analysis of match data
[0855] 4. The server receives and analyzes the live match feed
[0856] The server receives live video and audio feeds of the match in real time, allowing it to detect events that occur during the match (e.g., goals, falls) and input them into the generative model.
[0857] AI-powered voice generation
[0858] 5. The server generates real-time audio using the generative model
[0859] The server uses a generative model to generate cheers and other supportive sounds based on detected events and local audio feeds.
[0860] 6. The server streams the generated audio to the device
[0861] The server encodes the generated audio data and streams it to the communication terminal, which plays the audio in real time.
[0862] Remote cheering bat data transmission
[0863] 7. User uses a remote rooting device
[0864] Users cheer by shaking the remote cheering device, and vibration and motion data from the device is sent to the user's device.
[0865] 8. The device sends the support data to the server
[0866] The user's device sends the data received from the remote support device to the server, including vibration data and motion data.
[0867] Cheering sound generation and output
[0868] 9. The server analyzes the cheering data and generates voice
[0869] The server analyzes the cheering data and generates cheering sounds using a generative model, which are then played from the sound equipment in the stadium.
[0870] 10. Record audio from the viewing location and send it to your device
[0871] Microphones installed in the viewing area capture audio and transmit it to a server, which then streams it to the communication device and plays it back on the device.
[0872] Using Emotion Data with an Emotion Engine
[0873] 11. The emotion engine recognizes emotions on the user device
[0874] The emotion engine recognizes the user's emotions by analyzing their facial expressions and tone of voice.
[0875] The recognized emotion data is sent to the server in real time.
[0876] 12. Analyze emotional data and reflect it in speech generation
[0877] The server analyzes the received emotion data and dynamically changes the intensity and content of the generated voice based on the user's emotion. For example, if the user is excited, it generates louder cheers.
[0878] Specific examples
[0879] For example, consider the case where User A is watching a soccer match remotely. When User A launches the app and selects a match, the app connects to the server and streams the match video and AI-generated cheering sounds to the device.
[0880] When User A shakes the remote cheering device during a match, the vibration and motion data from the device is sent to the server via the terminal. The server analyzes the data and generates cheering sounds using a generation AI. The generated cheering sounds are played over the stadium's sound system, allowing on-site spectators to hear the cheers.
[0881] Furthermore, the emotion engine recognizes User A's emotions, and if User A is in an excited state, for example, louder cheers and cheering sounds are generated. This allows User A to feel an even stronger sense of presence and unity with the local area.
[0882] A microphone installed at the viewing location captures the sounds from the venue in real time and sends them to User A's device via a server. This allows User A to experience the cheers and excitement of the venue in real time through earphones. This system allows remote spectators to feel as if they are actually there, without feeling any physical distance.
[0883] The processing flow will be explained below.
[0884] Step 1:
[0885] The server initializes the generated AI model.
[0886] On startup, the server loads the generative AI model, sets the necessary parameters, and optimizes it using training data, ready to generate game cheers and sounds.
[0887] Step 2:
[0888] The user launches the app and selects a match
[0889] Users launch the app on their smartphone or tablet and select the game they want to watch. When the user clicks the "Start Watching" button, a connection request is sent from the device to the server.
[0890] Step 3:
[0891] The device sends a connection request to the server
[0892] The user's device sends a connection request to the server's API endpoint, which includes authentication information.
[0893] Step 4:
[0894] The server establishes the connection and performs authentication
[0895] The server receives the connection request and verifies the user's authentication information. If authentication is successful, the server establishes a session and notifies the terminal that the connection is successful.
[0896] Step 5:
[0897] The server receives a live feed of the match
[0898] The server receives live video and audio feeds of the match in real time and prepares to analyze the progress of the match.
[0899] Step 6:
[0900] The server analyzes the match data
[0901] The server analyzes the game footage in real time to detect specific events (e.g., goals, falls), and sends the detected event information to the generative AI model.
[0902] Step 7:
[0903] The server generates the voice using the generative AI model
[0904] The server inputs the detected event information into a generative AI model to generate cheering and cheering sounds in real time, and the generated audio data is encoded.
[0905] Step 8:
[0906] The server streams the generated audio to the device
[0907] The server transmits the encoded audio data in streaming format to the user's device, which plays the audio in real time.
[0908] Step 9:
[0909] User uses remote rooting device
[0910] Users can show their support by shaking the remote cheering device, and vibration and motion data from the device are sent to the user's device.
[0911] Step 10:
[0912] The device sends the support data to the server.
[0913] The user's device then sends the received device data, including vibration and motion data, to the server.
[0914] Step 11:
[0915] The server analyzes the support data
[0916] The server analyzes the received cheering device data and issues instructions to the generation AI model based on the results, which then generates the cheering sound.
[0917] Step 12:
[0918] The server generates cheering sounds using AI
[0919] The server generates cheering sounds based on the analysis results, and the generated audio data is encoded and sent to the stadium's sound system.
[0920] Step 13:
[0921] Play cheering sounds from the sound system
[0922] The sound equipment receives the cheering sound data sent from the server and plays it back in real time, allowing on-site spectators to feel the remote cheering.
[0923] Step 14:
[0924] Record audio from the viewing location and send it to your device
[0925] Microphones installed in the viewing area capture audio in real time and transmit it to a server, which encodes the captured audio for streaming to user devices.
[0926] Step 15:
[0927] The device plays audio from the viewing area
[0928] The user's device receives the streamed audio data and plays it back in real time, allowing the user to experience the cheers and atmosphere of the venue in real time.
[0929] Step 16:
[0930] An emotion engine recognizes emotions on the user's device
[0931] The user's device analyzes the user's facial expressions and tone of voice through a camera and microphone, and recognizes emotions using an emotion engine.
[0932] Step 17:
[0933] The device sends emotion data to the server.
[0934] The user's device transmits the recognized emotion data to the server in real time.
[0935] Step 18:
[0936] The server analyzes the emotional data and reflects it in the speech generation.
[0937] The server analyzes the emotion data and dynamically changes the intensity and content of the generated voices based on the user's emotions, for example, generating louder cheers if the user is excited.
[0938] Specific examples
[0939] For example, consider the case where User A is watching a soccer match remotely. When User A launches the app and selects a match, the app connects to the server and streams the match video and AI-generated cheering sounds to the device.
[0940] When User A shakes the remote cheering device during a match, the vibration and motion data from the device is sent to the server via the terminal. The server analyzes the data and generates cheering sounds using a generation AI. The generated cheering sounds are played over the stadium's sound system, allowing on-site spectators to hear the cheers.
[0941] Furthermore, the emotion engine recognizes User A's emotions, and if User A is in an excited state, for example, louder cheers and cheering sounds are generated. This allows User A to feel an even stronger sense of presence and unity with the local area.
[0942] A microphone installed at the viewing location captures the sounds from the venue in real time and sends them to User A's device via a server. This allows User A to experience the cheers and excitement of the venue in real time through earphones. This system allows remote spectators to feel as if they are actually there, without feeling any physical distance.
[0943] Example 2
[0944] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0945] Current remote viewing systems have difficulty fully reproducing the cheers and excitement of the stadium, meaning that remote spectators cannot fully experience the sense of presence and unity of being at the stadium. Furthermore, there is a lack of technology to effectively generate cheering sounds that utilize spectators' emotions and input signals from remote cheering devices.
[0946] The specific processing by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes a means for initializing the generative model and setting necessary parameters, a means for receiving live video and audio feeds of the game and detecting specific events, a means for transmitting vibration and motion data from the remote cheering device to the server via the communication terminal, a means for using an emotion engine that recognizes the user's emotions, and a means for transmitting the user's emotion data from the communication terminal to the server and reflecting it in sound generation. This allows the cheers and excitement of the viewing location to be reproduced in real time, allowing remote spectators to feel the presence of the venue. Furthermore, by utilizing the spectator's input signals and emotion data, more personalized cheering sounds can be generated.
[0947] A "generative model" is an algorithm that uses machine learning and artificial intelligence techniques to generate new data or patterns based on specific input data.
[0948] "Communication terminal" refers to a device such as a smartphone, tablet, or PC that can send and receive data via the Internet or other networks.
[0949] "Input signals" refer to data and commands sent from spectators or devices to the server, including vibration data, motion data from remote cheering devices, and user emotional data.
[0950] A "server" is a computer system that provides specific services or functions over a network, including data processing, storage, and management.
[0951] "Remote cheering device" refers to a device used by spectators to cheer from a remote location. Specifically, it includes stick-shaped or handheld devices that transmit vibration and motion data to a server.
[0952] An "emotion engine" refers to software or algorithms that analyze a user's facial expressions, tone of voice, physical movements, etc. to recognize emotions, and then generate appropriate responses in real time based on that data.
[0953] "Sound equipment" refers to speakers and sound systems that output the generated cheering sounds and cheers and reproduce them audibly.
[0954] "Live Game Video and Audio Feed" means a data stream that transmits real-time video and audio from the location where the Game is being played.
[0955] "User Emotion Data" refers to data that indicates the user's emotional state analyzed by the emotion engine. This data is used in the speech generation process using the generative model.
[0956] The present invention is a system that uses a generative model to reproduce the cheers and excitement of the spectator's location, providing a realistic experience for remote spectators. This system is realized by a server, communication terminals, remote cheering devices, audio equipment, and an emotion engine.
[0957] Hardware and software used
[0958] The system includes the following hardware and software:
[0959] 1. Server:
[0960] Computer system for executing generative AI models
[0961] Storage for caching past match data and audience reaction data
[0962] Network equipment for receiving and analyzing live video and audio feeds
[0963] 2. Communication terminal:
[0964] Smartphones, tablets, and computers used by spectators
[0965] Connectivity for receiving and sending data from remote rooted devices to the server
[0966] Speakers and earphones for playing audio data from the server
[0967] 3. Remote rooting devices:
[0968] A handheld device that spectators wave to cheer on the game.
[0969] A function that generates vibration and motion data and sends it to a communication device
[0970] 4. Sound equipment:
[0971] A speaker system that outputs cheering sounds and cheers within the stadium
[0972] 5. Emotion Engine:
[0973] Software that recognizes emotions by analyzing the user's facial expressions and tone of voice
[0974] A function to send the recognized emotion data to the server
[0975] Data processing and calculation
[0976] The main processes of this system are:
[0977] Initialize the generative AI model:
[0978] The server loads the generative AI model and sets the necessary parameters, and caches past match data and audience reaction data to prepare for real-time processing.
[0979] Receive live video and audio feeds:
[0980] The server receives live video and audio feeds of the match in real time and detects specific events.
[0981] Voice generation:
[0982] The server uses a generative model to generate cheers and sounds based on the detected events and streams them to the communication device.
[0983] Processing input data from a remote rooting device:
[0984] Vibration and motion data from the remote cheering device is transmitted to a server via a communication terminal, and cheering sounds are generated.
[0985] Use of emotion data:
[0986] The system analyzes the user's facial expressions and tone of voice, and sends the emotional data recognized by the emotion engine to the server, which then dynamically changes the intensity and content of the generated voice based on this data.
[0987] Specific examples
[0988] For example, consider the case where User A is watching a soccer match remotely. When User A launches the app and selects a match, the app connects to the server and streams the match video and AI-generated cheering sounds to the device.
[0989] When User A shakes the remote cheering device during a match, the vibration and motion data from the device is sent to the server via the terminal. The server analyzes the data and generates cheering sounds using a generation AI. The generated cheering sounds are played through the sound system at the stadium, so that on-site spectators can also hear the cheers.
[0990] Furthermore, the emotion engine recognizes User A's emotions, and if User A is in an excited state, for example, louder cheers and cheering sounds are generated. This allows User A to feel an even stronger sense of presence and unity with the local area.
[0991] Prompt Sentence Examples
[0992] An example of a prompt to be input to a generative AI model is written as follows:
[0993] "When a goal is scored during a game, how do we generate audio that synchronizes with the cheers of the crowd?"
[0994] "How to change the intensity of the cheering sounds generated when the user is recognized as excited"
[0995] "What kind of cheering sound should be generated based on the vibration data received from the remote cheering device?"
[0996] This allows the entire system to work together, providing spectators with a realistic experience in real time.
[0997] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0998] The flow of this system's program processing
[0999] Step 1:
[1000] Initializing a generative AI model
[1001] Server loads the model: When the server starts up, it loads the generative AI model from disk and loads it into memory, making the model immediately available for use.
[1002] Server sets parameters: The server sets the hyperparameters required for the model to operate, including the learning rate and batch size. The set parameters are important because they directly affect the model's performance.
[1003] Cache past data: The server caches past match data and audience reaction data, which speeds up real-time processing. Cached data improves response time when an event occurs.
[1004] Input: Generative AI model and configuration parameters from disk
[1005] Output: A generative AI model that can be run in memory
[1006] Step 2:
[1007] Launch the app and select a match
[1008] User launches app: The spectator launches the dedicated app on their smartphone or tablet. The app displays the home screen and shows a list of available matches.
[1009] User selects a game: The user selects the game they want to watch from the list. This selection information is sent from the communication device to the server, which then prepares the corresponding live feed.
[1010] Server accepts connection request: The server accepts the user's connection request and performs authentication. If authentication is successful, the server sends a connection establishment notification to the user's terminal.
[1011] Input: User's match selection information
[1012] Output: Connection establishment and match information on the server side
[1013] Step 3:
[1014] Receiving and analyzing live feeds
[1015] Server receives live video and audio: The server receives live video and audio feeds from the stadium in real time. These feeds are transmitted via a communications network.
[1016] Server detects specific events: The server analyzes and detects specific events that occur during the match (e.g., goals, falls). This data is fed into a generative AI model and used as the basis for generating cheering sounds and cheers.
[1017] Input: Live video and audio feed
[1018] Output: Detected event data
[1019] Step 4:
[1020] Real-time voice generation
[1021] Server generates sounds: Cheering and cheering sounds are generated by a generative AI model based on detected event data. The generated sounds change dynamically depending on the situation of the match.
[1022] The server encodes the generated audio: the encoded audio data is prepared for streaming, providing high-quality audio to the user in real time.
[1023] Input: Detected event data
[1024] Output: Generated cheering and cheering sound data
[1025] Step 5:
[1026] Streaming generated audio
[1027] The server streams audio data to the device: The server sends encoded audio data in real time to the user's communication device, which receives the data and plays it through its built-in speaker or earphones.
[1028] Input: Generated audio data
[1029] Output: Streamed audio data
[1030] Step 6:
[1031] Sending data from a remote rooting device
[1032] The user operates the device: The user performs cheering actions using the remote cheering device. The device generates vibration and motion data and transmits it to the communication terminal.
[1033] The terminal transmits the data to the server: The communication terminal transmits the data received from the device to the server, which analyzes the data and generates additional cheering sounds.
[1034] Input: Vibration and motion data from the remote rooting device
[1035] Output: Data sent to the server
[1036] Step 7:
[1037] Cheering sound generation
[1038] The server analyzes the data: vibration and motion data is analyzed, and a generative AI model is used to generate cheering sounds, which are then played over the on-site sound system.
[1039] Input: Vibration and motion data
[1040] Output: Generated cheer sound
[1041] Step 8:
[1042] Audio recording from the viewing area
[1043] The server captures the audio: Microphones installed in the viewing area capture the audio from the venue and send it to the server.
[1044] The server streams the audio to the device: The captured audio is sent to the communication device and played back in real time. This process is important for conveying the cheers and realism of the actual event to the user.
[1045] Input: Audio data from microphones installed in the viewing area
[1046] Output: Audio data streamed to the user's device
[1047] Step 9:
[1048] Analysis and use of emotional data
[1049] User device collects emotional data: The emotion engine analyzes the user's facial expressions and tone of voice to collect emotional data.
[1050] The device sends emotional data to the server: The analyzed emotional data is sent to the server in real time. The server analyzes this data and dynamically changes the intensity and content of the generated voice.
[1051] Input: User facial expressions and tone of voice
[1052] Output: Emotion data sent to the server.
[1053] (Application example 2)
[1054] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[1055] Current remote viewing systems have the problem of being unable to fully reproduce the sense of presence and unity that users feel at a real viewing location. In particular, it is difficult to experience the cheers and sounds of the fans at the venue in real time, and they are unable to generate cheering sounds that reflect the emotions of the remote spectators, limiting the viewing experience. Furthermore, they lack the functionality to generate cheering sounds using data from remote cheering devices. Technology that solves these issues is needed.
[1056] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1057] In this invention, the server includes means for reproducing the cheers and excitement of the spectator location using a generative model, means for transmitting input signals from spectators using a communication terminal to the server, means for analyzing the input signals, generating cheering sounds, and outputting them from a sound device, means for recording audio from the spectator location and transmitting the audio to the communication terminal, means for playing the audio on the communication terminal, means for recognizing the emotions of spectators using an emotion engine and transmitting the emotion data to the server in real time, and means for dynamically changing the intensity and content of the cheering sounds based on the emotion data. This allows remote spectators to experience a sense of realism similar to that of being at the venue, and cheering sounds that correspond to their emotions can be generated.
[1058] A "generative model" is a machine learning model that uses artificial intelligence to generate new data.
[1059] A "spectator venue" is the location where a sport or event actually takes place and where spectators physically gather.
[1060] "Cheers and enthusiasm" refers to the cheers and support emitted by spectators at the viewing venue, and are the sounds and atmosphere that indicate the excitement of the event.
[1061] A "communication terminal" is an electronic device used to send and receive data over the Internet, including smartphones and tablets.
[1062] "Input signals from spectators" refer to data that spectators send to the server via their communication terminals, and are signals that reflect the actions and emotions of the spectators.
[1063] A "server" is a computer system that processes and manages data over a network.
[1064] "Cheering sounds" are sounds generated based on input signals from spectators and generative models, and include sounds of cheering and cheering.
[1065] "Audio device" refers to equipment for outputting sound, including speakers and earphones.
[1066] An "emotion engine" is software that analyzes and recognizes the user's emotions from their facial expressions and movements.
[1067] "Real time" is a time concept that refers to instantaneous reaction and processing without delay.
[1068] A "remote cheering device" is a device that spectators use to cheer from home or another location, and has the ability to transmit vibration and motion data to a server.
[1069] The system for implementing this invention includes a generative AI model, an emotion engine, a communication terminal, a server, and an audio device. The operation of the entire system is as follows.
[1070] First, the user launches the remote viewing application on their communication device (smartphone or tablet). Within the application, the user selects the game or event they want to watch. The communication device then sends the selection information to the server, which then receives it.
[1071] The server receives live game feeds (video and audio) and inputs them into a generative AI model to generate cheering and support sounds for the spectators in real time. The generated audio is encoded, streamed to a communication device, and played back to the user in real time. The generative AI model then detects specific events (e.g., a goal or a foul) and dynamically generates audio based on those events.
[1072] Furthermore, the communication terminal is equipped with an emotion engine that recognizes the user's emotions. The emotion engine analyzes the user's facial expressions and tone of voice and transmits the recognized emotion data to the server in real time. The server analyzes this emotion data and dynamically changes the intensity and content of the cheering sounds according to the user's emotions. For example, if the user is excited, a louder cheer will be generated.
[1073] Users can also use a remote cheering device. Vibration and motion data from the remote cheering device is sent to a communication terminal, which then transmits the data to a server. The server analyzes the data and generates cheering sounds using a generative AI model, which are then played on the on-site sound equipment.
[1074] As a concrete example, consider the case where a user is watching a soccer match remotely. When the user launches the application and selects a match, live video and cheering sounds generated by a generative AI are streamed to the communication terminal. When the user shakes the remote cheering device during the match, vibration and motion data from the device are sent to the server via the communication terminal. The server analyzes the data and generates cheering sounds using a generative AI model. These cheering sounds are then played over the sound equipment at the viewing location.
[1075] An example of this prompt would be:
[1076] "A user is watching a soccer match remotely. The emotion engine on the user's smartphone recognizes the user's facial expressions and tone of voice, and the generation AI generates the cheering sounds of the local crowd in real time accordingly. As the user's excitement increases, the generation AI generates louder cheers and cheering sounds, providing a realistic spectator experience."
[1077] As described above, using this system allows remote spectators to experience the same sense of realism as if they were in person, and cheering sounds can be generated to suit their emotions, greatly improving the sense of realism and unity of remote spectators.
[1078] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1079] Step 1:
[1080] The server initializes the generated AI model.
[1081] Input: Past match data and audience reaction data
[1082] Output: Initialized generative AI model
[1083] At startup, the server loads the generative AI model, sets the necessary parameters, reads the model's training data, and prepares it for real-time processing.
[1084] Step 2:
[1085] The user starts the application on the communication terminal and selects the game they want to watch.
[1086] Input: User's match selection information
[1087] Output: Match selection information sent to server
[1088] The user launches the application and selects the game they want to watch. The selection information is sent to the server, which receives it.
[1089] Step 3:
[1090] The server receives a live feed of the game and uses a generative AI model to generate the cheers and support sounds of the spectators.
[1091] Input: Live video and audio feed
[1092] Output: Generated cheers and cheering sounds
[1093] The server receives live video and audio feeds of the game, feeds them into a generative AI model, and generates cheers and cheering sounds in real time, which are then encoded and streamed to communication devices.
[1094] Step 4:
[1095] The communication terminal reproduces the generated cheers and cheering sounds in real time.
[1096] Input: Audio data streamed from the server
[1097] Output: Real-time cheers and cheering sounds played
[1098] The communication terminal decodes the received audio data and plays it back in real time, allowing the user to experience realistic audio.
[1099] Step 5:
[1100] The emotion engine recognizes the user's emotions and transmits the emotion data to the server in real time.
[1101] Input: User facial expressions and tone of voice
[1102] Output: Emotion data sent to the server
[1103] The emotion engine analyzes the user's facial expressions and tone of voice on the communication device to recognize their emotions, and the recognized emotion data is sent to the server.
[1104] Step 6:
[1105] The server analyzes the emotional data and uses a generative AI model to dynamically change the intensity and content of the cheering sounds.
[1106] Input: User emotion data
[1107] Output: Cheering sounds with dynamically changing intensity and content
[1108] The server analyzes the received emotional data and dynamically changes the intensity and content of the cheering sounds based on the user's emotions using a generative AI model. For example, if the user is excited, a louder cheer will be generated.
[1109] Step 7:
[1110] The user uses the remote support device and transmits the vibration data and motion data to the communication terminal.
[1111] Input: Vibration and motion data from the remote rooting device
[1112] Output: Data sent to the communication device
[1113] Users cheer by shaking the remote cheering device, and vibration and motion data from the device are sent to the communication terminal.
[1114] Step 8:
[1115] The communication terminal transmits the data received from the remote support device to the server.
[1116] Input: Vibration and motion data from the remote rooting device
[1117] Output: Device data sent to the server
[1118] The communication terminal transmits the received vibration data and motion data to the server.
[1119] Step 9:
[1120] The server analyzes the data from the remote cheering device and generates cheering sounds using a generative AI model.
[1121] Input: Vibration and motion data
[1122] Output: Generated cheer sound
[1123] The server analyzes the data from the remote cheering device and generates cheering sounds using a generative AI model. The generated cheering sounds are then played over the sound equipment at the viewing area.
[1124] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1125] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1126] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[1127] [Third embodiment]
[1128] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[1129] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[1130] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1131] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[1132] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1133] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1134] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1135] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1136] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1137] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1138] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1139] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[1140] This system uses a generative model to recreate the cheers and excitement of the spectator's location, providing a realistic experience for remote spectators. The system is primarily composed of a server, communication terminals, remote cheering devices, and audio equipment.
[1141] Program processing and explanation
[1142] Initializing a generative AI model
[1143] 1. The server initializes the generative AI model
[1144] On startup, the server loads the generative AI model and sets the necessary parameters, so the model is ready to generate cheers and other acoustic patterns for the spectator area.
[1145] Launching the app and connecting
[1146] 2. The user launches the app and selects a match.
[1147] Users launch the app on their smartphone or other communication device and select the game they want to watch. Once the selection is complete, the device sends a connection request to the server.
[1148] 3. The server establishes the connection
[1149] The server receives a connection request from the user's device, establishes the connection through an authentication process, notifies the device that the connection is successful, and the user can begin watching the game.
[1150] Real-time analysis of match data
[1151] 4. The server receives and analyzes the live match feed
[1152] The server receives live video and audio feeds of the match in real time, allowing it to detect events that occur during the match (e.g., goals scored, falls) and input them into the generative model.
[1153] AI-powered voice generation
[1154] 5. The server generates real-time audio using the generative model
[1155] The server uses a generative model to generate cheers and other supportive sounds based on detected events and local audio feeds.
[1156] 6. The server streams the generated audio
[1157] The server encodes the generated audio data and streams it to the communication terminal, which plays the audio in real time.
[1158] Remote cheering bat data transmission
[1159] 7. User uses a remote rooting device
[1160] The user starts cheering by shaking the remote cheering device, which then transmits vibration and motion data to the communication terminal.
[1161] 8. The device sends the support data to the server
[1162] The communication terminal sends the received data from the cheering device to the server, which analyzes the data and uses it to generate cheering sounds.
[1163] Cheering sound generation and output
[1164] 9. The server analyzes the cheering data and generates voice
[1165] The server analyzes the cheering data and generates cheering sounds using a generative model, which are then played from the sound equipment in the stadium.
[1166] 10. Record audio from the viewing area and send it to a communication device
[1167] Microphones installed in the viewing area capture audio and transmit it to a server, which then streams it to the communication device and plays it back on the device.
[1168] Specific examples
[1169] For example, consider the case where User A is watching a soccer match remotely. When User A launches the app and selects a match, the app connects to the server and streams the match video and AI-generated cheering sounds to the device.
[1170] When User A shakes the remote cheering device during a match, the vibration and motion data from the device is sent to the server via the terminal. The server analyzes the data and generates cheering sounds using a generation AI. The generated cheering sounds are played over the stadium's sound system, allowing on-site spectators to hear the cheers.
[1171] Furthermore, microphones installed in the viewing area capture the sounds from the venue in real time and transmit them via a server to User A's device, allowing User A to experience the cheers and excitement of the crowd in real time through earphones.
[1172] The processing flow will be explained below.
[1173] Step 1:
[1174] The server initializes the generated AI model.
[1175] At startup, the server loads the AI model and configures the model's parameters, which includes loading the model's training data.
[1176] The server caches past match data and audience reaction data into the AI model, preparing it for real-time processing.
[1177] Step 2:
[1178] The user launches the app and selects a match
[1179] The user launches the app on their smartphone or tablet.
[1180] Select the game you want to watch from the app's main screen and click the "Start Watching" button.
[1181] Step 3:
[1182] The device sends a connection request to the server
[1183] The user's terminal sends a request to start watching to the server.
[1184] The device connects to the server's API endpoint and sends authentication information.
[1185] Step 4:
[1186] The server establishes the connection and performs authentication
[1187] The server receives the connection request and authenticates the user.
[1188] Once the authentication is complete, the server establishes a session and notifies the terminal of the successful connection.
[1189] Step 5:
[1190] The server receives a live feed of the match
[1191] The server receives live video and audio feeds of the match in real time.
[1192] The server analyzes this data and prepares to detect important events (e.g. goals, falls).
[1193] Step 6:
[1194] The server detects an important event
[1195] The server analyzes the video and audio of the match in real time to detect specific events.
[1196] When a specific event such as a goal or fall is detected, that information is passed on to the generating AI.
[1197] Step 7:
[1198] The server generates the voice using AI
[1199] The server uses a generative AI model to generate cheers and sounds based on the detected events.
[1200] The generated audio data is encoded in real time.
[1201] Step 8:
[1202] The server streams the generated audio to the device
[1203] The server streams the encoded audio data to the user terminal.
[1204] The user terminal plays back the received audio in real time.
[1205] Step 9:
[1206] User uses remote rooting device
[1207] Users cheer by waving the remote cheering device.
[1208] Vibration and motion data from the device is sent to the user's terminal.
[1209] Step 10:
[1210] The device sends the support data to the server.
[1211] The user's terminal transmits the data received from the remote support device to the server.
[1212] The data sent to the server includes vibration data and motion data.
[1213] Step 11:
[1214] The server analyzes the support data
[1215] The server analyzes the received support data in real time.
[1216] Based on the analysis results, the server determines what kind of cheering sound to generate.
[1217] Step 12:
[1218] The server generates cheering sounds using AI
[1219] The server uses a generative AI model to generate cheering sounds, including applause and cheers.
[1220] The generated cheering sound is encoded and transmitted to an audio device.
[1221] Step 13:
[1222] Play cheering sounds from the sound system
[1223] The cheering sound data transmitted from the server is received by the sound device.
[1224] The sound system will then play the received cheering sounds in real time, allowing on-site spectators to feel the cheering from the remote area.
[1225] Step 14:
[1226] Record audio from the viewing location and send it to your device
[1227] The server captures audio from microphones installed at the viewing area and transmits it to the server in real time.
[1228] The server encodes the captured audio for streaming to the user terminal.
[1229] Step 15:
[1230] The device plays audio from the viewing area
[1231] The user terminal receives the streamed audio data and plays it back in real time.
[1232] Users can experience the cheers and atmosphere of the venue in real time.
[1233] Example 1
[1234] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1235] The aim is to solve the problem that it is difficult for remote spectators to experience the sense of presence and cheers of being at the venue in real time, and that cheers from remote cheering devices are not easily conveyed to on-site spectators.
[1236] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1237] In this invention, the server includes a means for recreating the cheers and excitement of the spectator location using a generative model, a means for receiving and analyzing a live feed of the game and detecting events, and a means for encoding audio data and streaming it to a communication terminal, allowing remote spectators to experience realistic cheers in real time and for cheering via cheering devices to be reflected at the venue.
[1238] A "generative model" is an artificial intelligence technology that generates new data and information based on learned algorithms and data.
[1239] "Spectator venue" refers to a local venue or stadium where a sporting event, concert, etc. is held.
[1240] "Cheering and excitement" refers to the vocal and emotional excitement of spectators expressing their excitement and support for a game or event.
[1241] A "communication terminal" is an electronic device that can send and receive data over a network, such as a smartphone, tablet PC, or personal computer.
[1242] A "server" is a high performance computing device for storing, processing, and distributing data over a network.
[1243] An "input signal" is an electrical signal that contains instructions or data from a user or device.
[1244] "Cheering sounds" are sounds and sound effects made by spectators when cheering on a game or event.
[1245] "Sound equipment" means equipment such as speakers and amplifiers for reproducing sound.
[1246] "Recording" is the act of recording sound in digital or analog form.
[1247] "Streaming" refers to the technology of continuously transmitting and playing data in real time over the Internet.
[1248] A "remote support device" is a device that enables support from a remote location, and has the function of transmitting vibration data and motion data mainly to a communication terminal.
[1249] This system uses a generative model to recreate the cheers and excitement of the spectator's location, providing a realistic experience for remote spectators. The system mainly consists of a server, communication terminals, remote cheering devices, and audio equipment.
[1250] The server first initializes the generative AI model by loading a pre-trained model file into memory using a machine learning library such as TensorFlow or PyTorch. It then sets the necessary parameters and prepares the model for operation, ready to generate cheers and other acoustic patterns for the spectators.
[1251] Users launch a dedicated app on their communication device, such as a smartphone or tablet, and select the game they want to watch. Once a game is selected, the device sends the selection information to the server and requests a connection to the server. The server receives the connection request from the user device and performs an authentication process based on the authentication information. If authentication is successful, the session begins and the user can begin remote viewing.
[1252] The server receives live video and audio feeds of the game in real time. It then analyzes the data to detect goals and important events. The results of this analysis are converted into a format that can be used by the generative AI model and fed into it. The resulting audio data is then encoded and streamed to the user's device.
[1253] Users begin cheering by shaking the remote cheering device. This device transmits vibration and motion data to a communication terminal. The communication terminal then transmits this data to a server, which then generates cheering sounds based on that data. The generated cheering sounds are played from the sound equipment in the stadium, allowing on-site spectators to hear the cheers.
[1254] Furthermore, microphones installed in the viewing area capture the sound from the venue in real time and transmit it to a server, which then streams the sound to communication devices, allowing users to experience the cheers and excitement of the venue in real time.
[1255] As a concrete example, consider the case where User A is watching a soccer match remotely. When User A launches the app and selects a match, the app connects to the server, and game footage and cheering sounds generated by AI are streamed to the device. When User A shakes the remote cheering device during the match, the vibration and motion data from the device are sent to the server, and cheering sounds are generated and played from the on-site sound equipment. In addition, audio captured at the venue is sent to the device in real time, allowing User A to enjoy the immersive experience through earphones.
[1256] Examples of prompts for a generative AI model include:
[1257] "In order to generate audio that reproduces the cheers and excitement of the viewing area, we provide the following information:
[1258] 1. The crowd's reaction when a goal is scored
[1259] 2. The crowd booing when the fall happened
[1260] 3. The crowd roaring during halftime
[1261] Based on this, generate realistic cheer sounds in real time.
[1262] This system will enable remote spectators to experience the excitement and cheers of being at the venue in real time, improving the viewing experience.
[1263] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1264] Step 1: Initializing the generative AI model
[1265] The server initializes the generated AI model.
[1266] Input: trained model file, required parameters
[1267] Output: Initialized generative AI model
[1268] What it does: When the server starts up, it uses machine learning libraries like TensorFlow and PyTorch to load a trained generative AI model into memory, reads the model's parameters (e.g., settings for generating voices), and prepares the model for operation, ready to generate cheers and other acoustic patterns for the spectators.
[1269] Step 2: Launch the app and connect
[1270] The user launches the app and selects a match
[1271] Input: Select the match you want to watch
[1272] Output: A connection request is sent
[1273] Specific operation: The user launches the dedicated app on their smartphone or tablet and selects the game they want to watch from a list on the main screen. Once a game is selected, the app sends a connection request to the server with the selected information.
[1274] The server establishes the connection
[1275] Input: Connection request, user credentials
[1276] Output: Connection established notification
[1277] Specific operation: The server receives a connection request from a user terminal and performs an authentication process based on authentication information (e.g., user ID and password). If authentication is successful, the server starts a session and notifies the communication terminal that the connection has been established.
[1278] Step 3: Real-time analysis of match data
[1279] The server receives and analyzes the live match feed
[1280] Input: Live video and audio feed of the match
[1281] Output: Event data (scoring scenes, falls, etc.)
[1282] How it works: The server receives live video and audio feeds from live game broadcasts in real time. It analyzes the input video and audio data to detect goals and important events. The analysis results are then converted into a format that can be used by the generative AI model.
[1283] Step 4: AI-powered voice generation
[1284] The server generates real-time audio using the generative model
[1285] Input: Event data, local audio feed
[1286] Output: Generated audio data
[1287] How it works: The server uses a generative AI model to generate realistic cheering and cheering sounds based on the detected event data and local audio feeds. The generated audio data is then encoded and prepared for transmission to the communication device.
[1288] The server streams the generated audio
[1289] Input: Generated audio data
[1290] Output: Audio stream
[1291] Specific operation: The server streams the encoded audio data to the user's communication device in real time, allowing the user to play back the realistic audio generated on the device.
[1292] Step 5: Send data to your remote rooting device
[1293] User uses remote rooting device
[1294] Input: Vibration data and motion data of the remote rooting device
[1295] Output: Send data to a communication terminal
[1296] Specific operation: A user expresses his / her support by shaking a remote support device (e.g., a Bluetooth-connected support bat). This device transmits vibration data and motion data to a communication terminal.
[1297] The device sends the support data to the server.
[1298] Input: Vibration data and motion data of the remote rooting device
[1299] Output: Send data to the server
[1300] Specific operation: The communication device sends the received vibration and motion data to the server, which receives this data, analyzes it, and uses it to generate cheering sounds.
[1301] Step 6: Generate and output cheer sounds
[1302] The server analyzes the cheering data and generates voice
[1303] Input: Remote rooting device data
[1304] Output: Generated cheer sound
[1305] Specific operation: The server analyzes the received data and generates sounds according to the intensity and frequency of the cheering. The generated cheering sounds are sent to the sound equipment in the stadium and played back.
[1306] Record audio from the viewing location and send it to a communication device
[1307] Input: Audio data captured at the viewing location
[1308] Output: Audio stream to communication device
[1309] How it works: Microphones installed in the viewing area capture audio in real time. The captured audio data is sent to a server and then encoded and transmitted to a communication device. Users can play this audio on their device, allowing them to experience the realism of being at the venue.
[1310] (Application example 1)
[1311] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1312] Conventional remote viewing systems make it difficult to experience the cheers and excitement of the fans in real time, and they lack the interactivity of cheering. Furthermore, there is a lack of a way to efficiently analyze data from remote cheering devices and generate realistic cheering sounds. There is a need to solve these issues and provide a more realistic viewing experience.
[1313] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1314] In this invention, the server includes means for reproducing the cheers and excitement of the spectator location using a generative model, means for transmitting input signals from spectators to the server using a communication device, means for analyzing the input signals to generate cheering sounds and outputting them from a sound device, means for recording audio from the spectator location and transmitting it to the communication device, means for playing the audio on the communication device, means for performing real-time event analysis and generating cheers using a generative model, and means for acquiring vibration data and motion data from the cheering device and transmitting it to the server. This provides remote spectators with the same sense of realism as if they were watching the game in person, allowing users to send their cheers interactively through their devices.
[1315] A "generative model" is a model that uses artificial intelligence technology to generate specific patterns, sounds, images, etc.
[1316] A "spectator location" is a local location where entertainment or competition, such as a sporting event or concert, takes place.
[1317] "Cheers" are the cheers and excited voices that spectators make at the viewing venue in response to a game or performance.
[1318] "Excitement" refers to the atmosphere of excitement and enthusiasm created by the audience.
[1319] A "communication device" is a device for sending and receiving information, including smartphones, tablets, and computers.
[1320] A "spectator" is someone who watches events such as sports or concerts remotely or in person.
[1321] An "input signal" is data or instructions sent by a spectator through a communication device.
[1322] "Cheering sounds" are sounds generated in response to the cheering actions of spectators.
[1323] "Audio device" means a device for reproducing sound, including speakers and headphones.
[1324] "Real-time event analysis" is a process that instantly analyzes the status and behavior of an event currently in progress.
[1325] "Vibration data" is data recorded of the movement of the support device when it vibrates.
[1326] "Motion data" is data that records the movement of the support device.
[1327] A "server" is a computer that manages and provides information over a network.
[1328] The system embodying this invention uses a generative AI model to recreate the cheers and excitement of the spectator's location, providing a sense of realism to remote spectators. The system is primarily composed of a server, communication equipment, a remote cheering device, and an audio device.
[1329] 1. System Program
[1330] The system's program is built around a generative AI model and mainly performs the following processes:
[1331] 1. Initializing the generative AI model
[1332] The server loads the generative AI model and sets the necessary parameters, so the model is ready to generate cheers and other sound patterns for the spectator area.
[1333] 2. Connecting communication devices
[1334] The user starts up the communication device and selects the event they want to watch. The communication device sends a connection request to the server, and the server establishes the connection. If the connection is successful, the user can start watching.
[1335] 3. Event Data Analysis
[1336] The server receives live video and audio feeds of the event in real time, detects and analyzes specific events, and feeds the results into a generative AI model.
[1337] 4. Real-time speech generation
[1338] The server uses a generative AI model to generate cheers and cheering sounds in real time, and this generated audio data is sent to a communication device and played back in real time.
[1339] 5. Use of remote support devices
[1340] When a user shakes the remote cheering device, the device transmits vibration and motion data to the communication device, which then analyzes the data and generates cheering sounds.
[1341] 6. Data transmission and reception between the server and communication device
[1342] The server plays the generated cheering sounds from the on-site sound equipment and records the audio from the viewing location and transmits it to the communication device, which then plays back the audio, allowing the user to experience the cheers and realism of the venue.
[1343] 2. Hardware and Software Details
[1344] Hardware
[1345] Server: Run the system on a high-performance server (e.g., AWS EC2).
[1346] Communication devices: Common devices such as smartphones and tablets.
[1347] Remote rooting device: An IoT device for capturing user actions.
[1348] Sound device: An audio playback device such as a speaker or headphones.
[1349] software
[1350] Generative AI model: For example, we use OpenAI's GPT-3 model for speech generation.
[1351] Communication and Data Processing: Build WebSocket server and client functions using Python.
[1352] 3. Specific Examples
[1353] For example, if a user wants to watch a live concert remotely, the process would be as follows:
[1354] 1. A user turns on a communication device and selects a live concert. The communication device connects to a server and receives a real-time live feed.
[1355] 2. The server analyzes the live feed and detects specific events (e.g., highlights or excitement).
[1356] 3. The server uses a generative AI model to generate cheers and cheering sounds based on these events.
[1357] 4. When the user shakes the remote cheering device, the motion data is sent to the server via the communication device, and a cheering sound is generated.
[1358] 5. The generated audio is played on the communication device and also played on-site.
[1359] Prompt Sentence Examples
[1360] 1. Initializing the generative AI model
[1361] "Load a generative model and generate cheers and cheers in real time."
[1362] 2. User's live concert viewing
[1363] "Users simply launch the app, select a specific live concert, and watch it. The live feed, including the sounds of cheering and cheering from the venue, is streamed to their smartphone."
[1364] 3. Use a remote support device
[1365] "When a user remotely shakes a cheering device, the vibration and motion data from the device is sent to the server, which generates cheering sounds that can be played by other spectators watching in real time."
[1366] In this way, this invention allows users to enjoy a realistic viewing experience even from a remote location.
[1367] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1368] Step 1: Initializing the generative AI model
[1369] The server loads the generative AI model and sets the necessary parameters, so it is ready to generate cheers and other sound patterns for the spectator area. It takes the model's reference path as input and the generative AI model initialization as output.
[1370] Step 2: Connecting communication devices
[1371] A user starts up a communication device and selects the event they want to watch. The communication device sends a connection request to the server, and the server establishes the connection. The input is the event information selected by the user, and the output is the established connection.
[1372] Step 3: Receive live feeds and parse event data
[1373] The server receives live video and audio feeds of the event in real time. It detects and analyzes specific events (e.g., goals scored). This process has the live feed data as input and the analyzed event data as output.
[1374] Step 4: Real-time speech generation
[1375] The server uses a generative AI model to generate cheers and cheering sounds in real time. This generated audio data is sent to a communication device. The input is event data, and the output is generated audio data.
[1376] Step 5: Encode and transmit the audio data
[1377] The server encodes the generated audio data and streams it to the communication device, with the generated audio data as input and the encoded audio data as output.
[1378] Step 6: Real-time audio playback
[1379] The communication device plays back the received audio data in real time, with the encoded audio data as input and the audio being played back as output.
[1380] Step 7: Send data to your remote rooting device
[1381] The user cheers by shaking the remote cheering device. The device collects vibration and motion data and sends it to the communication device. The motion data is input and sent to the communication device as output.
[1382] Step 8: Sending remote rooting data to the server
[1383] The communication device receives data from the remote support device and transmits it to the server, which has the remote support data as input and transmits the data to the server as output.
[1384] Step 9: Analyzing cheering data and generating cheering sounds
[1385] The server analyzes the remote cheering data and generates cheering sounds using a generative model. The remote cheering data is input, and the generated cheering sounds are output.
[1386] Step 10: Playing the cheering sounds and sending the audio feed
[1387] The server plays the generated cheering sounds from an audio device and records the audio from the viewing location and sends it to a communication device.The generated cheering sounds and on-site audio are input, and the audio is played back and an audio feed is sent as output.
[1388] Step 11: Playback of local audio via communication device
[1389] The communication device plays back the received local audio. It has received audio data as input and plays back the audio as output.
[1390] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1391] This invention is a system that uses a generative model to reproduce the cheers and excitement of the spectator's location, providing a realistic experience for remote spectators. This system is realized by combining a server, communication terminals, remote cheering devices, audio equipment, and an emotion engine that recognizes the user's emotions.
[1392] Program processing and explanation
[1393] Initializing a generative AI model
[1394] 1. The server initializes the generative AI model
[1395] On startup, the server loads the generative AI model and sets the necessary parameters, which includes loading the model's training data.
[1396] The server caches past match data and audience reaction data into the AI model, preparing it for real-time processing.
[1397] Launching the app and connecting
[1398] 2. The user launches the app and selects a match.
[1399] Users launch the app on their smartphone or tablet and select the game they want to watch. Once the selection is complete, the device sends a connection request to the server.
[1400] 3. The server establishes the connection
[1401] The server receives a connection request from the user's device, establishes the connection through an authentication process, notifies the device that the connection is successful, and the user can begin watching the game.
[1402] Real-time analysis of match data
[1403] 4. The server receives and analyzes the live match feed
[1404] The server receives live video and audio feeds of the match in real time, allowing it to detect events that occur during the match (e.g., goals, falls) and input them into the generative model.
[1405] AI-powered voice generation
[1406] 5. The server generates real-time audio using the generative model
[1407] The server uses a generative model to generate cheers and other supportive sounds based on detected events and local audio feeds.
[1408] 6. The server streams the generated audio to the device
[1409] The server encodes the generated audio data and streams it to the communication terminal, which plays the audio in real time.
[1410] Remote cheering bat data transmission
[1411] 7. User uses a remote rooting device
[1412] Users cheer by shaking the remote cheering device, and vibration and motion data from the device is sent to the user's device.
[1413] 8. The device sends the support data to the server
[1414] The user's device sends the data received from the remote support device to the server, including vibration data and motion data.
[1415] Cheering sound generation and output
[1416] 9. The server analyzes the cheering data and generates voice
[1417] The server analyzes the cheering data and generates cheering sounds using a generative model, which are then played from the sound equipment in the stadium.
[1418] 10. Record audio from the viewing location and send it to your device
[1419] Microphones installed in the viewing area capture audio and transmit it to a server, which then streams it to the communication device and plays it back on the device.
[1420] Using Emotion Data with an Emotion Engine
[1421] 11. The emotion engine recognizes emotions on the user device
[1422] The emotion engine recognizes the user's emotions by analyzing their facial expressions and tone of voice.
[1423] The recognized emotion data is sent to the server in real time.
[1424] 12. Analyze emotional data and reflect it in speech generation
[1425] The server analyzes the received emotion data and dynamically changes the intensity and content of the generated voice based on the user's emotion. For example, if the user is excited, it generates louder cheers.
[1426] Specific examples
[1427] For example, consider the case where User A is watching a soccer match remotely. When User A launches the app and selects a match, the app connects to the server and streams the match video and AI-generated cheering sounds to the device.
[1428] When User A shakes the remote cheering device during a match, the vibration and motion data from the device is sent to the server via the terminal. The server analyzes the data and generates cheering sounds using a generation AI. The generated cheering sounds are played over the stadium's sound system, allowing on-site spectators to hear the cheers.
[1429] Furthermore, the emotion engine recognizes User A's emotions, and if User A is in an excited state, for example, louder cheers and cheering sounds are generated. This allows User A to feel an even stronger sense of presence and unity with the local area.
[1430] A microphone installed at the viewing location captures the sounds from the venue in real time and sends them to User A's device via a server. This allows User A to experience the cheers and excitement of the venue in real time through earphones. This system allows remote spectators to feel as if they are actually there, without feeling any physical distance.
[1431] The processing flow will be explained below.
[1432] Step 1:
[1433] The server initializes the generated AI model.
[1434] On startup, the server loads the generative AI model, sets the necessary parameters, and optimizes it using training data, ready to generate game cheers and sounds.
[1435] Step 2:
[1436] The user launches the app and selects a match
[1437] Users launch the app on their smartphone or tablet and select the game they want to watch. When the user clicks the "Start Watching" button, a connection request is sent from the device to the server.
[1438] Step 3:
[1439] The device sends a connection request to the server
[1440] The user's device sends a connection request to the server's API endpoint, which includes authentication information.
[1441] Step 4:
[1442] The server establishes the connection and performs authentication
[1443] The server receives the connection request and verifies the user's authentication information. If authentication is successful, the server establishes a session and notifies the terminal that the connection is successful.
[1444] Step 5:
[1445] The server receives a live feed of the match
[1446] The server receives live video and audio feeds of the match in real time and prepares to analyze the progress of the match.
[1447] Step 6:
[1448] The server analyzes the match data
[1449] The server analyzes the game footage in real time to detect specific events (e.g., goals, falls), and sends the detected event information to the generative AI model.
[1450] Step 7:
[1451] The server generates the voice using the generative AI model
[1452] The server inputs the detected event information into a generative AI model to generate cheering and cheering sounds in real time, and the generated audio data is encoded.
[1453] Step 8:
[1454] The server streams the generated audio to the device
[1455] The server transmits the encoded audio data in streaming format to the user's device, which plays the audio in real time.
[1456] Step 9:
[1457] User uses remote rooting device
[1458] Users can show their support by shaking the remote cheering device, and vibration and motion data from the device are sent to the user's device.
[1459] Step 10:
[1460] The device sends the support data to the server.
[1461] The user's device then sends the received device data, including vibration and motion data, to the server.
[1462] Step 11:
[1463] The server analyzes the support data
[1464] The server analyzes the received cheering device data and issues instructions to the generation AI model based on the results, which then generates the cheering sound.
[1465] Step 12:
[1466] The server generates cheering sounds using AI
[1467] The server generates cheering sounds based on the analysis results, and the generated audio data is encoded and sent to the stadium's sound system.
[1468] Step 13:
[1469] Play cheering sounds from the sound system
[1470] The sound equipment receives the cheering sound data sent from the server and plays it back in real time, allowing on-site spectators to feel the remote cheering.
[1471] Step 14:
[1472] Record audio from the viewing location and send it to your device
[1473] Microphones installed in the viewing area capture audio in real time and transmit it to a server, which encodes the captured audio for streaming to user devices.
[1474] Step 15:
[1475] The device plays audio from the viewing area
[1476] The user's device receives the streamed audio data and plays it back in real time, allowing the user to experience the cheers and atmosphere of the venue in real time.
[1477] Step 16:
[1478] An emotion engine recognizes emotions on the user's device
[1479] The user's device analyzes the user's facial expressions and tone of voice through a camera and microphone, and recognizes emotions using an emotion engine.
[1480] Step 17:
[1481] The device sends emotion data to the server.
[1482] The user's device transmits the recognized emotion data to the server in real time.
[1483] Step 18:
[1484] The server analyzes the emotional data and reflects it in the speech generation.
[1485] The server analyzes the emotion data and dynamically changes the intensity and content of the generated voices based on the user's emotions, for example, generating louder cheers if the user is excited.
[1486] Specific examples
[1487] For example, consider the case where User A is watching a soccer match remotely. When User A launches the app and selects a match, the app connects to the server and streams the match video and AI-generated cheering sounds to the device.
[1488] When User A shakes the remote cheering device during a match, the vibration and motion data from the device is sent to the server via the terminal. The server analyzes the data and generates cheering sounds using a generation AI. The generated cheering sounds are played over the stadium's sound system, allowing on-site spectators to hear the cheers.
[1489] Furthermore, the emotion engine recognizes User A's emotions, and if User A is in an excited state, for example, louder cheers and cheering sounds are generated. This allows User A to feel an even stronger sense of presence and unity with the local area.
[1490] A microphone installed at the viewing location captures the sounds from the venue in real time and sends them to User A's device via a server. This allows User A to experience the cheers and excitement of the venue in real time through earphones. This system allows remote spectators to feel as if they are actually there, without feeling any physical distance.
[1491] Example 2
[1492] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1493] Current remote viewing systems have difficulty fully reproducing the cheers and excitement of the stadium, meaning that remote spectators cannot fully experience the sense of presence and unity of being at the stadium. Furthermore, there is a lack of technology to effectively generate cheering sounds that utilize spectators' emotions and input signals from remote cheering devices.
[1494] The specific processing by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes a means for initializing the generative model and setting necessary parameters, a means for receiving live video and audio feeds of the game and detecting specific events, a means for transmitting vibration and motion data from the remote cheering device to the server via the communication terminal, a means for using an emotion engine that recognizes the user's emotions, and a means for transmitting the user's emotion data from the communication terminal to the server and reflecting it in sound generation. This allows the cheers and excitement of the viewing location to be reproduced in real time, allowing remote spectators to feel the presence of the venue. Furthermore, by utilizing the spectator's input signals and emotion data, more personalized cheering sounds can be generated.
[1495] A "generative model" is an algorithm that uses machine learning and artificial intelligence techniques to generate new data or patterns based on specific input data.
[1496] "Communication terminal" refers to a device such as a smartphone, tablet, or PC that can send and receive data via the Internet or other networks.
[1497] "Input signals" refer to data and commands sent from spectators or devices to the server, including vibration data, motion data from remote cheering devices, and user emotional data.
[1498] A "server" is a computer system that provides specific services or functions over a network, including data processing, storage, and management.
[1499] "Remote cheering device" refers to a device used by spectators to cheer from a remote location. Specifically, it includes stick-shaped or handheld devices that transmit vibration and motion data to a server.
[1500] An "emotion engine" refers to software or algorithms that analyze a user's facial expressions, tone of voice, physical movements, etc. to recognize emotions, and then generate appropriate responses in real time based on that data.
[1501] "Sound equipment" refers to speakers and sound systems that output the generated cheering sounds and cheers and reproduce them audibly.
[1502] "Live Game Video and Audio Feed" means a data stream that transmits real-time video and audio from the location where the Game is being played.
[1503] "User Emotion Data" refers to data that indicates the user's emotional state analyzed by the emotion engine. This data is used in the speech generation process using the generative model.
[1504] The present invention is a system that uses a generative model to reproduce the cheers and excitement of the spectator's location, providing a realistic experience for remote spectators. This system is realized by a server, communication terminals, remote cheering devices, audio equipment, and an emotion engine.
[1505] Hardware and software used
[1506] The system includes the following hardware and software:
[1507] 1. Server:
[1508] Computer system for executing generative AI models
[1509] Storage for caching past match data and audience reaction data
[1510] Network equipment for receiving and analyzing live video and audio feeds
[1511] 2. Communication terminal:
[1512] Smartphones, tablets, and computers used by spectators
[1513] Connectivity for receiving and sending data from remote rooted devices to the server
[1514] Speakers and earphones for playing audio data from the server
[1515] 3. Remote rooting devices:
[1516] A handheld device that spectators wave to cheer on the game.
[1517] A function that generates vibration and motion data and sends it to a communication device
[1518] 4. Sound equipment:
[1519] A speaker system that outputs cheering sounds and cheers within the stadium
[1520] 5. Emotion Engine:
[1521] Software that recognizes emotions by analyzing the user's facial expressions and tone of voice
[1522] A function to send the recognized emotion data to the server
[1523] Data processing and calculation
[1524] The main processes of this system are:
[1525] Initialize the generative AI model:
[1526] The server loads the generative AI model and sets the necessary parameters, and caches past match data and audience reaction data to prepare for real-time processing.
[1527] Receive live video and audio feeds:
[1528] The server receives live video and audio feeds of the match in real time and detects specific events.
[1529] Voice generation:
[1530] The server uses a generative model to generate cheers and sounds based on the detected events and streams them to the communication device.
[1531] Processing input data from a remote rooting device:
[1532] Vibration and motion data from the remote cheering device is transmitted to a server via a communication terminal, and cheering sounds are generated.
[1533] Use of emotion data:
[1534] The system analyzes the user's facial expressions and tone of voice, and sends the emotional data recognized by the emotion engine to the server, which then dynamically changes the intensity and content of the generated voice based on this data.
[1535] Specific examples
[1536] For example, consider the case where User A is watching a soccer match remotely. When User A launches the app and selects a match, the app connects to the server and streams the match video and AI-generated cheering sounds to the device.
[1537] When User A shakes the remote cheering device during a match, the vibration and motion data from the device is sent to the server via the terminal. The server analyzes the data and generates cheering sounds using a generation AI. The generated cheering sounds are played through the sound system at the stadium, so that on-site spectators can also hear the cheers.
[1538] Furthermore, the emotion engine recognizes User A's emotions, and if User A is in an excited state, for example, louder cheers and cheering sounds are generated. This allows User A to feel an even stronger sense of presence and unity with the local area.
[1539] Prompt Sentence Examples
[1540] An example of a prompt to be input to a generative AI model is written as follows:
[1541] "When a goal is scored during a game, how do we generate audio that synchronizes with the cheers of the crowd?"
[1542] "How to change the intensity of the cheering sounds generated when the user is recognized as excited"
[1543] "What kind of cheering sound should be generated based on the vibration data received from the remote cheering device?"
[1544] This allows the entire system to work together, providing spectators with a realistic experience in real time.
[1545] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1546] The flow of this system's program processing
[1547] Step 1:
[1548] Initializing a generative AI model
[1549] Server loads the model: When the server starts up, it loads the generative AI model from disk and loads it into memory, making the model immediately available for use.
[1550] Server sets parameters: The server sets the hyperparameters required for the model to operate, including the learning rate and batch size. The set parameters are important because they directly affect the model's performance.
[1551] Cache past data: The server caches past match data and audience reaction data, which speeds up real-time processing. Cached data improves response time when an event occurs.
[1552] Input: Generative AI model and configuration parameters from disk
[1553] Output: A generative AI model that can be run in memory
[1554] Step 2:
[1555] Launch the app and select a match
[1556] User launches app: The spectator launches the dedicated app on their smartphone or tablet. The app displays the home screen and shows a list of available matches.
[1557] User selects a game: The user selects the game they want to watch from the list. This selection information is sent from the communication device to the server, which then prepares the corresponding live feed.
[1558] Server accepts connection request: The server accepts the user's connection request and performs authentication. If authentication is successful, the server sends a connection establishment notification to the user's terminal.
[1559] Input: User's match selection information
[1560] Output: Connection establishment and match information on the server side
[1561] Step 3:
[1562] Receiving and analyzing live feeds
[1563] Server receives live video and audio: The server receives live video and audio feeds from the stadium in real time. These feeds are transmitted via a communications network.
[1564] Server detects specific events: The server analyzes and detects specific events that occur during the match (e.g., goals, falls). This data is fed into a generative AI model and used as the basis for generating cheering sounds and cheers.
[1565] Input: Live video and audio feed
[1566] Output: Detected event data
[1567] Step 4:
[1568] Real-time voice generation
[1569] Server generates sounds: Cheering and cheering sounds are generated by a generative AI model based on detected event data. The generated sounds change dynamically depending on the situation of the match.
[1570] The server encodes the generated audio: the encoded audio data is prepared for streaming, providing high-quality audio to the user in real time.
[1571] Input: Detected event data
[1572] Output: Generated cheering and cheering sound data
[1573] Step 5:
[1574] Streaming generated audio
[1575] The server streams audio data to the device: The server sends encoded audio data in real time to the user's communication device, which receives the data and plays it through its built-in speaker or earphones.
[1576] Input: Generated audio data
[1577] Output: Streamed audio data
[1578] Step 6:
[1579] Sending data from a remote rooting device
[1580] The user operates the device: The user performs cheering actions using the remote cheering device. The device generates vibration and motion data and transmits it to the communication terminal.
[1581] The terminal transmits the data to the server: The communication terminal transmits the data received from the device to the server, which analyzes the data and generates additional cheering sounds.
[1582] Input: Vibration and motion data from the remote rooting device
[1583] Output: Data sent to the server
[1584] Step 7:
[1585] Cheering sound generation
[1586] The server analyzes the data: vibration and motion data is analyzed, and a generative AI model is used to generate cheering sounds, which are then played over the on-site sound system.
[1587] Input: Vibration and motion data
[1588] Output: Generated cheer sound
[1589] Step 8:
[1590] Audio recording from the viewing area
[1591] The server captures the audio: Microphones installed in the viewing area capture the audio from the venue and send it to the server.
[1592] The server streams the audio to the device: The captured audio is sent to the communication device and played back in real time. This process is important for conveying the cheers and realism of the actual event to the user.
[1593] Input: Audio data from microphones installed in the viewing area
[1594] Output: Audio data streamed to the user's device
[1595] Step 9:
[1596] Analysis and use of emotional data
[1597] User device collects emotional data: The emotion engine analyzes the user's facial expressions and tone of voice to collect emotional data.
[1598] The device sends emotional data to the server: The analyzed emotional data is sent to the server in real time. The server analyzes this data and dynamically changes the intensity and content of the generated voice.
[1599] Input: User facial expressions and tone of voice
[1600] Output: Emotion data sent to the server.
[1601] (Application example 2)
[1602] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1603] Current remote viewing systems have the problem of being unable to fully reproduce the sense of presence and unity that users feel at a real viewing location. In particular, it is difficult to experience the cheers and sounds of the fans at the venue in real time, and they are unable to generate cheering sounds that reflect the emotions of the remote spectators, limiting the viewing experience. Furthermore, they lack the functionality to generate cheering sounds using data from remote cheering devices. Technology that solves these issues is needed.
[1604] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1605] In this invention, the server includes means for reproducing the cheers and excitement of the spectator location using a generative model, means for transmitting input signals from spectators using a communication terminal to the server, means for analyzing the input signals, generating cheering sounds, and outputting them from a sound device, means for recording audio from the spectator location and transmitting the audio to the communication terminal, means for playing the audio on the communication terminal, means for recognizing the emotions of spectators using an emotion engine and transmitting the emotion data to the server in real time, and means for dynamically changing the intensity and content of the cheering sounds based on the emotion data. This allows remote spectators to experience a sense of realism similar to that of being at the venue, and cheering sounds that correspond to their emotions can be generated.
[1606] A "generative model" is a machine learning model that uses artificial intelligence to generate new data.
[1607] A "spectator venue" is the location where a sport or event actually takes place and where spectators physically gather.
[1608] "Cheers and enthusiasm" refers to the cheers and support emitted by spectators at the viewing venue, and are the sounds and atmosphere that indicate the excitement of the event.
[1609] A "communication terminal" is an electronic device used to send and receive data over the Internet, including smartphones and tablets.
[1610] "Input signals from spectators" refer to data that spectators send to the server via their communication terminals, and are signals that reflect the actions and emotions of the spectators.
[1611] A "server" is a computer system that processes and manages data over a network.
[1612] "Cheering sounds" are sounds generated based on input signals from spectators and generative models, and include sounds of cheering and cheering.
[1613] "Audio device" refers to equipment for outputting sound, including speakers and earphones.
[1614] An "emotion engine" is software that analyzes and recognizes the user's emotions from their facial expressions and movements.
[1615] "Real time" is a time concept that refers to instantaneous reaction and processing without delay.
[1616] A "remote cheering device" is a device that spectators use to cheer from home or another location, and has the ability to transmit vibration and motion data to a server.
[1617] The system for implementing this invention includes a generative AI model, an emotion engine, a communication terminal, a server, and an audio device. The operation of the entire system is as follows.
[1618] First, the user launches the remote viewing application on their communication device (smartphone or tablet). Within the application, the user selects the game or event they want to watch. The communication device then sends the selection information to the server, which then receives it.
[1619] The server receives live game feeds (video and audio) and inputs them into a generative AI model to generate cheering and support sounds for the spectators in real time. The generated audio is encoded, streamed to a communication device, and played back to the user in real time. The generative AI model then detects specific events (e.g., a goal or a foul) and dynamically generates audio based on those events.
[1620] Furthermore, the communication terminal is equipped with an emotion engine that recognizes the user's emotions. The emotion engine analyzes the user's facial expressions and tone of voice and transmits the recognized emotion data to the server in real time. The server analyzes this emotion data and dynamically changes the intensity and content of the cheering sounds according to the user's emotions. For example, if the user is excited, a louder cheer will be generated.
[1621] Users can also use a remote cheering device. Vibration and motion data from the remote cheering device is sent to a communication terminal, which then transmits the data to a server. The server analyzes the data and generates cheering sounds using a generative AI model, which are then played on the on-site sound equipment.
[1622] As a concrete example, consider the case where a user is watching a soccer match remotely. When the user launches the application and selects a match, live video and cheering sounds generated by a generative AI are streamed to the communication terminal. When the user shakes the remote cheering device during the match, vibration and motion data from the device are sent to the server via the communication terminal. The server analyzes the data and generates cheering sounds using a generative AI model. These cheering sounds are then played over the sound equipment at the viewing location.
[1623] An example of this prompt would be:
[1624] "A user is watching a soccer match remotely. The emotion engine on the user's smartphone recognizes the user's facial expressions and tone of voice, and the generation AI generates the cheering sounds of the local crowd in real time accordingly. As the user's excitement increases, the generation AI generates louder cheers and cheering sounds, providing a realistic spectator experience."
[1625] As described above, using this system allows remote spectators to experience the same sense of realism as if they were in person, and cheering sounds can be generated to suit their emotions, greatly improving the sense of realism and unity of remote spectators.
[1626] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1627] Step 1:
[1628] The server initializes the generated AI model.
[1629] Input: Past match data and audience reaction data
[1630] Output: Initialized generative AI model
[1631] At startup, the server loads the generative AI model, sets the necessary parameters, reads the model's training data, and prepares it for real-time processing.
[1632] Step 2:
[1633] The user starts the application on the communication terminal and selects the game they want to watch.
[1634] Input: User's match selection information
[1635] Output: Match selection information sent to server
[1636] The user launches the application and selects the game they want to watch. The selection information is sent to the server, which receives it.
[1637] Step 3:
[1638] The server receives a live feed of the game and uses a generative AI model to generate the cheers and support sounds of the spectators.
[1639] Input: Live video and audio feed
[1640] Output: Generated cheers and cheering sounds
[1641] The server receives live video and audio feeds of the game, feeds them into a generative AI model, and generates cheers and cheering sounds in real time, which are then encoded and streamed to communication devices.
[1642] Step 4:
[1643] The communication terminal reproduces the generated cheers and cheering sounds in real time.
[1644] Input: Audio data streamed from the server
[1645] Output: Real-time cheers and cheering sounds played
[1646] The communication terminal decodes the received audio data and plays it back in real time, allowing the user to experience realistic audio.
[1647] Step 5:
[1648] The emotion engine recognizes the user's emotions and transmits the emotion data to the server in real time.
[1649] Input: User facial expressions and tone of voice
[1650] Output: Emotion data sent to the server
[1651] The emotion engine analyzes the user's facial expressions and tone of voice on the communication device to recognize their emotions, and the recognized emotion data is sent to the server.
[1652] Step 6:
[1653] The server analyzes the emotional data and uses a generative AI model to dynamically change the intensity and content of the cheering sounds.
[1654] Input: User emotion data
[1655] Output: Cheering sounds with dynamically changing intensity and content
[1656] The server analyzes the received emotional data and dynamically changes the intensity and content of the cheering sounds based on the user's emotions using a generative AI model. For example, if the user is excited, a louder cheer will be generated.
[1657] Step 7:
[1658] The user uses the remote support device and transmits the vibration data and motion data to the communication terminal.
[1659] Input: Vibration and motion data from the remote rooting device
[1660] Output: Data sent to the communication device
[1661] Users cheer by shaking the remote cheering device, and vibration and motion data from the device are sent to the communication terminal.
[1662] Step 8:
[1663] The communication terminal transmits the data received from the remote support device to the server.
[1664] Input: Vibration and motion data from the remote rooting device
[1665] Output: Device data sent to the server
[1666] The communication terminal transmits the received vibration data and motion data to the server.
[1667] Step 9:
[1668] The server analyzes the data from the remote cheering device and generates cheering sounds using a generative AI model.
[1669] Input: Vibration and motion data
[1670] Output: Generated cheer sound
[1671] The server analyzes the data from the remote cheering device and generates cheering sounds using a generative AI model. The generated cheering sounds are then played over the sound equipment at the viewing area.
[1672] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1673] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1674] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1675] [Fourth embodiment]
[1676] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1677] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1678] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1679] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1680] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1681] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1682] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1683] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1684] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1685] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1686] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1687] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1688] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1689] This system uses a generative model to recreate the cheers and excitement of the spectator's location, providing a realistic experience for remote spectators. The system is primarily composed of a server, communication terminals, remote cheering devices, and audio equipment.
[1690] Program processing and explanation
[1691] Initializing a generative AI model
[1692] 1. The server initializes the generative AI model
[1693] On startup, the server loads the generative AI model and sets the necessary parameters, so the model is ready to generate cheers and other acoustic patterns for the spectator area.
[1694] Launching the app and connecting
[1695] 2. The user launches the app and selects a match.
[1696] Users launch the app on their smartphone or other communication device and select the game they want to watch. Once the selection is complete, the device sends a connection request to the server.
[1697] 3. The server establishes the connection
[1698] The server receives a connection request from the user's device, establishes the connection through an authentication process, notifies the device that the connection is successful, and the user can begin watching the game.
[1699] Real-time analysis of match data
[1700] 4. The server receives and analyzes the live match feed
[1701] The server receives live video and audio feeds of the match in real time, allowing it to detect events that occur during the match (e.g., goals scored, falls) and input them into the generative model.
[1702] AI-powered voice generation
[1703] 5. The server generates real-time audio using the generative model
[1704] The server uses a generative model to generate cheers and other supportive sounds based on detected events and local audio feeds.
[1705] 6. The server streams the generated audio
[1706] The server encodes the generated audio data and streams it to the communication terminal, which plays the audio in real time.
[1707] Remote cheering bat data transmission
[1708] 7. User uses a remote rooting device
[1709] The user starts cheering by shaking the remote cheering device, which then transmits vibration and motion data to the communication terminal.
[1710] 8. The device sends the support data to the server
[1711] The communication terminal sends the received data from the cheering device to the server, which analyzes the data and uses it to generate cheering sounds.
[1712] Cheering sound generation and output
[1713] 9. The server analyzes the cheering data and generates voice
[1714] The server analyzes the cheering data and generates cheering sounds using a generative model, which are then played from the sound equipment in the stadium.
[1715] 10. Record audio from the viewing area and send it to a communication device
[1716] Microphones installed in the viewing area capture audio and transmit it to a server, which then streams it to the communication device and plays it back on the device.
[1717] Specific examples
[1718] For example, consider the case where User A is watching a soccer match remotely. When User A launches the app and selects a match, the app connects to the server and streams the match video and AI-generated cheering sounds to the device.
[1719] When User A shakes the remote cheering device during a match, the vibration and motion data from the device is sent to the server via the terminal. The server analyzes the data and generates cheering sounds using a generation AI. The generated cheering sounds are played over the stadium's sound system, allowing on-site spectators to hear the cheers.
[1720] Furthermore, microphones installed in the viewing area capture the sounds from the venue in real time and transmit them via a server to User A's device, allowing User A to experience the cheers and excitement of the crowd in real time through earphones.
[1721] The processing flow will be explained below.
[1722] Step 1:
[1723] The server initializes the generated AI model.
[1724] At startup, the server loads the AI model and configures the model's parameters, which includes loading the model's training data.
[1725] The server caches past match data and audience reaction data into the AI model, preparing it for real-time processing.
[1726] Step 2:
[1727] The user launches the app and selects a match
[1728] The user launches the app on their smartphone or tablet.
[1729] Select the game you want to watch from the app's main screen and click the "Start Watching" button.
[1730] Step 3:
[1731] The device sends a connection request to the server
[1732] The user's terminal sends a request to start watching to the server.
[1733] The device connects to the server's API endpoint and sends authentication information.
[1734] Step 4:
[1735] The server establishes the connection and performs authentication
[1736] The server receives the connection request and authenticates the user.
[1737] Once the authentication is complete, the server establishes a session and notifies the terminal of the successful connection.
[1738] Step 5:
[1739] The server receives a live feed of the match
[1740] The server receives live video and audio feeds of the match in real time.
[1741] The server analyzes this data and prepares to detect important events (e.g. goals, falls).
[1742] Step 6:
[1743] The server detects an important event
[1744] The server analyzes the video and audio of the match in real time to detect specific events.
[1745] When a specific event such as a goal or fall is detected, that information is passed on to the generating AI.
[1746] Step 7:
[1747] The server generates the voice using AI
[1748] The server uses a generative AI model to generate cheers and sounds based on the detected events.
[1749] The generated audio data is encoded in real time.
[1750] Step 8:
[1751] The server streams the generated audio to the device
[1752] The server streams the encoded audio data to the user terminal.
[1753] The user terminal plays back the received audio in real time.
[1754] Step 9:
[1755] User uses remote rooting device
[1756] Users cheer by waving the remote cheering device.
[1757] Vibration and motion data from the device is sent to the user's terminal.
[1758] Step 10:
[1759] The device sends the support data to the server.
[1760] The user's terminal transmits the data received from the remote support device to the server.
[1761] The data sent to the server includes vibration data and motion data.
[1762] Step 11:
[1763] The server analyzes the support data
[1764] The server analyzes the received support data in real time.
[1765] Based on the analysis results, the server determines what kind of cheering sound to generate.
[1766] Step 12:
[1767] The server generates cheering sounds using AI
[1768] The server uses a generative AI model to generate cheering sounds, including applause and cheers.
[1769] The generated cheering sound is encoded and transmitted to an audio device.
[1770] Step 13:
[1771] Play cheering sounds from the sound system
[1772] The cheering sound data transmitted from the server is received by the sound device.
[1773] The sound system will then play the received cheering sounds in real time, allowing on-site spectators to feel the cheering from the remote area.
[1774] Step 14:
[1775] Record audio from the viewing location and send it to your device
[1776] The server captures audio from microphones installed at the viewing area and transmits it to the server in real time.
[1777] The server encodes the captured audio for streaming to the user terminal.
[1778] Step 15:
[1779] The device plays audio from the viewing area
[1780] The user terminal receives the streamed audio data and plays it back in real time.
[1781] Users can experience the cheers and atmosphere of the venue in real time.
[1782] Example 1
[1783] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1784] The aim is to solve the problem that it is difficult for remote spectators to experience the sense of presence and cheers of being at the venue in real time, and that cheers from remote cheering devices are not easily conveyed to on-site spectators.
[1785] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1786] In this invention, the server includes a means for recreating the cheers and excitement of the spectator location using a generative model, a means for receiving and analyzing a live feed of the game and detecting events, and a means for encoding audio data and streaming it to a communication terminal, allowing remote spectators to experience realistic cheers in real time and for cheering via cheering devices to be reflected at the venue.
[1787] A "generative model" is an artificial intelligence technology that generates new data and information based on learned algorithms and data.
[1788] "Spectator venue" refers to a local venue or stadium where a sporting event, concert, etc. is held.
[1789] "Cheering and excitement" refers to the vocal and emotional excitement of spectators expressing their excitement and support for a game or event.
[1790] A "communication terminal" is an electronic device that can send and receive data over a network, such as a smartphone, tablet PC, or personal computer.
[1791] A "server" is a high performance computing device for storing, processing, and distributing data over a network.
[1792] An "input signal" is an electrical signal that contains instructions or data from a user or device.
[1793] "Cheering sounds" are sounds and sound effects made by spectators when cheering on a game or event.
[1794] "Sound equipment" means equipment such as speakers and amplifiers for reproducing sound.
[1795] "Recording" is the act of recording sound in digital or analog form.
[1796] "Streaming" refers to the technology of continuously transmitting and playing data in real time over the Internet.
[1797] A "remote support device" is a device that enables support from a remote location, and has the function of transmitting vibration data and motion data mainly to a communication terminal.
[1798] This system uses a generative model to recreate the cheers and excitement of the spectator's location, providing a realistic experience for remote spectators. The system mainly consists of a server, communication terminals, remote cheering devices, and audio equipment.
[1799] The server first initializes the generative AI model by loading a pre-trained model file into memory using a machine learning library such as TensorFlow or PyTorch. It then sets the necessary parameters and prepares the model for operation, ready to generate cheers and other acoustic patterns for the spectators.
[1800] Users launch a dedicated app on their communication device, such as a smartphone or tablet, and select the game they want to watch. Once a game is selected, the device sends the selection information to the server and requests a connection to the server. The server receives the connection request from the user device and performs an authentication process based on the authentication information. If authentication is successful, the session begins and the user can begin remote viewing.
[1801] The server receives live video and audio feeds of the game in real time. It then analyzes the data to detect goals and important events. The results of this analysis are converted into a format that can be used by the generative AI model and fed into it. The resulting audio data is then encoded and streamed to the user's device.
[1802] Users begin cheering by shaking the remote cheering device. This device transmits vibration and motion data to a communication terminal. The communication terminal then transmits this data to a server, which then generates cheering sounds based on that data. The generated cheering sounds are played from the sound equipment in the stadium, allowing on-site spectators to hear the cheers.
[1803] Furthermore, microphones installed in the viewing area capture the sound from the venue in real time and transmit it to a server, which then streams the sound to communication devices, allowing users to experience the cheers and excitement of the venue in real time.
[1804] As a concrete example, consider the case where User A is watching a soccer match remotely. When User A launches the app and selects a match, the app connects to the server, and game footage and cheering sounds generated by AI are streamed to the device. When User A shakes the remote cheering device during the match, the vibration and motion data from the device are sent to the server, and cheering sounds are generated and played from the on-site sound equipment. In addition, audio captured at the venue is sent to the device in real time, allowing User A to enjoy the immersive experience through earphones.
[1805] Examples of prompts for a generative AI model include:
[1806] "In order to generate audio that reproduces the cheers and excitement of the viewing area, we provide the following information:
[1807] 1. The crowd's reaction when a goal is scored
[1808] 2. The crowd booing when the fall happened
[1809] 3. The crowd roaring during halftime
[1810] Based on this, generate realistic cheer sounds in real time.
[1811] This system will enable remote spectators to experience the excitement and cheers of being at the venue in real time, improving the viewing experience.
[1812] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1813] Step 1: Initializing the generative AI model
[1814] The server initializes the generated AI model.
[1815] Input: trained model file, required parameters
[1816] Output: Initialized generative AI model
[1817] What it does: When the server starts up, it uses machine learning libraries like TensorFlow and PyTorch to load a trained generative AI model into memory, reads the model's parameters (e.g., settings for generating voices), and prepares the model for operation, ready to generate cheers and other acoustic patterns for the spectators.
[1818] Step 2: Launch the app and connect
[1819] The user launches the app and selects a match
[1820] Input: Select the match you want to watch
[1821] Output: A connection request is sent
[1822] Specific operation: The user launches the dedicated app on their smartphone or tablet and selects the game they want to watch from a list on the main screen. Once a game is selected, the app sends a connection request to the server with the selected information.
[1823] The server establishes the connection
[1824] Input: Connection request, user credentials
[1825] Output: Connection established notification
[1826] Specific operation: The server receives a connection request from a user terminal and performs an authentication process based on authentication information (e.g., user ID and password). If authentication is successful, the server starts a session and notifies the communication terminal that the connection has been established.
[1827] Step 3: Real-time analysis of match data
[1828] The server receives and analyzes the live match feed
[1829] Input: Live video and audio feed of the match
[1830] Output: Event data (scoring scenes, falls, etc.)
[1831] How it works: The server receives live video and audio feeds from live game broadcasts in real time. It analyzes the input video and audio data to detect goals and important events. The analysis results are then converted into a format that can be used by the generative AI model.
[1832] Step 4: AI-powered voice generation
[1833] The server generates real-time audio using the generative model
[1834] Input: Event data, local audio feed
[1835] Output: Generated audio data
[1836] How it works: The server uses a generative AI model to generate realistic cheering and cheering sounds based on the detected event data and local audio feeds. The generated audio data is then encoded and prepared for transmission to the communication device.
[1837] The server streams the generated audio
[1838] Input: Generated audio data
[1839] Output: Audio stream
[1840] Specific operation: The server streams the encoded audio data to the user's communication device in real time, allowing the user to play back the realistic audio generated on the device.
[1841] Step 5: Send data to your remote rooting device
[1842] User uses remote rooting device
[1843] Input: Vibration data and motion data of the remote rooting device
[1844] Output: Send data to a communication terminal
[1845] Specific operation: A user expresses his / her support by shaking a remote support device (e.g., a Bluetooth-connected support bat). This device transmits vibration data and motion data to a communication terminal.
[1846] The device sends the support data to the server.
[1847] Input: Vibration data and motion data of the remote rooting device
[1848] Output: Send data to the server
[1849] Specific operation: The communication device sends the received vibration and motion data to the server, which receives this data, analyzes it, and uses it to generate cheering sounds.
[1850] Step 6: Generate and output cheer sounds
[1851] The server analyzes the cheering data and generates voice
[1852] Input: Remote rooting device data
[1853] Output: Generated cheer sound
[1854] Specific operation: The server analyzes the received data and generates sounds according to the intensity and frequency of the cheering. The generated cheering sounds are sent to the sound equipment in the stadium and played back.
[1855] Record audio from the viewing location and send it to a communication device
[1856] Input: Audio data captured at the viewing location
[1857] Output: Audio stream to communication device
[1858] How it works: Microphones installed in the viewing area capture audio in real time. The captured audio data is sent to a server and then encoded and transmitted to a communication device. Users can play this audio on their device, allowing them to experience the realism of being at the venue.
[1859] (Application example 1)
[1860] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1861] Conventional remote viewing systems make it difficult to experience the cheers and excitement of the fans in real time, and they lack the interactivity of cheering. Furthermore, there is a lack of a way to efficiently analyze data from remote cheering devices and generate realistic cheering sounds. There is a need to solve these issues and provide a more realistic viewing experience.
[1862] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1863] In this invention, the server includes means for reproducing the cheers and excitement of the spectator location using a generative model, means for transmitting input signals from spectators to the server using a communication device, means for analyzing the input signals to generate cheering sounds and outputting them from a sound device, means for recording audio from the spectator location and transmitting it to the communication device, means for playing the audio on the communication device, means for performing real-time event analysis and generating cheers using a generative model, and means for acquiring vibration data and motion data from the cheering device and transmitting it to the server. This provides remote spectators with the same sense of realism as if they were watching the game in person, allowing users to send their cheers interactively through their devices.
[1864] A "generative model" is a model that uses artificial intelligence technology to generate specific patterns, sounds, images, etc.
[1865] A "spectator location" is a local location where entertainment or competition, such as a sporting event or concert, takes place.
[1866] "Cheers" are the cheers and excited voices that spectators make at the viewing venue in response to a game or performance.
[1867] "Excitement" refers to the atmosphere of excitement and enthusiasm created by the audience.
[1868] A "communication device" is a device for sending and receiving information, including smartphones, tablets, and computers.
[1869] A "spectator" is someone who watches events such as sports or concerts remotely or in person.
[1870] An "input signal" is data or instructions sent by a spectator through a communication device.
[1871] "Cheering sounds" are sounds generated in response to the cheering actions of spectators.
[1872] "Audio device" means a device for reproducing sound, including speakers and headphones.
[1873] "Real-time event analysis" is a process that instantly analyzes the status and behavior of an event currently in progress.
[1874] "Vibration data" is data recorded of the movement of the support device when it vibrates.
[1875] "Motion data" is data that records the movement of the support device.
[1876] A "server" is a computer that manages and provides information over a network.
[1877] The system embodying this invention uses a generative AI model to recreate the cheers and excitement of the spectator's location, providing a sense of realism to remote spectators. The system is primarily composed of a server, communication equipment, a remote cheering device, and an audio device.
[1878] 1. System Program
[1879] The system's program is built around a generative AI model and mainly performs the following processes:
[1880] 1. Initializing the generative AI model
[1881] The server loads the generative AI model and sets the necessary parameters, so the model is ready to generate cheers and other sound patterns for the spectator area.
[1882] 2. Connecting communication devices
[1883] The user starts up the communication device and selects the event they want to watch. The communication device sends a connection request to the server, and the server establishes the connection. If the connection is successful, the user can start watching.
[1884] 3. Event Data Analysis
[1885] The server receives live video and audio feeds of the event in real time, detects and analyzes specific events, and feeds the results into a generative AI model.
[1886] 4. Real-time speech generation
[1887] The server uses a generative AI model to generate cheers and cheering sounds in real time, and this generated audio data is sent to a communication device and played back in real time.
[1888] 5. Use of remote support devices
[1889] When a user shakes the remote cheering device, the device transmits vibration and motion data to the communication device, which then analyzes the data and generates cheering sounds.
[1890] 6. Data transmission and reception between the server and communication device
[1891] The server plays the generated cheering sounds from the on-site sound equipment and records the audio from the viewing location and transmits it to the communication device, which then plays back the audio, allowing the user to experience the cheers and realism of the venue.
[1892] 2. Hardware and Software Details
[1893] Hardware
[1894] Server: Run the system on a high-performance server (e.g., AWS EC2).
[1895] Communication devices: Common devices such as smartphones and tablets.
[1896] Remote rooting device: An IoT device for capturing user actions.
[1897] Sound device: An audio playback device such as a speaker or headphones.
[1898] software
[1899] Generative AI model: For example, we use OpenAI's GPT-3 model for speech generation.
[1900] Communication and Data Processing: Build WebSocket server and client functions using Python.
[1901] 3. Specific Examples
[1902] For example, if a user wants to watch a live concert remotely, the process would be as follows:
[1903] 1. A user turns on a communication device and selects a live concert. The communication device connects to a server and receives a real-time live feed.
[1904] 2. The server analyzes the live feed and detects specific events (e.g., highlights or excitement).
[1905] 3. The server uses a generative AI model to generate cheers and cheering sounds based on these events.
[1906] 4. When the user shakes the remote cheering device, the motion data is sent to the server via the communication device, and a cheering sound is generated.
[1907] 5. The generated audio is played on the communication device and also played on-site.
[1908] Prompt Sentence Examples
[1909] 1. Initializing the generative AI model
[1910] "Load a generative model and generate cheers and cheers in real time."
[1911] 2. User's live concert viewing
[1912] "Users simply launch the app, select a specific live concert, and watch it. The live feed, including the sounds of cheering and cheering from the venue, is streamed to their smartphone."
[1913] 3. Use a remote support device
[1914] "When a user remotely shakes a cheering device, the vibration and motion data from the device is sent to the server, which generates cheering sounds that can be played by other spectators watching in real time."
[1915] In this way, this invention allows users to enjoy a realistic viewing experience even from a remote location.
[1916] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1917] Step 1: Initializing the generative AI model
[1918] The server loads the generative AI model and sets the necessary parameters, so it is ready to generate cheers and other sound patterns for the spectator area. It takes the model's reference path as input and the generative AI model initialization as output.
[1919] Step 2: Connecting communication devices
[1920] A user starts up a communication device and selects the event they want to watch. The communication device sends a connection request to the server, and the server establishes the connection. The input is the event information selected by the user, and the output is the established connection.
[1921] Step 3: Receive live feeds and parse event data
[1922] The server receives live video and audio feeds of the event in real time. It detects and analyzes specific events (e.g., goals scored). This process has the live feed data as input and the analyzed event data as output.
[1923] Step 4: Real-time speech generation
[1924] The server uses a generative AI model to generate cheers and cheering sounds in real time. This generated audio data is sent to a communication device. The input is event data, and the output is generated audio data.
[1925] Step 5: Encode and transmit the audio data
[1926] The server encodes the generated audio data and streams it to the communication device, with the generated audio data as input and the encoded audio data as output.
[1927] Step 6: Real-time audio playback
[1928] The communication device plays back the received audio data in real time, with the encoded audio data as input and the audio being played back as output.
[1929] Step 7: Send data to your remote rooting device
[1930] The user cheers by shaking the remote cheering device. The device collects vibration and motion data and sends it to the communication device. The motion data is input and sent to the communication device as output.
[1931] Step 8: Sending remote rooting data to the server
[1932] The communication device receives data from the remote support device and transmits it to the server, which has the remote support data as input and transmits the data to the server as output.
[1933] Step 9: Analyzing cheering data and generating cheering sounds
[1934] The server analyzes the remote cheering data and generates cheering sounds using a generative model. The remote cheering data is input, and the generated cheering sounds are output.
[1935] Step 10: Playing the cheering sounds and sending the audio feed
[1936] The server plays the generated cheering sounds from an audio device and records the audio from the viewing location and sends it to a communication device.The generated cheering sounds and on-site audio are input, and the audio is played back and an audio feed is sent as output.
[1937] Step 11: Playback of local audio via communication device
[1938] The communication device plays back the received local audio. It has received audio data as input and plays back the audio as output.
[1939] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1940] This invention is a system that uses a generative model to reproduce the cheers and excitement of the spectator's location, providing a realistic experience for remote spectators. This system is realized by combining a server, communication terminals, remote cheering devices, audio equipment, and an emotion engine that recognizes the user's emotions.
[1941] Program processing and explanation
[1942] Initializing a generative AI model
[1943] 1. The server initializes the generative AI model
[1944] On startup, the server loads the generative AI model and sets the necessary parameters, which includes loading the model's training data.
[1945] The server caches past match data and audience reaction data into the AI model, preparing it for real-time processing.
[1946] Launching the app and connecting
[1947] 2. The user launches the app and selects a match.
[1948] Users launch the app on their smartphone or tablet and select the game they want to watch. Once the selection is complete, the device sends a connection request to the server.
[1949] 3. The server establishes the connection
[1950] The server receives a connection request from the user's device, establishes the connection through an authentication process, notifies the device that the connection is successful, and the user can begin watching the game.
[1951] Real-time analysis of match data
[1952] 4. The server receives and analyzes the live match feed
[1953] The server receives live video and audio feeds of the match in real time, allowing it to detect events that occur during the match (e.g., goals, falls) and input them into the generative model.
[1954] AI-powered voice generation
[1955] 5. The server generates real-time audio using the generative model
[1956] The server uses a generative model to generate cheers and other supportive sounds based on detected events and local audio feeds.
[1957] 6. The server streams the generated audio to the device
[1958] The server encodes the generated audio data and streams it to the communication terminal, which plays the audio in real time.
[1959] Remote cheering bat data transmission
[1960] 7. User uses a remote rooting device
[1961] Users cheer by shaking the remote cheering device, and vibration and motion data from the device is sent to the user's device.
[1962] 8. The device sends the support data to the server
[1963] The user's device sends the data received from the remote support device to the server, including vibration data and motion data.
[1964] Cheering sound generation and output
[1965] 9. The server analyzes the cheering data and generates voice
[1966] The server analyzes the cheering data and generates cheering sounds using a generative model, which are then played from the sound equipment in the stadium.
[1967] 10. Record audio from the viewing location and send it to your device
[1968] Microphones installed in the viewing area capture audio and transmit it to a server, which then streams it to the communication device and plays it back on the device.
[1969] Using Emotion Data with an Emotion Engine
[1970] 11. The emotion engine recognizes emotions on the user device
[1971] The emotion engine recognizes the user's emotions by analyzing their facial expressions and tone of voice.
[1972] The recognized emotion data is sent to the server in real time.
[1973] 12. Analyze emotional data and reflect it in speech generation
[1974] The server analyzes the received emotion data and dynamically changes the intensity and content of the generated voice based on the user's emotion. For example, if the user is excited, it generates louder cheers.
[1975] Specific examples
[1976] For example, consider the case where User A is watching a soccer match remotely. When User A launches the app and selects a match, the app connects to the server and streams the match video and AI-generated cheering sounds to the device.
[1977] When User A shakes the remote cheering device during a match, the vibration and motion data from the device is sent to the server via the terminal. The server analyzes the data and generates cheering sounds using a generation AI. The generated cheering sounds are played over the stadium's sound system, allowing on-site spectators to hear the cheers.
[1978] Furthermore, the emotion engine recognizes User A's emotions, and if User A is in an excited state, for example, louder cheers and cheering sounds are generated. This allows User A to feel an even stronger sense of presence and unity with the local area.
[1979] A microphone installed at the viewing location captures the sounds from the venue in real time and sends them to User A's device via a server. This allows User A to experience the cheers and excitement of the venue in real time through earphones. This system allows remote spectators to feel as if they are actually there, without feeling any physical distance.
[1980] The processing flow will be explained below.
[1981] Step 1:
[1982] The server initializes the generated AI model.
[1983] On startup, the server loads the generative AI model, sets the necessary parameters, and optimizes it using training data, ready to generate game cheers and sounds.
[1984] Step 2:
[1985] The user launches the app and selects a match
[1986] Users launch the app on their smartphone or tablet and select the game they want to watch. When the user clicks the "Start Watching" button, a connection request is sent from the device to the server.
[1987] Step 3:
[1988] The device sends a connection request to the server
[1989] The user's device sends a connection request to the server's API endpoint, which includes authentication information.
[1990] Step 4:
[1991] The server establishes the connection and performs authentication
[1992] The server receives the connection request and verifies the user's authentication information. If authentication is successful, the server establishes a session and notifies the terminal that the connection is successful.
[1993] Step 5:
[1994] The server receives a live feed of the match
[1995] The server receives live video and audio feeds of the match in real time and prepares to analyze the progress of the match.
[1996] Step 6:
[1997] The server analyzes the match data
[1998] The server analyzes the game footage in real time to detect specific events (e.g., goals, falls), and sends the detected event information to the generative AI model.
[1999] Step 7:
[2000] The server generates the voice using the generative AI model
[2001] The server inputs the detected event information into a generative AI model to generate cheering and cheering sounds in real time, and the generated audio data is encoded.
[2002] Step 8:
[2003] The server streams the generated audio to the device
[2004] The server transmits the encoded audio data in streaming format to the user's device, which plays the audio in real time.
[2005] Step 9:
[2006] User uses remote rooting device
[2007] Users can show their support by shaking the remote cheering device, and vibration and motion data from the device are sent to the user's device.
[2008] Step 10:
[2009] The device sends the support data to the server.
[2010] The user's device then sends the received device data, including vibration and motion data, to the server.
[2011] Step 11:
[2012] The server analyzes the support data
[2013] The server analyzes the received cheering device data and issues instructions to the generation AI model based on the results, which then generates the cheering sound.
[2014] Step 12:
[2015] The server generates cheering sounds using AI
[2016] The server generates cheering sounds based on the analysis results, and the generated audio data is encoded and sent to the stadium's sound system.
[2017] Step 13:
[2018] Play cheering sounds from the sound system
[2019] The sound equipment receives the cheering sound data sent from the server and plays it back in real time, allowing on-site spectators to feel the remote cheering.
[2020] Step 14:
[2021] Record audio from the viewing location and send it to your device
[2022] Microphones installed in the viewing area capture audio in real time and transmit it to a server, which encodes the captured audio for streaming to user devices.
[2023] Step 15:
[2024] The device plays audio from the viewing area
[2025] The user's device receives the streamed audio data and plays it back in real time, allowing the user to experience the cheers and atmosphere of the venue in real time.
[2026] Step 16:
[2027] An emotion engine recognizes emotions on the user's device
[2028] The user's device analyzes the user's facial expressions and tone of voice through a camera and microphone, and recognizes emotions using an emotion engine.
[2029] Step 17:
[2030] The device sends emotion data to the server.
[2031] The user's device transmits the recognized emotion data to the server in real time.
[2032] Step 18:
[2033] The server analyzes the emotional data and reflects it in the speech generation.
[2034] The server analyzes the emotion data and dynamically changes the intensity and content of the generated voices based on the user's emotions, for example, generating louder cheers if the user is excited.
[2035] Specific examples
[2036] For example, consider the case where User A is watching a soccer match remotely. When User A launches the app and selects a match, the app connects to the server and streams the match video and AI-generated cheering sounds to the device.
[2037] When User A shakes the remote cheering device during a match, the vibration and motion data from the device is sent to the server via the terminal. The server analyzes the data and generates cheering sounds using a generation AI. The generated cheering sounds are played over the stadium's sound system, allowing on-site spectators to hear the cheers.
[2038] Furthermore, the emotion engine recognizes User A's emotions, and if User A is in an excited state, for example, louder cheers and cheering sounds are generated. This allows User A to feel an even stronger sense of presence and unity with the local area.
[2039] A microphone installed at the viewing location captures the sounds from the venue in real time and sends them to User A's device via a server. This allows User A to experience the cheers and excitement of the venue in real time through earphones. This system allows remote spectators to feel as if they are actually there, without feeling any physical distance.
[2040] Example 2
[2041] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[2042] Current remote viewing systems have difficulty fully reproducing the cheers and excitement of the stadium, meaning that remote spectators cannot fully experience the sense of presence and unity of being at the stadium. Furthermore, there is a lack of technology to effectively generate cheering sounds that utilize spectators' emotions and input signals from remote cheering devices.
[2043] The specific processing by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes a means for initializing the generative model and setting necessary parameters, a means for receiving live video and audio feeds of the game and detecting specific events, a means for transmitting vibration and motion data from the remote cheering device to the server via the communication terminal, a means for using an emotion engine that recognizes the user's emotions, and a means for transmitting the user's emotion data from the communication terminal to the server and reflecting it in sound generation. This allows the cheers and excitement of the viewing location to be reproduced in real time, allowing remote spectators to feel the presence of the venue. Furthermore, by utilizing the spectator's input signals and emotion data, more personalized cheering sounds can be generated.
[2044] A "generative model" is an algorithm that uses machine learning and artificial intelligence techniques to generate new data or patterns based on specific input data.
[2045] "Communication terminal" refers to a device such as a smartphone, tablet, or PC that can send and receive data via the Internet or other networks.
[2046] "Input signals" refer to data and commands sent from spectators or devices to the server, including vibration data, motion data from remote cheering devices, and user emotional data.
[2047] A "server" is a computer system that provides specific services or functions over a network, including data processing, storage, and management.
[2048] "Remote cheering device" refers to a device used by spectators to cheer from a remote location. Specifically, it includes stick-shaped or handheld devices that transmit vibration and motion data to a server.
[2049] An "emotion engine" refers to software or algorithms that analyze a user's facial expressions, tone of voice, physical movements, etc. to recognize emotions, and then generate appropriate responses in real time based on that data.
[2050] "Sound equipment" refers to speakers and sound systems that output the generated cheering sounds and cheers and reproduce them audibly.
[2051] "Live Game Video and Audio Feed" means a data stream that transmits real-time video and audio from the location where the Game is being played.
[2052] "User Emotion Data" refers to data that indicates the user's emotional state analyzed by the emotion engine. This data is used in the speech generation process using the generative model.
[2053] The present invention is a system that uses a generative model to reproduce the cheers and excitement of the spectator's location, providing a realistic experience for remote spectators. This system is realized by a server, communication terminals, remote cheering devices, audio equipment, and an emotion engine.
[2054] Hardware and software used
[2055] The system includes the following hardware and software:
[2056] 1. Server:
[2057] Computer system for executing generative AI models
[2058] Storage for caching past match data and audience reaction data
[2059] Network equipment for receiving and analyzing live video and audio feeds
[2060] 2. Communication terminal:
[2061] Smartphones, tablets, and computers used by spectators
[2062] Connectivity for receiving and sending data from remote rooted devices to the server
[2063] Speakers and earphones for playing audio data from the server
[2064] 3. Remote rooting devices:
[2065] A handheld device that spectators wave to cheer on the game.
[2066] A function that generates vibration and motion data and sends it to a communication device
[2067] 4. Sound equipment:
[2068] A speaker system that outputs cheering sounds and cheers within the stadium
[2069] 5. Emotion Engine:
[2070] Software that recognizes emotions by analyzing the user's facial expressions and tone of voice
[2071] A function to send the recognized emotion data to the server
[2072] Data processing and calculation
[2073] The main processes of this system are:
[2074] Initialize the generative AI model:
[2075] The server loads the generative AI model and sets the necessary parameters, and caches past match data and audience reaction data to prepare for real-time processing.
[2076] Receive live video and audio feeds:
[2077] The server receives live video and audio feeds of the match in real time and detects specific events.
[2078] Voice generation:
[2079] The server uses a generative model to generate cheers and sounds based on the detected events and streams them to the communication device.
[2080] Processing input data from a remote rooting device:
[2081] Vibration and motion data from the remote cheering device is transmitted to a server via a communication terminal, and cheering sounds are generated.
[2082] Use of emotion data:
[2083] The system analyzes the user's facial expressions and tone of voice, and sends the emotional data recognized by the emotion engine to the server, which then dynamically changes the intensity and content of the generated voice based on this data.
[2084] Specific examples
[2085] For example, consider the case where User A is watching a soccer match remotely. When User A launches the app and selects a match, the app connects to the server and streams the match video and AI-generated cheering sounds to the device.
[2086] When User A shakes the remote cheering device during a match, the vibration and motion data from the device is sent to the server via the terminal. The server analyzes the data and generates cheering sounds using a generation AI. The generated cheering sounds are played through the sound system at the stadium, so that on-site spectators can also hear the cheers.
[2087] Furthermore, the emotion engine recognizes User A's emotions, and if User A is in an excited state, for example, louder cheers and cheering sounds are generated. This allows User A to feel an even stronger sense of presence and unity with the local area.
[2088] Prompt Sentence Examples
[2089] An example of a prompt to be input to a generative AI model is written as follows:
[2090] "When a goal is scored during a game, how do we generate audio that synchronizes with the cheers of the crowd?"
[2091] "How to change the intensity of the cheering sounds generated when the user is recognized as excited"
[2092] "What kind of cheering sound should be generated based on the vibration data received from the remote cheering device?"
[2093] This allows the entire system to work together, providing spectators with a realistic experience in real time.
[2094] The flow of the identification process in the second embodiment will be described with reference to FIG.
[2095] The flow of this system's program processing
[2096] Step 1:
[2097] Initializing a generative AI model
[2098] Server loads the model: When the server starts up, it loads the generative AI model from disk and loads it into memory, making the model immediately available for use.
[2099] Server sets parameters: The server sets the hyperparameters required for the model to operate, including the learning rate and batch size. The set parameters are important because they directly affect the model's performance.
[2100] Cache past data: The server caches past match data and audience reaction data, which speeds up real-time processing. Cached data improves response time when an event occurs.
[2101] Input: Generative AI model and configuration parameters from disk
[2102] Output: A generative AI model that can be run in memory
[2103] Step 2:
[2104] Launch the app and select a match
[2105] User launches app: The spectator launches the dedicated app on their smartphone or tablet. The app displays the home screen and shows a list of available matches.
[2106] User selects a game: The user selects the game they want to watch from the list. This selection information is sent from the communication device to the server, which then prepares the corresponding live feed.
[2107] Server accepts connection request: The server accepts the user's connection request and performs authentication. If authentication is successful, the server sends a connection establishment notification to the user's terminal.
[2108] Input: User's match selection information
[2109] Output: Connection establishment and match information on the server side
[2110] Step 3:
[2111] Receiving and analyzing live feeds
[2112] Server receives live video and audio: The server receives live video and audio feeds from the stadium in real time. These feeds are transmitted via a communications network.
[2113] Server detects specific events: The server analyzes and detects specific events that occur during the match (e.g., goals, falls). This data is fed into a generative AI model and used as the basis for generating cheering sounds and cheers.
[2114] Input: Live video and audio feed
[2115] Output: Detected event data
[2116] Step 4:
[2117] Real-time voice generation
[2118] Server generates sounds: Cheering and cheering sounds are generated by a generative AI model based on detected event data. The generated sounds change dynamically depending on the situation of the match.
[2119] The server encodes the generated audio: the encoded audio data is prepared for streaming, providing high-quality audio to the user in real time.
[2120] Input: Detected event data
[2121] Output: Generated cheering and cheering sound data
[2122] Step 5:
[2123] Streaming generated audio
[2124] The server streams audio data to the device: The server sends encoded audio data in real time to the user's communication device, which receives the data and plays it through its built-in speaker or earphones.
[2125] Input: Generated audio data
[2126] Output: Streamed audio data
[2127] Step 6:
[2128] Sending data from a remote rooting device
[2129] The user operates the device: The user performs cheering actions using the remote cheering device. The device generates vibration and motion data and transmits it to the communication terminal.
[2130] The terminal transmits the data to the server: The communication terminal transmits the data received from the device to the server, which analyzes the data and generates additional cheering sounds.
[2131] Input: Vibration and motion data from the remote rooting device
[2132] Output: Data sent to the server
[2133] Step 7:
[2134] Cheering sound generation
[2135] The server analyzes the data: vibration and motion data is analyzed, and a generative AI model is used to generate cheering sounds, which are then played over the on-site sound system.
[2136] Input: Vibration and motion data
[2137] Output: Generated cheer sound
[2138] Step 8:
[2139] Audio recording from the viewing area
[2140] The server captures the audio: Microphones installed in the viewing area capture the audio from the venue and send it to the server.
[2141] The server streams the audio to the device: The captured audio is sent to the communication device and played back in real time. This process is important for conveying the cheers and realism of the actual event to the user.
[2142] Input: Audio data from microphones installed in the viewing area
[2143] Output: Audio data streamed to the user's device
[2144] Step 9:
[2145] Analysis and use of emotional data
[2146] User device collects emotional data: The emotion engine analyzes the user's facial expressions and tone of voice to collect emotional data.
[2147] The device sends emotional data to the server: The analyzed emotional data is sent to the server in real time. The server analyzes this data and dynamically changes the intensity and content of the generated voice.
[2148] Input: User facial expressions and tone of voice
[2149] Output: Emotion data sent to the server.
[2150] (Application example 2)
[2151] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[2152] Current remote viewing systems have the problem of being unable to fully reproduce the sense of presence and unity that users feel at a real viewing location. In particular, it is difficult to experience the cheers and sounds of the fans at the venue in real time, and they are unable to generate cheering sounds that reflect the emotions of the remote spectators, limiting the viewing experience. Furthermore, they lack the functionality to generate cheering sounds using data from remote cheering devices. Technology that solves these issues is needed.
[2153] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[2154] In this invention, the server includes means for reproducing the cheers and excitement of the spectator location using a generative model, means for transmitting input signals from spectators using a communication terminal to the server, means for analyzing the input signals, generating cheering sounds, and outputting them from a sound device, means for recording audio from the spectator location and transmitting the audio to the communication terminal, means for playing the audio on the communication terminal, means for recognizing the emotions of spectators using an emotion engine and transmitting the emotion data to the server in real time, and means for dynamically changing the intensity and content of the cheering sounds based on the emotion data. This allows remote spectators to experience a sense of realism similar to that of being at the venue, and cheering sounds that correspond to their emotions can be generated.
[2155] A "generative model" is a machine learning model that uses artificial intelligence to generate new data.
[2156] A "spectator venue" is the location where a sport or event actually takes place and where spectators physically gather.
[2157] "Cheers and enthusiasm" refers to the cheers and support emitted by spectators at the viewing venue, and are the sounds and atmosphere that indicate the excitement of the event.
[2158] A "communication terminal" is an electronic device used to send and receive data over the Internet, including smartphones and tablets.
[2159] "Input signals from spectators" refer to data that spectators send to the server via their communication terminals, and are signals that reflect the actions and emotions of the spectators.
[2160] A "server" is a computer system that processes and manages data over a network.
[2161] "Cheering sounds" are sounds generated based on input signals from spectators and generative models, and include sounds of cheering and cheering.
[2162] "Audio device" refers to equipment for outputting sound, including speakers and earphones.
[2163] An "emotion engine" is software that analyzes and recognizes the user's emotions from their facial expressions and movements.
[2164] "Real time" is a time concept that refers to instantaneous reaction and processing without delay.
[2165] A "remote cheering device" is a device that spectators use to cheer from home or another location, and has the ability to transmit vibration and motion data to a server.
[2166] The system for implementing this invention includes a generative AI model, an emotion engine, a communication terminal, a server, and an audio device. The operation of the entire system is as follows.
[2167] First, the user launches the remote viewing application on their communication device (smartphone or tablet). Within the application, the user selects the game or event they want to watch. The communication device then sends the selection information to the server, which then receives it.
[2168] The server receives live game feeds (video and audio) and inputs them into a generative AI model to generate cheering and support sounds for the spectators in real time. The generated audio is encoded, streamed to a communication device, and played back to the user in real time. The generative AI model then detects specific events (e.g., a goal or a foul) and dynamically generates audio based on those events.
[2169] Furthermore, the communication terminal is equipped with an emotion engine that recognizes the user's emotions. The emotion engine analyzes the user's facial expressions and tone of voice and transmits the recognized emotion data to the server in real time. The server analyzes this emotion data and dynamically changes the intensity and content of the cheering sounds according to the user's emotions. For example, if the user is excited, a louder cheer will be generated.
[2170] Users can also use a remote cheering device. Vibration and motion data from the remote cheering device is sent to a communication terminal, which then transmits the data to a server. The server analyzes the data and generates cheering sounds using a generative AI model, which are then played on the on-site sound equipment.
[2171] As a concrete example, consider the case where a user is watching a soccer match remotely. When the user launches the application and selects a match, live video and cheering sounds generated by a generative AI are streamed to the communication terminal. When the user shakes the remote cheering device during the match, vibration and motion data from the device are sent to the server via the communication terminal. The server analyzes the data and generates cheering sounds using a generative AI model. These cheering sounds are then played over the sound equipment at the viewing location.
[2172] An example of this prompt would be:
[2173] "A user is watching a soccer match remotely. The emotion engine on the user's smartphone recognizes the user's facial expressions and tone of voice, and the generation AI generates the cheering sounds of the local crowd in real time accordingly. As the user's excitement increases, the generation AI generates louder cheers and cheering sounds, providing a realistic spectator experience."
[2174] As described above, using this system allows remote spectators to experience the same sense of realism as if they were in person, and cheering sounds can be generated to suit their emotions, greatly improving the sense of realism and unity of remote spectators.
[2175] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[2176] Step 1:
[2177] The server initializes the generated AI model.
[2178] Input: Past match data and audience reaction data
[2179] Output: Initialized generative AI model
[2180] At startup, the server loads the generative AI model, sets the necessary parameters, reads the model's training data, and prepares it for real-time processing.
[2181] Step 2:
[2182] The user starts the application on the communication terminal and selects the game they want to watch.
[2183] Input: User's match selection information
[2184] Output: Match selection information sent to server
[2185] The user launches the application and selects the game they want to watch. The selection information is sent to the server, which receives it.
[2186] Step 3:
[2187] The server receives a live feed of the game and uses a generative AI model to generate the cheers and support sounds of the spectators.
[2188] Input: Live video and audio feed
[2189] Output: Generated cheers and cheering sounds
[2190] The server receives live video and audio feeds of the game, feeds them into a generative AI model, and generates cheers and cheering sounds in real time, which are then encoded and streamed to communication devices.
[2191] Step 4:
[2192] The communication terminal reproduces the generated cheers and cheering sounds in real time.
[2193] Input: Audio data streamed from the server
[2194] Output: Real-time cheers and cheering sounds played
[2195] The communication terminal decodes the received audio data and plays it back in real time, allowing the user to experience realistic audio.
[2196] Step 5:
[2197] The emotion engine recognizes the user's emotions and transmits the emotion data to the server in real time.
[2198] Input: User facial expressions and tone of voice
[2199] Output: Emotion data sent to the server
[2200] The emotion engine analyzes the user's facial expressions and tone of voice on the communication device to recognize their emotions, and the recognized emotion data is sent to the server.
[2201] Step 6:
[2202] The server analyzes the emotional data and uses a generative AI model to dynamically change the intensity and content of the cheering sounds.
[2203] Input: User emotion data
[2204] Output: Cheering sounds with dynamically changing intensity and content
[2205] The server analyzes the received emotional data and dynamically changes the intensity and content of the cheering sounds based on the user's emotions using a generative AI model. For example, if the user is excited, a louder cheer will be generated.
[2206] Step 7:
[2207] The user uses the remote support device and transmits the vibration data and motion data to the communication terminal.
[2208] Input: Vibration and motion data from the remote rooting device
[2209] Output: Data sent to the communication device
[2210] Users cheer by shaking the remote cheering device, and vibration and motion data from the device are sent to the communication terminal.
[2211] Step 8:
[2212] The communication terminal transmits the data received from the remote support device to the server.
[2213] Input: Vibration and motion data from the remote rooting device
[2214] Output: Device data sent to the server
[2215] The communication terminal transmits the received vibration data and motion data to the server.
[2216] Step 9:
[2217] The server analyzes the data from the remote cheering device and generates cheering sounds using a generative AI model.
[2218] Input: Vibration and motion data
[2219] Output: Generated cheer sound
[2220] The server analyzes the data from the remote cheering device and generates cheering sounds using a generative AI model. The generated cheering sounds are then played over the sound equipment at the viewing area.
[2221] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[2222] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[2223] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[2224] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[2225] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[2226] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[2227] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[2228] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[2229] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[2230] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces ...
Claims
1. Using generative models, we will create a method to recreate the cheers and excitement of the spectators. a means for transmitting input signals from spectators to a server using a communication terminal; means for analyzing the input signal, generating cheering sounds, and outputting the cheering sounds from an audio device; a means for recording the sound of the spectator location and transmitting the sound to a communication terminal; a playback means in the communication terminal; A system including:
2. The system of claim 1 , further comprising: means for analyzing an audio feed from the viewing location to detect a particular event; and means for generating audio based on the event with a generative model.
3. The system of claim 1 , further comprising: means for transmitting vibration or motion data of a remote cheering device to the server when the spectator's input signal is from the device.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A