System
The system generates and streams live video of user-selected artists with AI, random set lists, and real-time interaction, overcoming existing limitations to provide an immersive virtual live experience.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-08-19
- Publication Date
- 2026-03-04
AI Technical Summary
Existing systems fail to provide an immersive and interactive virtual live experience for users due to limitations in generating live video of user-selected artists in real time, randomly determining set lists and performance orders, and reflecting user viewpoints and operations in real time, while also lacking high-quality, low-latency streaming capabilities.
A system that generates live video of a user-selected artist using AI technology, streams it in real time to a viewing device, randomly determines the set list and performance order, and reflects the user's viewpoint and interactions in a virtual reality environment, utilizing low-latency streaming protocols like WebRTC or HTTP Live Streaming.
Enables users to enjoy a realistic virtual live performance without time or location constraints, providing an immersive and interactive experience that mimics attending a live event.
Smart Images

Figure 2026035134000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] In modern society, there are an increasing number of people who want to attend concerts and live events but cannot find the time due to work or childcare, or who want to avoid the risk of infection. Additionally, there are challenges such as travel costs and time constraints that make it difficult to attend concerts by overseas artists. These factors limit access to live events, preventing many fans from enjoying a live experience in person. It is hoped that this problem can be solved and that a means for users to enjoy their private time to the fullest can be provided. [Means for solving the problem]
[0005] The present invention solves the above-mentioned problems by providing a system that includes a means for generating live video of an artist selected by the user, a means for streaming the generated live video to the user's viewing device, a means for randomly determining the set list and performance order of the live video, and a means for reflecting the user's viewpoint and operations in real time. Specifically, by using AI technology to generate the live video and a low-latency streaming protocol, the system allows users to feel as if they are actually attending a live concert. Furthermore, by randomly determining the set list and performance order, the system provides a different live experience each time, thereby sustaining the user's interest. In this way, users can enjoy an artist's live performance in a virtual reality environment without the constraints of time or location.
[0006] "User" refers to a person who uses the system to enjoy a virtual live experience.
[0007] "Artist" refers to the musician / performer providing the live footage.
[0008] "Live footage" refers to video data that recreates an artist's performance.
[0009] "Generating means" refers to the process and equipment that creates live footage from collected data.
[0010] "Delivery means" refers to the process and equipment that transmits the generated live video to a user's viewing device.
[0011] "Viewing device" refers to equipment such as a headset or computer that a user uses to view live footage.
[0012] "Random determination means" refers to processes and devices that use algorithms to randomly determine the set list and performance order for a live performance.
[0013] "Means for reflecting viewpoints and operations in real time" refers to the process and equipment that detects the user's viewpoint movements and interactions and immediately reflects them in the image.
[0014] "AI technology" refers to technology that uses artificial intelligence technology to generate live footage.
[0015] "Low latency streaming protocol" refers to a communication standard for delivering live video in real time without delay.
[0016] "Realism" refers to the feeling that users are participating in an actual live event.
[0017] A "virtual reality environment" refers to an environment in which users can enjoy experiences in an immersive virtual space. [Brief explanation of the drawings]
[0018] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0019] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0020] First, the terms used in the following description will be explained.
[0021] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0022] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0023] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0024] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0025] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0026] [First embodiment]
[0027] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0028] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0029] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0030] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0031] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0032] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0033] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0034] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0035] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0036] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0037] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0038] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0039] This invention relates to a system that generates live video footage of an artist selected by a user and distributes that video in real time, allowing users to enjoy the artist's live performance in a virtual reality environment.
[0040] System Overview
[0041] The system mainly consists of three elements: a server, a terminal (user's viewing device), and the user. The server is responsible for generating and distributing live video, while the terminal receives and displays the video. Users participate in the live broadcast through their terminal.
[0042] Server Processing
[0043] The server first collects live data from the artist, including video, audio, stage layout, and lighting data, which is obtained from the artist and production team via a dedicated API.
[0044] The server then uses AI technology to generate live video based on the collected data. For example, by using GAN (Generative Adversarial Networks) as a generative model, the video is synthesized in real time and lighting and effects are automatically placed.
[0045] Furthermore, when the server receives a request to start a live performance, it randomly generates a setlist from the songs in the database and shuffles the order of the songs to create a unique setlist. During this process, seamless transitions appropriate for each song are also automatically generated.
[0046] Finally, the server encodes, compresses, and prepares the generated live video for streaming, delivering it to the user's device using a low-latency streaming protocol (e.g., WebRTC or HTTP Live Streaming: HLS).
[0047] Terminal handling
[0048] The device receives live video data streamed from the server in real time. The received video data is decoded and converted into a format suitable for the VR headset. The device also detects the user's viewpoint movements and interactions in real time and reflects them in the video.
[0049] User behavior
[0050] First, users put on a VR headset and log in to the system. After logging in, they select their favorite artist or live event from the app interface. Once the selection is complete, a request is sent to the server.
[0051] Once the live viewing begins, users can freely move their viewpoint while wearing the VR headset, enjoying the atmosphere of the live venue. Various operations are also possible, such as changing the viewpoint, adjusting the volume, and using interactive elements (for example, recreating the sounds of applause and cheers in VR).
[0052] Specific examples
[0053] When a user selects a live event by their favorite artist and presses the start viewing button, the server receives the request and generates a virtual live video based on the collected data. The set list and performance order are also randomly determined. The generated live video is streamed from the server to the device. The device receives the video data, decodes it, and displays it on the VR headset. The user puts on the VR headset and watches the live performance in real time while moving their viewpoint.
[0054] In this way, the system of the present invention allows users to enjoy an artist's live performance in a virtual reality environment without time or place restrictions.
[0055] The processing flow will be explained below.
[0056] Server Processing
[0057] Step 1:
[0058] The server collects the artist's existing live footage, audio, stage layout and lighting data from a database, which is then retrieved from the artist and production team via a dedicated API.
[0059] Step 2:
[0060] The server uses AI technology to generate virtual live footage in real time based on the collected data. Complex video synthesis is performed using technologies such as GAN (Generative Adversarial Networks), and lighting and effects are automatically placed.
[0061] Step 3:
[0062] When the server receives a request to start a live performance, it randomly generates a setlist from the songs in the database and shuffles the order in which they are played, resulting in a unique setlist.
[0063] Step 4:
[0064] The server encodes and compresses the generated live video to prepare it for streaming, and delivers it to the device using a low-latency streaming protocol (e.g., WebRTC or HTTP Live Streaming: HLS).
[0065] Terminal handling
[0066] Step 1:
[0067] The device receives live video data streamed from the server in real time, with minimal buffering required to maintain low latency.
[0068] Step 2:
[0069] It decodes the received video data and converts it into a format suitable for VR headsets, allowing users to enjoy the video without interruption.
[0070] Step 3:
[0071] The device detects the user's viewpoint movements and operations in real time and reflects this information in the video display, allowing the user to freely look around the virtual space.
[0072] User behavior
[0073] Step 1:
[0074] Users simply put on a VR headset and log into the dedicated application, where their user information and viewing history are automatically loaded.
[0075] Step 2:
[0076] Users select their favorite artist or live event from the app's interface and press the Start Watch button, which sends a request to the server.
[0077] Step 3:
[0078] Once the live viewing begins, the user can freely move their viewpoint while wearing the VR headset, experiencing the atmosphere of the live venue. They can change the viewpoint, adjust the volume, and use interactive elements (for example, recreating the sounds of applause and cheers in VR).
[0079] This processing flow allows users to feel as if they are actually participating in a live event.
[0080] Example 1
[0081] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0082] Conventional live video streaming systems have struggled to generate live video of a user-selected artist in real time and stream it to a viewing device. In particular, they needed to randomly determine the set list and performance order of the live video, and reflect the user's viewpoint and operations in real time. Furthermore, the lack of technology for streaming high-quality video with low latency limited the user experience.
[0083] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0084] In this invention, the server includes: means for generating live video of a musician selected by the user; means for streaming the generated live video to the user's display device; means for randomly determining the songs and performance order of the live video; means for reflecting the user's viewpoint and operations in real time; means for collecting live data from musicians and production groups via a dedicated API; means for generating live video using AI technology based on the collected data; means for randomly generating a set list from the songs in the collected data and shuffling the performance order; means for encoding and compressing the generated live video and preparing it for streaming; means for streaming the video using a low-latency streaming protocol; and means for converting the decoded video into a format suitable for a visual device, thereby enabling users to enjoy high-quality live video in real time.
[0085] A "musician" is a person whose occupation is performing or composing music.
[0086] "Live footage" refers to footage of a specific event or performance that is filmed and distributed in real time.
[0087] A "display device" is a device for displaying images, and specifically includes VR headsets and computer monitors.
[0088] "Song List" refers to the list of songs that will be performed live.
[0089] "Performance order" refers to the order in which songs are performed at a live concert.
[0090] "Viewpoint" refers to the direction and position from which the user looks in the VR space.
[0091] An "interaction" is an action or input that a user takes to interact with a system.
[0092] "Musicians and production groups" refers to people and organizations involved in organizing live performances and producing videos.
[0093] A "dedicated API" is a predefined interface for accessing specific functionality or services.
[0094] "Live data" is information related to a live event, including video, audio, stage layout, lighting data, and the like.
[0095] "AI technology" refers to analytical and generation methods that use artificial intelligence, and specific examples include Generative Adversarial Networks (GAN).
[0096] "Encoding" is the process of converting digital data into a particular format.
[0097] "Compression" is a process for reducing the amount of data.
[0098] "Streaming distribution" is a technology that transmits video and audio in real time and plays them instantly on the receiving end.
[0099] A "low-latency streaming protocol" is a communication method for minimizing delays in the transmission of video and audio, and specific examples include WebRTC and HTTP Live Streaming (HLS).
[0100] "Decoding" is the process of returning encoded data to its original format.
[0101] "Visual equipment" means a device through which a user views an image, including a VR headset or other display device.
[0102] The present invention relates to a system that generates live video footage of a musician selected by a user and distributes the video in real time, allowing users to enjoy the musician's live performance in a virtual reality environment.
[0103] System Overview
[0104] The system mainly consists of three elements: a server, a terminal (user's viewing device), and a user. The server is responsible for generating and distributing live video, while the terminal receives and displays the video. Users participate in the live broadcast through their terminal.
[0105] Server Processing
[0106] The server first collects live data from musicians, including video, audio, stage layout, and lighting data, which is obtained from musicians and production groups via a dedicated API.
[0107] The server then uses AI technology to generate live video based on the collected data. Specifically, it uses Generative Adversarial Networks (GANs) to synthesize video in real time and automatically position lighting and effects. For example, the lighting position and color dynamically change as musicians move around the stage.
[0108] Furthermore, when the server receives a request to start a live performance, it randomly generates a setlist from the songs in the database, shuffles the order of the songs, and creates a unique setlist. During this process, it also automatically generates seamless transitions appropriate for each song. For example, it adds effects that make the transition between songs feel natural.
[0109] Finally, the server encodes and compresses the generated live video and prepares it for streaming, delivering it to the user's device using a low-latency streaming protocol (e.g., WebRTC or HTTP Live Streaming: HLS).
[0110] Terminal handling
[0111] The device receives live video data streamed from the server in real time. The received video data is decoded using codecs such as H.264 and VP9 and converted into a format suitable for VR headsets. The device also detects the user's viewpoint movements and interactions in real time and reflects them in the video.
[0112] User behavior
[0113] First, the user puts on the VR headset and logs in to the system. After logging in, they select their favorite musician or live event from the app interface. Once the selection is complete, a request is sent to the server.
[0114] Once the live viewing begins, users can freely move their viewpoint while wearing the VR headset, enjoying the atmosphere of the live venue. Various operations are also possible, such as changing the viewpoint, adjusting the volume, and using interactive elements (for example, recreating applause and cheers in virtual reality).
[0115] Specific examples
[0116] Suppose a user selects "Musician A's Live Event" on the app interface and presses the Start Viewing button. The server receives the request and collects Musician A's data (video, audio, stage layout, lighting data, etc.) through a dedicated API. Based on that data, AI technology (GAN) is used to generate live video and create a random set list and appropriate transitions. The generated video is encoded and delivered to the device using a low-latency streaming protocol. The device decodes the received data and displays the video on a VR headset. The user can watch the live performance in real time while changing their viewpoint. As a concrete example, the following prompt sentence can be input to the generative AI model:
[0117] "Generate live footage of any musician. Include random setlists and lighting effects."
[0118] In this way, users can enjoy live performances of musicians in a virtual reality environment without time or location constraints.
[0119] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0120] Step 1:
[0121] The user puts on the VR headset and logs into the system.
[0122] Specific operation: The user puts on the VR headset, launches the dedicated app, and logs in by entering their user ID and password.
[0123] Input: User ID, Password
[0124] Output: User authentication completed, individual settings loaded
[0125] Step 2:
[0126] Users select artists and live events.
[0127] What it does: Users use the touchpad or voice commands to select their favorite artists or live events from the app interface.
[0128] Input: Artist name, Live event name
[0129] Output: Send selection
[0130] Step 3:
[0131] The server collects live data.
[0132] Specific operation: Based on the selected artist and live event, the server collects live data such as video, audio, stage layout, and lighting data through a dedicated API.
[0133] Input: Selected artist name, live event name
[0134] Output: Collected live data (video, audio, stage layout, lighting data, etc.)
[0135] Step 4:
[0136] The server uses AI technology to generate live footage.
[0137] How it works: The server uses Generative Adversarial Networks (GAN) to generate real-time live video based on collected live data, with lighting and other effects automatically placed.
[0138] Input: Live data collected
[0139] Output: Generated live video
[0140] Step 5:
[0141] The server randomly generates the setlist and shuffles the order in which the songs are played.
[0142] How it works: The server randomly generates a setlist from the songs in the database, shuffles the order of the songs, and automatically generates seamless transitions between songs.
[0143] Input: Songs in the database
[0144] Output: Random setlist, unique playing order, seamless transitions
[0145] Step 6:
[0146] The server encodes and compresses the generated live video.
[0147] Specific operation: The server encodes the generated live video using codecs such as H.264 or VP9 and compresses the data.
[0148] Input: Generated live video
[0149] Output: Encoded video data
[0150] Step 7:
[0151] The server streams the video.
[0152] How it works: The server delivers video data using a low-latency streaming protocol (WebRTC or HTTP Live Streaming: HLS). The server monitors network conditions and adjusts the bitrate as needed.
[0153] Input: Encoded video data
[0154] Output: Streamed video data
[0155] Step 8:
[0156] The device receives and decodes the video.
[0157] How it works: The device receives video data streamed from the server in real time, decodes it using codecs such as H.264 and VP9, and then converts the decoded data into a format suitable for the VR headset.
[0158] Input: Streamed video data
[0159] Output: Decoded video data, displayed on a VR headset
[0160] Step 9:
[0161] Users watch live.
[0162] Specific operation: Users can enjoy the live performance while wearing a VR headset and freely moving their viewpoint. They can change the viewpoint, adjust the volume, and use interactive elements (such as applause and cheers). The user's viewpoint movement and interactions are reflected in the video in real time.
[0163] Input: User's viewpoint movement, operation input
[0164] Output: Real-time changing visual and audio experience
[0165] (Application example 1)
[0166] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0167] Today, many music fans desire to enjoy artists' live performances in real time, but are often unable to attend due to physical distance, location, or even time constraints. Furthermore, existing online live streaming systems lack the immersive and interactive feel, making it difficult to provide an experience equivalent to a live performance. Furthermore, few systems are able to reflect user actions in real time and allow for free viewpoint movement.
[0168] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0169] In this invention, the server includes a means for generating live video of an artist selected by the user, a means for distributing the generated live video to the user's viewing device, a means for randomly determining the set list and performance order of the live video, a means for reflecting the user's viewpoint and operations in real time, and a means for reproducing interactive elements in a virtual reality environment. This allows users to transcend physical limitations and enjoy a realistic virtual live performance in real time, even from the comfort of their own home.
[0170] "Means for generating live video of an artist selected by a user" refers to technical means for generating live video content of a specific artist designated by a user.
[0171] "Means for delivering the generated live video to the user's viewing device" refers to the technical means for transmitting the generated live video data to the viewing terminal used by the user.
[0172] "Means for randomly determining the set list and performance order of live video" refers to a technical means for randomly determining the list of songs to be performed live and their order.
[0173] "Means for reflecting the user's viewpoint and operations in real time" refers to technical means for reflecting the user's visual viewpoint movements and interactive operations in the video in real time.
[0174] "Means for reproducing interactive elements in a virtual reality environment" refers to the technical means for reproducing interactive elements, such as applause and cheers, that can be experienced by users within a virtual reality environment.
[0175] "AI technology" refers to technology that utilizes artificial intelligence and is primarily used to generate live footage and determine set lists.
[0176] A "low-latency streaming protocol" is a communication protocol for transmitting and receiving data with minimal delay, and is a technology that is particularly applicable to live video distribution, which requires real-time performance.
[0177] This invention relates to a system that generates live video footage of an artist selected by the user and distributes it in real time. The system mainly consists of three elements: a server, a terminal (user's viewing device), and the user. This system allows users to enjoy the artist's live performance in a virtual reality environment.
[0178] System Overview
[0179] In this system, the server is responsible for generating and distributing live video, and the terminal receives and displays the video. Users participate in the live broadcast through the terminal.
[0180] Server Processing
[0181] The server first collects the artist's live performance data. This data includes video, audio, stage layout, lighting, and other data. This data is obtained from the artist and production team via a dedicated API. Next, the server uses AI technology to generate live video based on the collected data. For example, by using a generative adversarial network (GAN) as a generative model, it synthesizes video in real time and automatically places lighting and effects. Furthermore, when the server receives a request to start a live performance, it randomly generates a setlist from a list of songs in its database and shuffles the performance order to create a unique setlist. During this process, seamless transitions appropriate for each song are also automatically generated. Finally, the server encodes and compresses the generated live video and prepares it for streaming. The video is then distributed to the user's device using a low-latency streaming protocol (e.g., WebRTC or HTTP Live Streaming: HLS).
[0182] Terminal handling
[0183] The device receives live video data streamed from the server in real time. The received video data is decoded and converted into a format suitable for the VR headset. The device also detects the user's viewpoint movements and interactions in real time and reflects them in the video. This includes using Unity to process interactive elements such as viewpoint movements and clapping sounds in real time.
[0184] User behavior
[0185] First, the user puts on the VR headset and logs into the system. After logging in, they select their favorite artist or live event from the app interface. Once the selection is complete, a request is sent to the server. Once viewing of the live performance begins, the user can freely move their viewpoint while wearing the VR headset, enjoying the atmosphere of the live venue. Various operations are also possible, such as changing the viewpoint, adjusting the volume, and using interactive elements (for example, recreating applause and cheers in VR).
[0186] Specific examples
[0187] When a user selects a live event by their favorite artist and presses the start viewing button, the server receives the request and generates a virtual live video based on the collected data. The set list and performance order are also randomly determined. The generated live video is streamed from the server to the device. The device receives the video data, decodes it, and displays it on the VR headset. The user puts on the VR headset and watches the live performance in real time while moving their viewpoint. In this way, the system of the present invention allows users to enjoy an artist's live performance in a virtual reality environment without the constraints of time or place.
[0188] Prompt Sentence Examples
[0189] "To generate a virtual live video of an artist selected by the user and deliver it with low latency, synthesize the collected live data (video, audio, stage layout) in real time and stream it using WebRTC. Reflect real-time viewpoint changes and interactive elements so that users can enjoy the virtual live performance."
[0190] This allows users to enjoy an experience that feels as if they are at an actual live concert venue in a VR environment.
[0191] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0192] Step 1:
[0193] The server receives the user request
[0194] Input: The user selects the artist's live performance they want to watch and submits a request.
[0195] Processing: The server receives the request data from the user, which may include the artist name, the date and time of the show, and the user's authentication information.
[0196] Output: Request data to be used in the next step.
[0197] Step 2:
[0198] Live Data Collection
[0199] Input: Request data to generate artist gigs.
[0200] Processing: The server collects live data (video, audio, stage layout, lighting data, etc.) from artists and production teams via a dedicated API.
[0201] Output: Live data collected.
[0202] Step 3:
[0203] Live video generation
[0204] Input: Live data collected.
[0205] Processing: The server uses a generative AI model (e.g., GAN) to generate live footage in real time, including compositing the footage and adding lighting and effects.
[0206] Output: The generated live video data.
[0207] Step 4:
[0208] Setlist and performance order randomly determined
[0209] Input: Live video data and artist song list.
[0210] Processing: The server retrieves the song list from the database, randomly generates a setlist, shuffles the order of the songs, and automatically creates seamless transitions between each song.
[0211] Output: Live video data with setlist and transitions.
[0212] Step 5:
[0213] Encoding and preparing live video for distribution
[0214] Input: Live video data with setlist and transitions.
[0215] Processing: The server encodes the video in H264 format and prepares the data for distribution using a low-latency streaming protocol such as WebRTC.
[0216] Output: Encoded live video data.
[0217] Step 6:
[0218] Live video streaming
[0219] Input: Encoded live video data.
[0220] Processing: The server streams the prepared video data to the user's device in real time.
[0221] Output: Live video data sent to user terminal.
[0222] Step 7:
[0223] Receiving and decoding live video data
[0224] Input: Live video data streamed from the server.
[0225] Processing: The device decodes the received video data and converts it into a format suitable for the VR headset.
[0226] Output: Live video data that can be displayed on a VR headset.
[0227] Step 8:
[0228] Reflects user viewpoint movements and operations
[0229] Input: User gaze movement data and interaction data.
[0230] Processing: The device uses Unity to reflect interactive elements such as the user's viewpoint and clapping sounds in the video in real time.
[0231] Output: Live video data reflecting the interaction.
[0232] Step 9:
[0233] User's live viewing experience
[0234] Input: Live video data reflecting viewpoint movement.
[0235] Processing: Users can watch the live stream through a VR headset, freely moving their viewpoint in real time, and can also adjust the volume and use interactive elements (such as applause and cheering).
[0236] Output: An immersive virtual live experience.
[0237] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0238] The present invention relates to a system that generates live video of an artist selected by the user and distributes the video in real time. The present invention can also be combined with an emotion engine that recognizes the user's emotions to further improve the user experience.
[0239] System configuration
[0240] The system mainly consists of a server, a terminal (user's viewing device), and an emotion engine. The server is responsible for generating and distributing live video, the terminal is responsible for receiving and displaying the video, and the emotion engine recognizes user emotions in real time.
[0241] Server Processing
[0242] The server first collects live data from the artist, including video, audio, stage layout, and lighting data, which is obtained from the artist and production team via a dedicated API.
[0243] The server then uses AI technology to generate live video based on the collected data. Generative Adversarial Networks (GANs) and other generative models are used to synthesize the video in real time. Lighting and effects are also automatically placed.
[0244] Furthermore, when the server receives a request to start a live performance, it randomly generates a setlist from the songs in the database and shuffles the order in which they are played, resulting in a unique setlist.
[0245] The server encodes and compresses the generated live video, prepares it for streaming, and delivers it to the user's device using a low-latency streaming protocol (e.g., WebRTC or HTTP Live Streaming: HLS).
[0246] Terminal handling
[0247] The device receives live video data streamed from the server in real time. The received video data is decoded and converted into a format suitable for the VR headset. The device also detects the user's viewpoint movements and interactions in real time and reflects them in the video display.
[0248] Emotion engine processing
[0249] The emotion engine analyzes the user's facial expressions and voice to recognize emotions in real time. This data is sent to a server and used to dynamically change the live video production, set list, and interactive elements (such as specific lighting and sound effects).
[0250] User behavior
[0251] First, users put on a VR headset and log in to the system. After logging in, they select their favorite artist or live event from the app interface. Once the selection is complete, a request is sent to the server.
[0252] Once the live viewing begins, users can freely move their viewpoint while wearing the VR headset, enjoying the atmosphere of the live venue. Various operations are also possible, such as changing the viewpoint, adjusting the volume, and using interactive elements (for example, recreating applause and cheers in VR).
[0253] The emotion engine detects the user's emotions and changes the live video presentation in real time. For example, if the user becomes excited, the lighting and effects will be enhanced to further enhance the user's experience.
[0254] Specific examples
[0255] For example, when a user selects a live event by their favorite artist and presses the start button, the server receives the request and generates a virtual live video based on the collected data. The set list and performance order are also randomly determined. Furthermore, an emotion engine analyzes the user's emotions in real time and sends that data to the server. The server then dynamically adjusts the live video production and set list based on this data.
[0256] In this way, the system of the present invention allows users to enjoy live performances by artists in a virtual reality environment without the constraints of time or place, and by using an emotion engine, the user's experience can be further enriched.
[0257] The processing flow will be explained below.
[0258] Server Processing
[0259] Step 1:
[0260] The server collects artists' existing live footage, audio sources, stage layout, and lighting data from a database, which is then retrieved from the artists and production teams via a dedicated API.
[0261] Step 2:
[0262] Based on the collected data, the server uses AI technology to generate virtual live footage. For example, it uses a generative model based on GAN (Generative Adversarial Networks) to synthesize the footage in real time. Lighting and effects are also automatically placed.
[0263] Step 3:
[0264] When the server receives a request to start a live performance, it randomly generates a setlist from the songs in its database and shuffles the order in which they are played, resulting in a unique setlist.
[0265] Step 4:
[0266] The server encodes and compresses the generated live video, preparing it for streaming, and delivers it to the user's device using a low-latency streaming protocol (e.g., WebRTC or HTTP Live Streaming: HLS).
[0267] Step 5:
[0268] The server receives data from the emotion engine and adjusts the live video production (lighting, effects, set list changes, etc.) in real time based on the user's emotions.
[0269] Terminal handling
[0270] Step 1:
[0271] The device receives live video data streamed from the server in real time, with minimal buffering required to maintain low latency.
[0272] Step 2:
[0273] It decodes the received video data and converts it into a format suitable for VR headsets, allowing users to enjoy the video without interruption.
[0274] Step 3:
[0275] The device detects the user's viewpoint movements and operations in real time and reflects this information in the video display, allowing the user to freely look around the virtual space.
[0276] Step 4:
[0277] The device captures the user's facial expressions, voice, and other data and sends it to the emotion engine, which analyzes the data and generates real-time emotion data.
[0278] Emotion engine processing
[0279] Step 1:
[0280] The emotion engine receives the user's facial expressions and voice data sent from the device and analyzes it in real time.
[0281] Step 2:
[0282] Based on the analysis results, the system recognizes the user's emotions and sends the data to the server. For example, if the user is excited, the emotional data is transmitted to the server.
[0283] User behavior
[0284] Step 1:
[0285] Users put on a VR headset and log in to a dedicated application, where their user information and viewing history are automatically loaded.
[0286] Step 2:
[0287] Users select their favorite artist or live event from the app's interface and press the Start Watch button, which sends a request to the server.
[0288] Step 3:
[0289] Once the live viewing begins, users can freely move their viewpoint while wearing the VR headset and enjoy the atmosphere of the live venue. They can change the viewpoint, adjust the volume, and use interactive elements (for example, recreating applause and cheers in VR).
[0290] Step 4:
[0291] The emotion engine analyzes the user's facial expressions and voice, and the results are sent to the server. If the user is excited, the lighting and effects will be enhanced, further enriching the experience.
[0292] In this way, the system of the present invention not only allows users to enjoy live performances by artists in a virtual reality environment without time or place constraints, but also uses an emotion engine to provide a dynamic live experience that reflects the user's real-time emotions.
[0293] Example 2
[0294] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0295] Conventional live video streaming systems allow users to watch live footage of artists of their choice, but the experience relies on limited visual and auditory information. As a result, real-time video production is not tailored to the user's emotions or interactions, resulting in a lack of depth of experience. Furthermore, because the set list and performance order are fixed, watching the same live footage repeatedly can lose its freshness, creating a problem.
[0296] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0297] In this invention, the server includes means for generating live video of an artist selected by the user, means for delivering the generated live video to the user's viewing device, means for randomly determining the set list and performance order of the live video, means for reflecting the user's viewpoint and operations in real time, and means for recognizing the user's emotions in real time and dynamically changing the direction of the live video, thereby enabling users to enjoy a fresh and realistic live experience that is individually customized.
[0298] "User-selected artists" are musical artists selected by users of the system according to their own preferences.
[0299] "Live footage" is video that visually and aurally captures a musical artist's performance or live performance.
[0300] "Means of generation" is a general term for methods and devices that use specific algorithms or AI technology to create live footage.
[0301] A "viewing device" is a device that a user uses to view video and audio, such as a VR headset or smartphone.
[0302] "Means of distribution" refers to the methods and technologies for sending the generated live video to the user's viewing device via a network such as the Internet.
[0303] A "set list" is a list of songs that will be performed live.
[0304] The "play order" is the rule that determines the order in which songs in a set list are played.
[0305] A "random determination means" is a method or device that randomly determines the set list or performance order using a specific algorithm or random number generation technology.
[0306] "Means for reflecting viewpoints and operations in real time" refers to methods and technologies that instantly detect the user's viewpoint movements and interactions and reflect them in the video display.
[0307] "Means for recognizing emotions in real time and dynamically changing the presentation of live video" refers to methods and technologies that analyze emotions from the user's facial expressions, voice, etc., and instantly change the presentation of live video based on the results.
[0308] "AI technology" refers to any technology that uses artificial intelligence, including machine learning and deep learning in particular.
[0309] "Low latency streaming protocol" is a protocol that minimizes data delays and is a technology that enables real-time communication.
[0310] This invention relates to a system that generates live video of an artist selected by the user and distributes the video in real time. Furthermore, this invention can further enhance the user experience by incorporating an emotion engine that recognizes the user's emotions. This system consists of three main components: a server, a terminal, and an emotion engine.
[0311] System configuration
[0312] The system includes a server, a terminal (user's viewing device), and an emotion engine. The server is responsible for generating and distributing live video, and the terminal receives and displays the video. The emotion engine recognizes user emotions in real time.
[0313] Server Processing
[0314] The server first collects live data from the artist. This data includes video, audio, stage layout, and lighting data. This data is obtained from the artist and production team via a dedicated API. Based on the collected data, the server uses AI technology (such as Generative Adversarial Networks: GAN) to generate live video. During this generation process, lighting and effects are also automatically placed.
[0315] Furthermore, when the server receives a request to start a live performance, it randomly generates a setlist from the song list in the database and shuffles the performance order to create a unique setlist.The generated live video is then encoded, compressed, and distributed to the user's device using a low-latency streaming protocol (e.g., WebRTC or HTTP Live Streaming: HLS).
[0316] Terminal handling
[0317] The device receives live video data streamed from the server in real time. The received video data is decoded and converted into a format suitable for the VR headset. The device also detects the user's viewpoint movements and interactions in real time and reflects them in the video display.
[0318] Emotion engine processing
[0319] The emotion engine analyzes the user's facial expressions and voice to recognize emotions in real time. This data is sent to a server and used to dynamically change the live video production, set list, and interactive elements (such as specific lighting and sound effects).
[0320] User behavior
[0321] First, users put on a VR headset and log in to the system. After logging in, they select their favorite artist or live event from the app interface. Once the selection is complete, a request is sent to the server. Once viewing of the live performance begins, users can freely move their viewpoint while wearing the VR headset, enjoying the atmosphere of the live venue. They can also change the viewpoint, adjust the volume, and use interactive elements (for example, recreating applause and cheers in VR). The emotion engine detects the user's emotions, which then changes the live video presentation in real time. For example, if the user becomes excited, lighting and effects are enhanced, further improving the user experience.
[0322] Specific examples
[0323] For example, when a user selects a live event by their favorite artist and presses the start button, the server receives the request and uses AI technology to generate a virtual live video based on the collected data. The set list and performance order are also randomly determined. After the live video is generated, the server encodes the video and streams it to the user with low latency. When the emotion engine detects the user's excitement, the server dynamically changes the live performance by enhancing lighting and effects.
[0324] Prompt Sentence Examples
[0325] Below are some example prompts to input to a generative AI model:
[0326] "Please explain a system that generates a virtual live video of a live event of an artist selected by the user based on related data, and dynamically changes the live performance by analyzing the user's emotions in real time using an emotion engine."
[0327] In this way, the system of the present invention allows users to enjoy live performances of artists in a virtual reality environment regardless of location. Furthermore, the use of an emotion engine can further enrich the user experience.
[0328] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0329] Step 1:
[0330] Live Data Collection
[0331] The server collects live data from artists, including video, audio, stage layout, and lighting data, and this data is obtained through a dedicated API.
[0332] Input: API requests from artists and production teams
[0333] Output: Live video, audio, stage layout, lighting data
[0334] Specific behavior:
[0335] The server sends a request to an API endpoint.
[0336] Receive and parse the JSON data returned as a response.
[0337] The acquired data is stored in an internal database.
[0338] Step 2:
[0339] Live video generation
[0340] The server uses AI techniques (such as Generative Adversarial Networks: GAN) to generate live footage based on the collected data, taking into account stage layout and lighting data during the generation process.
[0341] Input: Live video, audio, stage layout, lighting data
[0342] Output: Generated live video
[0343] Specific behavior:
[0344] The server inputs the collected data into the AI model.
[0345] The AI model generates images based on the specified parameters.
[0346] Lighting and effects are automatically added to the generated video data.
[0347] Step 3:
[0348] Handling Go Live Requests
[0349] The server receives a request from the user to start a live stream, which includes an artist and event identifier.
[0350] Input: User's request to go live
[0351] Output: Information about a specific artist or event
[0352] Specific behavior:
[0353] The server receives the request and loads the corresponding artist and event data.
[0354] Prepare to move on to the next step.
[0355] Step 4:
[0356] Deciding on the set list and performance order
[0357] The server randomly generates a setlist from the songs in the database and shuffles the order in which they are played.
[0358] Input: Song list
[0359] Output: Randomly determined setlist and playing order
[0360] Specific behavior:
[0361] The server retrieves the song list from the database.
[0362] Use a random algorithm to generate setlists and shuffle song order.
[0363] Step 5:
[0364] Encode video and prepare for streaming
[0365] The server encodes, compresses, and converts the generated live video into a streaming format, preparing it for distribution using a low-latency streaming protocol (e.g., WebRTC or HTTP Live Streaming: HLS).
[0366] Input: Generated live video
[0367] Output: Video encoded for streaming
[0368] Specific behavior:
[0369] The server encodes the video data in a specific codec.
[0370] The encoded data is sent to a streaming server.
[0371] Step 6:
[0372] Live video streaming
[0373] The server streams the encoded live video to the user's device in real time.
[0374] Input: Encoded live video
[0375] Output: Video streamed to user device
[0376] Specific behavior:
[0377] The server checks the streaming protocol settings and transmits the encoded video data in real time.
[0378] Monitor your streaming performance and make adjustments as needed.
[0379] Step 7:
[0380] Handling user gaze movements and interactions
[0381] The device receives live video data streamed from the server in real time, decodes it, and converts it into a format suitable for the VR headset. It also detects the user's viewpoint movements and interactions and reflects them in the video display.
[0382] Input: Streaming live video data, user's viewpoint movement and interaction information
[0383] Output: Image displayed on VR headset
[0384] Specific behavior:
[0385] The terminal decodes the received video data.
[0386] It detects the user's viewpoint movement and adjusts the displayed content in real time.
[0387] Processes user interaction information (e.g., button presses).
[0388] Step 8:
[0389] Emotion analysis using an emotion engine
[0390] The emotion engine analyzes the user's facial expressions and voice to recognize their emotions in real time, and this emotion data is sent to the server.
[0391] Input: User's facial expressions and voice data
[0392] Output: Parsed emotion data
[0393] Specific behavior:
[0394] The emotion engine captures the user's camera footage and microphone audio.
[0395] The acquired data is analyzed in real time to recognize the emotional state.
[0396] Step 9:
[0397] Dynamic changes to live performances
[0398] The server dynamically changes the live video presentation based on the emotional data received from the emotion engine. For example, if the user is excited, the lighting and effects will be enhanced.
[0399] Input: Emotion data received from the emotion engine
[0400] Output: Dynamically modified live video rendition
[0401] Specific behavior:
[0402] The server analyzes the emotion data and determines the corresponding change in presentation.
[0403] Implementing real-time performance changes to improve the user experience.
[0404] (Application example 2)
[0405] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0406] Conventional live video streaming systems make it difficult for users to experience the content interactively in real time, limiting the viewing experience. Furthermore, they lack the ability to dynamically change the presentation to reflect user emotions and reactions, making it difficult to maximize user satisfaction. Furthermore, there are technical challenges in delivering high-quality video with low latency.
[0407] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for generating live video of an artist selected by the user, means for delivering the generated live video to the user's viewing device, means for randomly determining the set list and performance order of the live video, means for reflecting the user's viewpoint and operations in real time, means for recognizing the user's emotions and dynamically changing the production and set list of the live video based on that data, and means for delivering the live video using a low-latency streaming protocol. This allows users to enjoy a more interactive and personalized live experience in real time.
[0408] "User" means any person or entity that uses the System to view live video.
[0409] "Artists" are musicians or performers who are the subject of live footage.
[0410] "Live footage" refers to footage that visually records an artist's performance.
[0411] "Viewing Device" means the hardware device through which a User views the live feed. Examples include smartphones and head-mounted displays.
[0412] A "set list" is a list of songs that will be played at a live concert.
[0413] "Performance order" refers to the order in which songs are performed on the set list.
[0414] "Real time" refers to the time in which an event that occurs is processed and reflected immediately without delay.
[0415] "Emotion recognition" is the process of analyzing a user's facial expressions and voice to identify their emotional state.
[0416] "Production" refers to the visual and auditory effects, such as lighting, effects, and sound, used in live footage.
[0417] "Dynamic change" refers to the ability of the system to make immediate changes in response to specific conditions or situations.
[0418] A "low latency streaming protocol" is a communication protocol that minimizes delays in sending and receiving data.
[0419] "AI technology" refers to all technologies based on artificial intelligence, and in this invention, it particularly refers to technologies that perform generative models and emotion recognition.
[0420] A "prompt sentence" is an input sentence that gives instructions to a generative AI model.
[0421] This invention relates to a system that generates live video of an artist selected by the user and distributes the video in real time. Furthermore, by combining this invention with an emotion engine that recognizes the user's emotions, the user experience can be further improved.
[0422] System configuration
[0423] The system mainly consists of a server, a terminal (user's viewing device), and an emotion engine. The server is responsible for generating and distributing live video, the terminal is responsible for receiving and displaying the video, and the emotion engine recognizes user emotions in real time.
[0424] Server Processing
[0425] The server first collects live data from the artist. This data includes video, audio, stage layout, and lighting data. This data is obtained from the artist and production team via a dedicated API. The server then uses AI technology to generate live video based on the collected data. Generative Adversarial Networks (GAN) and other generative models are used to synthesize video in real time. Lighting and effects are also automatically placed.
[0426] Furthermore, when the server receives a request to start a live performance, it randomly generates a setlist from the song list in the database and shuffles the performance order. This process creates a unique setlist. The server then encodes and compresses the generated live video and prepares it for streaming. The video is then delivered to the user's device using a low-latency streaming protocol (e.g., WebRTC or HTTP Live Streaming: HLS).
[0427] Terminal handling
[0428] The device receives live video data streamed from the server in real time. The received video data is decoded and converted into a format suitable for the VR headset. The device also detects the user's viewpoint movements and interactions in real time and reflects them in the video display.
[0429] Emotion engine processing
[0430] The emotion engine analyzes the user's facial expressions, voice, etc. to recognize emotions in real time. It uses specific APIs such as OpenAI® API and Microsoft® Azure® Face API. This data is sent to a server and used to dynamically change the live video production, set list, and interactive elements (such as specific lighting and sound effects).
[0431] User behavior
[0432] First, the user puts on a VR headset and logs into the system. After logging in, they select their favorite artist or live event from the app's interface. Once the selection is complete, a request is sent to the server. Once viewing of the live performance begins, the user can freely move their viewpoint while wearing the VR headset, enjoying the atmosphere of the live venue. Various operations are also possible, such as changing the viewpoint, adjusting the volume, and using interactive elements (for example, recreating applause and cheers in VR).
[0433] The emotion engine detects the user's emotions and changes the live video presentation in real time. For example, if the user becomes excited, the lighting and effects will be enhanced to further enhance the user's experience.
[0434] Specific examples
[0435] For example, when a user selects a live event by their favorite artist and presses the start button, the server receives the request and generates a virtual live video based on the collected data. The set list and performance order are also randomly determined. Furthermore, an emotion engine analyzes the user's emotions in real time and sends that data to the server. The server then dynamically adjusts the live video production and set list based on this data.
[0436] Examples of prompts include:
[0437] "Please implement an algorithm that detects a smile on the user's face, sends that information to the server, and enhances the lighting effects on the live video."
[0438] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0439] Step 1:
[0440] The server collects live data from artists. As input, it receives live video, audio, stage layout, and lighting data provided by artists and production teams via a dedicated API. This data is integrated within the server and prepared as the basis for generating live video. The output is an integrated live data set.
[0441] Step 2:
[0442] The server generates live video using a generative AI model (e.g., GAN). It uses the live data collected in step 1 as input. The AI model takes video, audio, stage layout, and lighting data as input and generates a synthesized live video in real time. The output is the generated live video.
[0443] Step 3:
[0444] When the server receives a request to start a live performance, it randomly generates a setlist from the song list in the database and shuffles the playing order. It uses the request to start a live performance and the song list in the database as input. It uses a random algorithm to generate the setlist and the playing order. The output is the randomly determined setlist and its playing order.
[0445] Step 4:
[0446] The server encodes, compresses, and prepares the generated live video for streaming. It uses the live video generated in step 2 and the setlist determined in step 3 as input. It prepares the video for streaming using a low-latency streaming protocol (WebRTC or HLS). The output is the encoded live video ready for streaming.
[0447] Step 5:
[0448] The device receives live video data streamed from the server in real time. As input, it receives the streaming data sent from the server. The device decodes and converts it into a format suitable for the VR headset. The output is the decoded and converted video data.
[0449] Step 6:
[0450] The device detects the user's viewpoint movements and operations in real time. Sensor information from the VR headset is used as input. Based on this, viewpoint movement and operation data is collected in real time and reflected in the video display. The output is a video display that reflects the viewpoint and operations.
[0451] Step 7:
[0452] The emotion engine analyzes the user's facial expressions and voice to recognize emotions in real time. The input is the user's facial expression data and voice data. Analysis is performed using OpenAI's API or Microsoft Azure's Face API to generate emotion data. The output is the analyzed emotion data.
[0453] Step 8:
[0454] The server dynamically changes the live video production and set list based on the emotional data sent from the emotion engine. The emotional data from the emotion engine is used as input. Based on the emotional data, lighting effects and production are adjusted, and the set list is also changed as necessary. The output is live video and production that adapts to the user's emotions.
[0455] Step 9:
[0456] A user puts on a VR headset, logs in to the system, and selects their favorite artist or live event from the app interface. The user's selection information is entered into the app as input. Once the selection is complete, a live viewing request is sent to the server. The output is a live viewing request.
[0457] Step 10:
[0458] While watching a live stream, the user changes the viewpoint, adjusts the volume, and uses interactive elements (such as recreating applause and cheers). The input is information about the user's viewpoint, volume adjustment, and use of interactive elements. Based on this, the viewing experience is dynamically adjusted. The output is a viewing experience that corresponds to the user's actions.
[0459] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0460] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (registered trademark) (Internet search engine).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0461] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0462] [Second embodiment]
[0463] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0464] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0465] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0466] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0467] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0468] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0469] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0470] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0471] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0472] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0473] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0474] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0475] This invention relates to a system that generates live video footage of an artist selected by a user and distributes that video in real time, allowing users to enjoy the artist's live performance in a virtual reality environment.
[0476] System Overview
[0477] The system mainly consists of three elements: a server, a terminal (user's viewing device), and the user. The server is responsible for generating and distributing live video, while the terminal receives and displays the video. Users participate in the live broadcast through their terminal.
[0478] Server Processing
[0479] The server first collects live data from the artist, including video, audio, stage layout, and lighting data, which is obtained from the artist and production team via a dedicated API.
[0480] The server then uses AI technology to generate live video based on the collected data. For example, by using GAN (Generative Adversarial Networks) as a generative model, the video is synthesized in real time and lighting and effects are automatically placed.
[0481] Furthermore, when the server receives a request to start a live performance, it randomly generates a setlist from the songs in the database and shuffles the order of the songs to create a unique setlist. During this process, seamless transitions appropriate for each song are also automatically generated.
[0482] Finally, the server encodes, compresses, and prepares the generated live video for streaming, delivering it to the user's device using a low-latency streaming protocol (e.g., WebRTC or HTTP Live Streaming: HLS).
[0483] Terminal handling
[0484] The device receives live video data streamed from the server in real time. The received video data is decoded and converted into a format suitable for the VR headset. The device also detects the user's viewpoint movements and interactions in real time and reflects them in the video.
[0485] User behavior
[0486] First, users put on a VR headset and log in to the system. After logging in, they select their favorite artist or live event from the app interface. Once the selection is complete, a request is sent to the server.
[0487] Once the live viewing begins, users can freely move their viewpoint while wearing the VR headset, enjoying the atmosphere of the live venue. Various operations are also possible, such as changing the viewpoint, adjusting the volume, and using interactive elements (for example, recreating the sounds of applause and cheers in VR).
[0488] Specific examples
[0489] When a user selects a live event by their favorite artist and presses the start viewing button, the server receives the request and generates a virtual live video based on the collected data. The set list and performance order are also randomly determined. The generated live video is streamed from the server to the device. The device receives the video data, decodes it, and displays it on the VR headset. The user puts on the VR headset and watches the live performance in real time while moving their viewpoint.
[0490] In this way, the system of the present invention allows users to enjoy an artist's live performance in a virtual reality environment without time or place restrictions.
[0491] The processing flow will be explained below.
[0492] Server Processing
[0493] Step 1:
[0494] The server collects the artist's existing live footage, audio, stage layout and lighting data from a database, which is then retrieved from the artist and production team via a dedicated API.
[0495] Step 2:
[0496] The server uses AI technology to generate virtual live footage in real time based on the collected data. Complex video synthesis is performed using technologies such as GAN (Generative Adversarial Networks), and lighting and effects are automatically placed.
[0497] Step 3:
[0498] When the server receives a request to start a live performance, it randomly generates a setlist from the songs in the database and shuffles the order in which they are played, resulting in a unique setlist.
[0499] Step 4:
[0500] The server encodes and compresses the generated live video to prepare it for streaming, and delivers it to the device using a low-latency streaming protocol (e.g., WebRTC or HTTP Live Streaming: HLS).
[0501] Terminal handling
[0502] Step 1:
[0503] The device receives live video data streamed from the server in real time, with minimal buffering required to maintain low latency.
[0504] Step 2:
[0505] It decodes the received video data and converts it into a format suitable for VR headsets, allowing users to enjoy the video without interruption.
[0506] Step 3:
[0507] The device detects the user's viewpoint movements and operations in real time and reflects this information in the video display, allowing the user to freely look around the virtual space.
[0508] User behavior
[0509] Step 1:
[0510] Users simply put on a VR headset and log into the dedicated application, where their user information and viewing history are automatically loaded.
[0511] Step 2:
[0512] Users select their favorite artist or live event from the app's interface and press the Start Watch button, which sends a request to the server.
[0513] Step 3:
[0514] Once the live viewing begins, the user can freely move their viewpoint while wearing the VR headset, experiencing the atmosphere of the live venue. They can change the viewpoint, adjust the volume, and use interactive elements (for example, recreating the sounds of applause and cheers in VR).
[0515] This processing flow allows users to feel as if they are actually participating in a live event.
[0516] Example 1
[0517] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0518] Conventional live video streaming systems have struggled to generate live video of a user-selected artist in real time and stream it to a viewing device. In particular, they needed to randomly determine the set list and performance order of the live video, and reflect the user's viewpoint and operations in real time. Furthermore, the lack of technology for streaming high-quality video with low latency limited the user experience.
[0519] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0520] In this invention, the server includes: means for generating live video of a musician selected by the user; means for streaming the generated live video to the user's display device; means for randomly determining the songs and performance order of the live video; means for reflecting the user's viewpoint and operations in real time; means for collecting live data from musicians and production groups via a dedicated API; means for generating live video using AI technology based on the collected data; means for randomly generating a set list from the songs in the collected data and shuffling the performance order; means for encoding and compressing the generated live video and preparing it for streaming; means for streaming the video using a low-latency streaming protocol; and means for converting the decoded video into a format suitable for a visual device, thereby enabling users to enjoy high-quality live video in real time.
[0521] A "musician" is a person whose occupation is performing or composing music.
[0522] "Live footage" refers to footage of a specific event or performance that is filmed and distributed in real time.
[0523] A "display device" is a device for displaying images, and specifically includes VR headsets and computer monitors.
[0524] "Song List" refers to the list of songs that will be performed live.
[0525] "Performance order" refers to the order in which songs are performed at a live concert.
[0526] "Viewpoint" refers to the direction and position from which the user looks in the VR space.
[0527] An "interaction" is an action or input that a user takes to interact with a system.
[0528] "Musicians and production groups" refers to people and organizations involved in organizing live performances and producing videos.
[0529] A "dedicated API" is a predefined interface for accessing specific functionality or services.
[0530] "Live data" is information related to a live event, including video, audio, stage layout, lighting data, and the like.
[0531] "AI technology" refers to analytical and generation methods that use artificial intelligence, and specific examples include Generative Adversarial Networks (GAN).
[0532] "Encoding" is the process of converting digital data into a particular format.
[0533] "Compression" is a process for reducing the amount of data.
[0534] "Streaming distribution" is a technology that transmits video and audio in real time and plays them instantly on the receiving end.
[0535] A "low-latency streaming protocol" is a communication method for minimizing delays in the transmission of video and audio, and specific examples include WebRTC and HTTP Live Streaming (HLS).
[0536] "Decoding" is the process of returning encoded data to its original format.
[0537] "Visual equipment" means a device through which a user views an image, including a VR headset or other display device.
[0538] The present invention relates to a system that generates live video footage of a musician selected by a user and distributes the video in real time, allowing users to enjoy the musician's live performance in a virtual reality environment.
[0539] System Overview
[0540] The system mainly consists of three elements: a server, a terminal (user's viewing device), and a user. The server is responsible for generating and distributing live video, while the terminal receives and displays the video. Users participate in the live broadcast through their terminal.
[0541] Server Processing
[0542] The server first collects live data from musicians, including video, audio, stage layout, and lighting data, which is obtained from musicians and production groups via a dedicated API.
[0543] The server then uses AI technology to generate live video based on the collected data. Specifically, it uses Generative Adversarial Networks (GANs) to synthesize video in real time and automatically position lighting and effects. For example, the lighting position and color dynamically change as musicians move around the stage.
[0544] Furthermore, when the server receives a request to start a live performance, it randomly generates a setlist from the songs in the database, shuffles the order of the songs, and creates a unique setlist. During this process, it also automatically generates seamless transitions appropriate for each song. For example, it adds effects that make the transition between songs feel natural.
[0545] Finally, the server encodes and compresses the generated live video and prepares it for streaming, delivering it to the user's device using a low-latency streaming protocol (e.g., WebRTC or HTTP Live Streaming: HLS).
[0546] Terminal handling
[0547] The device receives live video data streamed from the server in real time. The received video data is decoded using codecs such as H.264 and VP9 and converted into a format suitable for VR headsets. The device also detects the user's viewpoint movements and interactions in real time and reflects them in the video.
[0548] User behavior
[0549] First, the user puts on the VR headset and logs in to the system. After logging in, they select their favorite musician or live event from the app interface. Once the selection is complete, a request is sent to the server.
[0550] Once the live viewing begins, users can freely move their viewpoint while wearing the VR headset, enjoying the atmosphere of the live venue. Various operations are also possible, such as changing the viewpoint, adjusting the volume, and using interactive elements (for example, recreating applause and cheers in virtual reality).
[0551] Specific examples
[0552] Suppose a user selects "Musician A's Live Event" on the app interface and presses the Start Viewing button. The server receives the request and collects Musician A's data (video, audio, stage layout, lighting data, etc.) through a dedicated API. Based on that data, AI technology (GAN) is used to generate live video and create a random set list and appropriate transitions. The generated video is encoded and delivered to the device using a low-latency streaming protocol. The device decodes the received data and displays the video on a VR headset. The user can watch the live performance in real time while changing their viewpoint. As a concrete example, the following prompt sentence can be input to the generative AI model:
[0553] "Generate live footage of any musician. Include random setlists and lighting effects."
[0554] In this way, users can enjoy live performances of musicians in a virtual reality environment without time or location constraints.
[0555] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0556] Step 1:
[0557] The user puts on the VR headset and logs into the system.
[0558] Specific operation: The user puts on the VR headset, launches the dedicated app, and logs in by entering their user ID and password.
[0559] Input: User ID, Password
[0560] Output: User authentication completed, individual settings loaded
[0561] Step 2:
[0562] Users select artists and live events.
[0563] What it does: Users use the touchpad or voice commands to select their favorite artists or live events from the app interface.
[0564] Input: Artist name, Live event name
[0565] Output: Send selection
[0566] Step 3:
[0567] The server collects live data.
[0568] Specific operation: Based on the selected artist and live event, the server collects live data such as video, audio, stage layout, and lighting data through a dedicated API.
[0569] Input: Selected artist name, live event name
[0570] Output: Collected live data (video, audio, stage layout, lighting data, etc.)
[0571] Step 4:
[0572] The server uses AI technology to generate live footage.
[0573] How it works: The server uses Generative Adversarial Networks (GAN) to generate real-time live video based on collected live data, with lighting and other effects automatically placed.
[0574] Input: Live data collected
[0575] Output: Generated live video
[0576] Step 5:
[0577] The server randomly generates the setlist and shuffles the order in which the songs are played.
[0578] How it works: The server randomly generates a setlist from the songs in the database, shuffles the order of the songs, and automatically generates seamless transitions between songs.
[0579] Input: Songs in the database
[0580] Output: Random setlist, unique playing order, seamless transitions
[0581] Step 6:
[0582] The server encodes and compresses the generated live video.
[0583] Specific operation: The server encodes the generated live video using codecs such as H.264 or VP9 and compresses the data.
[0584] Input: Generated live video
[0585] Output: Encoded video data
[0586] Step 7:
[0587] The server streams the video.
[0588] How it works: The server delivers video data using a low-latency streaming protocol (WebRTC or HTTP Live Streaming: HLS). The server monitors network conditions and adjusts the bitrate as needed.
[0589] Input: Encoded video data
[0590] Output: Streamed video data
[0591] Step 8:
[0592] The device receives and decodes the video.
[0593] How it works: The device receives video data streamed from the server in real time, decodes it using codecs such as H.264 and VP9, and then converts the decoded data into a format suitable for the VR headset.
[0594] Input: Streamed video data
[0595] Output: Decoded video data, displayed on a VR headset
[0596] Step 9:
[0597] Users watch live.
[0598] Specific operation: Users can enjoy the live performance while wearing a VR headset and freely moving their viewpoint. They can change the viewpoint, adjust the volume, and use interactive elements (such as applause and cheers). The user's viewpoint movement and interactions are reflected in the video in real time.
[0599] Input: User's viewpoint movement, operation input
[0600] Output: Real-time changing visual and audio experience
[0601] (Application example 1)
[0602] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0603] Today, many music fans desire to enjoy artists' live performances in real time, but are often unable to attend due to physical distance, location, or even time constraints. Furthermore, existing online live streaming systems lack the immersive and interactive feel, making it difficult to provide an experience equivalent to a live performance. Furthermore, few systems are able to reflect user actions in real time and allow for free viewpoint movement.
[0604] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0605] In this invention, the server includes a means for generating live video of an artist selected by the user, a means for distributing the generated live video to the user's viewing device, a means for randomly determining the set list and performance order of the live video, a means for reflecting the user's viewpoint and operations in real time, and a means for reproducing interactive elements in a virtual reality environment. This allows users to transcend physical limitations and enjoy a realistic virtual live performance in real time, even from the comfort of their own home.
[0606] "Means for generating live video of an artist selected by a user" refers to technical means for generating live video content of a specific artist designated by a user.
[0607] "Means for delivering the generated live video to the user's viewing device" refers to the technical means for transmitting the generated live video data to the viewing terminal used by the user.
[0608] "Means for randomly determining the set list and performance order of live video" refers to a technical means for randomly determining the list of songs to be performed live and their order.
[0609] "Means for reflecting the user's viewpoint and operations in real time" refers to technical means for reflecting the user's visual viewpoint movements and interactive operations in the video in real time.
[0610] "Means for reproducing interactive elements in a virtual reality environment" refers to the technical means for reproducing interactive elements, such as applause and cheers, that can be experienced by users within a virtual reality environment.
[0611] "AI technology" refers to technology that utilizes artificial intelligence and is primarily used to generate live footage and determine set lists.
[0612] A "low-latency streaming protocol" is a communication protocol for transmitting and receiving data with minimal delay, and is a technology that is particularly applicable to live video distribution, which requires real-time performance.
[0613] This invention relates to a system that generates live video footage of an artist selected by the user and distributes it in real time. The system mainly consists of three elements: a server, a terminal (user's viewing device), and the user. This system allows users to enjoy the artist's live performance in a virtual reality environment.
[0614] System Overview
[0615] In this system, the server is responsible for generating and distributing live video, and the terminal receives and displays the video. Users participate in the live broadcast through the terminal.
[0616] Server Processing
[0617] The server first collects the artist's live performance data. This data includes video, audio, stage layout, lighting, and other data. This data is obtained from the artist and production team via a dedicated API. Next, the server uses AI technology to generate live video based on the collected data. For example, by using a generative adversarial network (GAN) as a generative model, it synthesizes video in real time and automatically places lighting and effects. Furthermore, when the server receives a request to start a live performance, it randomly generates a setlist from a list of songs in its database and shuffles the performance order to create a unique setlist. During this process, seamless transitions appropriate for each song are also automatically generated. Finally, the server encodes and compresses the generated live video and prepares it for streaming. The video is then distributed to the user's device using a low-latency streaming protocol (e.g., WebRTC or HTTP Live Streaming: HLS).
[0618] Terminal handling
[0619] The device receives live video data streamed from the server in real time. The received video data is decoded and converted into a format suitable for the VR headset. The device also detects the user's viewpoint movements and interactions in real time and reflects them in the video. This includes using Unity to process interactive elements such as viewpoint movements and clapping sounds in real time.
[0620] User behavior
[0621] First, the user puts on the VR headset and logs into the system. After logging in, they select their favorite artist or live event from the app interface. Once the selection is complete, a request is sent to the server. Once viewing of the live performance begins, the user can freely move their viewpoint while wearing the VR headset, enjoying the atmosphere of the live venue. Various operations are also possible, such as changing the viewpoint, adjusting the volume, and using interactive elements (for example, recreating applause and cheers in VR).
[0622] Specific examples
[0623] When a user selects a live event by their favorite artist and presses the start viewing button, the server receives the request and generates a virtual live video based on the collected data. The set list and performance order are also randomly determined. The generated live video is streamed from the server to the device. The device receives the video data, decodes it, and displays it on the VR headset. The user puts on the VR headset and watches the live performance in real time while moving their viewpoint. In this way, the system of the present invention allows users to enjoy an artist's live performance in a virtual reality environment without the constraints of time or place.
[0624] Prompt Sentence Examples
[0625] "To generate a virtual live video of an artist selected by the user and deliver it with low latency, synthesize the collected live data (video, audio, stage layout) in real time and stream it using WebRTC. Reflect real-time viewpoint changes and interactive elements so that users can enjoy the virtual live performance."
[0626] This allows users to enjoy an experience that feels as if they are at an actual live concert venue in a VR environment.
[0627] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0628] Step 1:
[0629] The server receives the user request
[0630] Input: The user selects the artist's live performance they want to watch and submits a request.
[0631] Processing: The server receives the request data from the user, which may include the artist name, the date and time of the show, and the user's authentication information.
[0632] Output: Request data to be used in the next step.
[0633] Step 2:
[0634] Live Data Collection
[0635] Input: Request data to generate artist gigs.
[0636] Processing: The server collects live data (video, audio, stage layout, lighting data, etc.) from artists and production teams via a dedicated API.
[0637] Output: Live data collected.
[0638] Step 3:
[0639] Live video generation
[0640] Input: Live data collected.
[0641] Processing: The server uses a generative AI model (e.g., GAN) to generate live footage in real time, including compositing the footage and adding lighting and effects.
[0642] Output: The generated live video data.
[0643] Step 4:
[0644] Setlist and performance order randomly determined
[0645] Input: Live video data and artist song list.
[0646] Processing: The server retrieves the song list from the database, randomly generates a setlist, shuffles the order of the songs, and automatically creates seamless transitions between each song.
[0647] Output: Live video data with setlist and transitions.
[0648] Step 5:
[0649] Encoding and preparing live video for distribution
[0650] Input: Live video data with setlist and transitions.
[0651] Processing: The server encodes the video in H264 format and prepares the data for distribution using a low-latency streaming protocol such as WebRTC.
[0652] Output: Encoded live video data.
[0653] Step 6:
[0654] Live video streaming
[0655] Input: Encoded live video data.
[0656] Processing: The server streams the prepared video data to the user's device in real time.
[0657] Output: Live video data sent to user terminal.
[0658] Step 7:
[0659] Receiving and decoding live video data
[0660] Input: Live video data streamed from the server.
[0661] Processing: The device decodes the received video data and converts it into a format suitable for the VR headset.
[0662] Output: Live video data that can be displayed on a VR headset.
[0663] Step 8:
[0664] Reflects user viewpoint movements and operations
[0665] Input: User gaze movement data and interaction data.
[0666] Processing: The device uses Unity to reflect interactive elements such as the user's viewpoint and clapping sounds in the video in real time.
[0667] Output: Live video data reflecting the interaction.
[0668] Step 9:
[0669] User's live viewing experience
[0670] Input: Live video data reflecting viewpoint movement.
[0671] Processing: Users can watch the live stream through a VR headset, freely moving their viewpoint in real time, and can also adjust the volume and use interactive elements (such as applause and cheering).
[0672] Output: An immersive virtual live experience.
[0673] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0674] The present invention relates to a system that generates live video of an artist selected by the user and distributes the video in real time. The present invention can also be combined with an emotion engine that recognizes the user's emotions to further improve the user experience.
[0675] System configuration
[0676] The system mainly consists of a server, a terminal (user's viewing device), and an emotion engine. The server is responsible for generating and distributing live video, the terminal is responsible for receiving and displaying the video, and the emotion engine recognizes user emotions in real time.
[0677] Server Processing
[0678] The server first collects live data from the artist, including video, audio, stage layout, and lighting data, which is obtained from the artist and production team via a dedicated API.
[0679] The server then uses AI technology to generate live video based on the collected data. Generative Adversarial Networks (GANs) and other generative models are used to synthesize the video in real time. Lighting and effects are also automatically placed.
[0680] Furthermore, when the server receives a request to start a live performance, it randomly generates a setlist from the songs in the database and shuffles the order in which they are played, resulting in a unique setlist.
[0681] The server encodes and compresses the generated live video, prepares it for streaming, and delivers it to the user's device using a low-latency streaming protocol (e.g., WebRTC or HTTP Live Streaming: HLS).
[0682] Terminal handling
[0683] The device receives live video data streamed from the server in real time. The received video data is decoded and converted into a format suitable for the VR headset. The device also detects the user's viewpoint movements and interactions in real time and reflects them in the video display.
[0684] Emotion engine processing
[0685] The emotion engine analyzes the user's facial expressions and voice to recognize emotions in real time. This data is sent to a server and used to dynamically change the live video production, set list, and interactive elements (such as specific lighting and sound effects).
[0686] User behavior
[0687] First, users put on a VR headset and log in to the system. After logging in, they select their favorite artist or live event from the app interface. Once the selection is complete, a request is sent to the server.
[0688] Once the live viewing begins, users can freely move their viewpoint while wearing the VR headset, enjoying the atmosphere of the live venue. Various operations are also possible, such as changing the viewpoint, adjusting the volume, and using interactive elements (for example, recreating applause and cheers in VR).
[0689] The emotion engine detects the user's emotions and changes the live video presentation in real time. For example, if the user becomes excited, the lighting and effects will be enhanced to further enhance the user's experience.
[0690] Specific examples
[0691] For example, when a user selects a live event by their favorite artist and presses the start button, the server receives the request and generates a virtual live video based on the collected data. The set list and performance order are also randomly determined. Furthermore, an emotion engine analyzes the user's emotions in real time and sends that data to the server. The server then dynamically adjusts the live video production and set list based on this data.
[0692] In this way, the system of the present invention allows users to enjoy live performances by artists in a virtual reality environment without the constraints of time or place, and by using an emotion engine, the user's experience can be further enriched.
[0693] The processing flow will be explained below.
[0694] Server Processing
[0695] Step 1:
[0696] The server collects artists' existing live footage, audio sources, stage layout, and lighting data from a database, which is then retrieved from the artists and production teams via a dedicated API.
[0697] Step 2:
[0698] Based on the collected data, the server uses AI technology to generate virtual live footage. For example, it uses a generative model based on GAN (Generative Adversarial Networks) to synthesize the footage in real time. Lighting and effects are also automatically placed.
[0699] Step 3:
[0700] When the server receives a request to start a live performance, it randomly generates a setlist from the songs in its database and shuffles the order in which they are played, resulting in a unique setlist.
[0701] Step 4:
[0702] The server encodes and compresses the generated live video, preparing it for streaming, and delivers it to the user's device using a low-latency streaming protocol (e.g., WebRTC or HTTP Live Streaming: HLS).
[0703] Step 5:
[0704] The server receives data from the emotion engine and adjusts the live video production (lighting, effects, set list changes, etc.) in real time based on the user's emotions.
[0705] Terminal handling
[0706] Step 1:
[0707] The device receives live video data streamed from the server in real time, with minimal buffering required to maintain low latency.
[0708] Step 2:
[0709] It decodes the received video data and converts it into a format suitable for VR headsets, allowing users to enjoy the video without interruption.
[0710] Step 3:
[0711] The device detects the user's viewpoint movements and operations in real time and reflects this information in the video display, allowing the user to freely look around the virtual space.
[0712] Step 4:
[0713] The device captures the user's facial expressions, voice, and other data and sends it to the emotion engine, which analyzes the data and generates real-time emotion data.
[0714] Emotion engine processing
[0715] Step 1:
[0716] The emotion engine receives the user's facial expressions and voice data sent from the device and analyzes it in real time.
[0717] Step 2:
[0718] Based on the analysis results, the system recognizes the user's emotions and sends the data to the server. For example, if the user is excited, the emotional data is transmitted to the server.
[0719] User behavior
[0720] Step 1:
[0721] Users put on a VR headset and log in to a dedicated application, where their user information and viewing history are automatically loaded.
[0722] Step 2:
[0723] Users select their favorite artist or live event from the app's interface and press the Start Watch button, which sends a request to the server.
[0724] Step 3:
[0725] Once the live viewing begins, users can freely move their viewpoint while wearing the VR headset and enjoy the atmosphere of the live venue. They can change the viewpoint, adjust the volume, and use interactive elements (for example, recreating applause and cheers in VR).
[0726] Step 4:
[0727] The emotion engine analyzes the user's facial expressions and voice, and the results are sent to the server. If the user is excited, the lighting and effects will be enhanced, further enriching the experience.
[0728] In this way, the system of the present invention not only allows users to enjoy live performances by artists in a virtual reality environment without time or place constraints, but also uses an emotion engine to provide a dynamic live experience that reflects the user's real-time emotions.
[0729] Example 2
[0730] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0731] Conventional live video streaming systems allow users to watch live footage of artists of their choice, but the experience relies on limited visual and auditory information. As a result, real-time video production is not tailored to the user's emotions or interactions, resulting in a lack of depth of experience. Furthermore, because the set list and performance order are fixed, watching the same live footage repeatedly can lose its freshness, creating a problem.
[0732] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0733] In this invention, the server includes means for generating live video of an artist selected by the user, means for delivering the generated live video to the user's viewing device, means for randomly determining the set list and performance order of the live video, means for reflecting the user's viewpoint and operations in real time, and means for recognizing the user's emotions in real time and dynamically changing the direction of the live video, thereby enabling users to enjoy a fresh and realistic live experience that is individually customized.
[0734] "User-selected artists" are musical artists selected by users of the system according to their own preferences.
[0735] "Live footage" is video that visually and aurally captures a musical artist's performance or live performance.
[0736] "Means of generation" is a general term for methods and devices that use specific algorithms or AI technology to create live footage.
[0737] A "viewing device" is a device that a user uses to view video and audio, such as a VR headset or smartphone.
[0738] "Means of distribution" refers to the methods and technologies for sending the generated live video to the user's viewing device via a network such as the Internet.
[0739] A "set list" is a list of songs that will be performed live.
[0740] The "play order" is the rule that determines the order in which songs in a set list are played.
[0741] A "random determination means" is a method or device that randomly determines the set list or performance order using a specific algorithm or random number generation technology.
[0742] "Means for reflecting viewpoints and operations in real time" refers to methods and technologies that instantly detect the user's viewpoint movements and interactions and reflect them in the video display.
[0743] "Means for recognizing emotions in real time and dynamically changing the presentation of live video" refers to methods and technologies that analyze emotions from the user's facial expressions, voice, etc., and instantly change the presentation of live video based on the results.
[0744] "AI technology" refers to any technology that uses artificial intelligence, including machine learning and deep learning in particular.
[0745] "Low latency streaming protocol" is a protocol that minimizes data delays and is a technology that enables real-time communication.
[0746] This invention relates to a system that generates live video of an artist selected by the user and distributes the video in real time. Furthermore, this invention can further enhance the user experience by incorporating an emotion engine that recognizes the user's emotions. This system consists of three main components: a server, a terminal, and an emotion engine.
[0747] System configuration
[0748] The system includes a server, a terminal (user's viewing device), and an emotion engine. The server is responsible for generating and distributing live video, and the terminal receives and displays the video. The emotion engine recognizes user emotions in real time.
[0749] Server Processing
[0750] The server first collects live data from the artist. This data includes video, audio, stage layout, and lighting data. This data is obtained from the artist and production team via a dedicated API. Based on the collected data, the server uses AI technology (such as Generative Adversarial Networks: GAN) to generate live video. During this generation process, lighting and effects are also automatically placed.
[0751] Furthermore, when the server receives a request to start a live performance, it randomly generates a setlist from the song list in the database and shuffles the performance order to create a unique setlist.The generated live video is then encoded, compressed, and distributed to the user's device using a low-latency streaming protocol (e.g., WebRTC or HTTP Live Streaming: HLS).
[0752] Terminal handling
[0753] The device receives live video data streamed from the server in real time. The received video data is decoded and converted into a format suitable for the VR headset. The device also detects the user's viewpoint movements and interactions in real time and reflects them in the video display.
[0754] Emotion engine processing
[0755] The emotion engine analyzes the user's facial expressions and voice to recognize emotions in real time. This data is sent to a server and used to dynamically change the live video production, set list, and interactive elements (such as specific lighting and sound effects).
[0756] User behavior
[0757] First, users put on a VR headset and log in to the system. After logging in, they select their favorite artist or live event from the app interface. Once the selection is complete, a request is sent to the server. Once viewing of the live performance begins, users can freely move their viewpoint while wearing the VR headset, enjoying the atmosphere of the live venue. They can also change the viewpoint, adjust the volume, and use interactive elements (for example, recreating applause and cheers in VR). The emotion engine detects the user's emotions, which then changes the live video presentation in real time. For example, if the user becomes excited, lighting and effects are enhanced, further improving the user experience.
[0758] Specific examples
[0759] For example, when a user selects a live event by their favorite artist and presses the start button, the server receives the request and uses AI technology to generate a virtual live video based on the collected data. The set list and performance order are also randomly determined. After the live video is generated, the server encodes the video and streams it to the user with low latency. When the emotion engine detects the user's excitement, the server dynamically changes the live performance by enhancing lighting and effects.
[0760] Prompt Sentence Examples
[0761] Below are some example prompts to input to a generative AI model:
[0762] "Please explain a system that generates a virtual live video of a live event of an artist selected by the user based on related data, and dynamically changes the live performance by analyzing the user's emotions in real time using an emotion engine."
[0763] In this way, the system of the present invention allows users to enjoy live performances of artists in a virtual reality environment regardless of location. Furthermore, the use of an emotion engine can further enrich the user experience.
[0764] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0765] Step 1:
[0766] Live Data Collection
[0767] The server collects live data from artists, including video, audio, stage layout, and lighting data, and this data is obtained through a dedicated API.
[0768] Input: API requests from artists and production teams
[0769] Output: Live video, audio, stage layout, lighting data
[0770] Specific behavior:
[0771] The server sends a request to an API endpoint.
[0772] Receive and parse the JSON data returned as a response.
[0773] The acquired data is stored in an internal database.
[0774] Step 2:
[0775] Live video generation
[0776] The server uses AI techniques (such as Generative Adversarial Networks: GAN) to generate live footage based on the collected data, taking into account stage layout and lighting data during the generation process.
[0777] Input: Live video, audio, stage layout, lighting data
[0778] Output: Generated live video
[0779] Specific behavior:
[0780] The server inputs the collected data into the AI model.
[0781] The AI model generates images based on the specified parameters.
[0782] Lighting and effects are automatically added to the generated video data.
[0783] Step 3:
[0784] Handling Go Live Requests
[0785] The server receives a request from the user to start a live stream, which includes an artist and event identifier.
[0786] Input: User's request to go live
[0787] Output: Information about a specific artist or event
[0788] Specific behavior:
[0789] The server receives the request and loads the corresponding artist and event data.
[0790] Prepare to move on to the next step.
[0791] Step 4:
[0792] Deciding on the set list and performance order
[0793] The server randomly generates a setlist from the songs in the database and shuffles the order in which they are played.
[0794] Input: Song list
[0795] Output: Randomly determined setlist and playing order
[0796] Specific behavior:
[0797] The server retrieves the song list from the database.
[0798] Use a random algorithm to generate setlists and shuffle song order.
[0799] Step 5:
[0800] Encode video and prepare for streaming
[0801] The server encodes, compresses, and converts the generated live video into a streaming format, preparing it for distribution using a low-latency streaming protocol (e.g., WebRTC or HTTP Live Streaming: HLS).
[0802] Input: Generated live video
[0803] Output: Video encoded for streaming
[0804] Specific behavior:
[0805] The server encodes the video data in a specific codec.
[0806] The encoded data is sent to a streaming server.
[0807] Step 6:
[0808] Live video streaming
[0809] The server streams the encoded live video to the user's device in real time.
[0810] Input: Encoded live video
[0811] Output: Video streamed to user device
[0812] Specific behavior:
[0813] The server checks the streaming protocol settings and transmits the encoded video data in real time.
[0814] Monitor your streaming performance and make adjustments as needed.
[0815] Step 7:
[0816] Handling user gaze movements and interactions
[0817] The device receives live video data streamed from the server in real time, decodes it, and converts it into a format suitable for the VR headset. It also detects the user's viewpoint movements and interactions and reflects them in the video display.
[0818] Input: Streaming live video data, user's viewpoint movement and interaction information
[0819] Output: Image displayed on VR headset
[0820] Specific behavior:
[0821] The terminal decodes the received video data.
[0822] It detects the user's viewpoint movement and adjusts the displayed content in real time.
[0823] Processes user interaction information (e.g., button presses).
[0824] Step 8:
[0825] Emotion analysis using an emotion engine
[0826] The emotion engine analyzes the user's facial expressions and voice to recognize their emotions in real time, and this emotion data is sent to the server.
[0827] Input: User's facial expressions and voice data
[0828] Output: Parsed emotion data
[0829] Specific behavior:
[0830] The emotion engine captures the user's camera footage and microphone audio.
[0831] The acquired data is analyzed in real time to recognize the emotional state.
[0832] Step 9:
[0833] Dynamic changes to live performances
[0834] The server dynamically changes the live video presentation based on the emotional data received from the emotion engine. For example, if the user is excited, the lighting and effects will be enhanced.
[0835] Input: Emotion data received from the emotion engine
[0836] Output: Dynamically modified live video rendition
[0837] Specific behavior:
[0838] The server analyzes the emotion data and determines the corresponding change in presentation.
[0839] Implementing real-time performance changes to improve the user experience.
[0840] (Application example 2)
[0841] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0842] Conventional live video streaming systems make it difficult for users to experience the content interactively in real time, limiting the viewing experience. Furthermore, they lack the ability to dynamically change the presentation to reflect user emotions and reactions, making it difficult to maximize user satisfaction. Furthermore, there are technical challenges in delivering high-quality video with low latency.
[0843] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for generating live video of an artist selected by the user, means for delivering the generated live video to the user's viewing device, means for randomly determining the set list and performance order of the live video, means for reflecting the user's viewpoint and operations in real time, means for recognizing the user's emotions and dynamically changing the production and set list of the live video based on that data, and means for delivering the live video using a low-latency streaming protocol. This allows users to enjoy a more interactive and personalized live experience in real time.
[0844] "User" means any person or entity that uses the System to view live video.
[0845] "Artists" are musicians or performers who are the subject of live footage.
[0846] "Live footage" refers to footage that visually records an artist's performance.
[0847] "Viewing Device" means the hardware device through which a User views the live feed. Examples include smartphones and head-mounted displays.
[0848] A "set list" is a list of songs that will be played at a live concert.
[0849] "Performance order" refers to the order in which songs are performed on the set list.
[0850] "Real time" refers to the time in which an event that occurs is processed and reflected immediately without delay.
[0851] "Emotion recognition" is the process of analyzing a user's facial expressions and voice to identify their emotional state.
[0852] "Production" refers to the visual and auditory effects, such as lighting, effects, and sound, used in live footage.
[0853] "Dynamic change" refers to the ability of the system to make immediate changes in response to specific conditions or situations.
[0854] A "low latency streaming protocol" is a communication protocol that minimizes delays in sending and receiving data.
[0855] "AI technology" refers to all technologies based on artificial intelligence, and in this invention, it particularly refers to technologies that perform generative models and emotion recognition.
[0856] A "prompt sentence" is an input sentence that gives instructions to a generative AI model.
[0857] This invention relates to a system that generates live video of an artist selected by the user and distributes the video in real time. Furthermore, by combining this invention with an emotion engine that recognizes the user's emotions, the user experience can be further improved.
[0858] System configuration
[0859] The system mainly consists of a server, a terminal (user's viewing device), and an emotion engine. The server is responsible for generating and distributing live video, the terminal is responsible for receiving and displaying the video, and the emotion engine recognizes user emotions in real time.
[0860] Server Processing
[0861] The server first collects live data from the artist. This data includes video, audio, stage layout, and lighting data. This data is obtained from the artist and production team via a dedicated API. The server then uses AI technology to generate live video based on the collected data. Generative Adversarial Networks (GAN) and other generative models are used to synthesize video in real time. Lighting and effects are also automatically placed.
[0862] Furthermore, when the server receives a request to start a live performance, it randomly generates a setlist from the song list in the database and shuffles the performance order. This process creates a unique setlist. The server then encodes and compresses the generated live video and prepares it for streaming. The video is then delivered to the user's device using a low-latency streaming protocol (e.g., WebRTC or HTTP Live Streaming: HLS).
[0863] Terminal handling
[0864] The device receives live video data streamed from the server in real time. The received video data is decoded and converted into a format suitable for the VR headset. The device also detects the user's viewpoint movements and interactions in real time and reflects them in the video display.
[0865] Emotion engine processing
[0866] The emotion engine analyzes the user's facial expressions, voice, and other data to recognize emotions in real time. It uses specific APIs, such as OpenAI's API or Microsoft Azure's Face API. This data is sent to a server and used to dynamically change the live video presentation, set list, and interactive elements (such as specific lighting and sound effects).
[0867] User behavior
[0868] First, the user puts on a VR headset and logs into the system. After logging in, they select their favorite artist or live event from the app's interface. Once the selection is complete, a request is sent to the server. Once viewing of the live performance begins, the user can freely move their viewpoint while wearing the VR headset, enjoying the atmosphere of the live venue. Various operations are also possible, such as changing the viewpoint, adjusting the volume, and using interactive elements (for example, recreating applause and cheers in VR).
[0869] The emotion engine detects the user's emotions and changes the live video presentation in real time. For example, if the user becomes excited, the lighting and effects will be enhanced to further enhance the user's experience.
[0870] Specific examples
[0871] For example, when a user selects a live event by their favorite artist and presses the start button, the server receives the request and generates a virtual live video based on the collected data. The set list and performance order are also randomly determined. Furthermore, an emotion engine analyzes the user's emotions in real time and sends that data to the server. The server then dynamically adjusts the live video production and set list based on this data.
[0872] Examples of prompts include:
[0873] "Please implement an algorithm that detects a smile on the user's face, sends that information to the server, and enhances the lighting effects on the live video."
[0874] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0875] Step 1:
[0876] The server collects live data from artists. As input, it receives live video, audio, stage layout, and lighting data provided by artists and production teams via a dedicated API. This data is integrated within the server and prepared as the basis for generating live video. The output is an integrated live data set.
[0877] Step 2:
[0878] The server generates live video using a generative AI model (e.g., GAN). It uses the live data collected in step 1 as input. The AI model takes video, audio, stage layout, and lighting data as input and generates a synthesized live video in real time. The output is the generated live video.
[0879] Step 3:
[0880] When the server receives a request to start a live performance, it randomly generates a setlist from the song list in the database and shuffles the playing order. It uses the request to start a live performance and the song list in the database as input. It uses a random algorithm to generate the setlist and the playing order. The output is the randomly determined setlist and its playing order.
[0881] Step 4:
[0882] The server encodes, compresses, and prepares the generated live video for streaming. It uses the live video generated in step 2 and the setlist determined in step 3 as input. It prepares the video for streaming using a low-latency streaming protocol (WebRTC or HLS). The output is the encoded live video ready for streaming.
[0883] Step 5:
[0884] The device receives live video data streamed from the server in real time. As input, it receives the streaming data sent from the server. The device decodes and converts it into a format suitable for the VR headset. The output is the decoded and converted video data.
[0885] Step 6:
[0886] The device detects the user's viewpoint movements and operations in real time. Sensor information from the VR headset is used as input. Based on this, viewpoint movement and operation data is collected in real time and reflected in the video display. The output is a video display that reflects the viewpoint and operations.
[0887] Step 7:
[0888] The emotion engine analyzes the user's facial expressions and voice to recognize emotions in real time. The input is the user's facial expression data and voice data. Analysis is performed using OpenAI's API or Microsoft Azure's Face API to generate emotion data. The output is the analyzed emotion data.
[0889] Step 8:
[0890] The server dynamically changes the live video production and set list based on the emotional data sent from the emotion engine. The emotional data from the emotion engine is used as input. Based on the emotional data, lighting effects and production are adjusted, and the set list is also changed as necessary. The output is live video and production that adapts to the user's emotions.
[0891] Step 9:
[0892] A user puts on a VR headset, logs in to the system, and selects their favorite artist or live event from the app interface. The user's selection information is entered into the app as input. Once the selection is complete, a live viewing request is sent to the server. The output is a live viewing request.
[0893] Step 10:
[0894] While watching a live stream, the user changes the viewpoint, adjusts the volume, and uses interactive elements (such as recreating applause and cheers). The input is information about the user's viewpoint, volume adjustment, and use of interactive elements. Based on this, the viewing experience is dynamically adjusted. The output is a viewing experience that corresponds to the user's actions.
[0895] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0896] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0897] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0898] [Third embodiment]
[0899] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0900] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0901] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0902] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0903] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0904] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0905] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0906] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0907] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0908] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0909] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0910] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0911] This invention relates to a system that generates live video footage of an artist selected by a user and distributes that video in real time, allowing users to enjoy the artist's live performance in a virtual reality environment.
[0912] System Overview
[0913] The system mainly consists of three elements: a server, a terminal (user's viewing device), and the user. The server is responsible for generating and distributing live video, while the terminal receives and displays the video. Users participate in the live broadcast through their terminal.
[0914] Server Processing
[0915] The server first collects live data from the artist, including video, audio, stage layout, and lighting data, which is obtained from the artist and production team via a dedicated API.
[0916] The server then uses AI technology to generate live video based on the collected data. For example, by using GAN (Generative Adversarial Networks) as a generative model, the video is synthesized in real time and lighting and effects are automatically placed.
[0917] Furthermore, when the server receives a request to start a live performance, it randomly generates a setlist from the songs in the database and shuffles the order of the songs to create a unique setlist. During this process, seamless transitions appropriate for each song are also automatically generated.
[0918] Finally, the server encodes, compresses, and prepares the generated live video for streaming, delivering it to the user's device using a low-latency streaming protocol (e.g., WebRTC or HTTP Live Streaming: HLS).
[0919] Terminal handling
[0920] The device receives live video data streamed from the server in real time. The received video data is decoded and converted into a format suitable for the VR headset. The device also detects the user's viewpoint movements and interactions in real time and reflects them in the video.
[0921] User behavior
[0922] First, users put on a VR headset and log in to the system. After logging in, they select their favorite artist or live event from the app interface. Once the selection is complete, a request is sent to the server.
[0923] Once the live viewing begins, users can freely move their viewpoint while wearing the VR headset, enjoying the atmosphere of the live venue. Various operations are also possible, such as changing the viewpoint, adjusting the volume, and using interactive elements (for example, recreating the sounds of applause and cheers in VR).
[0924] Specific examples
[0925] When a user selects a live event by their favorite artist and presses the start viewing button, the server receives the request and generates a virtual live video based on the collected data. The set list and performance order are also randomly determined. The generated live video is streamed from the server to the device. The device receives the video data, decodes it, and displays it on the VR headset. The user puts on the VR headset and watches the live performance in real time while moving their viewpoint.
[0926] In this way, the system of the present invention allows users to enjoy an artist's live performance in a virtual reality environment without time or place restrictions.
[0927] The processing flow will be explained below.
[0928] Server Processing
[0929] Step 1:
[0930] The server collects the artist's existing live footage, audio, stage layout and lighting data from a database, which is then retrieved from the artist and production team via a dedicated API.
[0931] Step 2:
[0932] The server uses AI technology to generate virtual live footage in real time based on the collected data. Complex video synthesis is performed using technologies such as GAN (Generative Adversarial Networks), and lighting and effects are automatically placed.
[0933] Step 3:
[0934] When the server receives a request to start a live performance, it randomly generates a setlist from the songs in the database and shuffles the order in which they are played, resulting in a unique setlist.
[0935] Step 4:
[0936] The server encodes and compresses the generated live video to prepare it for streaming, and delivers it to the device using a low-latency streaming protocol (e.g., WebRTC or HTTP Live Streaming: HLS).
[0937] Terminal handling
[0938] Step 1:
[0939] The device receives live video data streamed from the server in real time, with minimal buffering required to maintain low latency.
[0940] Step 2:
[0941] It decodes the received video data and converts it into a format suitable for VR headsets, allowing users to enjoy the video without interruption.
[0942] Step 3:
[0943] The device detects the user's viewpoint movements and operations in real time and reflects this information in the video display, allowing the user to freely look around the virtual space.
[0944] User behavior
[0945] Step 1:
[0946] Users simply put on a VR headset and log into the dedicated application, where their user information and viewing history are automatically loaded.
[0947] Step 2:
[0948] Users select their favorite artist or live event from the app's interface and press the Start Watch button, which sends a request to the server.
[0949] Step 3:
[0950] Once the live viewing begins, the user can freely move their viewpoint while wearing the VR headset, experiencing the atmosphere of the live venue. They can change the viewpoint, adjust the volume, and use interactive elements (for example, recreating the sounds of applause and cheers in VR).
[0951] This processing flow allows users to feel as if they are actually participating in a live event.
[0952] Example 1
[0953] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0954] Conventional live video streaming systems have struggled to generate live video of a user-selected artist in real time and stream it to a viewing device. In particular, they needed to randomly determine the set list and performance order of the live video, and reflect the user's viewpoint and operations in real time. Furthermore, the lack of technology for streaming high-quality video with low latency limited the user experience.
[0955] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0956] In this invention, the server includes: means for generating live video of a musician selected by the user; means for streaming the generated live video to the user's display device; means for randomly determining the songs and performance order of the live video; means for reflecting the user's viewpoint and operations in real time; means for collecting live data from musicians and production groups via a dedicated API; means for generating live video using AI technology based on the collected data; means for randomly generating a set list from the songs in the collected data and shuffling the performance order; means for encoding and compressing the generated live video and preparing it for streaming; means for streaming the video using a low-latency streaming protocol; and means for converting the decoded video into a format suitable for a visual device, thereby enabling users to enjoy high-quality live video in real time.
[0957] A "musician" is a person whose occupation is performing or composing music.
[0958] "Live footage" refers to footage of a specific event or performance that is filmed and distributed in real time.
[0959] A "display device" is a device for displaying images, and specifically includes VR headsets and computer monitors.
[0960] "Song List" refers to the list of songs that will be performed live.
[0961] "Performance order" refers to the order in which songs are performed at a live concert.
[0962] "Viewpoint" refers to the direction and position from which the user looks in the VR space.
[0963] An "interaction" is an action or input that a user takes to interact with a system.
[0964] "Musicians and production groups" refers to people and organizations involved in organizing live performances and producing videos.
[0965] A "dedicated API" is a predefined interface for accessing specific functionality or services.
[0966] "Live data" is information related to a live event, including video, audio, stage layout, lighting data, and the like.
[0967] "AI technology" refers to analytical and generation methods that use artificial intelligence, and specific examples include Generative Adversarial Networks (GAN).
[0968] "Encoding" is the process of converting digital data into a particular format.
[0969] "Compression" is a process for reducing the amount of data.
[0970] "Streaming distribution" is a technology that transmits video and audio in real time and plays them instantly on the receiving end.
[0971] A "low-latency streaming protocol" is a communication method for minimizing delays in the transmission of video and audio, and specific examples include WebRTC and HTTP Live Streaming (HLS).
[0972] "Decoding" is the process of returning encoded data to its original format.
[0973] "Visual equipment" means a device through which a user views an image, including a VR headset or other display device.
[0974] The present invention relates to a system that generates live video footage of a musician selected by a user and distributes the video in real time, allowing users to enjoy the musician's live performance in a virtual reality environment.
[0975] System Overview
[0976] The system mainly consists of three elements: a server, a terminal (user's viewing device), and a user. The server is responsible for generating and distributing live video, while the terminal receives and displays the video. Users participate in the live broadcast through their terminal.
[0977] Server Processing
[0978] The server first collects live data from musicians, including video, audio, stage layout, and lighting data, which is obtained from musicians and production groups via a dedicated API.
[0979] The server then uses AI technology to generate live video based on the collected data. Specifically, it uses Generative Adversarial Networks (GANs) to synthesize video in real time and automatically position lighting and effects. For example, the lighting position and color dynamically change as musicians move around the stage.
[0980] Furthermore, when the server receives a request to start a live performance, it randomly generates a setlist from the songs in the database, shuffles the order of the songs, and creates a unique setlist. During this process, it also automatically generates seamless transitions appropriate for each song. For example, it adds effects that make the transition between songs feel natural.
[0981] Finally, the server encodes and compresses the generated live video and prepares it for streaming, delivering it to the user's device using a low-latency streaming protocol (e.g., WebRTC or HTTP Live Streaming: HLS).
[0982] Terminal handling
[0983] The device receives live video data streamed from the server in real time. The received video data is decoded using codecs such as H.264 and VP9 and converted into a format suitable for VR headsets. The device also detects the user's viewpoint movements and interactions in real time and reflects them in the video.
[0984] User behavior
[0985] First, the user puts on the VR headset and logs in to the system. After logging in, they select their favorite musician or live event from the app interface. Once the selection is complete, a request is sent to the server.
[0986] Once the live viewing begins, users can freely move their viewpoint while wearing the VR headset, enjoying the atmosphere of the live venue. Various operations are also possible, such as changing the viewpoint, adjusting the volume, and using interactive elements (for example, recreating applause and cheers in virtual reality).
[0987] Specific examples
[0988] Suppose a user selects "Musician A's Live Event" on the app interface and presses the Start Viewing button. The server receives the request and collects Musician A's data (video, audio, stage layout, lighting data, etc.) through a dedicated API. Based on that data, AI technology (GAN) is used to generate live video and create a random set list and appropriate transitions. The generated video is encoded and delivered to the device using a low-latency streaming protocol. The device decodes the received data and displays the video on a VR headset. The user can watch the live performance in real time while changing their viewpoint. As a concrete example, the following prompt sentence can be input to the generative AI model:
[0989] "Generate live footage of any musician. Include random setlists and lighting effects."
[0990] In this way, users can enjoy live performances of musicians in a virtual reality environment without time or location constraints.
[0991] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0992] Step 1:
[0993] The user puts on the VR headset and logs into the system.
[0994] Specific operation: The user puts on the VR headset, launches the dedicated app, and logs in by entering their user ID and password.
[0995] Input: User ID, Password
[0996] Output: User authentication completed, individual settings loaded
[0997] Step 2:
[0998] Users select artists and live events.
[0999] What it does: Users use the touchpad or voice commands to select their favorite artists or live events from the app interface.
[1000] Input: Artist name, Live event name
[1001] Output: Send selection
[1002] Step 3:
[1003] The server collects live data.
[1004] Specific operation: Based on the selected artist and live event, the server collects live data such as video, audio, stage layout, and lighting data through a dedicated API.
[1005] Input: Selected artist name, live event name
[1006] Output: Collected live data (video, audio, stage layout, lighting data, etc.)
[1007] Step 4:
[1008] The server uses AI technology to generate live footage.
[1009] How it works: The server uses Generative Adversarial Networks (GAN) to generate real-time live video based on collected live data, with lighting and other effects automatically placed.
[1010] Input: Live data collected
[1011] Output: Generated live video
[1012] Step 5:
[1013] The server randomly generates the setlist and shuffles the order in which the songs are played.
[1014] How it works: The server randomly generates a setlist from the songs in the database, shuffles the order of the songs, and automatically generates seamless transitions between songs.
[1015] Input: Songs in the database
[1016] Output: Random setlist, unique playing order, seamless transitions
[1017] Step 6:
[1018] The server encodes and compresses the generated live video.
[1019] Specific operation: The server encodes the generated live video using codecs such as H.264 or VP9 and compresses the data.
[1020] Input: Generated live video
[1021] Output: Encoded video data
[1022] Step 7:
[1023] The server streams the video.
[1024] How it works: The server delivers video data using a low-latency streaming protocol (WebRTC or HTTP Live Streaming: HLS). The server monitors network conditions and adjusts the bitrate as needed.
[1025] Input: Encoded video data
[1026] Output: Streamed video data
[1027] Step 8:
[1028] The device receives and decodes the video.
[1029] How it works: The device receives video data streamed from the server in real time, decodes it using codecs such as H.264 and VP9, and then converts the decoded data into a format suitable for the VR headset.
[1030] Input: Streamed video data
[1031] Output: Decoded video data, displayed on a VR headset
[1032] Step 9:
[1033] Users watch live.
[1034] Specific operation: Users can enjoy the live performance while wearing a VR headset and freely moving their viewpoint. They can change the viewpoint, adjust the volume, and use interactive elements (such as applause and cheers). The user's viewpoint movement and interactions are reflected in the video in real time.
[1035] Input: User's viewpoint movement, operation input
[1036] Output: Real-time changing visual and audio experience
[1037] (Application example 1)
[1038] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1039] Today, many music fans desire to enjoy artists' live performances in real time, but are often unable to attend due to physical distance, location, or even time constraints. Furthermore, existing online live streaming systems lack the immersive and interactive feel, making it difficult to provide an experience equivalent to a live performance. Furthermore, few systems are able to reflect user actions in real time and allow for free viewpoint movement.
[1040] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1041] In this invention, the server includes a means for generating live video of an artist selected by the user, a means for distributing the generated live video to the user's viewing device, a means for randomly determining the set list and performance order of the live video, a means for reflecting the user's viewpoint and operations in real time, and a means for reproducing interactive elements in a virtual reality environment. This allows users to transcend physical limitations and enjoy a realistic virtual live performance in real time, even from the comfort of their own home.
[1042] "Means for generating live video of an artist selected by a user" refers to technical means for generating live video content of a specific artist designated by a user.
[1043] "Means for delivering the generated live video to the user's viewing device" refers to the technical means for transmitting the generated live video data to the viewing terminal used by the user.
[1044] "Means for randomly determining the set list and performance order of live video" refers to a technical means for randomly determining the list of songs to be performed live and their order.
[1045] "Means for reflecting the user's viewpoint and operations in real time" refers to technical means for reflecting the user's visual viewpoint movements and interactive operations in the video in real time.
[1046] "Means for reproducing interactive elements in a virtual reality environment" refers to the technical means for reproducing interactive elements, such as applause and cheers, that can be experienced by users within a virtual reality environment.
[1047] "AI technology" refers to technology that utilizes artificial intelligence and is primarily used to generate live footage and determine set lists.
[1048] A "low-latency streaming protocol" is a communication protocol for transmitting and receiving data with minimal delay, and is a technology that is particularly applicable to live video distribution, which requires real-time performance.
[1049] This invention relates to a system that generates live video footage of an artist selected by the user and distributes it in real time. The system mainly consists of three elements: a server, a terminal (user's viewing device), and the user. This system allows users to enjoy the artist's live performance in a virtual reality environment.
[1050] System Overview
[1051] In this system, the server is responsible for generating and distributing live video, and the terminal receives and displays the video. Users participate in the live broadcast through the terminal.
[1052] Server Processing
[1053] The server first collects the artist's live performance data. This data includes video, audio, stage layout, lighting, and other data. This data is obtained from the artist and production team via a dedicated API. Next, the server uses AI technology to generate live video based on the collected data. For example, by using a generative adversarial network (GAN) as a generative model, it synthesizes video in real time and automatically places lighting and effects. Furthermore, when the server receives a request to start a live performance, it randomly generates a setlist from a list of songs in its database and shuffles the performance order to create a unique setlist. During this process, seamless transitions appropriate for each song are also automatically generated. Finally, the server encodes and compresses the generated live video and prepares it for streaming. The video is then distributed to the user's device using a low-latency streaming protocol (e.g., WebRTC or HTTP Live Streaming: HLS).
[1054] Terminal handling
[1055] The device receives live video data streamed from the server in real time. The received video data is decoded and converted into a format suitable for the VR headset. The device also detects the user's viewpoint movements and interactions in real time and reflects them in the video. This includes using Unity to process interactive elements such as viewpoint movements and clapping sounds in real time.
[1056] User behavior
[1057] First, the user puts on the VR headset and logs into the system. After logging in, they select their favorite artist or live event from the app interface. Once the selection is complete, a request is sent to the server. Once viewing of the live performance begins, the user can freely move their viewpoint while wearing the VR headset, enjoying the atmosphere of the live venue. Various operations are also possible, such as changing the viewpoint, adjusting the volume, and using interactive elements (for example, recreating applause and cheers in VR).
[1058] Specific examples
[1059] When a user selects a live event by their favorite artist and presses the start viewing button, the server receives the request and generates a virtual live video based on the collected data. The set list and performance order are also randomly determined. The generated live video is streamed from the server to the device. The device receives the video data, decodes it, and displays it on the VR headset. The user puts on the VR headset and watches the live performance in real time while moving their viewpoint. In this way, the system of the present invention allows users to enjoy an artist's live performance in a virtual reality environment without the constraints of time or place.
[1060] Prompt Sentence Examples
[1061] "To generate a virtual live video of an artist selected by the user and deliver it with low latency, synthesize the collected live data (video, audio, stage layout) in real time and stream it using WebRTC. Reflect real-time viewpoint changes and interactive elements so that users can enjoy the virtual live performance."
[1062] This allows users to enjoy an experience that feels as if they are at an actual live concert venue in a VR environment.
[1063] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1064] Step 1:
[1065] The server receives the user request
[1066] Input: The user selects the artist's live performance they want to watch and submits a request.
[1067] Processing: The server receives the request data from the user, which may include the artist name, the date and time of the show, and the user's authentication information.
[1068] Output: Request data to be used in the next step.
[1069] Step 2:
[1070] Live Data Collection
[1071] Input: Request data to generate artist gigs.
[1072] Processing: The server collects live data (video, audio, stage layout, lighting data, etc.) from artists and production teams via a dedicated API.
[1073] Output: Live data collected.
[1074] Step 3:
[1075] Live video generation
[1076] Input: Live data collected.
[1077] Processing: The server uses a generative AI model (e.g., GAN) to generate live footage in real time, including compositing the footage and adding lighting and effects.
[1078] Output: The generated live video data.
[1079] Step 4:
[1080] Setlist and performance order randomly determined
[1081] Input: Live video data and artist song list.
[1082] Processing: The server retrieves the song list from the database, randomly generates a setlist, shuffles the order of the songs, and automatically creates seamless transitions between each song.
[1083] Output: Live video data with setlist and transitions.
[1084] Step 5:
[1085] Encoding and preparing live video for distribution
[1086] Input: Live video data with setlist and transitions.
[1087] Processing: The server encodes the video in H264 format and prepares the data for distribution using a low-latency streaming protocol such as WebRTC.
[1088] Output: Encoded live video data.
[1089] Step 6:
[1090] Live video streaming
[1091] Input: Encoded live video data.
[1092] Processing: The server streams the prepared video data to the user's device in real time.
[1093] Output: Live video data sent to user terminal.
[1094] Step 7:
[1095] Receiving and decoding live video data
[1096] Input: Live video data streamed from the server.
[1097] Processing: The device decodes the received video data and converts it into a format suitable for the VR headset.
[1098] Output: Live video data that can be displayed on a VR headset.
[1099] Step 8:
[1100] Reflects user viewpoint movements and operations
[1101] Input: User gaze movement data and interaction data.
[1102] Processing: The device uses Unity to reflect interactive elements such as the user's viewpoint and clapping sounds in the video in real time.
[1103] Output: Live video data reflecting the interaction.
[1104] Step 9:
[1105] User's live viewing experience
[1106] Input: Live video data reflecting viewpoint movement.
[1107] Processing: Users can watch the live stream through a VR headset, freely moving their viewpoint in real time, and can also adjust the volume and use interactive elements (such as applause and cheering).
[1108] Output: An immersive virtual live experience.
[1109] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1110] The present invention relates to a system that generates live video of an artist selected by the user and distributes the video in real time. The present invention can also be combined with an emotion engine that recognizes the user's emotions to further improve the user experience.
[1111] System configuration
[1112] The system mainly consists of a server, a terminal (user's viewing device), and an emotion engine. The server is responsible for generating and distributing live video, the terminal is responsible for receiving and displaying the video, and the emotion engine recognizes user emotions in real time.
[1113] Server Processing
[1114] The server first collects live data from the artist, including video, audio, stage layout, and lighting data, which is obtained from the artist and production team via a dedicated API.
[1115] The server then uses AI technology to generate live video based on the collected data. Generative Adversarial Networks (GANs) and other generative models are used to synthesize the video in real time. Lighting and effects are also automatically placed.
[1116] Furthermore, when the server receives a request to start a live performance, it randomly generates a setlist from the songs in the database and shuffles the order in which they are played, resulting in a unique setlist.
[1117] The server encodes and compresses the generated live video, prepares it for streaming, and delivers it to the user's device using a low-latency streaming protocol (e.g., WebRTC or HTTP Live Streaming: HLS).
[1118] Terminal handling
[1119] The device receives live video data streamed from the server in real time. The received video data is decoded and converted into a format suitable for the VR headset. The device also detects the user's viewpoint movements and interactions in real time and reflects them in the video display.
[1120] Emotion engine processing
[1121] The emotion engine analyzes the user's facial expressions and voice to recognize emotions in real time. This data is sent to a server and used to dynamically change the live video production, set list, and interactive elements (such as specific lighting and sound effects).
[1122] User behavior
[1123] First, users put on a VR headset and log in to the system. After logging in, they select their favorite artist or live event from the app interface. Once the selection is complete, a request is sent to the server.
[1124] Once the live viewing begins, users can freely move their viewpoint while wearing the VR headset, enjoying the atmosphere of the live venue. Various operations are also possible, such as changing the viewpoint, adjusting the volume, and using interactive elements (for example, recreating applause and cheers in VR).
[1125] The emotion engine detects the user's emotions and changes the live video presentation in real time. For example, if the user becomes excited, the lighting and effects will be enhanced to further enhance the user's experience.
[1126] Specific examples
[1127] For example, when a user selects a live event by their favorite artist and presses the start button, the server receives the request and generates a virtual live video based on the collected data. The set list and performance order are also randomly determined. Furthermore, an emotion engine analyzes the user's emotions in real time and sends that data to the server. The server then dynamically adjusts the live video production and set list based on this data.
[1128] In this way, the system of the present invention allows users to enjoy live performances by artists in a virtual reality environment without the constraints of time or place, and by using an emotion engine, the user's experience can be further enriched.
[1129] The processing flow will be explained below.
[1130] Server Processing
[1131] Step 1:
[1132] The server collects artists' existing live footage, audio sources, stage layout, and lighting data from a database, which is then retrieved from the artists and production teams via a dedicated API.
[1133] Step 2:
[1134] Based on the collected data, the server uses AI technology to generate virtual live footage. For example, it uses a generative model based on GAN (Generative Adversarial Networks) to synthesize the footage in real time. Lighting and effects are also automatically placed.
[1135] Step 3:
[1136] When the server receives a request to start a live performance, it randomly generates a setlist from the songs in its database and shuffles the order in which they are played, resulting in a unique setlist.
[1137] Step 4:
[1138] The server encodes and compresses the generated live video, preparing it for streaming, and delivers it to the user's device using a low-latency streaming protocol (e.g., WebRTC or HTTP Live Streaming: HLS).
[1139] Step 5:
[1140] The server receives data from the emotion engine and adjusts the live video production (lighting, effects, set list changes, etc.) in real time based on the user's emotions.
[1141] Terminal handling
[1142] Step 1:
[1143] The device receives live video data streamed from the server in real time, with minimal buffering required to maintain low latency.
[1144] Step 2:
[1145] It decodes the received video data and converts it into a format suitable for VR headsets, allowing users to enjoy the video without interruption.
[1146] Step 3:
[1147] The device detects the user's viewpoint movements and operations in real time and reflects this information in the video display, allowing the user to freely look around the virtual space.
[1148] Step 4:
[1149] The device captures the user's facial expressions, voice, and other data and sends it to the emotion engine, which analyzes the data and generates real-time emotion data.
[1150] Emotion engine processing
[1151] Step 1:
[1152] The emotion engine receives the user's facial expressions and voice data sent from the device and analyzes it in real time.
[1153] Step 2:
[1154] Based on the analysis results, the system recognizes the user's emotions and sends the data to the server. For example, if the user is excited, the emotional data is transmitted to the server.
[1155] User behavior
[1156] Step 1:
[1157] Users put on a VR headset and log in to a dedicated application, where their user information and viewing history are automatically loaded.
[1158] Step 2:
[1159] Users select their favorite artist or live event from the app's interface and press the Start Watch button, which sends a request to the server.
[1160] Step 3:
[1161] Once the live viewing begins, users can freely move their viewpoint while wearing the VR headset and enjoy the atmosphere of the live venue. They can change the viewpoint, adjust the volume, and use interactive elements (for example, recreating applause and cheers in VR).
[1162] Step 4:
[1163] The emotion engine analyzes the user's facial expressions and voice, and the results are sent to the server. If the user is excited, the lighting and effects will be enhanced, further enriching the experience.
[1164] In this way, the system of the present invention not only allows users to enjoy live performances by artists in a virtual reality environment without time or place constraints, but also uses an emotion engine to provide a dynamic live experience that reflects the user's real-time emotions.
[1165] Example 2
[1166] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1167] Conventional live video streaming systems allow users to watch live footage of artists of their choice, but the experience relies on limited visual and auditory information. As a result, real-time video production is not tailored to the user's emotions or interactions, resulting in a lack of depth of experience. Furthermore, because the set list and performance order are fixed, watching the same live footage repeatedly can lose its freshness, creating a problem.
[1168] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1169] In this invention, the server includes means for generating live video of an artist selected by the user, means for delivering the generated live video to the user's viewing device, means for randomly determining the set list and performance order of the live video, means for reflecting the user's viewpoint and operations in real time, and means for recognizing the user's emotions in real time and dynamically changing the direction of the live video, thereby enabling users to enjoy a fresh and realistic live experience that is individually customized.
[1170] "User-selected artists" are musical artists selected by users of the system according to their own preferences.
[1171] "Live footage" is video that visually and aurally captures a musical artist's performance or live performance.
[1172] "Means of generation" is a general term for methods and devices that use specific algorithms or AI technology to create live footage.
[1173] A "viewing device" is a device that a user uses to view video and audio, such as a VR headset or smartphone.
[1174] "Means of distribution" refers to the methods and technologies for sending the generated live video to the user's viewing device via a network such as the Internet.
[1175] A "set list" is a list of songs that will be performed live.
[1176] The "play order" is the rule that determines the order in which songs in a set list are played.
[1177] A "random determination means" is a method or device that randomly determines the set list or performance order using a specific algorithm or random number generation technology.
[1178] "Means for reflecting viewpoints and operations in real time" refers to methods and technologies that instantly detect the user's viewpoint movements and interactions and reflect them in the video display.
[1179] "Means for recognizing emotions in real time and dynamically changing the presentation of live video" refers to methods and technologies that analyze emotions from the user's facial expressions, voice, etc., and instantly change the presentation of live video based on the results.
[1180] "AI technology" refers to any technology that uses artificial intelligence, including machine learning and deep learning in particular.
[1181] "Low latency streaming protocol" is a protocol that minimizes data delays and is a technology that enables real-time communication.
[1182] This invention relates to a system that generates live video of an artist selected by the user and distributes the video in real time. Furthermore, this invention can further enhance the user experience by incorporating an emotion engine that recognizes the user's emotions. This system consists of three main components: a server, a terminal, and an emotion engine.
[1183] System configuration
[1184] The system includes a server, a terminal (user's viewing device), and an emotion engine. The server is responsible for generating and distributing live video, and the terminal receives and displays the video. The emotion engine recognizes user emotions in real time.
[1185] Server Processing
[1186] The server first collects live data from the artist. This data includes video, audio, stage layout, and lighting data. This data is obtained from the artist and production team via a dedicated API. Based on the collected data, the server uses AI technology (such as Generative Adversarial Networks: GAN) to generate live video. During this generation process, lighting and effects are also automatically placed.
[1187] Furthermore, when the server receives a request to start a live performance, it randomly generates a setlist from the song list in the database and shuffles the performance order to create a unique setlist.The generated live video is then encoded, compressed, and distributed to the user's device using a low-latency streaming protocol (e.g., WebRTC or HTTP Live Streaming: HLS).
[1188] Terminal handling
[1189] The device receives live video data streamed from the server in real time. The received video data is decoded and converted into a format suitable for the VR headset. The device also detects the user's viewpoint movements and interactions in real time and reflects them in the video display.
[1190] Emotion engine processing
[1191] The emotion engine analyzes the user's facial expressions and voice to recognize emotions in real time. This data is sent to a server and used to dynamically change the live video production, set list, and interactive elements (such as specific lighting and sound effects).
[1192] User behavior
[1193] First, users put on a VR headset and log in to the system. After logging in, they select their favorite artist or live event from the app interface. Once the selection is complete, a request is sent to the server. Once viewing of the live performance begins, users can freely move their viewpoint while wearing the VR headset, enjoying the atmosphere of the live venue. They can also change the viewpoint, adjust the volume, and use interactive elements (for example, recreating applause and cheers in VR). The emotion engine detects the user's emotions, which then changes the live video presentation in real time. For example, if the user becomes excited, lighting and effects are enhanced, further improving the user experience.
[1194] Specific examples
[1195] For example, when a user selects a live event by their favorite artist and presses the start button, the server receives the request and uses AI technology to generate a virtual live video based on the collected data. The set list and performance order are also randomly determined. After the live video is generated, the server encodes the video and streams it to the user with low latency. When the emotion engine detects the user's excitement, the server dynamically changes the live performance by enhancing lighting and effects.
[1196] Prompt Sentence Examples
[1197] Below are some example prompts to input to a generative AI model:
[1198] "Please explain a system that generates a virtual live video of a live event of an artist selected by the user based on related data, and dynamically changes the live performance by analyzing the user's emotions in real time using an emotion engine."
[1199] In this way, the system of the present invention allows users to enjoy live performances of artists in a virtual reality environment regardless of location. Furthermore, the use of an emotion engine can further enrich the user experience.
[1200] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1201] Step 1:
[1202] Live Data Collection
[1203] The server collects live data from artists, including video, audio, stage layout, and lighting data, and this data is obtained through a dedicated API.
[1204] Input: API requests from artists and production teams
[1205] Output: Live video, audio, stage layout, lighting data
[1206] Specific behavior:
[1207] The server sends a request to an API endpoint.
[1208] Receive and parse the JSON data returned as a response.
[1209] The acquired data is stored in an internal database.
[1210] Step 2:
[1211] Live video generation
[1212] The server uses AI techniques (such as Generative Adversarial Networks: GAN) to generate live footage based on the collected data, taking into account stage layout and lighting data during the generation process.
[1213] Input: Live video, audio, stage layout, lighting data
[1214] Output: Generated live video
[1215] Specific behavior:
[1216] The server inputs the collected data into the AI model.
[1217] The AI model generates images based on the specified parameters.
[1218] Lighting and effects are automatically added to the generated video data.
[1219] Step 3:
[1220] Handling Go Live Requests
[1221] The server receives a request from the user to start a live stream, which includes an artist and event identifier.
[1222] Input: User's request to go live
[1223] Output: Information about a specific artist or event
[1224] Specific behavior:
[1225] The server receives the request and loads the corresponding artist and event data.
[1226] Prepare to move on to the next step.
[1227] Step 4:
[1228] Deciding on the set list and performance order
[1229] The server randomly generates a setlist from the songs in the database and shuffles the order in which they are played.
[1230] Input: Song list
[1231] Output: Randomly determined setlist and playing order
[1232] Specific behavior:
[1233] The server retrieves the song list from the database.
[1234] Use a random algorithm to generate setlists and shuffle song order.
[1235] Step 5:
[1236] Encode video and prepare for streaming
[1237] The server encodes, compresses, and converts the generated live video into a streaming format, preparing it for distribution using a low-latency streaming protocol (e.g., WebRTC or HTTP Live Streaming: HLS).
[1238] Input: Generated live video
[1239] Output: Video encoded for streaming
[1240] Specific behavior:
[1241] The server encodes the video data in a specific codec.
[1242] The encoded data is sent to a streaming server.
[1243] Step 6:
[1244] Live video streaming
[1245] The server streams the encoded live video to the user's device in real time.
[1246] Input: Encoded live video
[1247] Output: Video streamed to user device
[1248] Specific behavior:
[1249] The server checks the streaming protocol settings and transmits the encoded video data in real time.
[1250] Monitor your streaming performance and make adjustments as needed.
[1251] Step 7:
[1252] Handling user gaze movements and interactions
[1253] The device receives live video data streamed from the server in real time, decodes it, and converts it into a format suitable for the VR headset. It also detects the user's viewpoint movements and interactions and reflects them in the video display.
[1254] Input: Streaming live video data, user's viewpoint movement and interaction information
[1255] Output: Image displayed on VR headset
[1256] Specific behavior:
[1257] The terminal decodes the received video data.
[1258] It detects the user's viewpoint movement and adjusts the displayed content in real time.
[1259] Processes user interaction information (e.g., button presses).
[1260] Step 8:
[1261] Emotion analysis using an emotion engine
[1262] The emotion engine analyzes the user's facial expressions and voice to recognize their emotions in real time, and this emotion data is sent to the server.
[1263] Input: User's facial expressions and voice data
[1264] Output: Parsed emotion data
[1265] Specific behavior:
[1266] The emotion engine captures the user's camera footage and microphone audio.
[1267] The acquired data is analyzed in real time to recognize the emotional state.
[1268] Step 9:
[1269] Dynamic changes to live performances
[1270] The server dynamically changes the live video presentation based on the emotional data received from the emotion engine. For example, if the user is excited, the lighting and effects will be enhanced.
[1271] Input: Emotion data received from the emotion engine
[1272] Output: Dynamically modified live video rendition
[1273] Specific behavior:
[1274] The server analyzes the emotion data and determines the corresponding change in presentation.
[1275] Implementing real-time performance changes to improve the user experience.
[1276] (Application example 2)
[1277] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1278] Conventional live video streaming systems make it difficult for users to experience the content interactively in real time, limiting the viewing experience. Furthermore, they lack the ability to dynamically change the presentation to reflect user emotions and reactions, making it difficult to maximize user satisfaction. Furthermore, there are technical challenges in delivering high-quality video with low latency.
[1279] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for generating live video of an artist selected by the user, means for delivering the generated live video to the user's viewing device, means for randomly determining the set list and performance order of the live video, means for reflecting the user's viewpoint and operations in real time, means for recognizing the user's emotions and dynamically changing the production and set list of the live video based on that data, and means for delivering the live video using a low-latency streaming protocol. This allows users to enjoy a more interactive and personalized live experience in real time.
[1280] "User" means any person or entity that uses the System to view live video.
[1281] "Artists" are musicians or performers who are the subject of live footage.
[1282] "Live footage" refers to footage that visually records an artist's performance.
[1283] "Viewing Device" means the hardware device through which a User views the live feed. Examples include smartphones and head-mounted displays.
[1284] A "set list" is a list of songs that will be played at a live concert.
[1285] "Performance order" refers to the order in which songs are performed on the set list.
[1286] "Real time" refers to the time in which an event that occurs is processed and reflected immediately without delay.
[1287] "Emotion recognition" is the process of analyzing a user's facial expressions and voice to identify their emotional state.
[1288] "Production" refers to the visual and auditory effects, such as lighting, effects, and sound, used in live footage.
[1289] "Dynamic change" refers to the ability of the system to make immediate changes in response to specific conditions or situations.
[1290] A "low latency streaming protocol" is a communication protocol that minimizes delays in sending and receiving data.
[1291] "AI technology" refers to all technologies based on artificial intelligence, and in this invention, it particularly refers to technologies that perform generative models and emotion recognition.
[1292] A "prompt sentence" is an input sentence that gives instructions to a generative AI model.
[1293] This invention relates to a system that generates live video of an artist selected by the user and distributes the video in real time. Furthermore, by combining this invention with an emotion engine that recognizes the user's emotions, the user experience can be further improved.
[1294] System configuration
[1295] The system mainly consists of a server, a terminal (user's viewing device), and an emotion engine. The server is responsible for generating and distributing live video, the terminal is responsible for receiving and displaying the video, and the emotion engine recognizes user emotions in real time.
[1296] Server Processing
[1297] The server first collects live data from the artist. This data includes video, audio, stage layout, and lighting data. This data is obtained from the artist and production team via a dedicated API. The server then uses AI technology to generate live video based on the collected data. Generative Adversarial Networks (GAN) and other generative models are used to synthesize video in real time. Lighting and effects are also automatically placed.
[1298] Furthermore, when the server receives a request to start a live performance, it randomly generates a setlist from the song list in the database and shuffles the performance order. This process creates a unique setlist. The server then encodes and compresses the generated live video and prepares it for streaming. The video is then delivered to the user's device using a low-latency streaming protocol (e.g., WebRTC or HTTP Live Streaming: HLS).
[1299] Terminal handling
[1300] The device receives live video data streamed from the server in real time. The received video data is decoded and converted into a format suitable for the VR headset. The device also detects the user's viewpoint movements and interactions in real time and reflects them in the video display.
[1301] Emotion engine processing
[1302] The emotion engine analyzes the user's facial expressions, voice, and other data to recognize emotions in real time. It uses specific APIs, such as OpenAI's API or Microsoft Azure's Face API. This data is sent to a server and used to dynamically change the live video presentation, set list, and interactive elements (such as specific lighting and sound effects).
[1303] User behavior
[1304] First, the user puts on a VR headset and logs into the system. After logging in, they select their favorite artist or live event from the app's interface. Once the selection is complete, a request is sent to the server. Once viewing of the live performance begins, the user can freely move their viewpoint while wearing the VR headset, enjoying the atmosphere of the live venue. Various operations are also possible, such as changing the viewpoint, adjusting the volume, and using interactive elements (for example, recreating applause and cheers in VR).
[1305] The emotion engine detects the user's emotions and changes the live video presentation in real time. For example, if the user becomes excited, the lighting and effects will be enhanced to further enhance the user's experience.
[1306] Specific examples
[1307] For example, when a user selects a live event by their favorite artist and presses the start button, the server receives the request and generates a virtual live video based on the collected data. The set list and performance order are also randomly determined. Furthermore, an emotion engine analyzes the user's emotions in real time and sends that data to the server. The server then dynamically adjusts the live video production and set list based on this data.
[1308] Examples of prompts include:
[1309] "Please implement an algorithm that detects a smile on the user's face, sends that information to the server, and enhances the lighting effects on the live video."
[1310] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1311] Step 1:
[1312] The server collects live data from artists. As input, it receives live video, audio, stage layout, and lighting data provided by artists and production teams via a dedicated API. This data is integrated within the server and prepared as the basis for generating live video. The output is an integrated live data set.
[1313] Step 2:
[1314] The server generates live video using a generative AI model (e.g., GAN). It uses the live data collected in step 1 as input. The AI model takes video, audio, stage layout, and lighting data as input and generates a synthesized live video in real time. The output is the generated live video.
[1315] Step 3:
[1316] When the server receives a request to start a live performance, it randomly generates a setlist from the song list in the database and shuffles the playing order. It uses the request to start a live performance and the song list in the database as input. It uses a random algorithm to generate the setlist and the playing order. The output is the randomly determined setlist and its playing order.
[1317] Step 4:
[1318] The server encodes, compresses, and prepares the generated live video for streaming. It uses the live video generated in step 2 and the setlist determined in step 3 as input. It prepares the video for streaming using a low-latency streaming protocol (WebRTC or HLS). The output is the encoded live video ready for streaming.
[1319] Step 5:
[1320] The device receives live video data streamed from the server in real time. As input, it receives the streaming data sent from the server. The device decodes and converts it into a format suitable for the VR headset. The output is the decoded and converted video data.
[1321] Step 6:
[1322] The device detects the user's viewpoint movements and operations in real time. Sensor information from the VR headset is used as input. Based on this, viewpoint movement and operation data is collected in real time and reflected in the video display. The output is a video display that reflects the viewpoint and operations.
[1323] Step 7:
[1324] The emotion engine analyzes the user's facial expressions and voice to recognize emotions in real time. The input is the user's facial expression data and voice data. Analysis is performed using OpenAI's API or Microsoft Azure's Face API to generate emotion data. The output is the analyzed emotion data.
[1325] Step 8:
[1326] The server dynamically changes the live video production and set list based on the emotional data sent from the emotion engine. The emotional data from the emotion engine is used as input. Based on the emotional data, lighting effects and production are adjusted, and the set list is also changed as necessary. The output is live video and production that adapts to the user's emotions.
[1327] Step 9:
[1328] A user puts on a VR headset, logs in to the system, and selects their favorite artist or live event from the app interface. The user's selection information is entered into the app as input. Once the selection is complete, a live viewing request is sent to the server. The output is a live viewing request.
[1329] Step 10:
[1330] While watching a live stream, the user changes the viewpoint, adjusts the volume, and uses interactive elements (such as recreating applause and cheers). The input is information about the user's viewpoint, volume adjustment, and use of interactive elements. Based on this, the viewing experience is dynamically adjusted. The output is a viewing experience that corresponds to the user's actions.
[1331] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1332] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1333] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1334] [Fourth embodiment]
[1335] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1336] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1337] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1338] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1339] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1340] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1341] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1342] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1343] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1344] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1345] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1346] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1347] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1348] This invention relates to a system that generates live video footage of an artist selected by a user and distributes that video in real time, allowing users to enjoy the artist's live performance in a virtual reality environment.
[1349] System Overview
[1350] The system mainly consists of three elements: a server, a terminal (user's viewing device), and the user. The server is responsible for generating and distributing live video, while the terminal receives and displays the video. Users participate in the live broadcast through their terminal.
[1351] Server Processing
[1352] The server first collects live data from the artist, including video, audio, stage layout, and lighting data, which is obtained from the artist and production team via a dedicated API.
[1353] The server then uses AI technology to generate live video based on the collected data. For example, by using GAN (Generative Adversarial Networks) as a generative model, the video is synthesized in real time and lighting and effects are automatically placed.
[1354] Furthermore, when the server receives a request to start a live performance, it randomly generates a setlist from the songs in the database and shuffles the order of the songs to create a unique setlist. During this process, seamless transitions appropriate for each song are also automatically generated.
[1355] Finally, the server encodes, compresses, and prepares the generated live video for streaming, delivering it to the user's device using a low-latency streaming protocol (e.g., WebRTC or HTTP Live Streaming: HLS).
[1356] Terminal handling
[1357] The device receives live video data streamed from the server in real time. The received video data is decoded and converted into a format suitable for the VR headset. The device also detects the user's viewpoint movements and interactions in real time and reflects them in the video.
[1358] User behavior
[1359] First, users put on a VR headset and log in to the system. After logging in, they select their favorite artist or live event from the app interface. Once the selection is complete, a request is sent to the server.
[1360] Once the live viewing begins, users can freely move their viewpoint while wearing the VR headset, enjoying the atmosphere of the live venue. Various operations are also possible, such as changing the viewpoint, adjusting the volume, and using interactive elements (for example, recreating the sounds of applause and cheers in VR).
[1361] Specific examples
[1362] When a user selects a live event by their favorite artist and presses the start viewing button, the server receives the request and generates a virtual live video based on the collected data. The set list and performance order are also randomly determined. The generated live video is streamed from the server to the device. The device receives the video data, decodes it, and displays it on the VR headset. The user puts on the VR headset and watches the live performance in real time while moving their viewpoint.
[1363] In this way, the system of the present invention allows users to enjoy an artist's live performance in a virtual reality environment without time or place restrictions.
[1364] The processing flow will be explained below.
[1365] Server Processing
[1366] Step 1:
[1367] The server collects the artist's existing live footage, audio, stage layout and lighting data from a database, which is then retrieved from the artist and production team via a dedicated API.
[1368] Step 2:
[1369] The server uses AI technology to generate virtual live footage in real time based on the collected data. Complex video synthesis is performed using technologies such as GAN (Generative Adversarial Networks), and lighting and effects are automatically placed.
[1370] Step 3:
[1371] When the server receives a request to start a live performance, it randomly generates a setlist from the songs in the database and shuffles the order in which they are played, resulting in a unique setlist.
[1372] Step 4:
[1373] The server encodes and compresses the generated live video to prepare it for streaming, and delivers it to the device using a low-latency streaming protocol (e.g., WebRTC or HTTP Live Streaming: HLS).
[1374] Terminal handling
[1375] Step 1:
[1376] The device receives live video data streamed from the server in real time, with minimal buffering required to maintain low latency.
[1377] Step 2:
[1378] It decodes the received video data and converts it into a format suitable for VR headsets, allowing users to enjoy the video without interruption.
[1379] Step 3:
[1380] The device detects the user's viewpoint movements and operations in real time and reflects this information in the video display, allowing the user to freely look around the virtual space.
[1381] User behavior
[1382] Step 1:
[1383] Users simply put on a VR headset and log into the dedicated application, where their user information and viewing history are automatically loaded.
[1384] Step 2:
[1385] Users select their favorite artist or live event from the app's interface and press the Start Watch button, which sends a request to the server.
[1386] Step 3:
[1387] Once the live viewing begins, the user can freely move their viewpoint while wearing the VR headset, experiencing the atmosphere of the live venue. They can change the viewpoint, adjust the volume, and use interactive elements (for example, recreating the sounds of applause and cheers in VR).
[1388] This processing flow allows users to feel as if they are actually participating in a live event.
[1389] Example 1
[1390] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1391] Conventional live video streaming systems have struggled to generate live video of a user-selected artist in real time and stream it to a viewing device. In particular, they needed to randomly determine the set list and performance order of the live video, and reflect the user's viewpoint and operations in real time. Furthermore, the lack of technology for streaming high-quality video with low latency limited the user experience.
[1392] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1393] In this invention, the server includes: means for generating live video of a musician selected by the user; means for streaming the generated live video to the user's display device; means for randomly determining the songs and performance order of the live video; means for reflecting the user's viewpoint and operations in real time; means for collecting live data from musicians and production groups via a dedicated API; means for generating live video using AI technology based on the collected data; means for randomly generating a set list from the songs in the collected data and shuffling the performance order; means for encoding and compressing the generated live video and preparing it for streaming; means for streaming the video using a low-latency streaming protocol; and means for converting the decoded video into a format suitable for a visual device, thereby enabling users to enjoy high-quality live video in real time.
[1394] A "musician" is a person whose occupation is performing or composing music.
[1395] "Live footage" refers to footage of a specific event or performance that is filmed and distributed in real time.
[1396] A "display device" is a device for displaying images, and specifically includes VR headsets and computer monitors.
[1397] "Song List" refers to the list of songs that will be performed live.
[1398] "Performance order" refers to the order in which songs are performed at a live concert.
[1399] "Viewpoint" refers to the direction and position from which the user looks in the VR space.
[1400] An "interaction" is an action or input that a user takes to interact with a system.
[1401] "Musicians and production groups" refers to people and organizations involved in organizing live performances and producing videos.
[1402] A "dedicated API" is a predefined interface for accessing specific functionality or services.
[1403] "Live data" is information related to a live event, including video, audio, stage layout, lighting data, and the like.
[1404] "AI technology" refers to analytical and generation methods that use artificial intelligence, and specific examples include Generative Adversarial Networks (GAN).
[1405] "Encoding" is the process of converting digital data into a particular format.
[1406] "Compression" is a process for reducing the amount of data.
[1407] "Streaming distribution" is a technology that transmits video and audio in real time and plays them instantly on the receiving end.
[1408] A "low-latency streaming protocol" is a communication method for minimizing delays in the transmission of video and audio, and specific examples include WebRTC and HTTP Live Streaming (HLS).
[1409] "Decoding" is the process of returning encoded data to its original format.
[1410] "Visual equipment" means a device through which a user views an image, including a VR headset or other display device.
[1411] The present invention relates to a system that generates live video footage of a musician selected by a user and distributes the video in real time, allowing users to enjoy the musician's live performance in a virtual reality environment.
[1412] System Overview
[1413] The system mainly consists of three elements: a server, a terminal (user's viewing device), and a user. The server is responsible for generating and distributing live video, while the terminal receives and displays the video. Users participate in the live broadcast through their terminal.
[1414] Server Processing
[1415] The server first collects live data from musicians, including video, audio, stage layout, and lighting data, which is obtained from musicians and production groups via a dedicated API.
[1416] The server then uses AI technology to generate live video based on the collected data. Specifically, it uses Generative Adversarial Networks (GANs) to synthesize video in real time and automatically position lighting and effects. For example, the lighting position and color dynamically change as musicians move around the stage.
[1417] Furthermore, when the server receives a request to start a live performance, it randomly generates a setlist from the songs in the database, shuffles the order of the songs, and creates a unique setlist. During this process, it also automatically generates seamless transitions appropriate for each song. For example, it adds effects that make the transition between songs feel natural.
[1418] Finally, the server encodes and compresses the generated live video and prepares it for streaming, delivering it to the user's device using a low-latency streaming protocol (e.g., WebRTC or HTTP Live Streaming: HLS).
[1419] Terminal handling
[1420] The device receives live video data streamed from the server in real time. The received video data is decoded using codecs such as H.264 and VP9 and converted into a format suitable for VR headsets. The device also detects the user's viewpoint movements and interactions in real time and reflects them in the video.
[1421] User behavior
[1422] First, the user puts on the VR headset and logs in to the system. After logging in, they select their favorite musician or live event from the app interface. Once the selection is complete, a request is sent to the server.
[1423] Once the live viewing begins, users can freely move their viewpoint while wearing the VR headset, enjoying the atmosphere of the live venue. Various operations are also possible, such as changing the viewpoint, adjusting the volume, and using interactive elements (for example, recreating applause and cheers in virtual reality).
[1424] Specific examples
[1425] Suppose a user selects "Musician A's Live Event" on the app interface and presses the Start Viewing button. The server receives the request and collects Musician A's data (video, audio, stage layout, lighting data, etc.) through a dedicated API. Based on that data, AI technology (GAN) is used to generate live video and create a random set list and appropriate transitions. The generated video is encoded and delivered to the device using a low-latency streaming protocol. The device decodes the received data and displays the video on a VR headset. The user can watch the live performance in real time while changing their viewpoint. As a concrete example, the following prompt sentence can be input to the generative AI model:
[1426] "Generate live footage of any musician. Include random setlists and lighting effects."
[1427] In this way, users can enjoy live performances of musicians in a virtual reality environment without time or location constraints.
[1428] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1429] Step 1:
[1430] The user puts on the VR headset and logs into the system.
[1431] Specific operation: The user puts on the VR headset, launches the dedicated app, and logs in by entering their user ID and password.
[1432] Input: User ID, Password
[1433] Output: User authentication completed, individual settings loaded
[1434] Step 2:
[1435] Users select artists and live events.
[1436] What it does: Users use the touchpad or voice commands to select their favorite artists or live events from the app interface.
[1437] Input: Artist name, Live event name
[1438] Output: Send selection
[1439] Step 3:
[1440] The server collects live data.
[1441] Specific operation: Based on the selected artist and live event, the server collects live data such as video, audio, stage layout, and lighting data through a dedicated API.
[1442] Input: Selected artist name, live event name
[1443] Output: Collected live data (video, audio, stage layout, lighting data, etc.)
[1444] Step 4:
[1445] The server uses AI technology to generate live footage.
[1446] How it works: The server uses Generative Adversarial Networks (GAN) to generate real-time live video based on collected live data, with lighting and other effects automatically placed.
[1447] Input: Live data collected
[1448] Output: Generated live video
[1449] Step 5:
[1450] The server randomly generates the setlist and shuffles the order in which the songs are played.
[1451] How it works: The server randomly generates a setlist from the songs in the database, shuffles the order of the songs, and automatically generates seamless transitions between songs.
[1452] Input: Songs in the database
[1453] Output: Random setlist, unique playing order, seamless transitions
[1454] Step 6:
[1455] The server encodes and compresses the generated live video.
[1456] Specific operation: The server encodes the generated live video using codecs such as H.264 or VP9 and compresses the data.
[1457] Input: Generated live video
[1458] Output: Encoded video data
[1459] Step 7:
[1460] The server streams the video.
[1461] How it works: The server delivers video data using a low-latency streaming protocol (WebRTC or HTTP Live Streaming: HLS). The server monitors network conditions and adjusts the bitrate as needed.
[1462] Input: Encoded video data
[1463] Output: Streamed video data
[1464] Step 8:
[1465] The device receives and decodes the video.
[1466] How it works: The device receives video data streamed from the server in real time, decodes it using codecs such as H.264 and VP9, and then converts the decoded data into a format suitable for the VR headset.
[1467] Input: Streamed video data
[1468] Output: Decoded video data, displayed on a VR headset
[1469] Step 9:
[1470] Users watch live.
[1471] Specific operation: Users can enjoy the live performance while wearing a VR headset and freely moving their viewpoint. They can change the viewpoint, adjust the volume, and use interactive elements (such as applause and cheers). The user's viewpoint movement and interactions are reflected in the video in real time.
[1472] Input: User's viewpoint movement, operation input
[1473] Output: Real-time changing visual and audio experience
[1474] (Application example 1)
[1475] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1476] Today, many music fans desire to enjoy artists' live performances in real time, but are often unable to attend due to physical distance, location, or even time constraints. Furthermore, existing online live streaming systems lack the immersive and interactive feel, making it difficult to provide an experience equivalent to a live performance. Furthermore, few systems are able to reflect user actions in real time and allow for free viewpoint movement.
[1477] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1478] In this invention, the server includes a means for generating live video of an artist selected by the user, a means for distributing the generated live video to the user's viewing device, a means for randomly determining the set list and performance order of the live video, a means for reflecting the user's viewpoint and operations in real time, and a means for reproducing interactive elements in a virtual reality environment. This allows users to transcend physical limitations and enjoy a realistic virtual live performance in real time, even from the comfort of their own home.
[1479] "Means for generating live video of an artist selected by a user" refers to technical means for generating live video content of a specific artist designated by a user.
[1480] "Means for delivering the generated live video to the user's viewing device" refers to the technical means for transmitting the generated live video data to the viewing terminal used by the user.
[1481] "Means for randomly determining the set list and performance order of live video" refers to a technical means for randomly determining the list of songs to be performed live and their order.
[1482] "Means for reflecting the user's viewpoint and operations in real time" refers to technical means for reflecting the user's visual viewpoint movements and interactive operations in the video in real time.
[1483] "Means for reproducing interactive elements in a virtual reality environment" refers to the technical means for reproducing interactive elements, such as applause and cheers, that can be experienced by users within a virtual reality environment.
[1484] "AI technology" refers to technology that utilizes artificial intelligence and is primarily used to generate live footage and determine set lists.
[1485] A "low-latency streaming protocol" is a communication protocol for transmitting and receiving data with minimal delay, and is a technology that is particularly applicable to live video distribution, which requires real-time performance.
[1486] This invention relates to a system that generates live video footage of an artist selected by the user and distributes it in real time. The system mainly consists of three elements: a server, a terminal (user's viewing device), and the user. This system allows users to enjoy the artist's live performance in a virtual reality environment.
[1487] System Overview
[1488] In this system, the server is responsible for generating and distributing live video, and the terminal receives and displays the video. Users participate in the live broadcast through the terminal.
[1489] Server Processing
[1490] The server first collects the artist's live performance data. This data includes video, audio, stage layout, lighting, and other data. This data is obtained from the artist and production team via a dedicated API. Next, the server uses AI technology to generate live video based on the collected data. For example, by using a generative adversarial network (GAN) as a generative model, it synthesizes video in real time and automatically places lighting and effects. Furthermore, when the server receives a request to start a live performance, it randomly generates a setlist from a list of songs in its database and shuffles the performance order to create a unique setlist. During this process, seamless transitions appropriate for each song are also automatically generated. Finally, the server encodes and compresses the generated live video and prepares it for streaming. The video is then distributed to the user's device using a low-latency streaming protocol (e.g., WebRTC or HTTP Live Streaming: HLS).
[1491] Terminal handling
[1492] The device receives live video data streamed from the server in real time. The received video data is decoded and converted into a format suitable for the VR headset. The device also detects the user's viewpoint movements and interactions in real time and reflects them in the video. This includes using Unity to process interactive elements such as viewpoint movements and clapping sounds in real time.
[1493] User behavior
[1494] First, the user puts on the VR headset and logs into the system. After logging in, they select their favorite artist or live event from the app interface. Once the selection is complete, a request is sent to the server. Once viewing of the live performance begins, the user can freely move their viewpoint while wearing the VR headset, enjoying the atmosphere of the live venue. Various operations are also possible, such as changing the viewpoint, adjusting the volume, and using interactive elements (for example, recreating applause and cheers in VR).
[1495] Specific examples
[1496] When a user selects a live event by their favorite artist and presses the start viewing button, the server receives the request and generates a virtual live video based on the collected data. The set list and performance order are also randomly determined. The generated live video is streamed from the server to the device. The device receives the video data, decodes it, and displays it on the VR headset. The user puts on the VR headset and watches the live performance in real time while moving their viewpoint. In this way, the system of the present invention allows users to enjoy an artist's live performance in a virtual reality environment without the constraints of time or place.
[1497] Prompt Sentence Examples
[1498] "To generate a virtual live video of an artist selected by the user and deliver it with low latency, synthesize the collected live data (video, audio, stage layout) in real time and stream it using WebRTC. Reflect real-time viewpoint changes and interactive elements so that users can enjoy the virtual live performance."
[1499] This allows users to enjoy an experience that feels as if they are at an actual live concert venue in a VR environment.
[1500] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1501] Step 1:
[1502] The server receives the user request
[1503] Input: The user selects the artist's live performance they want to watch and submits a request.
[1504] Processing: The server receives the request data from the user, which may include the artist name, the date and time of the show, and the user's authentication information.
[1505] Output: Request data to be used in the next step.
[1506] Step 2:
[1507] Live Data Collection
[1508] Input: Request data to generate artist gigs.
[1509] Processing: The server collects live data (video, audio, stage layout, lighting data, etc.) from artists and production teams via a dedicated API.
[1510] Output: Live data collected.
[1511] Step 3:
[1512] Live video generation
[1513] Input: Live data collected.
[1514] Processing: The server uses a generative AI model (e.g., GAN) to generate live footage in real time, including compositing the footage and adding lighting and effects.
[1515] Output: The generated live video data.
[1516] Step 4:
[1517] Setlist and performance order randomly determined
[1518] Input: Live video data and artist song list.
[1519] Processing: The server retrieves the song list from the database, randomly generates a setlist, shuffles the order of the songs, and automatically creates seamless transitions between each song.
[1520] Output: Live video data with setlist and transitions.
[1521] Step 5:
[1522] Encoding and preparing live video for distribution
[1523] Input: Live video data with setlist and transitions.
[1524] Processing: The server encodes the video in H264 format and prepares the data for distribution using a low-latency streaming protocol such as WebRTC.
[1525] Output: Encoded live video data.
[1526] Step 6:
[1527] Live video streaming
[1528] Input: Encoded live video data.
[1529] Processing: The server streams the prepared video data to the user's device in real time.
[1530] Output: Live video data sent to user terminal.
[1531] Step 7:
[1532] Receiving and decoding live video data
[1533] Input: Live video data streamed from the server.
[1534] Processing: The device decodes the received video data and converts it into a format suitable for the VR headset.
[1535] Output: Live video data that can be displayed on a VR headset.
[1536] Step 8:
[1537] Reflects user viewpoint movements and operations
[1538] Input: User gaze movement data and interaction data.
[1539] Processing: The device uses Unity to reflect interactive elements such as the user's viewpoint and clapping sounds in the video in real time.
[1540] Output: Live video data reflecting the interaction.
[1541] Step 9:
[1542] User's live viewing experience
[1543] Input: Live video data reflecting viewpoint movement.
[1544] Processing: Users can watch the live stream through a VR headset, freely moving their viewpoint in real time, and can also adjust the volume and use interactive elements (such as applause and cheering).
[1545] Output: An immersive virtual live experience.
[1546] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1547] The present invention relates to a system that generates live video of an artist selected by the user and distributes the video in real time. The present invention can also be combined with an emotion engine that recognizes the user's emotions to further improve the user experience.
[1548] System configuration
[1549] The system mainly consists of a server, a terminal (user's viewing device), and an emotion engine. The server is responsible for generating and distributing live video, the terminal is responsible for receiving and displaying the video, and the emotion engine recognizes user emotions in real time.
[1550] Server Processing
[1551] The server first collects live data from the artist, including video, audio, stage layout, and lighting data, which is obtained from the artist and production team via a dedicated API.
[1552] The server then uses AI technology to generate live video based on the collected data. Generative Adversarial Networks (GANs) and other generative models are used to synthesize the video in real time. Lighting and effects are also automatically placed.
[1553] Furthermore, when the server receives a request to start a live performance, it randomly generates a setlist from the songs in the database and shuffles the order in which they are played, resulting in a unique setlist.
[1554] The server encodes and compresses the generated live video, prepares it for streaming, and delivers it to the user's device using a low-latency streaming protocol (e.g., WebRTC or HTTP Live Streaming: HLS).
[1555] Terminal handling
[1556] The device receives live video data streamed from the server in real time. The received video data is decoded and converted into a format suitable for the VR headset. The device also detects the user's viewpoint movements and interactions in real time and reflects them in the video display.
[1557] Emotion engine processing
[1558] The emotion engine analyzes the user's facial expressions and voice to recognize emotions in real time. This data is sent to a server and used to dynamically change the live video production, set list, and interactive elements (such as specific lighting and sound effects).
[1559] User behavior
[1560] First, users put on a VR headset and log in to the system. After logging in, they select their favorite artist or live event from the app interface. Once the selection is complete, a request is sent to the server.
[1561] Once the live viewing begins, users can freely move their viewpoint while wearing the VR headset, enjoying the atmosphere of the live venue. Various operations are also possible, such as changing the viewpoint, adjusting the volume, and using interactive elements (for example, recreating applause and cheers in VR).
[1562] The emotion engine detects the user's emotions and changes the live video presentation in real time. For example, if the user becomes excited, the lighting and effects will be enhanced to further enhance the user's experience.
[1563] Specific examples
[1564] For example, when a user selects a live event by their favorite artist and presses the start button, the server receives the request and generates a virtual live video based on the collected data. The set list and performance order are also randomly determined. Furthermore, an emotion engine analyzes the user's emotions in real time and sends that data to the server. The server then dynamically adjusts the live video production and set list based on this data.
[1565] In this way, the system of the present invention allows users to enjoy live performances by artists in a virtual reality environment without the constraints of time or place, and by using an emotion engine, the user's experience can be further enriched.
[1566] The processing flow will be explained below.
[1567] Server Processing
[1568] Step 1:
[1569] The server collects artists' existing live footage, audio sources, stage layout, and lighting data from a database, which is then retrieved from the artists and production teams via a dedicated API.
[1570] Step 2:
[1571] Based on the collected data, the server uses AI technology to generate virtual live footage. For example, it uses a generative model based on GAN (Generative Adversarial Networks) to synthesize the footage in real time. Lighting and effects are also automatically placed.
[1572] Step 3:
[1573] When the server receives a request to start a live performance, it randomly generates a setlist from the songs in its database and shuffles the order in which they are played, resulting in a unique setlist.
[1574] Step 4:
[1575] The server encodes and compresses the generated live video, preparing it for streaming, and delivers it to the user's device using a low-latency streaming protocol (e.g., WebRTC or HTTP Live Streaming: HLS).
[1576] Step 5:
[1577] The server receives data from the emotion engine and adjusts the live video production (lighting, effects, set list changes, etc.) in real time based on the user's emotions.
[1578] Terminal handling
[1579] Step 1:
[1580] The device receives live video data streamed from the server in real time, with minimal buffering required to maintain low latency.
[1581] Step 2:
[1582] It decodes the received video data and converts it into a format suitable for VR headsets, allowing users to enjoy the video without interruption.
[1583] Step 3:
[1584] The device detects the user's viewpoint movements and operations in real time and reflects this information in the video display, allowing the user to freely look around the virtual space.
[1585] Step 4:
[1586] The device captures the user's facial expressions, voice, and other data and sends it to the emotion engine, which analyzes the data and generates real-time emotion data.
[1587] Emotion engine processing
[1588] Step 1:
[1589] The emotion engine receives the user's facial expressions and voice data sent from the device and analyzes it in real time.
[1590] Step 2:
[1591] Based on the analysis results, the system recognizes the user's emotions and sends the data to the server. For example, if the user is excited, the emotional data is transmitted to the server.
[1592] User behavior
[1593] Step 1:
[1594] Users put on a VR headset and log in to a dedicated application, where their user information and viewing history are automatically loaded.
[1595] Step 2:
[1596] Users select their favorite artist or live event from the app's interface and press the Start Watch button, which sends a request to the server.
[1597] Step 3:
[1598] Once the live viewing begins, users can freely move their viewpoint while wearing the VR headset and enjoy the atmosphere of the live venue. They can change the viewpoint, adjust the volume, and use interactive elements (for example, recreating applause and cheers in VR).
[1599] Step 4:
[1600] The emotion engine analyzes the user's facial expressions and voice, and the results are sent to the server. If the user is excited, the lighting and effects will be enhanced, further enriching the experience.
[1601] In this way, the system of the present invention not only allows users to enjoy live performances by artists in a virtual reality environment without time or place constraints, but also uses an emotion engine to provide a dynamic live experience that reflects the user's real-time emotions.
[1602] Example 2
[1603] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1604] Conventional live video streaming systems allow users to watch live footage of artists of their choice, but the experience relies on limited visual and auditory information. As a result, real-time video production is not tailored to the user's emotions or interactions, resulting in a lack of depth of experience. Furthermore, because the set list and performance order are fixed, watching the same live footage repeatedly can lose its freshness, creating a problem.
[1605] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1606] In this invention, the server includes means for generating live video of an artist selected by the user, means for delivering the generated live video to the user's viewing device, means for randomly determining the set list and performance order of the live video, means for reflecting the user's viewpoint and operations in real time, and means for recognizing the user's emotions in real time and dynamically changing the direction of the live video, thereby enabling users to enjoy a fresh and realistic live experience that is individually customized.
[1607] "User-selected artists" are musical artists selected by users of the system according to their own preferences.
[1608] "Live footage" is video that visually and aurally captures a musical artist's performance or live performance.
[1609] "Means of generation" is a general term for methods and devices that use specific algorithms or AI technology to create live footage.
[1610] A "viewing device" is a device that a user uses to view video and audio, such as a VR headset or smartphone.
[1611] "Means of distribution" refers to the methods and technologies for sending the generated live video to the user's viewing device via a network such as the Internet.
[1612] A "set list" is a list of songs that will be performed live.
[1613] The "play order" is the rule that determines the order in which songs in a set list are played.
[1614] A "random determination means" is a method or device that randomly determines the set list or performance order using a specific algorithm or random number generation technology.
[1615] "Means for reflecting viewpoints and operations in real time" refers to methods and technologies that instantly detect the user's viewpoint movements and interactions and reflect them in the video display.
[1616] "Means for recognizing emotions in real time and dynamically changing the presentation of live video" refers to methods and technologies that analyze emotions from the user's facial expressions, voice, etc., and instantly change the presentation of live video based on the results.
[1617] "AI technology" refers to any technology that uses artificial intelligence, including machine learning and deep learning in particular.
[1618] "Low latency streaming protocol" is a protocol that minimizes data delays and is a technology that enables real-time communication.
[1619] This invention relates to a system that generates live video of an artist selected by the user and distributes the video in real time. Furthermore, this invention can further enhance the user experience by incorporating an emotion engine that recognizes the user's emotions. This system consists of three main components: a server, a terminal, and an emotion engine.
[1620] System configuration
[1621] The system includes a server, a terminal (user's viewing device), and an emotion engine. The server is responsible for generating and distributing live video, and the terminal receives and displays the video. The emotion engine recognizes user emotions in real time.
[1622] Server Processing
[1623] The server first collects live data from the artist. This data includes video, audio, stage layout, and lighting data. This data is obtained from the artist and production team via a dedicated API. Based on the collected data, the server uses AI technology (such as Generative Adversarial Networks: GAN) to generate live video. During this generation process, lighting and effects are also automatically placed.
[1624] Furthermore, when the server receives a request to start a live performance, it randomly generates a setlist from the song list in the database and shuffles the performance order to create a unique setlist.The generated live video is then encoded, compressed, and distributed to the user's device using a low-latency streaming protocol (e.g., WebRTC or HTTP Live Streaming: HLS).
[1625] Terminal handling
[1626] The device receives live video data streamed from the server in real time. The received video data is decoded and converted into a format suitable for the VR headset. The device also detects the user's viewpoint movements and interactions in real time and reflects them in the video display.
[1627] Emotion engine processing
[1628] The emotion engine analyzes the user's facial expressions and voice to recognize emotions in real time. This data is sent to a server and used to dynamically change the live video production, set list, and interactive elements (such as specific lighting and sound effects).
[1629] User behavior
[1630] First, users put on a VR headset and log in to the system. After logging in, they select their favorite artist or live event from the app interface. Once the selection is complete, a request is sent to the server. Once viewing of the live performance begins, users can freely move their viewpoint while wearing the VR headset, enjoying the atmosphere of the live venue. They can also change the viewpoint, adjust the volume, and use interactive elements (for example, recreating applause and cheers in VR). The emotion engine detects the user's emotions, which then changes the live video presentation in real time. For example, if the user becomes excited, lighting and effects are enhanced, further improving the user experience.
[1631] Specific examples
[1632] For example, when a user selects a live event by their favorite artist and presses the start button, the server receives the request and uses AI technology to generate a virtual live video based on the collected data. The set list and performance order are also randomly determined. After the live video is generated, the server encodes the video and streams it to the user with low latency. When the emotion engine detects the user's excitement, the server dynamically changes the live performance by enhancing lighting and effects.
[1633] Prompt Sentence Examples
[1634] Below are some example prompts to input to a generative AI model:
[1635] "Please explain a system that generates a virtual live video of a live event of an artist selected by the user based on related data, and dynamically changes the live performance by analyzing the user's emotions in real time using an emotion engine."
[1636] In this way, the system of the present invention allows users to enjoy live performances of artists in a virtual reality environment regardless of location. Furthermore, the use of an emotion engine can further enrich the user experience.
[1637] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1638] Step 1:
[1639] Live Data Collection
[1640] The server collects live data from artists, including video, audio, stage layout, and lighting data, and this data is obtained through a dedicated API.
[1641] Input: API requests from artists and production teams
[1642] Output: Live video, audio, stage layout, lighting data
[1643] Specific behavior:
[1644] The server sends a request to an API endpoint.
[1645] Receive and parse the JSON data returned as a response.
[1646] The acquired data is stored in an internal database.
[1647] Step 2:
[1648] Live video generation
[1649] The server uses AI techniques (such as Generative Adversarial Networks: GAN) to generate live footage based on the collected data, taking into account stage layout and lighting data during the generation process.
[1650] Input: Live video, audio, stage layout, lighting data
[1651] Output: Generated live video
[1652] Specific behavior:
[1653] The server inputs the collected data into the AI model.
[1654] The AI model generates images based on the specified parameters.
[1655] Lighting and effects are automatically added to the generated video data.
[1656] Step 3:
[1657] Handling Go Live Requests
[1658] The server receives a request from the user to start a live stream, which includes an artist and event identifier.
[1659] Input: User's request to go live
[1660] Output: Information about a specific artist or event
[1661] Specific behavior:
[1662] The server receives the request and loads the corresponding artist and event data.
[1663] Prepare to move on to the next step.
[1664] Step 4:
[1665] Deciding on the set list and performance order
[1666] The server randomly generates a setlist from the songs in the database and shuffles the order in which they are played.
[1667] Input: Song list
[1668] Output: Randomly determined setlist and playing order
[1669] Specific behavior:
[1670] The server retrieves the song list from the database.
[1671] Use a random algorithm to generate setlists and shuffle song order.
[1672] Step 5:
[1673] Encode video and prepare for streaming
[1674] The server encodes, compresses, and converts the generated live video into a streaming format, preparing it for distribution using a low-latency streaming protocol (e.g., WebRTC or HTTP Live Streaming: HLS).
[1675] Input: Generated live video
[1676] Output: Video encoded for streaming
[1677] Specific behavior:
[1678] The server encodes the video data in a specific codec.
[1679] The encoded data is sent to a streaming server.
[1680] Step 6:
[1681] Live video streaming
[1682] The server streams the encoded live video to the user's device in real time.
[1683] Input: Encoded live video
[1684] Output: Video streamed to user device
[1685] Specific behavior:
[1686] The server checks the streaming protocol settings and transmits the encoded video data in real time.
[1687] Monitor your streaming performance and make adjustments as needed.
[1688] Step 7:
[1689] Handling user gaze movements and interactions
[1690] The device receives live video data streamed from the server in real time, decodes it, and converts it into a format suitable for the VR headset. It also detects the user's viewpoint movements and interactions and reflects them in the video display.
[1691] Input: Streaming live video data, user's viewpoint movement and interaction information
[1692] Output: Image displayed on VR headset
[1693] Specific behavior:
[1694] The terminal decodes the received video data.
[1695] It detects the user's viewpoint movement and adjusts the displayed content in real time.
[1696] Processes user interaction information (e.g., button presses).
[1697] Step 8:
[1698] Emotion analysis using an emotion engine
[1699] The emotion engine analyzes the user's facial expressions and voice to recognize their emotions in real time, and this emotion data is sent to the server.
[1700] Input: User's facial expressions and voice data
[1701] Output: Parsed emotion data
[1702] Specific behavior:
[1703] The emotion engine captures the user's camera footage and microphone audio.
[1704] The acquired data is analyzed in real time to recognize the emotional state.
[1705] Step 9:
[1706] Dynamic changes to live performances
[1707] The server dynamically changes the live video presentation based on the emotional data received from the emotion engine. For example, if the user is excited, the lighting and effects will be enhanced.
[1708] Input: Emotion data received from the emotion engine
[1709] Output: Dynamically modified live video rendition
[1710] Specific behavior:
[1711] The server analyzes the emotion data and determines the corresponding change in presentation.
[1712] Implementing real-time performance changes to improve the user experience.
[1713] (Application example 2)
[1714] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1715] Conventional live video streaming systems make it difficult for users to experience the content interactively in real time, limiting the viewing experience. Furthermore, they lack the ability to dynamically change the presentation to reflect user emotions and reactions, making it difficult to maximize user satisfaction. Furthermore, there are technical challenges in delivering high-quality video with low latency.
[1716] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for generating live video of an artist selected by the user, means for delivering the generated live video to the user's viewing device, means for randomly determining the set list and performance order of the live video, means for reflecting the user's viewpoint and operations in real time, means for recognizing the user's emotions and dynamically changing the production and set list of the live video based on that data, and means for delivering the live video using a low-latency streaming protocol. This allows users to enjoy a more interactive and personalized live experience in real time.
[1717] "User" means any person or entity that uses the System to view live video.
[1718] "Artists" are musicians or performers who are the subject of live footage.
[1719] "Live footage" refers to footage that visually records an artist's performance.
[1720] "Viewing Device" means the hardware device through which a User views the live feed. Examples include smartphones and head-mounted displays.
[1721] A "set list" is a list of songs that will be played at a live concert.
[1722] "Performance order" refers to the order in which songs are performed on the set list.
[1723] "Real time" refers to the time in which an event that occurs is processed and reflected immediately without delay.
[1724] "Emotion recognition" is the process of analyzing a user's facial expressions and voice to identify their emotional state.
[1725] "Production" refers to the visual and auditory effects, such as lighting, effects, and sound, used in live footage.
[1726] "Dynamic change" refers to the ability of the system to make immediate changes in response to specific conditions or situations.
[1727] A "low latency streaming protocol" is a communication protocol that minimizes delays in sending and receiving data.
[1728] "AI technology" refers to all technologies based on artificial intelligence, and in this invention, it particularly refers to technologies that perform generative models and emotion recognition.
[1729] A "prompt sentence" is an input sentence that gives instructions to a generative AI model.
[1730] This invention relates to a system that generates live video of an artist selected by the user and distributes the video in real time. Furthermore, by combining this invention with an emotion engine that recognizes the user's emotions, the user experience can be further improved.
[1731] System configuration
[1732] The system mainly consists of a server, a terminal (user's viewing device), and an emotion engine. The server is responsible for generating and distributing live video, the terminal is responsible for receiving and displaying the video, and the emotion engine recognizes user emotions in real time.
[1733] Server Processing
[1734] The server first collects live data from the artist. This data includes video, audio, stage layout, and lighting data. This data is obtained from the artist and production team via a dedicated API. The server then uses AI technology to generate live video based on the collected data. Generative Adversarial Networks (GAN) and other generative models are used to synthesize video in real time. Lighting and effects are also automatically placed.
[1735] Furthermore, when the server receives a request to start a live performance, it randomly generates a setlist from the song list in the database and shuffles the performance order. This process creates a unique setlist. The server then encodes and compresses the generated live video and prepares it for streaming. The video is then delivered to the user's device using a low-latency streaming protocol (e.g., WebRTC or HTTP Live Streaming: HLS).
[1736] Terminal handling
[1737] The device receives live video data streamed from the server in real time. The received video data is decoded and converted into a format suitable for the VR headset. The device also detects the user's viewpoint movements and interactions in real time and reflects them in the video display.
[1738] Emotion engine processing
[1739] The emotion engine analyzes the user's facial expressions, voice, and other data to recognize emotions in real time. It uses specific APIs, such as OpenAI's API or Microsoft Azure's Face API. This data is sent to a server and used to dynamically change the live video presentation, set list, and interactive elements (such as specific lighting and sound effects).
[1740] User behavior
[1741] First, the user puts on a VR headset and logs into the system. After logging in, they select their favorite artist or live event from the app's interface. Once the selection is complete, a request is sent to the server. Once viewing of the live performance begins, the user can freely move their viewpoint while wearing the VR headset, enjoying the atmosphere of the live venue. Various operations are also possible, such as changing the viewpoint, adjusting the volume, and using interactive elements (for example, recreating applause and cheers in VR).
[1742] The emotion engine detects the user's emotions and changes the live video presentation in real time. For example, if the user becomes excited, the lighting and effects will be enhanced to further enhance the user's experience.
[1743] Specific examples
[1744] For example, when a user selects a live event by their favorite artist and presses the start button, the server receives the request and generates a virtual live video based on the collected data. The set list and performance order are also randomly determined. Furthermore, an emotion engine analyzes the user's emotions in real time and sends that data to the server. The server then dynamically adjusts the live video production and set list based on this data.
[1745] Examples of prompts include:
[1746] "Please implement an algorithm that detects a smile on the user's face, sends that information to the server, and enhances the lighting effects on the live video."
[1747] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1748] Step 1:
[1749] The server collects live data from artists. As input, it receives live video, audio, stage layout, and lighting data provided by artists and production teams via a dedicated API. This data is integrated within the server and prepared as the basis for generating live video. The output is an integrated live data set.
[1750] Step 2:
[1751] The server generates live video using a generative AI model (e.g., GAN). It uses the live data collected in step 1 as input. The AI model takes video, audio, stage layout, and lighting data as input and generates a synthesized live video in real time. The output is the generated live video.
[1752] Step 3:
[1753] When the server receives a request to start a live performance, it randomly generates a setlist from the song list in the database and shuffles the playing order. It uses the request to start a live performance and the song list in the database as input. It uses a random algorithm to generate the setlist and the playing order. The output is the randomly determined setlist and its playing order.
[1754] Step 4:
[1755] The server encodes, compresses, and prepares the generated live video for streaming. It uses the live video generated in step 2 and the setlist determined in step 3 as input. It prepares the video for streaming using a low-latency streaming protocol (WebRTC or HLS). The output is the encoded live video ready for streaming.
[1756] Step 5:
[1757] The device receives live video data streamed from the server in real time. As input, it receives the streaming data sent from the server. The device decodes and converts it into a format suitable for the VR headset. The output is the decoded and converted video data.
[1758] Step 6:
[1759] The device detects the user's viewpoint movements and operations in real time. Sensor information from the VR headset is used as input. Based on this, viewpoint movement and operation data is collected in real time and reflected in the video display. The output is a video display that reflects the viewpoint and operations.
[1760] Step 7:
[1761] The emotion engine analyzes the user's facial expressions and voice to recognize emotions in real time. The input is the user's facial expression data and voice data. Analysis is performed using OpenAI's API or Microsoft Azure's Face API to generate emotion data. The output is the analyzed emotion data.
[1762] Step 8:
[1763] The server dynamically changes the live video production and set list based on the emotional data sent from the emotion engine. The emotional data from the emotion engine is used as input. Based on the emotional data, lighting effects and production are adjusted, and the set list is also changed as necessary. The output is live video and production that adapts to the user's emotions.
[1764] Step 9:
[1765] A user puts on a VR headset, logs in to the system, and selects their favorite artist or live event from the app interface. The user's selection information is entered into the app as input. Once the selection is complete, a live viewing request is sent to the server. The output is a live viewing request.
[1766] Step 10:
[1767] While watching a live stream, the user changes the viewpoint, adjusts the volume, and uses interactive elements (such as recreating applause and cheers). The input is information about the user's viewpoint, volume adjustment, and use of interactive elements. Based on this, the viewing experience is dynamically adjusted. The output is a viewing experience that corresponds to the user's actions.
[1768] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1769] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1770] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1771] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1772] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1773] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1774] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1775] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1776] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1777] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1778] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1779] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1780] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1781] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1782] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1783] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1784] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1785] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1786] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1787] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1788] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1789] The following is further disclosed regarding the above embodiment.
[1790] (Claim 1)
[1791] means for generating live footage of an artist selected by the user;
[1792] A means for delivering the generated live video to a user's viewing device;
[1793] A means for randomly determining the set list and performance order of the live video;
[1794] A means of reflecting the user's viewpoint and operations in real time;
[1795] A system including:
[1796] (Claim 2)
[1797] The system of claim 1, which utilizes AI technology in generating live video.
[1798] (Claim 3)
[1799] 10. The system of claim 1, wherein the system delivers live video using a low-latency streaming protocol.
[1800] "Example 1"
[1801] (Claim 1)
[1802] means for generating live footage of a musician selected by the user;
[1803] means for delivering the generated live video to a user's display device;
[1804] A means for randomly determining the songs and performance order of the live video;
[1805] A means of reflecting the user's viewpoint and operations in real time;
[1806] A means to collect live data from musicians and production groups through a dedicated API,
[1807] A means of generating live video using AI technology based on collected data, and
[1808] A means to randomly generate a set list from the songs in the collected data and shuffle the order of performance;
[1809] a means for encoding and compressing the generated live video and preparing it for streaming;
[1810] means for delivering video using a low latency streaming protocol;
[1811] means for converting the decoded video into a format suitable for a visual device;
[1812] A system including:
[1813] (Claim 2)
[1814] The system of claim 1, which utilizes AI technology in generating live video.
[1815] (Claim 3)
[1816] 10. The system of claim 1, wherein the system delivers live video using a low-latency streaming protocol.
[1817] "Application Example 1"
[1818] (Claim 1)
[1819] means for generating live footage of an artist selected by the user;
[1820] A means for delivering the generated live video to a user's viewing device;
[1821] A means for randomly determining the set list and performance order of the live video;
[1822] A means of reflecting the user's viewpoint and operations in real time;
[1823] a means for recreating the interactive element in a virtual reality environment;
[1824] A system including:
[1825] (Claim 2)
[1826] The system of claim 1, which utilizes AI technology in generating live video.
[1827] (Claim 3)
[1828] 10. The system of claim 1, wherein the system delivers live video using a low-latency streaming protocol.
[1829] "Example 2: Combining Emotion Engines"
[1830] (Claim 1)
[1831] means for generating live footage of an artist selected by the user;
[1832] A means for delivering the generated live video to a user's viewing device;
[1833] A means for randomly determining the set list and performance order of the live video;
[1834] A means of reflecting the user's viewpoint and operations in real time;
[1835] A means to recognize user emotions in real time and dynamically change the direction of live video,
[1836] A system including:
[1837] (Claim 2)
[1838] The system of claim 1, which utilizes AI technology in generating live video.
[1839] (Claim 3)
[1840] 10. The system of claim 1, wherein the system delivers live video using a low-latency streaming protocol.
[1841] "Application example 2 when combining emotion engines"
[1842] (Claim 1)
[1843] means for generating live footage of an artist selected by the user;
[1844] A means for delivering the generated live video to a user's viewing device;
[1845] A means for randomly determining the set list and performance order of the live video;
[1846] A means of reflecting the user's viewpoint and operations in real time;
[1847] A means to recognize the user's emotions and dynamically change the performance and set list of the live video based on that data;
[1848] A means for delivering live video using a low latency streaming protocol;
[1849] A system including:
[1850] (Claim 2)
[1851] The system of claim 1, which utilizes AI technology in generating live video.
[1852] (Claim 3)
[1853] The system of claim 1, which utilizes a specific API to recognize user emotions. [Explanation of symbols]
[1854] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. means for generating live footage of an artist selected by the user; A means for delivering the generated live video to a user's viewing device; and A means for randomly determining the set list and performance order of the live video; A means of reflecting the user's viewpoint and operations in real time; A system including:
2. The system of claim 1, which utilizes AI technology in generating live video.
3. The system of claim 1 , wherein the live video is delivered using a low-latency streaming protocol.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A