system

The system addresses ticketing and accessibility issues by providing real-time event experiences through VR, enhancing user engagement and revenue generation.

JP2026015067APending Publication Date: 2026-01-29SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024116541
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-19
Publication Date
2026-01-29

AI Technical Summary

Technical Problem

Users are unable to purchase tickets for sporting events or live music events due to limited seating, face difficulties traveling to distant venues, and have challenges attending real events due to physical disabilities, while organizers lack effective revenue sources.

Method used

A system that captures real-time video from multiple cameras, processes it using AI to generate supplemental information, and streams it to a VR device, allowing users to select viewpoints and communicate with others, providing an immersive event experience from home.

Benefits of technology

Enables users to enjoy events in real-time from home with enhanced interaction and detailed information, increasing user satisfaction and revenue streams for organizers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026015067000001_ABST
    Figure 2026015067000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system comprising: means for obtaining real-time video from a plurality of cameras; means for transferring the obtained video to a streaming server; means for receiving video from a viewpoint selected by a user using a VR device; means for analyzing the video using AI to generate supplemental information; means for providing the supplemental information to the users; and means for facilitating communication between the users.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] The present invention aims to solve the problems faced by users who cannot purchase tickets for sporting events or live music events due to limited seating, users who have difficulty traveling to distant venues, and users who cannot attend real events due to physical disabilities.The invention also aims to secure and increase revenue sources for organizers. [Means for solving the problem]

[0005] The present invention provides a means for acquiring real-time video from multiple cameras and transmitting the video to a streaming server. The user can then select a viewpoint using a VR device and receive video from the selected viewpoint. The system also includes a means for analyzing the video data using AI to generate supplemental information, and a means for providing the generated supplemental information to the user. Furthermore, the system includes a means for supporting communication between users, allowing users to enjoy a real-life event experience even from home.

[0006] A "camera" is a device for capturing video or images.

[0007] "Real-time video" refers to video data that shows the situation of an event that is currently occurring on the spot with almost no time lag.

[0008] A "streaming server" is a central system for continuously distributing video data over the Internet.

[0009] "User" refers to any individual or organization that uses this system.

[0010] "VR device" refers to devices such as headsets and goggles for displaying virtual reality content.

[0011] "Viewpoints" refer to camera positions taken from different locations and angles of the event.

[0012] "Video data" refers to digital information of images and videos captured by a camera.

[0013] "AI" stands for artificial intelligence and refers to the technology that analyzes video data and generates supplementary information based on the results.

[0014] "Supplementary information" refers to the results of AI analysis and additional data that are intended to help users understand and enjoy the content.

[0015] "Means of communication" refers to functions within the system that enable users to exchange messages and make voice calls. [Brief explanation of the drawings]

[0016] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14]FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0017] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0018] First, the terms used in the following description will be explained.

[0019] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0020] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0021] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0022] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0023] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0024] [First embodiment]

[0025] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0026] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0027] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0028] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0029] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0030] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0031] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0032] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0033] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0034] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0035] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0036] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0037] The system of the present invention is designed to enable users to enjoy events such as sports and live music in real time from the comfort of their own homes using a VR device. Specific means for implementing this system are described below.

[0038] System Configuration

[0039] The system mainly consists of the following components:

[0040] Multiple Cameras

[0041] Streaming Server

[0042] User's device

[0043] VR device

[0044] AI analysis server

[0045] System Operation

[0046] Camera installation and video acquisition

[0047] The server sets up multiple cameras at the event venue, each configured to capture footage from a different perspective. For example, in the case of a sporting event, cameras may be placed around the goal, from the athlete's perspective, or in the spectator seats.

[0048] The cameras capture video in real time and send it to a streaming server, which allows the server to simultaneously obtain video data from multiple cameras.

[0049] Streaming video

[0050] The server processes the acquired video data in real time and transfers it to the user's device via a streaming server, where the user can view the streaming video through a VR device.

[0051] User Interface

[0052] The user's device installs an application, through which viewpoint selection and other operations are performed. The user can select the camera viewpoint they want to see from the viewpoint selection screen displayed within the application. For example, in the case of a music concert, they can select the viewpoint of a specific instrument performer.

[0053] AI-based analysis and information provision

[0054] The AI ​​analysis server analyzes the video data being streamed in real time. For example, in the case of a sporting event, the AI ​​analyzes the movements of players and the position of the ball and generates supplementary information, which can provide information such as highlights of plays and tactical analysis.

[0055] This supplementary information is displayed in real time on the user's device, allowing users to understand and enjoy the event more in-depth, including the progress of the match, set lists of the songs being performed, and lyric information.

[0056] Communication Features

[0057] Users can also communicate with other users who are watching the same event. The server provides chat message and voice chat functions, allowing users to interact with each other. This communication function allows users to enjoy a more interactive experience.

[0058] Specific examples

[0059] Example 1: Watching a soccer game

[0060] 1. The server sets up multiple cameras in the stadium, such as from the goalkeeper's perspective, the spectator's perspective, and the player's perspective.

[0061] 2. The user puts on the VR device at home and launches the application.

[0062] 3. The user selects the goalkeeper perspective within the app and receives streaming video from that perspective.

[0063] 4. The AI ​​analysis server analyzes player movements and generates important play and tactical information.

[0064] 5. The analysis results are displayed in real time on the user's device, allowing them to chat with other viewers.

[0065] Example 2: Watching a live music concert

[0066] 1. The server sets up cameras on multiple musicians on stage and in the audience.

[0067] 2. Through the application, users select the perspective of their favorite instrumentalist.

[0068] 3. The server streams the selected viewpoint to the user's device.

[0069] 4. The AI ​​analysis server provides set lists and lyric information for the songs being performed in real time.

[0070] 5. Users can enjoy the live stream while chatting with other fans.

[0071] In this way, the system of the present invention allows users to enjoy a real event experience from the comfort of their own home. Furthermore, AI analysis and community functions provide new added value that differs from traditional viewing and appreciation methods.

[0072] The processing flow will be explained below.

[0073] Step 1:

[0074] The server performs an initialization procedure for multiple cameras, specifically obtaining the ID of each camera and configuring it to establish a streaming connection, allowing real-time video capture from each camera.

[0075] Step 2:

[0076] The camera acquires video in real time from the location where it is installed. Specifically, the camera captures frames (image data) at regular intervals and stores the video data in a buffer.

[0077] Step 3:

[0078] The camera transmits the captured video data to a server. Specifically, the camera uploads frame data to the server in real time via a network using a streaming protocol.

[0079] Step 4:

[0080] The server processes the received video data to relay it to the streaming server. Specifically, it arranges each received frame in order and sends it to the streaming server. The streaming server manages the video data and prepares it to be provided in response to user requests.

[0081] Step 5:

[0082] The user puts on the VR device and starts the application. Specifically, the application displays the user interface and opens a viewpoint selection screen.

[0083] Step 6:

[0084] The user operates the application interface to select the desired viewpoint. Specifically, the user selects, for example, the "goalkeeper's viewpoint" from the list displayed on the viewpoint selection screen.

[0085] Step 7:

[0086] The device receives the user's selection and requests the video data of that viewpoint from the server. Specifically, the device sends the ID of the selected viewpoint to the server and requests transmission of the corresponding stream.

[0087] Step 8:

[0088] The server acquires the video stream from the requested viewpoint and delivers it to the terminal. Specifically, it selects the corresponding camera stream and starts the process of transferring it to the user's terminal.

[0089] Step 9:

[0090] The AI ​​analysis server analyzes the video data received in real time, specifically recognizing important events in the video (such as player movements and ball position) and generating analysis results.

[0091] Step 10:

[0092] The server sends the AI ​​analysis results to the user's device, specifically, sending data including generated supplementary information to the device in real time.

[0093] Step 11:

[0094] The device displays the analysis results to the user, overlaying the analyzed information on the screen so that the user can instantly access important information.

[0095] Step 12:

[0096] Users can communicate with other users while watching an event by using the chat and voice chat functions within the application to communicate via text messages and voice.

[0097] Step 13:

[0098] The server manages and relays chat and voice chat messages between users, specifically, it delivers received messages to the correct destination and establishes communication between users.

[0099] Example 1

[0100] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0101] Modern event experiences are limited by physical constraints, making it difficult for users in distant locations to participate in events in real time. Furthermore, there are few ways to quickly grasp important moments and detailed information about an event, and there are limited ways to share it with other users through communication. This leads to a decline in overall user satisfaction.

[0102] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0103] In this invention, the server includes means for acquiring real-time video from a plurality of image acquisition devices, means for transferring the acquired video to a data processing device, means for a user to select a viewpoint using a virtual reality device and receive video from that viewpoint, means for analyzing the video data using artificial intelligence and generating supplemental information, and means for providing the supplemental information to users. This allows users who are far away to participate in an event in real time, quickly grasp important scenes and detailed information, and share the event experience through communication with other users.

[0104] An "image capture device" is a device that is installed at an event venue and captures images in real time from different viewpoints.

[0105] "Data processing device" means a device that receives video data transmitted from the image acquisition device and processes and encodes it in real time.

[0106] "User" refers to a person who uses the system to view and interact with events in real time.

[0107] A "virtual reality device" is a device such as a headset or goggles that a user wears to experience virtual reality.

[0108] "Artificial intelligence" refers to advanced algorithms and software used to analyze video data and generate supplemental information.

[0109] "Supplementary information" is information generated from the results of analysis by artificial intelligence, and is intended to provide users with additional understanding and interest.

[0110] The "viewpoint selection interface" is a function of the application that allows the user to select and manipulate different camera viewpoints.

[0111] The "communication function" is a function that allows users to communicate with other users in real time via messages and voice.

[0112] A "streaming server" is a server that transfers processed video data to a user's terminal in real time.

[0113] A "user terminal" is a device such as a PC, smartphone, or tablet operated by the user, which functions in conjunction with the virtual reality device.

[0114] The system of this invention is designed to enable users to watch events such as sports and live music in real time from the comfort of their own homes using a virtual reality device (hereinafter referred to as a VR device). Specific means for implementing this system are described below.

[0115] System Configuration

[0116] The system mainly consists of the following components:

[0117] Multiple image capture devices (cameras)

[0118] Data processing device (streaming server and AI analysis server)

[0119] User's device

[0120] VR device

[0121] System Operation

[0122] Camera installation and video acquisition

[0123] The server installs multiple image capture devices at the event venue. These image capture devices are positioned so that they capture video from different viewpoints. For example, in the case of a sporting event, image capture devices are placed around the goal area, from the players' viewpoints, and in the spectator seats. The image capture devices capture video in real time and send the video to a data processing device. This allows the server to simultaneously acquire video data from multiple image capture devices.

[0124] Streaming video

[0125] The server processes the acquired video data in real time and transfers it to the user's device via a streaming server. The user can then view this streaming video through a VR device.

[0126] User Interface

[0127] A dedicated application is installed on the user's device, and this application allows for viewpoint selection and other operations. The user can select the image capture device viewpoint they want to see from the viewpoint selection interface. For example, in the case of a live music performance, they can select the viewpoint of a specific instrument performer.

[0128] AI-based analysis and information provision

[0129] The AI ​​analysis server analyzes the video data being streamed in real time. For example, in the case of a sporting event, it analyzes the movements of the players and the position of the ball and generates supplementary information. This supplementary information is displayed in real time on the user's device, allowing the user to visually obtain information such as the progress of the game, highlights, and tactical analysis.

[0130] Communication Features

[0131] Users can communicate with other users who are watching the same event. The server provides chat message and voice chat functions, supporting users to interact with each other through these. This communication function allows users to enjoy a more interactive experience.

[0132] Specific examples

[0133] Example 1: Watching a soccer game

[0134] 1. The server installs multiple image capture devices at the stadium's goalkeeper's viewpoint, the spectator's viewpoint, and the player's viewpoint.

[0135] 2. The user puts on the VR device at home and launches the application.

[0136] 3. The user selects the goalkeeper's perspective within the app and receives streaming video from that perspective.

[0137] 4. The AI ​​analysis server analyzes player movements and generates important play and tactical information.

[0138] 5. The analysis results are displayed in real time on the user's device, allowing them to chat with other viewers.

[0139] Example 2: Watching a live music concert

[0140] 1. The server installs image capture devices on multiple instrument performers on stage and in the audience seats.

[0141] 2. Through the application, users can select the perspective of their favorite instrumentalist.

[0142] 3. The server streams the video from the selected viewpoint to the user's device.

[0143] 4. The AI ​​analysis server provides set lists and lyric information for the songs being performed in real time.

[0144] 5. Users can enjoy the live performance while chatting with other fans.

[0145] Example prompts for generative AI models

[0146] While watching a soccer match from the goalkeeper's point of view, generate sentences that provide analysis information on player movements and tactics.

[0147] To enjoy live music, generate text that provides perspective footage of specific instrumentalists and set list information.

[0148] In this way, the system of the present invention allows users to enjoy a real event experience from the comfort of their own home. Furthermore, AI analysis and community functions provide new added value that differs from traditional viewing and appreciation methods.

[0149] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0150] System program processing flow

[0151] Step 1: Installing the image capture device and acquiring images

[0152] The server installs multiple image capture devices (cameras) at the event venue. These cameras are positioned to capture images in real time from different viewpoints. Specifically, in a sporting event, cameras are placed around the goal, from the players' viewpoint, and in the spectator seats. The input from the cameras is real-time images, and the output is captured image data.

[0153] Step 2: Transferring video data

[0154] The camera transmits the captured real-time video to a data processing device (streaming server). In this step, the video data input from the camera is transferred to the streaming server via the network. The output is multiple video data stored in the streaming server.

[0155] Step 3: Processing the video data

[0156] The server processes the received video data in real time. Specifically, it compresses and encodes the video data and converts it into a format suitable for streaming. The input is the video data sent from the camera, and the output is the encoded streaming video data.

[0157] Step 4: Streaming the video data

[0158] The server transfers the processed video data to the user's device via a streaming server. The input is encoded video data, and the output is a video stream that can be played back in real time on the user's device.

[0159] Step 5: Launch the application and select a viewpoint

[0160] The user launches a dedicated application on the device. After the user completes login authentication, a viewpoint selection screen is displayed. Here, the user can select the viewpoint they want to view. The input is the viewpoint selected by the user, and the output is video data based on the selected viewpoint.

[0161] Step 6: Receiving and displaying viewpoint images

[0162] The user's device receives the video from the selected viewpoint and displays it on the VR device. The input is video data sent from the streaming server, and the output is real-time video displayed on the user's VR device.

[0163] Step 7: Analyze the video data

[0164] The AI ​​analytics server analyzes the video data received in real time. Specifically, it identifies player movements and ball position and generates important highlights and tactical information. The input is streaming video data, and the output is supplemental information based on the analysis.

[0165] Step 8: Provide supporting information

[0166] The AI ​​analysis server sends the generated supplementary information to the user's device, which displays it in real time. The input is the analyzed supplementary information, and the output is the information displayed on the user's device.

[0167] Step 9: Providing communication features

[0168] Users communicate with other users through chat and voice chat functions within the application. This function is provided by a server. The input is the user's message or voice data, and the output is the communication data transmitted to other users.

[0169] Through this step-by-step process, users can watch events in real time from the comfort of their own homes and enjoy a variety of interactive experiences.

[0170] (Application example 1)

[0171] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0172] Conventional live event viewing systems make it difficult for users to enjoy the event from multiple angles in real time or to interact with other viewers. They also often fail to obtain important scenes or specific information in real time while watching. This limits the user experience and prevents the full appeal of the event.

[0173] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0174] In this invention, the server includes means for acquiring real-time video from multiple cameras, means for transferring the acquired video to a streaming server, means for a user to select a viewpoint using a VR device and receive video from that viewpoint, means for analyzing video data using AI and generating supplemental information, means for providing the supplemental information to the user, means for supporting communication between users, means for interacting in a virtual space using smart glasses or a head-mounted display, means for providing real-time chat and voice calls with other users, means for automatically detecting highlights of an event from video in real time using a generative AI model, and means for providing the generated highlights to the user based on prompt text. This allows users to enjoy an event from multiple viewpoints in an immersive way, and enables interactive interaction with other viewers and the acquisition of important scenes and supplemental information through real-time analysis using AI.

[0175] "Multiple cameras" refers to multiple imaging devices installed at an event venue to capture footage in real time from different perspectives.

[0176] A "streaming server" refers to a server system that processes acquired video data in real time and distributes it to the user's device.

[0177] "User terminal" refers to a computer device that allows a user to receive event footage, select viewpoints, display supplementary information, and communicate.

[0178] "VR device" refers to devices such as head-mounted displays and smart glasses that users wear to experience virtual reality.

[0179] "AI analysis server" refers to a server system that runs artificial intelligence algorithms to analyze acquired video data and generate supplemental information.

[0180] "Supplementary information" refers to additional information such as highlights of the performance, tactical information, and set lists of songs performed, which is generated by AI analysis based on streaming video.

[0181] "Communication function" refers to the function that allows users to interact with other viewers through chat messages and voice calls while watching an event.

[0182] "Smart glasses" refers to a glasses-type device that users can wear to view images in a virtual space.

[0183] A "head-mounted display" refers to a display device worn on the head that allows users to intuitively experience a virtual space.

[0184] "Virtual space" refers to a digitally constructed simulated environment that users experience visually and aurally through VR devices.

[0185] "Real-time chat" refers to a means of communication that allows users to exchange text messages in real time.

[0186] "Voice Call" means a means by which Users can engage in real-time voice communication.

[0187] "Generative AI model" refers to a machine learning algorithm that automatically generates highlights and supplemental information based on input data.

[0188] "Prompt sentence" refers to text input provided to a generative AI model to generate a specific output.

[0189] This invention is a system that allows users to enjoy events such as sports and live music in real time from home using a VR device. The system consists of multiple cameras, a streaming server, a user's device, a VR device, and an AI analysis server.

[0190] System Configuration

[0191] Hardware Configuration

[0192] Multiple cameras: Installed at the event venue, capturing footage in real time from different perspectives.

[0193] Streaming server: A server system that processes acquired video data in real time and distributes it to user devices. Protocols used include WebRTC and RTMP.

[0194] User's device: A computer device used to receive the video, select viewpoints, display supplementary information, and communicate. Smartphones and PCs are often used.

[0195] VR equipment: A device such as a head-mounted display or smart glasses worn by a user to create a virtual reality experience.

[0196] AI analysis server: A server system that runs artificial intelligence algorithms to analyze video data in real time and generate supplementary information. It uses TensorFlow and OpenCV.

[0197] Software Configuration

[0198] Viewpoint switching function: A function that receives a user's viewpoint switching request and sends the corresponding camera video stream to the user's device.

[0199] Real-time AI analysis: The AI ​​analysis server receives the acquired video data and generates important scenes and supplementary information in real time. Analysis is performed using TensorFlow and OpenCV.

[0200] Communication feature: Supports real-time chat and voice calls between users. Firebase and Agora.io SDK are used.

[0201] Generative AI model: Applies machine learning algorithms to generate key scenes and supplementary information from video footage. Generates appropriate output based on prompts.

[0202] Operation explanation

[0203] Viewpoint switching

[0204] When a user selects a viewpoint, the server uses a WebRTC client to acquire the video stream from the selected camera and transmit it to the user's device. For example, in a live music concert, it is possible to select the viewpoint of a specific instrumentalist.

[0205] Real-time AI analysis

[0206] The video data sent to the streaming server is analyzed by an AI analysis server, which uses TensorFlow models to detect important plays during sporting events or lyrics of live music concerts in real time and notify users.

[0207] Communication Features

[0208] The communication feature uses the Firebase real-time database to manage text messages and the Agora.io SDK to provide voice calls, allowing users to interact with other viewers in real time within the VR space.

[0209] Specific examples

[0210] Example 1: Watching a sports game

[0211] 1. The server sets up cameras at multiple viewpoints in the stadium.

[0212] 2. The user puts on the VR device at home and launches the application.

[0213] 3. The user selects the viewpoint of their choice and receives streaming video from that viewpoint.

[0214] 4. The AI ​​analysis server analyzes players' movements and generates important play and tactical information in real time, which is then provided to users.

[0215] Example 2: Watching a live music concert

[0216] 1. The server sets up cameras on multiple musicians on stage and in the audience.

[0217] 2. The user selects the perspective of their favorite instrument performer and receives streaming footage from that perspective.

[0218] 3. The AI ​​analysis server provides set lists and lyric information for the songs being performed in real time.

[0219] Prompt Sentence Examples

[0220] Here is an example prompt for automatic event highlight detection using an AI model:

[0221] Input: "Auto-detect key plays from this footage and create highlights."

[0222] Output: Automatically highlighted video clip

[0223] This invention allows users to enjoy a more immersive multi-perspective experience, interact with other users, and experience the event in greater depth thanks to real-time analysis by AI.

[0224] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0225] Step 1:

[0226] Multiple cameras are installed at the event venue, each capturing video in real time from a different viewpoint. The server acquires the video data from the cameras and transfers it to the streaming server.

[0227] Input: Live video from the event venue

[0228] Data processing: Video capture and transfer to streaming server

[0229] Output: Raw video data stored on a streaming server

[0230] Step 2:

[0231] The streaming server receives the acquired video data and delivers the video data to the user's terminal in real time.

[0232] Input: Raw video data stored on a streaming server

[0233] Data processing: video data compression and transfer

[0234] Output: Real-time video sent to user device

[0235] Step 3:

[0236] The user selects a viewpoint through the terminal, which then sends a viewpoint switching request to the server, which then delivers the video stream of the selected viewpoint to the user terminal.

[0237] Input: User's request to switch viewpoint

[0238] Data processing: Selecting the video stream corresponding to the viewpoint

[0239] Output: Video of the selected viewpoint sent to the user's device

[0240] Step 4:

[0241] The video data sent to the streaming server is analyzed in real time by an AI analysis server. A generative AI model is used for the analysis to extract important scenes and supplementary information. For example, important plays during a sporting event or lyrics from a live music concert can be automatically extracted.

[0242] Input: Raw video data sent from the streaming server

[0243] Data Computing: Real-time analytics with TensorFlow and OpenCV

[0244] Output: Analyzed important scenes and supplementary information

[0245] Step 5:

[0246] The user's device receives supplemental information sent from the AI ​​analysis server and displays it in real time, such as highlights of important performances and set lists of songs performed.

[0247] Input: Supplementary information sent from the AI ​​analysis server

[0248] Data processing: Converting supplementary information into a display format

[0249] Output: Supplementary information displayed on the user's terminal

[0250] Step 6:

[0251] Users can chat and talk to other users watching the same event in real time through their devices, and the server uses the Firebase real-time database and Agora.io SDK to support this functionality.

[0252] Input: User text messages and voice data

[0253] Data processing: real-time transmission of text messages, compression and transmission of voice data

[0254] Output: Messages and audio delivered to other users

[0255] Step 7:

[0256] When users use the generative AI model as part of real-time AI analysis, they can input prompts to automatically detect event highlights and specific scenes from video.

[0257] Input: Prompt text (e.g. "Automatically detect important plays from this footage and create highlights")

[0258] Data calculation: Execute a generative AI model based on the input prompt.

[0259] Output: Automatically generated highlight reels and clips of specific scenes

[0260] This allows users to enjoy an immersive multi-perspective experience, real-time interactive interactions, and detailed supplemental information about the event through AI analysis.

[0261] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0262] This invention provides a new user experience by combining a system that allows users to experience events such as sports and live music concerts in real time using a VR device from the comfort of their own home with an emotion engine that recognizes the user's emotions. Specific means for implementing this system are described below.

[0263] System Configuration

[0264] The system mainly consists of the following components:

[0265] Multiple Cameras

[0266] Streaming Server

[0267] User's device

[0268] VR device

[0269] AI analysis server

[0270] Emotion Engine

[0271] System Operation

[0272] Camera installation and video acquisition

[0273] The server installs multiple cameras at the event venue and captures video in real time from each camera. This video data is then transferred to the streaming server.

[0274] Streaming video

[0275] The server relays the acquired video to the streaming server in real time, allowing the user's device to receive the video data directly from the streaming server.

[0276] User Interface

[0277] The user's device performs viewpoint selection and other operations via the application. The user can select each camera viewpoint from the viewpoint selection screen displayed within the application.

[0278] AI-based analysis and information provision

[0279] The AI ​​analysis server analyzes the video data being streamed in real time, and based on the results of this analysis, users can receive supplemental information such as highlights of plays and tactical analysis.

[0280] Emotion recognition by emotion engine

[0281] The emotion engine recognizes the user's emotions in real time by analyzing their facial expressions and voice data. Based on the analysis results, the emotion engine generates supplementary information and effects according to the user's emotions. For example, if the user shows a surprised expression, it can respond by displaying a replay of the play at that moment.

[0282] Emotional Data Feedback

[0283] The emotional data recognized by the emotion engine is fed back to the AI ​​analysis server. Based on this data, the system adjusts the user experience across the entire system. For example, if data is obtained showing multiple users expressing excitement, the system can respond by providing more detailed explanatory information.

[0284] Communication Features

[0285] Users can chat and voice chat with other users while watching the event, and the system manages and relays these communications, facilitating interaction between users.

[0286] Specific examples

[0287] Example 1: Watching a soccer game

[0288] 1. The server sets up cameras at multiple viewpoints in the stadium, such as the goalkeeper's viewpoint, the spectator's viewpoint, and the player's viewpoint.

[0289] 2. The user puts on the VR device at home and launches the application.

[0290] 3. The user selects "Goalkeeper View" on the viewpoint selection screen and receives video from that viewpoint.

[0291] 4. The AI ​​analysis server analyzes player movements and ball trajectory to generate important play and tactical information.

[0292] 5. The emotion engine analyzes the user's facial expressions and voice, and when it recognizes moments of surprise or joy, it instantly displays replay footage and detailed commentary.

[0293] 6. Users can enjoy communicating with other viewers using the chat function.

[0294] Example 2: Watching a live music concert

[0295] 1. The server sets up cameras on the stage for multiple musicians and in the audience.

[0296] 2. Through the application, the user selects the perspective of, for example, a drummer.

[0297] 3. The server streams the selected viewpoint to the user's device.

[0298] 4. The AI ​​analysis server provides performance set lists and lyric information in real time.

[0299] 5. The emotion engine recognizes the user's emotions and responds by adding special effects when the user is emotional.

[0300] 6. Users can enjoy the live show while chatting with other fans.

[0301] This system allows users to enjoy live sports and music events from the comfort of their own home through an advanced VR experience, and the introduction of an emotion engine provides a more personalized experience.

[0302] The processing flow will be explained below.

[0303] Step 1:

[0304] The server will set up multiple cameras at the event venue, and will configure the initial settings of these cameras so that they can capture images from each viewpoint in real time.

[0305] Step 2:

[0306] The camera captures video in real time from its installed location, acquiring frames at regular intervals and generating video data.

[0307] Step 3:

[0308] The camera transmits the captured video data to a server. Specifically, the video data is transferred to a streaming server via a network in real time.

[0309] Step 4:

[0310] The server receives video data from multiple cameras and manages it in a streaming server, which prepares to provide the corresponding video stream in response to a viewpoint request from a user.

[0311] Step 5:

[0312] The user puts on the VR device and launches the application, which displays the main screen and offers viewpoint selection options.

[0313] Step 6:

[0314] The user selects the desired viewpoint on the viewpoint selection screen within the application. For example, the user selects the "goalkeeper viewpoint."

[0315] Step 7:

[0316] Based on the user's viewpoint selection, the terminal requests the video data of the viewpoint from the streaming server. Specifically, the terminal transmits the selected viewpoint ID to the streaming server and requests the corresponding stream.

[0317] Step 8:

[0318] The streaming server acquires the video stream of the requested viewpoint and transmits it to the terminal in real time, allowing the user to view the video from the selected viewpoint.

[0319] Step 9:

[0320] The AI ​​analysis server analyzes the streaming video data in real time to detect player movements and important events (e.g., goals, shots).

[0321] Step 10:

[0322] The AI ​​analysis server generates supplemental information based on the analysis results and sends it to the device, allowing users to receive the supplemental information in real time.

[0323] Step 11:

[0324] The device displays the AI ​​analysis results to the user, overlaying supplemental information on the screen in real time.

[0325] Step 12:

[0326] The emotion engine analyzes the user's facial expressions and voice data to recognize their emotions. Specifically, it captures the user's reactions through a camera and microphone and generates emotion data.

[0327] Step 13:

[0328] The emotion engine feeds the recognized emotion data back to the AI ​​analysis server, which then adjusts the overall system experience based on this data, for example, providing special effects or additional information to excited users.

[0329] Step 14:

[0330] Users can communicate with other users while watching an event, using chat and voice chat functions within the application to interact with other viewers.

[0331] Step 15:

[0332] The server manages and relays chat and voice chat messages between users, thereby supporting real-time communication between users.

[0333] This detailed processing flow allows users to enjoy events in real time from home and enjoy a personalized experience through the emotion engine.

[0334] Example 2

[0335] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0336] In conventional virtual reality event viewing systems, users are limited to selecting a viewpoint and viewing experience, and real-time emotion analysis and personalized experiences based on that analysis are not performed. Furthermore, communication functions between users are insufficient, making it insufficient to provide a richer user experience.

[0337] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0338] In this invention, the server includes means for acquiring real-time video from multiple image capture devices, means for transferring the acquired video to a data distribution server, means for allowing a user to select a viewpoint using a virtual reality device and receiving video from that viewpoint, means for analyzing video data using artificial intelligence and generating supplemental information, means for providing the supplemental information to the user, means for supporting communication between users, means for analyzing user emotions and generating supplemental information and effects based on the emotions, and means for feeding back the emotion data to the analysis server and adjusting the overall user experience. This enables the provision of personalized information and the addition of effects according to the user's emotions, further enhancing communication between users.

[0339] An "imaging device" is a device for capturing images in real time.

[0340] The "data distribution server" is a server that distributes acquired video data to user terminals in real time.

[0341] A "virtual reality device" is a device that allows users to immerse themselves in a virtual environment and select and control their viewpoint.

[0342] "Video data" refers to real-time video captured by an imaging device.

[0343] "Artificial intelligence" is a technology that analyzes video data and user emotional data to generate supplementary information and effects.

[0344] "Supplemental Information" is additional information provided to enhance the viewing or viewing experience.

[0345] "Effects" are visual and sound effects added to the video.

[0346] "User emotion" refers to the emotional state analyzed from the user's facial expressions and voice.

[0347] "Feedback" is the process of adjusting the overall system experience based on acquired data.

[0348] "Communication means" refers to a function that supports real-time communication between users.

[0349] This invention provides a new user experience by combining a system that allows users to experience events such as sports and live music concerts in real time using a VR device from the comfort of their own home with an emotion engine that recognizes the user's emotions. Specific means for implementing this system are described below.

[0350] System Configuration

[0351] The system mainly consists of the following components:

[0352] Multiple imaging devices (cameras)

[0353] Data distribution server

[0354] User's device

[0355] Virtual reality device (VR device)

[0356] AI analysis server

[0357] Emotion Engine

[0358] System Operation

[0359] Camera installation and video acquisition

[0360] The server installs multiple camera devices at the event venue, captures video in real time from each camera device, and transfers the video data to a data distribution server.

[0361] Streaming video

[0362] The server relays the acquired video to the data distribution server in real time, allowing the user's device to receive the video data directly from the data distribution server.

[0363] User Interface

[0364] The user's device performs viewpoint selection and other operations via the application. The user can select each camera viewpoint from the viewpoint selection screen displayed within the application.

[0365] AI-based analysis and information provision

[0366] The AI ​​analysis server analyzes the video data being streamed in real time, and based on the results of this analysis, users can receive supplemental information such as highlights of plays and tactical analysis.

[0367] Emotion recognition by emotion engine

[0368] The emotion engine recognizes the user's emotions in real time by analyzing their facial expressions and voice data. Based on the analysis results, the emotion engine generates supplementary information and effects according to the user's emotions. For example, if the user shows a surprised expression, it can respond by displaying a replay of the play at that moment.

[0369] Emotional Data Feedback

[0370] The emotional data recognized by the emotion engine is fed back to the AI ​​analysis server. Based on this data, the system adjusts the user experience across the entire system. For example, if data is obtained showing multiple users expressing excitement, the system can respond by providing more detailed explanatory information.

[0371] Communication Features

[0372] Users can chat and voice chat with other users while watching the event, and the system manages and relays these communications, facilitating interaction between users.

[0373] Specific examples

[0374] Example 1: Watching a soccer game

[0375] 1. The server sets up cameras at multiple viewpoints in the stadium, such as the goalkeeper's viewpoint, the spectator's viewpoint, and the player's viewpoint.

[0376] 2. The user puts on the VR device at home and launches the application.

[0377] 3. The user selects "Goalkeeper View" on the viewpoint selection screen and receives video from that viewpoint.

[0378] 4. The AI ​​analysis server analyzes player movements and ball trajectory to generate important play and tactical information.

[0379] 5. The emotion engine analyzes the user's facial expressions and voice, and when it recognizes moments of surprise or joy, it instantly displays replay footage and detailed commentary.

[0380] 6. Users can enjoy communicating with other viewers using the chat function.

[0381] Example 2: Watching a live music concert

[0382] 1. The server sets up cameras on the stage for multiple musicians and in the audience.

[0383] 2. Through the application, the user selects the perspective of, for example, a drummer.

[0384] 3. The server streams the selected viewpoint to the user's device.

[0385] 4. The AI ​​analysis server provides performance set lists and lyric information in real time.

[0386] 5. The emotion engine recognizes the user's emotions and responds by adding special effects when the user is emotional.

[0387] 6. Users can enjoy the live show while chatting with other fans.

[0388] Prompt Sentence Examples

[0389] Here are some examples of specific prompts for this system:

[0390] 1. Emotion Recognition in Sporting Events: Describe a process for displaying a replay and detailed commentary the moment the user expresses surprise.

[0391] 2. Personalizing Live Music Experience: What is your process for adding effects in real time based on the user's emotions?

[0392] This system allows users to enjoy live sports and music events from the comfort of their own home through an advanced VR experience, and the introduction of an emotion engine provides a more personalized experience.

[0393] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0394] Step 1: Camera installation and video acquisition

[0395] The server will set up multiple camera devices at the event venue. For example, at a sporting event, cameras will be installed to view the goalkeeper, the spectators, and the players, while at a music concert, cameras will be installed to view each musician on stage and in the spectators' seats. The server will acquire images from these cameras in real time.

[0396] Input: Video data from multiple cameras installed at the event venue.

[0397] Data processing: Converts the analog video signal from the camera into a digital signal and formats it for transmission to the streaming server.

[0398] Output: The formatted digital video data is transferred to a streaming server.

[0399] Step 2: Streaming the video

[0400] The server relays the captured video in real time to the data distribution server, which then distributes the video to the user's device.

[0401] Input: Digital video data transferred from the server.

[0402] Data processing: Packetization and streaming optimization for relaying and real-time delivery of digital video data.

[0403] Output: Streaming video data is sent to the user's device.

[0404] Step 3: User Interface

[0405] The user wears the VR device at home and starts the application. The user selects the desired camera viewpoint from the viewpoint selection screen within the application.

[0406] Input: Viewpoint selection information made by the user to the application.

[0407] Data processing: Based on the user's viewpoint selection, the video data of the corresponding camera viewpoint is extracted.

[0408] Output: Real-time video from the selected viewpoint is displayed in the user's VR device.

[0409] Step 4: AI-based video analysis and information provision

[0410] The AI ​​analysis server analyzes the video data being streamed in real time, for example, analyzing the movements of players and the trajectory of the ball, and then generates highlights of the play and tactical information based on that analysis.

[0411] Input: Real-time video data.

[0412] Data Processing: Analyze video data using generative AI models to extract and generate key events and tactical information.

[0413] Output: Analyzed highlights and tactical information are provided to the user.

[0414] Step 5: Emotion Recognition with the Emotion Engine

[0415] The emotion engine analyzes the user's facial expressions and voice data in real time to recognize their emotions. For example, if the user shows a surprised expression, it will display a replay of the play at that moment.

[0416] Input: User's facial and voice data.

[0417] Data processing: Using an AI model, the system analyzes the user's emotions and generates corresponding effects and supplementary information.

[0418] Output: Real-time effects and replay footage are displayed according to the user's emotions.

[0419] Step 6: Feedback of emotional data

[0420] The emotional data recognized by the emotion engine is fed back to the AI ​​analysis server. The system adjusts the user experience based on this data. For example, if multiple users show signs of excitement, more detailed commentary will be provided.

[0421] Input: Recognized user emotion data.

[0422] Data processing: Emotional data is sent to the analysis server and reflected in the user experience of the entire system.

[0423] Output: A tailored user experience based on sentiment data.

[0424] Step 7: Communication Functions

[0425] Users can chat and voice chat with other users while watching the event, and the system manages and relays these communications, facilitating interaction between users.

[0426] Input: Chat messages and voice data between users.

[0427] Data processing: Relaying and logging chat messages and voice data.

[0428] Output: A real-time communication experience.

[0429] (Application example 2)

[0430] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0431] In conventional online shopping systems, users are limited to visual information about products, which is far from the shopping experience of a physical store. Furthermore, when users express interest in a product or have a particular emotion, they are unable to provide dynamic information that responds to that interest. Furthermore, real-time communication between users is not adequately supported.

[0432] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for acquiring real-time video from multiple image capture devices, means for transferring the acquired video to a streaming server, means for a user to select a viewpoint using a virtual reality device and receive video from that viewpoint, means for analyzing video data using artificial intelligence and generating supplemental information, means for providing the supplemental information to the user, means for supporting communication between users, means for recognizing user emotions using an emotion engine and generating additional information and effects based on the recognition results, and means for feeding back emotion data to the artificial intelligence analysis server to adjust the user experience. This allows users to enjoy a realistic shopping experience from the comfort of their own home, as if they were in a physical store, and makes it possible to provide dynamic information about interesting products and add emotion-based effects.

[0433] A "camera" is a device that captures images of the real world and records them as digital data.

[0434] A "streaming server" is a server that distributes acquired video data in real time.

[0435] A "virtual reality device" is a device that allows a user to enjoy an experience in a 3D space, and generally includes a head-mounted display and smart glasses.

[0436] "Artificial intelligence" is a technology that analyzes large amounts of data and finds patterns and trends to make decisions and make predictions.

[0437] "Supplementary information" refers to additional information, commentary, promotions, etc. related to the video the user is watching.

[0438] "Communication support means" is a function that supports chat and voice chat between users, enabling smooth information exchange.

[0439] The "emotion engine" is a system that analyzes data such as the user's facial expressions and voice, and recognizes emotions in real time.

[0440] "Emotion data" is data that indicates the user's emotional state as analyzed by the emotion engine.

[0441] The "artificial intelligence analysis server" is a server that analyzes emotional data and video data and makes decisions such as providing information and generating effects.

[0442] The present invention is a system that provides a shopping experience from home that makes you feel as if you are actually in a store, and to achieve this, it combines multiple technologies such as a camera device, a streaming server, a virtual reality device, artificial intelligence, and an emotion engine. Specific embodiments of the system are described below.

[0443] System Configuration

[0444] Camera: Installed in multiple areas of the store, it captures high-resolution video in real time. Specific hardware used includes high-resolution cameras (e.g., Sony α7 series).

[0445] Streaming server: Used to distribute acquired video data to user devices in real time. The specific software used is AWS Elemental MediaLive.

[0446] Virtual reality devices: Devices that allow users to immerse themselves in virtual reality spaces. Specific devices include head-mounted displays (Oculus Rift) and smart glasses (Google Glass).

[0447] Artificial intelligence analysis server: Used to analyze video data and generate supplementary information and effects. Specific software includes Google Cloud Vision API and TensorFlow.

[0448] Emotion engine: Analyzes the user's facial expressions and voice to recognize emotions in real time. Specific software includes the Affectiva SDK.

[0449] System Operation

[0450] Video Acquisition and Streaming

[0451] The server captures real-time video from high-resolution cameras installed in the physical store and streams it to user devices using AWS Elemental MediaLive. Users can then use virtual reality devices to view the store's interior from a specified viewpoint in real time.

[0452] Supplementary Information and Emotion Recognition

[0453] The AI ​​analysis server uses Google Cloud Vision API and TensorFlow to analyze the captured video data and generate supplementary information useful to the user. Meanwhile, the emotion engine analyzes the user's facial expressions and voice to collect emotional data. For example, if the user smiles, a special effect corresponding to that emotion can be added.

[0454] Emotional Data Feedback

[0455] The emotion data collected by the emotion engine is fed back to the AI ​​analysis server, which can then tailor the user experience to be more personalized. For example, if a user shows a strong interest in a particular product, the server can display additional promotional information related to that product.

[0456] User Interactions

[0457] Users can select their viewpoint within the virtual reality space and enjoy shopping while viewing the store in real time. They can also use the chat function to communicate with other users and store staff in real time, allowing them to have an experience equivalent to shopping in a physical store, even from the comfort of their own home.

[0458] Specific examples

[0459] Example 1

[0460] If a user makes an inquisitive facial expression while browsing the fragrance section of a physical store, the emotion engine will recognize this and pop up a special promotion for related products.

[0461] Example 2

[0462] When a user is viewing the clothing section, the AI ​​analysis server recognizes tags within the products and provides detailed information and stock information for the relevant products in real time.

[0463] Prompt Sentence Examples

[0464] "View product information for the fragrance section"

[0465] "Please let me know about promotions for this product."

[0466] Add recommended products to your cart

[0467] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0468] Step 1:

[0469] The server acquires real-time video from multiple high-resolution camera devices installed in the physical store. Specifically, it collects video feeds from each camera and manages them centrally. The input is video data from the camera devices, and the output is video data for transfer to the streaming server.

[0470] Step 2:

[0471] The server transfers the acquired video data to the streaming server. AWS Elemental MediaLive is used to transfer the video data in real time. The input is the video data acquired in step 1, and the output is a streaming data stream. This allows the user's device to receive the video in real time.

[0472] Step 3:

[0473] The user's device uses a virtual reality device to view the video obtained from the streaming server. The user selects a viewpoint from the application and receives video from a specific camera based on that selection. The input is the user's viewpoint selection and streaming data, and the output is the video provided to the user.

[0474] Step 4:

[0475] The AI ​​analysis server uses Google Cloud Vision API and TensorFlow to analyze video data and generate supplemental information. Specifically, it performs object recognition and motion analysis within the video to generate related product information and promotions. The input is streaming data, and the output is supplemental information.

[0476] Step 5:

[0477] The user's device displays the generated supplementary information in real time. Specifically, the information is overlaid on the video being viewed. The input is supplementary information, and the output is a video with additional information provided to the user.

[0478] Step 6:

[0479] The emotion engine analyzes the user's facial expressions and voice to recognize emotions in real time. It uses the Affectiva SDK to analyze data acquired from the user's camera and microphone. The input is the user's facial expressions and voice data, and the output is emotional data.

[0480] Step 7:

[0481] The emotional data obtained by the emotion engine is fed back to the AI ​​analysis server, which uses this data to generate additional information and effects to personalize the user experience. The input is emotional data, and the output is personalized supplementary information and effects.

[0482] Step 8:

[0483] The user's device displays the generated personalized information, allowing the user to enjoy a personalized shopping experience. The input is personalized information, and the output is a customized video provided to the user.

[0484] Step 9:

[0485] The communication support means allows users to chat and voice chat with other users and store staff in real time. Specifically, it sends and receives text messages and voice data. The input is the user's message or voice data, and the output is the delivery of the message or voice to the other party.

[0486] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0487] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0488] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0489] [Second embodiment]

[0490] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0491] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0492] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0493] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0494] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0495] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0496] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0497] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0498] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0499] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0500] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0501] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0502] The system of the present invention is designed to enable users to enjoy events such as sports and live music in real time from the comfort of their own homes using a VR device. Specific means for implementing this system are described below.

[0503] System Configuration

[0504] The system mainly consists of the following components:

[0505] Multiple Cameras

[0506] Streaming Server

[0507] User's device

[0508] VR device

[0509] AI analysis server

[0510] System Operation

[0511] Camera installation and video acquisition

[0512] The server sets up multiple cameras at the event venue, each configured to capture footage from a different perspective. For example, in the case of a sporting event, cameras may be placed around the goal, from the athlete's perspective, or in the spectator seats.

[0513] The cameras capture video in real time and send it to a streaming server, which allows the server to simultaneously obtain video data from multiple cameras.

[0514] Streaming video

[0515] The server processes the acquired video data in real time and transfers it to the user's device via a streaming server, where the user can view the streaming video through a VR device.

[0516] User Interface

[0517] The user's device installs an application, through which viewpoint selection and other operations are performed. The user can select the camera viewpoint they want to see from the viewpoint selection screen displayed within the application. For example, in the case of a music concert, they can select the viewpoint of a specific instrument performer.

[0518] AI-based analysis and information provision

[0519] The AI ​​analysis server analyzes the video data being streamed in real time. For example, in the case of a sporting event, the AI ​​analyzes the movements of players and the position of the ball and generates supplementary information, which can provide information such as highlights of plays and tactical analysis.

[0520] This supplementary information is displayed in real time on the user's device, allowing users to understand and enjoy the event more in-depth, including the progress of the match, set lists of the songs being performed, and lyric information.

[0521] Communication Features

[0522] Users can also communicate with other users who are watching the same event. The server provides chat message and voice chat functions, allowing users to interact with each other. This communication function allows users to enjoy a more interactive experience.

[0523] Specific examples

[0524] Example 1: Watching a soccer game

[0525] 1. The server sets up multiple cameras in the stadium, such as from the goalkeeper's perspective, the spectator's perspective, and the player's perspective.

[0526] 2. The user puts on the VR device at home and launches the application.

[0527] 3. The user selects the goalkeeper perspective within the app and receives streaming video from that perspective.

[0528] 4. The AI ​​analysis server analyzes player movements and generates important play and tactical information.

[0529] 5. The analysis results are displayed in real time on the user's device, allowing them to chat with other viewers.

[0530] Example 2: Watching a live music concert

[0531] 1. The server sets up cameras on multiple musicians on stage and in the audience.

[0532] 2. Through the application, users select the perspective of their favorite instrumentalist.

[0533] 3. The server streams the selected viewpoint to the user's device.

[0534] 4. The AI ​​analysis server provides set lists and lyric information for the songs being performed in real time.

[0535] 5. Users can enjoy the live stream while chatting with other fans.

[0536] In this way, the system of the present invention allows users to enjoy a real event experience from the comfort of their own home. Furthermore, AI analysis and community functions provide new added value that differs from traditional viewing and appreciation methods.

[0537] The processing flow will be explained below.

[0538] Step 1:

[0539] The server performs an initialization procedure for multiple cameras, specifically obtaining the ID of each camera and configuring it to establish a streaming connection, allowing real-time video capture from each camera.

[0540] Step 2:

[0541] The camera acquires video in real time from the location where it is installed. Specifically, the camera captures frames (image data) at regular intervals and stores the video data in a buffer.

[0542] Step 3:

[0543] The camera transmits the captured video data to a server. Specifically, the camera uploads frame data to the server in real time via a network using a streaming protocol.

[0544] Step 4:

[0545] The server processes the received video data to relay it to the streaming server. Specifically, it arranges each received frame in order and sends it to the streaming server. The streaming server manages the video data and prepares it to be provided in response to user requests.

[0546] Step 5:

[0547] The user puts on the VR device and starts the application. Specifically, the application displays the user interface and opens a viewpoint selection screen.

[0548] Step 6:

[0549] The user operates the application interface to select the desired viewpoint. Specifically, the user selects, for example, the "goalkeeper's viewpoint" from the list displayed on the viewpoint selection screen.

[0550] Step 7:

[0551] The device receives the user's selection and requests the video data of that viewpoint from the server. Specifically, the device sends the ID of the selected viewpoint to the server and requests transmission of the corresponding stream.

[0552] Step 8:

[0553] The server acquires the video stream from the requested viewpoint and delivers it to the terminal. Specifically, it selects the corresponding camera stream and starts the process of transferring it to the user's terminal.

[0554] Step 9:

[0555] The AI ​​analysis server analyzes the video data received in real time, specifically recognizing important events in the video (such as player movements and ball position) and generating analysis results.

[0556] Step 10:

[0557] The server sends the AI ​​analysis results to the user's device, specifically, sending data including generated supplementary information to the device in real time.

[0558] Step 11:

[0559] The device displays the analysis results to the user, overlaying the analyzed information on the screen so that the user can instantly access important information.

[0560] Step 12:

[0561] Users can communicate with other users while watching an event by using the chat and voice chat functions within the application to communicate via text messages and voice.

[0562] Step 13:

[0563] The server manages and relays chat and voice chat messages between users, specifically, it delivers received messages to the correct destination and establishes communication between users.

[0564] Example 1

[0565] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0566] Modern event experiences are limited by physical constraints, making it difficult for users in distant locations to participate in events in real time. Furthermore, there are few ways to quickly grasp important moments and detailed information about an event, and there are limited ways to share it with other users through communication. This leads to a decline in overall user satisfaction.

[0567] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0568] In this invention, the server includes means for acquiring real-time video from a plurality of image acquisition devices, means for transferring the acquired video to a data processing device, means for a user to select a viewpoint using a virtual reality device and receive video from that viewpoint, means for analyzing the video data using artificial intelligence and generating supplemental information, and means for providing the supplemental information to users. This allows users who are far away to participate in an event in real time, quickly grasp important scenes and detailed information, and share the event experience through communication with other users.

[0569] An "image capture device" is a device that is installed at an event venue and captures images in real time from different viewpoints.

[0570] "Data processing device" means a device that receives video data transmitted from the image acquisition device and processes and encodes it in real time.

[0571] "User" refers to a person who uses the system to view and interact with events in real time.

[0572] A "virtual reality device" is a device such as a headset or goggles that a user wears to experience virtual reality.

[0573] "Artificial intelligence" refers to advanced algorithms and software used to analyze video data and generate supplemental information.

[0574] "Supplementary information" is information generated from the results of analysis by artificial intelligence, and is intended to provide users with additional understanding and interest.

[0575] The "viewpoint selection interface" is a function of the application that allows the user to select and manipulate different camera viewpoints.

[0576] The "communication function" is a function that allows users to communicate with other users in real time via messages and voice.

[0577] A "streaming server" is a server that transfers processed video data to a user's terminal in real time.

[0578] A "user terminal" is a device such as a PC, smartphone, or tablet operated by the user, which functions in conjunction with the virtual reality device.

[0579] The system of this invention is designed to enable users to watch events such as sports and live music in real time from the comfort of their own homes using a virtual reality device (hereinafter referred to as a VR device). Specific means for implementing this system are described below.

[0580] System Configuration

[0581] The system mainly consists of the following components:

[0582] Multiple image capture devices (cameras)

[0583] Data processing device (streaming server and AI analysis server)

[0584] User's device

[0585] VR device

[0586] System Operation

[0587] Camera installation and video acquisition

[0588] The server installs multiple image capture devices at the event venue. These image capture devices are positioned so that they capture video from different viewpoints. For example, in the case of a sporting event, image capture devices are placed around the goal area, from the players' viewpoints, and in the spectator seats. The image capture devices capture video in real time and send the video to a data processing device. This allows the server to simultaneously acquire video data from multiple image capture devices.

[0589] Streaming video

[0590] The server processes the acquired video data in real time and transfers it to the user's device via a streaming server. The user can then view this streaming video through a VR device.

[0591] User Interface

[0592] A dedicated application is installed on the user's device, and this application allows for viewpoint selection and other operations. The user can select the image capture device viewpoint they want to see from the viewpoint selection interface. For example, in the case of a live music performance, they can select the viewpoint of a specific instrument performer.

[0593] AI-based analysis and information provision

[0594] The AI ​​analysis server analyzes the video data being streamed in real time. For example, in the case of a sporting event, it analyzes the movements of the players and the position of the ball and generates supplementary information. This supplementary information is displayed in real time on the user's device, allowing the user to visually obtain information such as the progress of the game, highlights, and tactical analysis.

[0595] Communication Features

[0596] Users can communicate with other users who are watching the same event. The server provides chat message and voice chat functions, supporting users to interact with each other through these. This communication function allows users to enjoy a more interactive experience.

[0597] Specific examples

[0598] Example 1: Watching a soccer game

[0599] 1. The server installs multiple image capture devices at the stadium's goalkeeper's viewpoint, the spectator's viewpoint, and the player's viewpoint.

[0600] 2. The user puts on the VR device at home and launches the application.

[0601] 3. The user selects the goalkeeper's perspective within the app and receives streaming video from that perspective.

[0602] 4. The AI ​​analysis server analyzes player movements and generates important play and tactical information.

[0603] 5. The analysis results are displayed in real time on the user's device, allowing them to chat with other viewers.

[0604] Example 2: Watching a live music concert

[0605] 1. The server installs image capture devices on multiple instrument performers on stage and in the audience seats.

[0606] 2. Through the application, users can select the perspective of their favorite instrumentalist.

[0607] 3. The server streams the video from the selected viewpoint to the user's device.

[0608] 4. The AI ​​analysis server provides set lists and lyric information for the songs being performed in real time.

[0609] 5. Users can enjoy the live performance while chatting with other fans.

[0610] Example prompts for generative AI models

[0611] While watching a soccer match from the goalkeeper's point of view, generate sentences that provide analysis information on player movements and tactics.

[0612] To enjoy live music, generate text that provides perspective footage of specific instrumentalists and set list information.

[0613] In this way, the system of the present invention allows users to enjoy a real event experience from the comfort of their own home. Furthermore, AI analysis and community functions provide new added value that differs from traditional viewing and appreciation methods.

[0614] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0615] System program processing flow

[0616] Step 1: Installing the image capture device and acquiring images

[0617] The server installs multiple image capture devices (cameras) at the event venue. These cameras are positioned to capture images in real time from different viewpoints. Specifically, in a sporting event, cameras are placed around the goal, from the players' viewpoint, and in the spectator seats. The input from the cameras is real-time images, and the output is captured image data.

[0618] Step 2: Transferring video data

[0619] The camera transmits the captured real-time video to a data processing device (streaming server). In this step, the video data input from the camera is transferred to the streaming server via the network. The output is multiple video data stored in the streaming server.

[0620] Step 3: Processing the video data

[0621] The server processes the received video data in real time. Specifically, it compresses and encodes the video data and converts it into a format suitable for streaming. The input is the video data sent from the camera, and the output is the encoded streaming video data.

[0622] Step 4: Streaming the video data

[0623] The server transfers the processed video data to the user's device via a streaming server. The input is encoded video data, and the output is a video stream that can be played back in real time on the user's device.

[0624] Step 5: Launch the application and select a viewpoint

[0625] The user launches a dedicated application on the device. After the user completes login authentication, a viewpoint selection screen is displayed. Here, the user can select the viewpoint they want to view. The input is the viewpoint selected by the user, and the output is video data based on the selected viewpoint.

[0626] Step 6: Receiving and displaying viewpoint images

[0627] The user's device receives the video from the selected viewpoint and displays it on the VR device. The input is video data sent from the streaming server, and the output is real-time video displayed on the user's VR device.

[0628] Step 7: Analyze the video data

[0629] The AI ​​analytics server analyzes the video data received in real time. Specifically, it identifies player movements and ball position and generates important highlights and tactical information. The input is streaming video data, and the output is supplemental information based on the analysis.

[0630] Step 8: Provide supporting information

[0631] The AI ​​analysis server sends the generated supplementary information to the user's device, which displays it in real time. The input is the analyzed supplementary information, and the output is the information displayed on the user's device.

[0632] Step 9: Providing communication features

[0633] Users communicate with other users through chat and voice chat functions within the application. This function is provided by a server. The input is the user's message or voice data, and the output is the communication data transmitted to other users.

[0634] Through this step-by-step process, users can watch events in real time from the comfort of their own homes and enjoy a variety of interactive experiences.

[0635] (Application example 1)

[0636] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0637] Conventional live event viewing systems make it difficult for users to enjoy the event from multiple angles in real time or to interact with other viewers. They also often fail to obtain important scenes or specific information in real time while watching. This limits the user experience and prevents the full appeal of the event.

[0638] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0639] In this invention, the server includes means for acquiring real-time video from multiple cameras, means for transferring the acquired video to a streaming server, means for a user to select a viewpoint using a VR device and receive video from that viewpoint, means for analyzing video data using AI and generating supplemental information, means for providing the supplemental information to the user, means for supporting communication between users, means for interacting in a virtual space using smart glasses or a head-mounted display, means for providing real-time chat and voice calls with other users, means for automatically detecting highlights of an event from video in real time using a generative AI model, and means for providing the generated highlights to the user based on prompt text. This allows users to enjoy an event from multiple viewpoints in an immersive way, and enables interactive interaction with other viewers and the acquisition of important scenes and supplemental information through real-time analysis using AI.

[0640] "Multiple cameras" refers to multiple imaging devices installed at an event venue to capture footage in real time from different perspectives.

[0641] A "streaming server" refers to a server system that processes acquired video data in real time and distributes it to the user's device.

[0642] "User terminal" refers to a computer device that allows a user to receive event footage, select viewpoints, display supplementary information, and communicate.

[0643] "VR device" refers to devices such as head-mounted displays and smart glasses that users wear to experience virtual reality.

[0644] "AI analysis server" refers to a server system that runs artificial intelligence algorithms to analyze acquired video data and generate supplemental information.

[0645] "Supplementary information" refers to additional information such as highlights of the performance, tactical information, and set lists of songs performed, which is generated by AI analysis based on streaming video.

[0646] "Communication function" refers to the function that allows users to interact with other viewers through chat messages and voice calls while watching an event.

[0647] "Smart glasses" refers to a glasses-type device that users can wear to view images in a virtual space.

[0648] A "head-mounted display" refers to a display device worn on the head that allows users to intuitively experience a virtual space.

[0649] "Virtual space" refers to a digitally constructed simulated environment that users experience visually and aurally through VR devices.

[0650] "Real-time chat" refers to a means of communication that allows users to exchange text messages in real time.

[0651] "Voice Call" means a means by which Users can engage in real-time voice communication.

[0652] "Generative AI model" refers to a machine learning algorithm that automatically generates highlights and supplemental information based on input data.

[0653] "Prompt sentence" refers to text input provided to a generative AI model to generate a specific output.

[0654] This invention is a system that allows users to enjoy events such as sports and live music in real time from home using a VR device. The system consists of multiple cameras, a streaming server, a user's device, a VR device, and an AI analysis server.

[0655] System Configuration

[0656] Hardware Configuration

[0657] Multiple cameras: Installed at the event venue, capturing footage in real time from different perspectives.

[0658] Streaming server: A server system that processes acquired video data in real time and distributes it to user devices. Protocols used include WebRTC and RTMP.

[0659] User's device: A computer device used to receive the video, select viewpoints, display supplementary information, and communicate. Smartphones and PCs are often used.

[0660] VR equipment: A device such as a head-mounted display or smart glasses worn by a user to create a virtual reality experience.

[0661] AI analysis server: A server system that runs artificial intelligence algorithms to analyze video data in real time and generate supplementary information. It uses TensorFlow and OpenCV.

[0662] Software Configuration

[0663] Viewpoint switching function: A function that receives a user's viewpoint switching request and sends the corresponding camera video stream to the user's device.

[0664] Real-time AI analysis: The AI ​​analysis server receives the acquired video data and generates important scenes and supplementary information in real time. Analysis is performed using TensorFlow and OpenCV.

[0665] Communication feature: Supports real-time chat and voice calls between users. Firebase and Agora.io SDK are used.

[0666] Generative AI model: Applies machine learning algorithms to generate key scenes and supplementary information from video footage. Generates appropriate output based on prompts.

[0667] Operation explanation

[0668] Viewpoint switching

[0669] When a user selects a viewpoint, the server uses a WebRTC client to acquire the video stream from the selected camera and transmit it to the user's device. For example, in a live music concert, it is possible to select the viewpoint of a specific instrumentalist.

[0670] Real-time AI analysis

[0671] The video data sent to the streaming server is analyzed by an AI analysis server, which uses TensorFlow models to detect important plays during sporting events or lyrics of live music concerts in real time and notify users.

[0672] Communication Features

[0673] The communication feature uses the Firebase real-time database to manage text messages and the Agora.io SDK to provide voice calls, allowing users to interact with other viewers in real time within the VR space.

[0674] Specific examples

[0675] Example 1: Watching a sports game

[0676] 1. The server sets up cameras at multiple viewpoints in the stadium.

[0677] 2. The user puts on the VR device at home and launches the application.

[0678] 3. The user selects the viewpoint of their choice and receives streaming video from that viewpoint.

[0679] 4. The AI ​​analysis server analyzes players' movements and generates important play and tactical information in real time, which is then provided to users.

[0680] Example 2: Watching a live music concert

[0681] 1. The server sets up cameras on multiple musicians on stage and in the audience.

[0682] 2. The user selects the perspective of their favorite instrument performer and receives streaming footage from that perspective.

[0683] 3. The AI ​​analysis server provides set lists and lyric information for the songs being performed in real time.

[0684] Prompt Sentence Examples

[0685] Here is an example prompt for automatic event highlight detection using an AI model:

[0686] Input: "Auto-detect key plays from this footage and create highlights."

[0687] Output: Automatically highlighted video clip

[0688] This invention allows users to enjoy a more immersive multi-perspective experience, interact with other users, and experience the event in greater depth thanks to real-time analysis by AI.

[0689] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0690] Step 1:

[0691] Multiple cameras are installed at the event venue, each capturing video in real time from a different viewpoint. The server acquires the video data from the cameras and transfers it to the streaming server.

[0692] Input: Live video from the event venue

[0693] Data processing: Video capture and transfer to streaming server

[0694] Output: Raw video data stored on a streaming server

[0695] Step 2:

[0696] The streaming server receives the acquired video data and delivers the video data to the user's terminal in real time.

[0697] Input: Raw video data stored on a streaming server

[0698] Data processing: video data compression and transfer

[0699] Output: Real-time video sent to user device

[0700] Step 3:

[0701] The user selects a viewpoint through the terminal, which then sends a viewpoint switching request to the server, which then delivers the video stream of the selected viewpoint to the user terminal.

[0702] Input: User's request to switch viewpoint

[0703] Data processing: Selecting the video stream corresponding to the viewpoint

[0704] Output: Video of the selected viewpoint sent to the user's device

[0705] Step 4:

[0706] The video data sent to the streaming server is analyzed in real time by an AI analysis server. A generative AI model is used for the analysis to extract important scenes and supplementary information. For example, important plays during a sporting event or lyrics from a live music concert can be automatically extracted.

[0707] Input: Raw video data sent from the streaming server

[0708] Data Computing: Real-time analytics with TensorFlow and OpenCV

[0709] Output: Analyzed important scenes and supplementary information

[0710] Step 5:

[0711] The user's device receives supplemental information sent from the AI ​​analysis server and displays it in real time, such as highlights of important performances and set lists of songs performed.

[0712] Input: Supplementary information sent from the AI ​​analysis server

[0713] Data processing: Converting supplementary information into a display format

[0714] Output: Supplementary information displayed on the user's terminal

[0715] Step 6:

[0716] Users can chat and talk to other users watching the same event in real time through their devices, and the server uses the Firebase real-time database and Agora.io SDK to support this functionality.

[0717] Input: User text messages and voice data

[0718] Data processing: real-time transmission of text messages, compression and transmission of voice data

[0719] Output: Messages and audio delivered to other users

[0720] Step 7:

[0721] When users use the generative AI model as part of real-time AI analysis, they can input prompts to automatically detect event highlights and specific scenes from video.

[0722] Input: Prompt text (e.g. "Automatically detect important plays from this footage and create highlights")

[0723] Data calculation: Execute a generative AI model based on the input prompt.

[0724] Output: Automatically generated highlight reels and clips of specific scenes

[0725] This allows users to enjoy an immersive multi-perspective experience, real-time interactive interactions, and detailed supplemental information about the event through AI analysis.

[0726] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0727] This invention provides a new user experience by combining a system that allows users to experience events such as sports and live music concerts in real time using a VR device from the comfort of their own home with an emotion engine that recognizes the user's emotions. Specific means for implementing this system are described below.

[0728] System Configuration

[0729] The system mainly consists of the following components:

[0730] Multiple Cameras

[0731] Streaming Server

[0732] User's device

[0733] VR device

[0734] AI analysis server

[0735] Emotion Engine

[0736] System Operation

[0737] Camera installation and video acquisition

[0738] The server installs multiple cameras at the event venue and captures video in real time from each camera. This video data is then transferred to the streaming server.

[0739] Streaming video

[0740] The server relays the acquired video to the streaming server in real time, allowing the user's device to receive the video data directly from the streaming server.

[0741] User Interface

[0742] The user's device performs viewpoint selection and other operations via the application. The user can select each camera viewpoint from the viewpoint selection screen displayed within the application.

[0743] AI-based analysis and information provision

[0744] The AI ​​analysis server analyzes the video data being streamed in real time, and based on the results of this analysis, users can receive supplemental information such as highlights of plays and tactical analysis.

[0745] Emotion recognition by emotion engine

[0746] The emotion engine recognizes the user's emotions in real time by analyzing their facial expressions and voice data. Based on the analysis results, the emotion engine generates supplementary information and effects according to the user's emotions. For example, if the user shows a surprised expression, it can respond by displaying a replay of the play at that moment.

[0747] Emotional Data Feedback

[0748] The emotional data recognized by the emotion engine is fed back to the AI ​​analysis server. Based on this data, the system adjusts the user experience across the entire system. For example, if data is obtained showing multiple users expressing excitement, the system can respond by providing more detailed explanatory information.

[0749] Communication Features

[0750] Users can chat and voice chat with other users while watching the event, and the system manages and relays these communications, facilitating interaction between users.

[0751] Specific examples

[0752] Example 1: Watching a soccer game

[0753] 1. The server sets up cameras at multiple viewpoints in the stadium, such as the goalkeeper's viewpoint, the spectator's viewpoint, and the player's viewpoint.

[0754] 2. The user puts on the VR device at home and launches the application.

[0755] 3. The user selects "Goalkeeper View" on the viewpoint selection screen and receives video from that viewpoint.

[0756] 4. The AI ​​analysis server analyzes player movements and ball trajectory to generate important play and tactical information.

[0757] 5. The emotion engine analyzes the user's facial expressions and voice, and when it recognizes moments of surprise or joy, it instantly displays replay footage and detailed commentary.

[0758] 6. Users can enjoy communicating with other viewers using the chat function.

[0759] Example 2: Watching a live music concert

[0760] 1. The server sets up cameras on the stage for multiple musicians and in the audience.

[0761] 2. Through the application, the user selects the perspective of, for example, a drummer.

[0762] 3. The server streams the selected viewpoint to the user's device.

[0763] 4. The AI ​​analysis server provides performance set lists and lyric information in real time.

[0764] 5. The emotion engine recognizes the user's emotions and responds by adding special effects when the user is emotional.

[0765] 6. Users can enjoy the live show while chatting with other fans.

[0766] This system allows users to enjoy live sports and music events from the comfort of their own home through an advanced VR experience, and the introduction of an emotion engine provides a more personalized experience.

[0767] The processing flow will be explained below.

[0768] Step 1:

[0769] The server will set up multiple cameras at the event venue, and will configure the initial settings of these cameras so that they can capture images from each viewpoint in real time.

[0770] Step 2:

[0771] The camera captures video in real time from its installed location, acquiring frames at regular intervals and generating video data.

[0772] Step 3:

[0773] The camera transmits the captured video data to a server. Specifically, the video data is transferred to a streaming server via a network in real time.

[0774] Step 4:

[0775] The server receives video data from multiple cameras and manages it in a streaming server, which prepares to provide the corresponding video stream in response to a viewpoint request from a user.

[0776] Step 5:

[0777] The user puts on the VR device and launches the application, which displays the main screen and offers viewpoint selection options.

[0778] Step 6:

[0779] The user selects the desired viewpoint on the viewpoint selection screen within the application. For example, the user selects the "goalkeeper viewpoint."

[0780] Step 7:

[0781] Based on the user's viewpoint selection, the terminal requests the video data of the viewpoint from the streaming server. Specifically, the terminal transmits the selected viewpoint ID to the streaming server and requests the corresponding stream.

[0782] Step 8:

[0783] The streaming server acquires the video stream of the requested viewpoint and transmits it to the terminal in real time, allowing the user to view the video from the selected viewpoint.

[0784] Step 9:

[0785] The AI ​​analysis server analyzes the streaming video data in real time to detect player movements and important events (e.g., goals, shots).

[0786] Step 10:

[0787] The AI ​​analysis server generates supplemental information based on the analysis results and sends it to the device, allowing users to receive the supplemental information in real time.

[0788] Step 11:

[0789] The device displays the AI ​​analysis results to the user, overlaying supplemental information on the screen in real time.

[0790] Step 12:

[0791] The emotion engine analyzes the user's facial expressions and voice data to recognize their emotions. Specifically, it captures the user's reactions through a camera and microphone and generates emotion data.

[0792] Step 13:

[0793] The emotion engine feeds the recognized emotion data back to the AI ​​analysis server, which then adjusts the overall system experience based on this data, for example, providing special effects or additional information to excited users.

[0794] Step 14:

[0795] Users can communicate with other users while watching an event, using chat and voice chat functions within the application to interact with other viewers.

[0796] Step 15:

[0797] The server manages and relays chat and voice chat messages between users, thereby supporting real-time communication between users.

[0798] This detailed processing flow allows users to enjoy events in real time from home and enjoy a personalized experience through the emotion engine.

[0799] Example 2

[0800] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0801] In conventional virtual reality event viewing systems, users are limited to selecting a viewpoint and viewing experience, and real-time emotion analysis and personalized experiences based on that analysis are not performed. Furthermore, communication functions between users are insufficient, making it insufficient to provide a richer user experience.

[0802] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0803] In this invention, the server includes means for acquiring real-time video from multiple image capture devices, means for transferring the acquired video to a data distribution server, means for allowing a user to select a viewpoint using a virtual reality device and receiving video from that viewpoint, means for analyzing video data using artificial intelligence and generating supplemental information, means for providing the supplemental information to the user, means for supporting communication between users, means for analyzing user emotions and generating supplemental information and effects based on the emotions, and means for feeding back the emotion data to the analysis server and adjusting the overall user experience. This enables the provision of personalized information and the addition of effects according to the user's emotions, further enhancing communication between users.

[0804] An "imaging device" is a device for capturing images in real time.

[0805] The "data distribution server" is a server that distributes acquired video data to user terminals in real time.

[0806] A "virtual reality device" is a device that allows users to immerse themselves in a virtual environment and select and control their viewpoint.

[0807] "Video data" refers to real-time video captured by an imaging device.

[0808] "Artificial intelligence" is a technology that analyzes video data and user emotional data to generate supplementary information and effects.

[0809] "Supplemental Information" is additional information provided to enhance the viewing or viewing experience.

[0810] "Effects" are visual and sound effects added to the video.

[0811] "User emotion" refers to the emotional state analyzed from the user's facial expressions and voice.

[0812] "Feedback" is the process of adjusting the overall system experience based on acquired data.

[0813] "Communication means" refers to a function that supports real-time communication between users.

[0814] This invention provides a new user experience by combining a system that allows users to experience events such as sports and live music concerts in real time using a VR device from the comfort of their own home with an emotion engine that recognizes the user's emotions. Specific means for implementing this system are described below.

[0815] System Configuration

[0816] The system mainly consists of the following components:

[0817] Multiple imaging devices (cameras)

[0818] Data distribution server

[0819] User's device

[0820] Virtual reality device (VR device)

[0821] AI analysis server

[0822] Emotion Engine

[0823] System Operation

[0824] Camera installation and video acquisition

[0825] The server installs multiple camera devices at the event venue, captures video in real time from each camera device, and transfers the video data to a data distribution server.

[0826] Streaming video

[0827] The server relays the acquired video to the data distribution server in real time, allowing the user's device to receive the video data directly from the data distribution server.

[0828] User Interface

[0829] The user's device performs viewpoint selection and other operations via the application. The user can select each camera viewpoint from the viewpoint selection screen displayed within the application.

[0830] AI-based analysis and information provision

[0831] The AI ​​analysis server analyzes the video data being streamed in real time, and based on the results of this analysis, users can receive supplemental information such as highlights of plays and tactical analysis.

[0832] Emotion recognition by emotion engine

[0833] The emotion engine recognizes the user's emotions in real time by analyzing their facial expressions and voice data. Based on the analysis results, the emotion engine generates supplementary information and effects according to the user's emotions. For example, if the user shows a surprised expression, it can respond by displaying a replay of the play at that moment.

[0834] Emotional Data Feedback

[0835] The emotional data recognized by the emotion engine is fed back to the AI ​​analysis server. Based on this data, the system adjusts the user experience across the entire system. For example, if data is obtained showing multiple users expressing excitement, the system can respond by providing more detailed explanatory information.

[0836] Communication Features

[0837] Users can chat and voice chat with other users while watching the event, and the system manages and relays these communications, facilitating interaction between users.

[0838] Specific examples

[0839] Example 1: Watching a soccer game

[0840] 1. The server sets up cameras at multiple viewpoints in the stadium, such as the goalkeeper's viewpoint, the spectator's viewpoint, and the player's viewpoint.

[0841] 2. The user puts on the VR device at home and launches the application.

[0842] 3. The user selects "Goalkeeper View" on the viewpoint selection screen and receives video from that viewpoint.

[0843] 4. The AI ​​analysis server analyzes player movements and ball trajectory to generate important play and tactical information.

[0844] 5. The emotion engine analyzes the user's facial expressions and voice, and when it recognizes moments of surprise or joy, it instantly displays replay footage and detailed commentary.

[0845] 6. Users can enjoy communicating with other viewers using the chat function.

[0846] Example 2: Watching a live music concert

[0847] 1. The server sets up cameras on the stage for multiple musicians and in the audience.

[0848] 2. Through the application, the user selects the perspective of, for example, a drummer.

[0849] 3. The server streams the selected viewpoint to the user's device.

[0850] 4. The AI ​​analysis server provides performance set lists and lyric information in real time.

[0851] 5. The emotion engine recognizes the user's emotions and responds by adding special effects when the user is emotional.

[0852] 6. Users can enjoy the live show while chatting with other fans.

[0853] Prompt Sentence Examples

[0854] Here are some examples of specific prompts for this system:

[0855] 1. Emotion Recognition in Sporting Events: Describe a process for displaying a replay and detailed commentary the moment the user expresses surprise.

[0856] 2. Personalizing Live Music Experience: What is your process for adding effects in real time based on the user's emotions?

[0857] This system allows users to enjoy live sports and music events from the comfort of their own home through an advanced VR experience, and the introduction of an emotion engine provides a more personalized experience.

[0858] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0859] Step 1: Camera installation and video acquisition

[0860] The server will set up multiple camera devices at the event venue. For example, at a sporting event, cameras will be installed to view the goalkeeper, the spectators, and the players, while at a music concert, cameras will be installed to view each musician on stage and in the spectators' seats. The server will acquire images from these cameras in real time.

[0861] Input: Video data from multiple cameras installed at the event venue.

[0862] Data processing: Converts the analog video signal from the camera into a digital signal and formats it for transmission to the streaming server.

[0863] Output: The formatted digital video data is transferred to a streaming server.

[0864] Step 2: Streaming the video

[0865] The server relays the captured video in real time to the data distribution server, which then distributes the video to the user's device.

[0866] Input: Digital video data transferred from the server.

[0867] Data processing: Packetization and streaming optimization for relaying and real-time delivery of digital video data.

[0868] Output: Streaming video data is sent to the user's device.

[0869] Step 3: User Interface

[0870] The user wears the VR device at home and starts the application. The user selects the desired camera viewpoint from the viewpoint selection screen within the application.

[0871] Input: Viewpoint selection information made by the user to the application.

[0872] Data processing: Based on the user's viewpoint selection, the video data of the corresponding camera viewpoint is extracted.

[0873] Output: Real-time video from the selected viewpoint is displayed in the user's VR device.

[0874] Step 4: AI-based video analysis and information provision

[0875] The AI ​​analysis server analyzes the video data being streamed in real time, for example, analyzing the movements of players and the trajectory of the ball, and then generates highlights of the play and tactical information based on that analysis.

[0876] Input: Real-time video data.

[0877] Data Processing: Analyze video data using generative AI models to extract and generate key events and tactical information.

[0878] Output: Analyzed highlights and tactical information are provided to the user.

[0879] Step 5: Emotion Recognition with the Emotion Engine

[0880] The emotion engine analyzes the user's facial expressions and voice data in real time to recognize their emotions. For example, if the user shows a surprised expression, it will display a replay of the play at that moment.

[0881] Input: User's facial and voice data.

[0882] Data processing: Using an AI model, the system analyzes the user's emotions and generates corresponding effects and supplementary information.

[0883] Output: Real-time effects and replay footage are displayed according to the user's emotions.

[0884] Step 6: Feedback of emotional data

[0885] The emotional data recognized by the emotion engine is fed back to the AI ​​analysis server. The system adjusts the user experience based on this data. For example, if multiple users show signs of excitement, more detailed commentary will be provided.

[0886] Input: Recognized user emotion data.

[0887] Data processing: Emotional data is sent to the analysis server and reflected in the user experience of the entire system.

[0888] Output: A tailored user experience based on sentiment data.

[0889] Step 7: Communication Functions

[0890] Users can chat and voice chat with other users while watching the event, and the system manages and relays these communications, facilitating interaction between users.

[0891] Input: Chat messages and voice data between users.

[0892] Data processing: Relaying and logging chat messages and voice data.

[0893] Output: A real-time communication experience.

[0894] (Application example 2)

[0895] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0896] In conventional online shopping systems, users are limited to visual information about products, which is far from the shopping experience of a physical store. Furthermore, when users express interest in a product or have a particular emotion, they are unable to provide dynamic information that responds to that interest. Furthermore, real-time communication between users is not adequately supported.

[0897] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for acquiring real-time video from multiple image capture devices, means for transferring the acquired video to a streaming server, means for a user to select a viewpoint using a virtual reality device and receive video from that viewpoint, means for analyzing video data using artificial intelligence and generating supplemental information, means for providing the supplemental information to the user, means for supporting communication between users, means for recognizing user emotions using an emotion engine and generating additional information and effects based on the recognition results, and means for feeding back emotion data to the artificial intelligence analysis server to adjust the user experience. This allows users to enjoy a realistic shopping experience from the comfort of their own home, as if they were in a physical store, and makes it possible to provide dynamic information about interesting products and add emotion-based effects.

[0898] A "camera" is a device that captures images of the real world and records them as digital data.

[0899] A "streaming server" is a server that distributes acquired video data in real time.

[0900] A "virtual reality device" is a device that allows a user to enjoy an experience in a 3D space, and generally includes a head-mounted display and smart glasses.

[0901] "Artificial intelligence" is a technology that analyzes large amounts of data and finds patterns and trends to make decisions and make predictions.

[0902] "Supplementary information" refers to additional information, commentary, promotions, etc. related to the video the user is watching.

[0903] "Communication support means" is a function that supports chat and voice chat between users, enabling smooth information exchange.

[0904] The "emotion engine" is a system that analyzes data such as the user's facial expressions and voice, and recognizes emotions in real time.

[0905] "Emotion data" is data that indicates the user's emotional state as analyzed by the emotion engine.

[0906] The "artificial intelligence analysis server" is a server that analyzes emotional data and video data and makes decisions such as providing information and generating effects.

[0907] The present invention is a system that provides a shopping experience from home that makes you feel as if you are actually in a store, and to achieve this, it combines multiple technologies such as a camera device, a streaming server, a virtual reality device, artificial intelligence, and an emotion engine. Specific embodiments of the system are described below.

[0908] System Configuration

[0909] Camera: Installed in multiple areas of the store, it captures high-resolution video in real time. Specific hardware used includes high-resolution cameras (e.g., Sony α7 series).

[0910] Streaming server: Used to distribute acquired video data to user devices in real time. The specific software used is AWS Elemental MediaLive.

[0911] Virtual reality devices: Devices that allow users to immerse themselves in virtual reality spaces. Specific devices include head-mounted displays (Oculus Rift) and smart glasses (Google Glass).

[0912] Artificial intelligence analysis server: Used to analyze video data and generate supplementary information and effects. Specific software includes Google Cloud Vision API and TensorFlow.

[0913] Emotion engine: Analyzes the user's facial expressions and voice to recognize emotions in real time. Specific software includes the Affectiva SDK.

[0914] System Operation

[0915] Video Acquisition and Streaming

[0916] The server captures real-time video from high-resolution cameras installed in the physical store and streams it to user devices using AWS Elemental MediaLive. Users can then use virtual reality devices to view the store's interior from a specified viewpoint in real time.

[0917] Supplementary Information and Emotion Recognition

[0918] The AI ​​analysis server uses Google Cloud Vision API and TensorFlow to analyze the captured video data and generate supplementary information useful to the user. Meanwhile, the emotion engine analyzes the user's facial expressions and voice to collect emotional data. For example, if the user smiles, a special effect corresponding to that emotion can be added.

[0919] Emotional Data Feedback

[0920] The emotion data collected by the emotion engine is fed back to the AI ​​analysis server, which can then tailor the user experience to be more personalized. For example, if a user shows a strong interest in a particular product, the server can display additional promotional information related to that product.

[0921] User Interactions

[0922] Users can select their viewpoint within the virtual reality space and enjoy shopping while viewing the store in real time. They can also use the chat function to communicate with other users and store staff in real time, allowing them to have an experience equivalent to shopping in a physical store, even from the comfort of their own home.

[0923] Specific examples

[0924] Example 1

[0925] If a user makes an inquisitive facial expression while browsing the fragrance section of a physical store, the emotion engine will recognize this and pop up a special promotion for related products.

[0926] Example 2

[0927] When a user is viewing the clothing section, the AI ​​analysis server recognizes tags within the products and provides detailed information and stock information for the relevant products in real time.

[0928] Prompt Sentence Examples

[0929] "View product information for the fragrance section"

[0930] "Please let me know about promotions for this product."

[0931] Add recommended products to your cart

[0932] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0933] Step 1:

[0934] The server acquires real-time video from multiple high-resolution camera devices installed in the physical store. Specifically, it collects video feeds from each camera and manages them centrally. The input is video data from the camera devices, and the output is video data for transfer to the streaming server.

[0935] Step 2:

[0936] The server transfers the acquired video data to the streaming server. AWS Elemental MediaLive is used to transfer the video data in real time. The input is the video data acquired in step 1, and the output is a streaming data stream. This allows the user's device to receive the video in real time.

[0937] Step 3:

[0938] The user's device uses a virtual reality device to view the video obtained from the streaming server. The user selects a viewpoint from the application and receives video from a specific camera based on that selection. The input is the user's viewpoint selection and streaming data, and the output is the video provided to the user.

[0939] Step 4:

[0940] The AI ​​analysis server uses Google Cloud Vision API and TensorFlow to analyze video data and generate supplemental information. Specifically, it performs object recognition and motion analysis within the video to generate related product information and promotions. The input is streaming data, and the output is supplemental information.

[0941] Step 5:

[0942] The user's device displays the generated supplementary information in real time. Specifically, the information is overlaid on the video being viewed. The input is supplementary information, and the output is a video with additional information provided to the user.

[0943] Step 6:

[0944] The emotion engine analyzes the user's facial expressions and voice to recognize emotions in real time. It uses the Affectiva SDK to analyze data acquired from the user's camera and microphone. The input is the user's facial expressions and voice data, and the output is emotional data.

[0945] Step 7:

[0946] The emotional data obtained by the emotion engine is fed back to the AI ​​analysis server, which uses this data to generate additional information and effects to personalize the user experience. The input is emotional data, and the output is personalized supplementary information and effects.

[0947] Step 8:

[0948] The user's device displays the generated personalized information, allowing the user to enjoy a personalized shopping experience. The input is personalized information, and the output is a customized video provided to the user.

[0949] Step 9:

[0950] The communication support means allows users to chat and voice chat with other users and store staff in real time. Specifically, it sends and receives text messages and voice data. The input is the user's message or voice data, and the output is the delivery of the message or voice to the other party.

[0951] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0952] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0953] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0954] [Third embodiment]

[0955] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0956] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[0957] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0958] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0959] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0960] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0961] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0962] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0963] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0964] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0965] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0966] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0967] The system of the present invention is designed to enable users to enjoy events such as sports and live music in real time from the comfort of their own homes using a VR device. Specific means for implementing this system are described below.

[0968] System Configuration

[0969] The system mainly consists of the following components:

[0970] Multiple Cameras

[0971] Streaming Server

[0972] User's device

[0973] VR device

[0974] AI analysis server

[0975] System Operation

[0976] Camera installation and video acquisition

[0977] The server sets up multiple cameras at the event venue, each configured to capture footage from a different perspective. For example, in the case of a sporting event, cameras may be placed around the goal, from the athlete's perspective, or in the spectator seats.

[0978] The cameras capture video in real time and send it to a streaming server, which allows the server to simultaneously obtain video data from multiple cameras.

[0979] Streaming video

[0980] The server processes the acquired video data in real time and transfers it to the user's device via a streaming server, where the user can view the streaming video through a VR device.

[0981] User Interface

[0982] The user's device installs an application, through which viewpoint selection and other operations are performed. The user can select the camera viewpoint they want to see from the viewpoint selection screen displayed within the application. For example, in the case of a music concert, they can select the viewpoint of a specific instrument performer.

[0983] AI-based analysis and information provision

[0984] The AI ​​analysis server analyzes the video data being streamed in real time. For example, in the case of a sporting event, the AI ​​analyzes the movements of players and the position of the ball and generates supplementary information, which can provide information such as highlights of plays and tactical analysis.

[0985] This supplementary information is displayed in real time on the user's device, allowing users to understand and enjoy the event more in-depth, including the progress of the match, set lists of the songs being performed, and lyric information.

[0986] Communication Features

[0987] Users can also communicate with other users who are watching the same event. The server provides chat message and voice chat functions, allowing users to interact with each other. This communication function allows users to enjoy a more interactive experience.

[0988] Specific examples

[0989] Example 1: Watching a soccer game

[0990] 1. The server sets up multiple cameras in the stadium, such as from the goalkeeper's perspective, the spectator's perspective, and the player's perspective.

[0991] 2. The user puts on the VR device at home and launches the application.

[0992] 3. The user selects the goalkeeper perspective within the app and receives streaming video from that perspective.

[0993] 4. The AI ​​analysis server analyzes player movements and generates important play and tactical information.

[0994] 5. The analysis results are displayed in real time on the user's device, allowing them to chat with other viewers.

[0995] Example 2: Watching a live music concert

[0996] 1. The server sets up cameras on multiple musicians on stage and in the audience.

[0997] 2. Through the application, users select the perspective of their favorite instrumentalist.

[0998] 3. The server streams the selected viewpoint to the user's device.

[0999] 4. The AI ​​analysis server provides set lists and lyric information for the songs being performed in real time.

[1000] 5. Users can enjoy the live stream while chatting with other fans.

[1001] In this way, the system of the present invention allows users to enjoy a real event experience from the comfort of their own home. Furthermore, AI analysis and community functions provide new added value that differs from traditional viewing and appreciation methods.

[1002] The processing flow will be explained below.

[1003] Step 1:

[1004] The server performs an initialization procedure for multiple cameras, specifically obtaining the ID of each camera and configuring it to establish a streaming connection, allowing real-time video capture from each camera.

[1005] Step 2:

[1006] The camera acquires video in real time from the location where it is installed. Specifically, the camera captures frames (image data) at regular intervals and stores the video data in a buffer.

[1007] Step 3:

[1008] The camera transmits the captured video data to a server. Specifically, the camera uploads frame data to the server in real time via a network using a streaming protocol.

[1009] Step 4:

[1010] The server processes the received video data to relay it to the streaming server. Specifically, it arranges each received frame in order and sends it to the streaming server. The streaming server manages the video data and prepares it to be provided in response to user requests.

[1011] Step 5:

[1012] The user puts on the VR device and starts the application. Specifically, the application displays the user interface and opens a viewpoint selection screen.

[1013] Step 6:

[1014] The user operates the application interface to select the desired viewpoint. Specifically, the user selects, for example, the "goalkeeper's viewpoint" from the list displayed on the viewpoint selection screen.

[1015] Step 7:

[1016] The device receives the user's selection and requests the video data of that viewpoint from the server. Specifically, the device sends the ID of the selected viewpoint to the server and requests transmission of the corresponding stream.

[1017] Step 8:

[1018] The server acquires the video stream from the requested viewpoint and delivers it to the terminal. Specifically, it selects the corresponding camera stream and starts the process of transferring it to the user's terminal.

[1019] Step 9:

[1020] The AI ​​analysis server analyzes the video data received in real time, specifically recognizing important events in the video (such as player movements and ball position) and generating analysis results.

[1021] Step 10:

[1022] The server sends the AI ​​analysis results to the user's device, specifically, sending data including generated supplementary information to the device in real time.

[1023] Step 11:

[1024] The device displays the analysis results to the user, overlaying the analyzed information on the screen so that the user can instantly access important information.

[1025] Step 12:

[1026] Users can communicate with other users while watching an event by using the chat and voice chat functions within the application to communicate via text messages and voice.

[1027] Step 13:

[1028] The server manages and relays chat and voice chat messages between users, specifically, it delivers received messages to the correct destination and establishes communication between users.

[1029] Example 1

[1030] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1031] Modern event experiences are limited by physical constraints, making it difficult for users in distant locations to participate in events in real time. Furthermore, there are few ways to quickly grasp important moments and detailed information about an event, and there are limited ways to share it with other users through communication. This leads to a decline in overall user satisfaction.

[1032] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1033] In this invention, the server includes means for acquiring real-time video from a plurality of image acquisition devices, means for transferring the acquired video to a data processing device, means for a user to select a viewpoint using a virtual reality device and receive video from that viewpoint, means for analyzing the video data using artificial intelligence and generating supplemental information, and means for providing the supplemental information to users. This allows users who are far away to participate in an event in real time, quickly grasp important scenes and detailed information, and share the event experience through communication with other users.

[1034] An "image capture device" is a device that is installed at an event venue and captures images in real time from different viewpoints.

[1035] "Data processing device" means a device that receives video data transmitted from the image acquisition device and processes and encodes it in real time.

[1036] "User" refers to a person who uses the system to view and interact with events in real time.

[1037] A "virtual reality device" is a device such as a headset or goggles that a user wears to experience virtual reality.

[1038] "Artificial intelligence" refers to advanced algorithms and software used to analyze video data and generate supplemental information.

[1039] "Supplementary information" is information generated from the results of analysis by artificial intelligence, and is intended to provide users with additional understanding and interest.

[1040] The "viewpoint selection interface" is a function of the application that allows the user to select and manipulate different camera viewpoints.

[1041] The "communication function" is a function that allows users to communicate with other users in real time via messages and voice.

[1042] A "streaming server" is a server that transfers processed video data to a user's terminal in real time.

[1043] A "user terminal" is a device such as a PC, smartphone, or tablet operated by the user, which functions in conjunction with the virtual reality device.

[1044] The system of this invention is designed to enable users to watch events such as sports and live music in real time from the comfort of their own homes using a virtual reality device (hereinafter referred to as a VR device). Specific means for implementing this system are described below.

[1045] System Configuration

[1046] The system mainly consists of the following components:

[1047] Multiple image capture devices (cameras)

[1048] Data processing device (streaming server and AI analysis server)

[1049] User's device

[1050] VR device

[1051] System Operation

[1052] Camera installation and video acquisition

[1053] The server installs multiple image capture devices at the event venue. These image capture devices are positioned so that they capture video from different viewpoints. For example, in the case of a sporting event, image capture devices are placed around the goal area, from the players' viewpoints, and in the spectator seats. The image capture devices capture video in real time and send the video to a data processing device. This allows the server to simultaneously acquire video data from multiple image capture devices.

[1054] Streaming video

[1055] The server processes the acquired video data in real time and transfers it to the user's device via a streaming server. The user can then view this streaming video through a VR device.

[1056] User Interface

[1057] A dedicated application is installed on the user's device, and this application allows for viewpoint selection and other operations. The user can select the image capture device viewpoint they want to see from the viewpoint selection interface. For example, in the case of a live music performance, they can select the viewpoint of a specific instrument performer.

[1058] AI-based analysis and information provision

[1059] The AI ​​analysis server analyzes the video data being streamed in real time. For example, in the case of a sporting event, it analyzes the movements of the players and the position of the ball and generates supplementary information. This supplementary information is displayed in real time on the user's device, allowing the user to visually obtain information such as the progress of the game, highlights, and tactical analysis.

[1060] Communication Features

[1061] Users can communicate with other users who are watching the same event. The server provides chat message and voice chat functions, supporting users to interact with each other through these. This communication function allows users to enjoy a more interactive experience.

[1062] Specific examples

[1063] Example 1: Watching a soccer game

[1064] 1. The server installs multiple image capture devices at the stadium's goalkeeper's viewpoint, the spectator's viewpoint, and the player's viewpoint.

[1065] 2. The user puts on the VR device at home and launches the application.

[1066] 3. The user selects the goalkeeper's perspective within the app and receives streaming video from that perspective.

[1067] 4. The AI ​​analysis server analyzes player movements and generates important play and tactical information.

[1068] 5. The analysis results are displayed in real time on the user's device, allowing them to chat with other viewers.

[1069] Example 2: Watching a live music concert

[1070] 1. The server installs image capture devices on multiple instrument performers on stage and in the audience seats.

[1071] 2. Through the application, users can select the perspective of their favorite instrumentalist.

[1072] 3. The server streams the video from the selected viewpoint to the user's device.

[1073] 4. The AI ​​analysis server provides set lists and lyric information for the songs being performed in real time.

[1074] 5. Users can enjoy the live performance while chatting with other fans.

[1075] Example prompts for generative AI models

[1076] While watching a soccer match from the goalkeeper's point of view, generate sentences that provide analysis information on player movements and tactics.

[1077] To enjoy live music, generate text that provides perspective footage of specific instrumentalists and set list information.

[1078] In this way, the system of the present invention allows users to enjoy a real event experience from the comfort of their own home. Furthermore, AI analysis and community functions provide new added value that differs from traditional viewing and appreciation methods.

[1079] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1080] System program processing flow

[1081] Step 1: Installing the image capture device and acquiring images

[1082] The server installs multiple image capture devices (cameras) at the event venue. These cameras are positioned to capture images in real time from different viewpoints. Specifically, in a sporting event, cameras are placed around the goal, from the players' viewpoint, and in the spectator seats. The input from the cameras is real-time images, and the output is captured image data.

[1083] Step 2: Transferring video data

[1084] The camera transmits the captured real-time video to a data processing device (streaming server). In this step, the video data input from the camera is transferred to the streaming server via the network. The output is multiple video data stored in the streaming server.

[1085] Step 3: Processing the video data

[1086] The server processes the received video data in real time. Specifically, it compresses and encodes the video data and converts it into a format suitable for streaming. The input is the video data sent from the camera, and the output is the encoded streaming video data.

[1087] Step 4: Streaming the video data

[1088] The server transfers the processed video data to the user's device via a streaming server. The input is encoded video data, and the output is a video stream that can be played back in real time on the user's device.

[1089] Step 5: Launch the application and select a viewpoint

[1090] The user launches a dedicated application on the device. After the user completes login authentication, a viewpoint selection screen is displayed. Here, the user can select the viewpoint they want to view. The input is the viewpoint selected by the user, and the output is video data based on the selected viewpoint.

[1091] Step 6: Receiving and displaying viewpoint images

[1092] The user's device receives the video from the selected viewpoint and displays it on the VR device. The input is video data sent from the streaming server, and the output is real-time video displayed on the user's VR device.

[1093] Step 7: Analyze the video data

[1094] The AI ​​analytics server analyzes the video data received in real time. Specifically, it identifies player movements and ball position and generates important highlights and tactical information. The input is streaming video data, and the output is supplemental information based on the analysis.

[1095] Step 8: Provide supporting information

[1096] The AI ​​analysis server sends the generated supplementary information to the user's device, which displays it in real time. The input is the analyzed supplementary information, and the output is the information displayed on the user's device.

[1097] Step 9: Providing communication features

[1098] Users communicate with other users through chat and voice chat functions within the application. This function is provided by a server. The input is the user's message or voice data, and the output is the communication data transmitted to other users.

[1099] Through this step-by-step process, users can watch events in real time from the comfort of their own homes and enjoy a variety of interactive experiences.

[1100] (Application example 1)

[1101] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1102] Conventional live event viewing systems make it difficult for users to enjoy the event from multiple angles in real time or to interact with other viewers. They also often fail to obtain important scenes or specific information in real time while watching. This limits the user experience and prevents the full appeal of the event.

[1103] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1104] In this invention, the server includes means for acquiring real-time video from multiple cameras, means for transferring the acquired video to a streaming server, means for a user to select a viewpoint using a VR device and receive video from that viewpoint, means for analyzing video data using AI and generating supplemental information, means for providing the supplemental information to the user, means for supporting communication between users, means for interacting in a virtual space using smart glasses or a head-mounted display, means for providing real-time chat and voice calls with other users, means for automatically detecting highlights of an event from video in real time using a generative AI model, and means for providing the generated highlights to the user based on prompt text. This allows users to enjoy an event from multiple viewpoints in an immersive way, and enables interactive interaction with other viewers and the acquisition of important scenes and supplemental information through real-time analysis using AI.

[1105] "Multiple cameras" refers to multiple imaging devices installed at an event venue to capture footage in real time from different perspectives.

[1106] A "streaming server" refers to a server system that processes acquired video data in real time and distributes it to the user's device.

[1107] "User terminal" refers to a computer device that allows a user to receive event footage, select viewpoints, display supplementary information, and communicate.

[1108] "VR device" refers to devices such as head-mounted displays and smart glasses that users wear to experience virtual reality.

[1109] "AI analysis server" refers to a server system that runs artificial intelligence algorithms to analyze acquired video data and generate supplemental information.

[1110] "Supplementary information" refers to additional information such as highlights of the performance, tactical information, and set lists of songs performed, which is generated by AI analysis based on streaming video.

[1111] "Communication function" refers to the function that allows users to interact with other viewers through chat messages and voice calls while watching an event.

[1112] "Smart glasses" refers to a glasses-type device that users can wear to view images in a virtual space.

[1113] A "head-mounted display" refers to a display device worn on the head that allows users to intuitively experience a virtual space.

[1114] "Virtual space" refers to a digitally constructed simulated environment that users experience visually and aurally through VR devices.

[1115] "Real-time chat" refers to a means of communication that allows users to exchange text messages in real time.

[1116] "Voice Call" means a means by which Users can engage in real-time voice communication.

[1117] "Generative AI model" refers to a machine learning algorithm that automatically generates highlights and supplemental information based on input data.

[1118] "Prompt sentence" refers to text input provided to a generative AI model to generate a specific output.

[1119] This invention is a system that allows users to enjoy events such as sports and live music in real time from home using a VR device. The system consists of multiple cameras, a streaming server, a user's device, a VR device, and an AI analysis server.

[1120] System Configuration

[1121] Hardware Configuration

[1122] Multiple cameras: Installed at the event venue, capturing footage in real time from different perspectives.

[1123] Streaming server: A server system that processes acquired video data in real time and distributes it to user devices. Protocols used include WebRTC and RTMP.

[1124] User's device: A computer device used to receive the video, select viewpoints, display supplementary information, and communicate. Smartphones and PCs are often used.

[1125] VR equipment: A device such as a head-mounted display or smart glasses worn by a user to create a virtual reality experience.

[1126] AI analysis server: A server system that runs artificial intelligence algorithms to analyze video data in real time and generate supplementary information. It uses TensorFlow and OpenCV.

[1127] Software Configuration

[1128] Viewpoint switching function: A function that receives a user's viewpoint switching request and sends the corresponding camera video stream to the user's device.

[1129] Real-time AI analysis: The AI ​​analysis server receives the acquired video data and generates important scenes and supplementary information in real time. Analysis is performed using TensorFlow and OpenCV.

[1130] Communication feature: Supports real-time chat and voice calls between users. Firebase and Agora.io SDK are used.

[1131] Generative AI model: Applies machine learning algorithms to generate key scenes and supplementary information from video footage. Generates appropriate output based on prompts.

[1132] Operation explanation

[1133] Viewpoint switching

[1134] When a user selects a viewpoint, the server uses a WebRTC client to acquire the video stream from the selected camera and transmit it to the user's device. For example, in a live music concert, it is possible to select the viewpoint of a specific instrumentalist.

[1135] Real-time AI analysis

[1136] The video data sent to the streaming server is analyzed by an AI analysis server, which uses TensorFlow models to detect important plays during sporting events or lyrics of live music concerts in real time and notify users.

[1137] Communication Features

[1138] The communication feature uses the Firebase real-time database to manage text messages and the Agora.io SDK to provide voice calls, allowing users to interact with other viewers in real time within the VR space.

[1139] Specific examples

[1140] Example 1: Watching a sports game

[1141] 1. The server sets up cameras at multiple viewpoints in the stadium.

[1142] 2. The user puts on the VR device at home and launches the application.

[1143] 3. The user selects the viewpoint of their choice and receives streaming video from that viewpoint.

[1144] 4. The AI ​​analysis server analyzes players' movements and generates important play and tactical information in real time, which is then provided to users.

[1145] Example 2: Watching a live music concert

[1146] 1. The server sets up cameras on multiple musicians on stage and in the audience.

[1147] 2. The user selects the perspective of their favorite instrument performer and receives streaming footage from that perspective.

[1148] 3. The AI ​​analysis server provides set lists and lyric information for the songs being performed in real time.

[1149] Prompt Sentence Examples

[1150] Here is an example prompt for automatic event highlight detection using an AI model:

[1151] Input: "Auto-detect key plays from this footage and create highlights."

[1152] Output: Automatically highlighted video clip

[1153] This invention allows users to enjoy a more immersive multi-perspective experience, interact with other users, and experience the event in greater depth thanks to real-time analysis by AI.

[1154] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1155] Step 1:

[1156] Multiple cameras are installed at the event venue, each capturing video in real time from a different viewpoint. The server acquires the video data from the cameras and transfers it to the streaming server.

[1157] Input: Live video from the event venue

[1158] Data processing: Video capture and transfer to streaming server

[1159] Output: Raw video data stored on a streaming server

[1160] Step 2:

[1161] The streaming server receives the acquired video data and delivers the video data to the user's terminal in real time.

[1162] Input: Raw video data stored on a streaming server

[1163] Data processing: video data compression and transfer

[1164] Output: Real-time video sent to user device

[1165] Step 3:

[1166] The user selects a viewpoint through the terminal, which then sends a viewpoint switching request to the server, which then delivers the video stream of the selected viewpoint to the user terminal.

[1167] Input: User's request to switch viewpoint

[1168] Data processing: Selecting the video stream corresponding to the viewpoint

[1169] Output: Video of the selected viewpoint sent to the user's device

[1170] Step 4:

[1171] The video data sent to the streaming server is analyzed in real time by an AI analysis server. A generative AI model is used for the analysis to extract important scenes and supplementary information. For example, important plays during a sporting event or lyrics from a live music concert can be automatically extracted.

[1172] Input: Raw video data sent from the streaming server

[1173] Data Computing: Real-time analytics with TensorFlow and OpenCV

[1174] Output: Analyzed important scenes and supplementary information

[1175] Step 5:

[1176] The user's device receives supplemental information sent from the AI ​​analysis server and displays it in real time, such as highlights of important performances and set lists of songs performed.

[1177] Input: Supplementary information sent from the AI ​​analysis server

[1178] Data processing: Converting supplementary information into a display format

[1179] Output: Supplementary information displayed on the user's terminal

[1180] Step 6:

[1181] Users can chat and talk to other users watching the same event in real time through their devices, and the server uses the Firebase real-time database and Agora.io SDK to support this functionality.

[1182] Input: User text messages and voice data

[1183] Data processing: real-time transmission of text messages, compression and transmission of voice data

[1184] Output: Messages and audio delivered to other users

[1185] Step 7:

[1186] When users use the generative AI model as part of real-time AI analysis, they can input prompts to automatically detect event highlights and specific scenes from video.

[1187] Input: Prompt text (e.g. "Automatically detect important plays from this footage and create highlights")

[1188] Data calculation: Execute a generative AI model based on the input prompt.

[1189] Output: Automatically generated highlight reels and clips of specific scenes

[1190] This allows users to enjoy an immersive multi-perspective experience, real-time interactive interactions, and detailed supplemental information about the event through AI analysis.

[1191] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1192] This invention provides a new user experience by combining a system that allows users to experience events such as sports and live music concerts in real time using a VR device from the comfort of their own home with an emotion engine that recognizes the user's emotions. Specific means for implementing this system are described below.

[1193] System Configuration

[1194] The system mainly consists of the following components:

[1195] Multiple Cameras

[1196] Streaming Server

[1197] User's device

[1198] VR device

[1199] AI analysis server

[1200] Emotion Engine

[1201] System Operation

[1202] Camera installation and video acquisition

[1203] The server installs multiple cameras at the event venue and captures video in real time from each camera. This video data is then transferred to the streaming server.

[1204] Streaming video

[1205] The server relays the acquired video to the streaming server in real time, allowing the user's device to receive the video data directly from the streaming server.

[1206] User Interface

[1207] The user's device performs viewpoint selection and other operations via the application. The user can select each camera viewpoint from the viewpoint selection screen displayed within the application.

[1208] AI-based analysis and information provision

[1209] The AI ​​analysis server analyzes the video data being streamed in real time, and based on the results of this analysis, users can receive supplemental information such as highlights of plays and tactical analysis.

[1210] Emotion recognition by emotion engine

[1211] The emotion engine recognizes the user's emotions in real time by analyzing their facial expressions and voice data. Based on the analysis results, the emotion engine generates supplementary information and effects according to the user's emotions. For example, if the user shows a surprised expression, it can respond by displaying a replay of the play at that moment.

[1212] Emotional Data Feedback

[1213] The emotional data recognized by the emotion engine is fed back to the AI ​​analysis server. Based on this data, the system adjusts the user experience across the entire system. For example, if data is obtained showing multiple users expressing excitement, the system can respond by providing more detailed explanatory information.

[1214] Communication Features

[1215] Users can chat and voice chat with other users while watching the event, and the system manages and relays these communications, facilitating interaction between users.

[1216] Specific examples

[1217] Example 1: Watching a soccer game

[1218] 1. The server sets up cameras at multiple viewpoints in the stadium, such as the goalkeeper's viewpoint, the spectator's viewpoint, and the player's viewpoint.

[1219] 2. The user puts on the VR device at home and launches the application.

[1220] 3. The user selects "Goalkeeper View" on the viewpoint selection screen and receives video from that viewpoint.

[1221] 4. The AI ​​analysis server analyzes player movements and ball trajectory to generate important play and tactical information.

[1222] 5. The emotion engine analyzes the user's facial expressions and voice, and when it recognizes moments of surprise or joy, it instantly displays replay footage and detailed commentary.

[1223] 6. Users can enjoy communicating with other viewers using the chat function.

[1224] Example 2: Watching a live music concert

[1225] 1. The server sets up cameras on the stage for multiple musicians and in the audience.

[1226] 2. Through the application, the user selects the perspective of, for example, a drummer.

[1227] 3. The server streams the selected viewpoint to the user's device.

[1228] 4. The AI ​​analysis server provides performance set lists and lyric information in real time.

[1229] 5. The emotion engine recognizes the user's emotions and responds by adding special effects when the user is emotional.

[1230] 6. Users can enjoy the live show while chatting with other fans.

[1231] This system allows users to enjoy live sports and music events from the comfort of their own home through an advanced VR experience, and the introduction of an emotion engine provides a more personalized experience.

[1232] The processing flow will be explained below.

[1233] Step 1:

[1234] The server will set up multiple cameras at the event venue, and will configure the initial settings of these cameras so that they can capture images from each viewpoint in real time.

[1235] Step 2:

[1236] The camera captures video in real time from its installed location, acquiring frames at regular intervals and generating video data.

[1237] Step 3:

[1238] The camera transmits the captured video data to a server. Specifically, the video data is transferred to a streaming server via a network in real time.

[1239] Step 4:

[1240] The server receives video data from multiple cameras and manages it in a streaming server, which prepares to provide the corresponding video stream in response to a viewpoint request from a user.

[1241] Step 5:

[1242] The user puts on the VR device and launches the application, which displays the main screen and offers viewpoint selection options.

[1243] Step 6:

[1244] The user selects the desired viewpoint on the viewpoint selection screen within the application. For example, the user selects the "goalkeeper viewpoint."

[1245] Step 7:

[1246] Based on the user's viewpoint selection, the terminal requests the video data of the viewpoint from the streaming server. Specifically, the terminal transmits the selected viewpoint ID to the streaming server and requests the corresponding stream.

[1247] Step 8:

[1248] The streaming server acquires the video stream of the requested viewpoint and transmits it to the terminal in real time, allowing the user to view the video from the selected viewpoint.

[1249] Step 9:

[1250] The AI ​​analysis server analyzes the streaming video data in real time to detect player movements and important events (e.g., goals, shots).

[1251] Step 10:

[1252] The AI ​​analysis server generates supplemental information based on the analysis results and sends it to the device, allowing users to receive the supplemental information in real time.

[1253] Step 11:

[1254] The device displays the AI ​​analysis results to the user, overlaying supplemental information on the screen in real time.

[1255] Step 12:

[1256] The emotion engine analyzes the user's facial expressions and voice data to recognize their emotions. Specifically, it captures the user's reactions through a camera and microphone and generates emotion data.

[1257] Step 13:

[1258] The emotion engine feeds the recognized emotion data back to the AI ​​analysis server, which then adjusts the overall system experience based on this data, for example, providing special effects or additional information to excited users.

[1259] Step 14:

[1260] Users can communicate with other users while watching an event, using chat and voice chat functions within the application to interact with other viewers.

[1261] Step 15:

[1262] The server manages and relays chat and voice chat messages between users, thereby supporting real-time communication between users.

[1263] This detailed processing flow allows users to enjoy events in real time from home and enjoy a personalized experience through the emotion engine.

[1264] Example 2

[1265] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1266] In conventional virtual reality event viewing systems, users are limited to selecting a viewpoint and viewing experience, and real-time emotion analysis and personalized experiences based on that analysis are not performed. Furthermore, communication functions between users are insufficient, making it insufficient to provide a richer user experience.

[1267] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1268] In this invention, the server includes means for acquiring real-time video from multiple image capture devices, means for transferring the acquired video to a data distribution server, means for allowing a user to select a viewpoint using a virtual reality device and receiving video from that viewpoint, means for analyzing video data using artificial intelligence and generating supplemental information, means for providing the supplemental information to the user, means for supporting communication between users, means for analyzing user emotions and generating supplemental information and effects based on the emotions, and means for feeding back the emotion data to the analysis server and adjusting the overall user experience. This enables the provision of personalized information and the addition of effects according to the user's emotions, further enhancing communication between users.

[1269] An "imaging device" is a device for capturing images in real time.

[1270] The "data distribution server" is a server that distributes acquired video data to user terminals in real time.

[1271] A "virtual reality device" is a device that allows users to immerse themselves in a virtual environment and select and control their viewpoint.

[1272] "Video data" refers to real-time video captured by an imaging device.

[1273] "Artificial intelligence" is a technology that analyzes video data and user emotional data to generate supplementary information and effects.

[1274] "Supplemental Information" is additional information provided to enhance the viewing or viewing experience.

[1275] "Effects" are visual and sound effects added to the video.

[1276] "User emotion" refers to the emotional state analyzed from the user's facial expressions and voice.

[1277] "Feedback" is the process of adjusting the overall system experience based on acquired data.

[1278] "Communication means" refers to a function that supports real-time communication between users.

[1279] This invention provides a new user experience by combining a system that allows users to experience events such as sports and live music concerts in real time using a VR device from the comfort of their own home with an emotion engine that recognizes the user's emotions. Specific means for implementing this system are described below.

[1280] System Configuration

[1281] The system mainly consists of the following components:

[1282] Multiple imaging devices (cameras)

[1283] Data distribution server

[1284] User's device

[1285] Virtual reality device (VR device)

[1286] AI analysis server

[1287] Emotion Engine

[1288] System Operation

[1289] Camera installation and video acquisition

[1290] The server installs multiple camera devices at the event venue, captures video in real time from each camera device, and transfers the video data to a data distribution server.

[1291] Streaming video

[1292] The server relays the acquired video to the data distribution server in real time, allowing the user's device to receive the video data directly from the data distribution server.

[1293] User Interface

[1294] The user's device performs viewpoint selection and other operations via the application. The user can select each camera viewpoint from the viewpoint selection screen displayed within the application.

[1295] AI-based analysis and information provision

[1296] The AI ​​analysis server analyzes the video data being streamed in real time, and based on the results of this analysis, users can receive supplemental information such as highlights of plays and tactical analysis.

[1297] Emotion recognition by emotion engine

[1298] The emotion engine recognizes the user's emotions in real time by analyzing their facial expressions and voice data. Based on the analysis results, the emotion engine generates supplementary information and effects according to the user's emotions. For example, if the user shows a surprised expression, it can respond by displaying a replay of the play at that moment.

[1299] Emotional Data Feedback

[1300] The emotional data recognized by the emotion engine is fed back to the AI ​​analysis server. Based on this data, the system adjusts the user experience across the entire system. For example, if data is obtained showing multiple users expressing excitement, the system can respond by providing more detailed explanatory information.

[1301] Communication Features

[1302] Users can chat and voice chat with other users while watching the event, and the system manages and relays these communications, facilitating interaction between users.

[1303] Specific examples

[1304] Example 1: Watching a soccer game

[1305] 1. The server sets up cameras at multiple viewpoints in the stadium, such as the goalkeeper's viewpoint, the spectator's viewpoint, and the player's viewpoint.

[1306] 2. The user puts on the VR device at home and launches the application.

[1307] 3. The user selects "Goalkeeper View" on the viewpoint selection screen and receives video from that viewpoint.

[1308] 4. The AI ​​analysis server analyzes player movements and ball trajectory to generate important play and tactical information.

[1309] 5. The emotion engine analyzes the user's facial expressions and voice, and when it recognizes moments of surprise or joy, it instantly displays replay footage and detailed commentary.

[1310] 6. Users can enjoy communicating with other viewers using the chat function.

[1311] Example 2: Watching a live music concert

[1312] 1. The server sets up cameras on the stage for multiple musicians and in the audience.

[1313] 2. Through the application, the user selects the perspective of, for example, a drummer.

[1314] 3. The server streams the selected viewpoint to the user's device.

[1315] 4. The AI ​​analysis server provides performance set lists and lyric information in real time.

[1316] 5. The emotion engine recognizes the user's emotions and responds by adding special effects when the user is emotional.

[1317] 6. Users can enjoy the live show while chatting with other fans.

[1318] Prompt Sentence Examples

[1319] Here are some examples of specific prompts for this system:

[1320] 1. Emotion Recognition in Sporting Events: Describe a process for displaying a replay and detailed commentary the moment the user expresses surprise.

[1321] 2. Personalizing Live Music Experience: What is your process for adding effects in real time based on the user's emotions?

[1322] This system allows users to enjoy live sports and music events from the comfort of their own home through an advanced VR experience, and the introduction of an emotion engine provides a more personalized experience.

[1323] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1324] Step 1: Camera installation and video acquisition

[1325] The server will set up multiple camera devices at the event venue. For example, at a sporting event, cameras will be installed to view the goalkeeper, the spectators, and the players, while at a music concert, cameras will be installed to view each musician on stage and in the spectators' seats. The server will acquire images from these cameras in real time.

[1326] Input: Video data from multiple cameras installed at the event venue.

[1327] Data processing: Converts the analog video signal from the camera into a digital signal and formats it for transmission to the streaming server.

[1328] Output: The formatted digital video data is transferred to a streaming server.

[1329] Step 2: Streaming the video

[1330] The server relays the captured video in real time to the data distribution server, which then distributes the video to the user's device.

[1331] Input: Digital video data transferred from the server.

[1332] Data processing: Packetization and streaming optimization for relaying and real-time delivery of digital video data.

[1333] Output: Streaming video data is sent to the user's device.

[1334] Step 3: User Interface

[1335] The user wears the VR device at home and starts the application. The user selects the desired camera viewpoint from the viewpoint selection screen within the application.

[1336] Input: Viewpoint selection information made by the user to the application.

[1337] Data processing: Based on the user's viewpoint selection, the video data of the corresponding camera viewpoint is extracted.

[1338] Output: Real-time video from the selected viewpoint is displayed in the user's VR device.

[1339] Step 4: AI-based video analysis and information provision

[1340] The AI ​​analysis server analyzes the video data being streamed in real time, for example, analyzing the movements of players and the trajectory of the ball, and then generates highlights of the play and tactical information based on that analysis.

[1341] Input: Real-time video data.

[1342] Data Processing: Analyze video data using generative AI models to extract and generate key events and tactical information.

[1343] Output: Analyzed highlights and tactical information are provided to the user.

[1344] Step 5: Emotion Recognition with the Emotion Engine

[1345] The emotion engine analyzes the user's facial expressions and voice data in real time to recognize their emotions. For example, if the user shows a surprised expression, it will display a replay of the play at that moment.

[1346] Input: User's facial and voice data.

[1347] Data processing: Using an AI model, the system analyzes the user's emotions and generates corresponding effects and supplementary information.

[1348] Output: Real-time effects and replay footage are displayed according to the user's emotions.

[1349] Step 6: Feedback of emotional data

[1350] The emotional data recognized by the emotion engine is fed back to the AI ​​analysis server. The system adjusts the user experience based on this data. For example, if multiple users show signs of excitement, more detailed commentary will be provided.

[1351] Input: Recognized user emotion data.

[1352] Data processing: Emotional data is sent to the analysis server and reflected in the user experience of the entire system.

[1353] Output: A tailored user experience based on sentiment data.

[1354] Step 7: Communication Functions

[1355] Users can chat and voice chat with other users while watching the event, and the system manages and relays these communications, facilitating interaction between users.

[1356] Input: Chat messages and voice data between users.

[1357] Data processing: Relaying and logging chat messages and voice data.

[1358] Output: A real-time communication experience.

[1359] (Application example 2)

[1360] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1361] In conventional online shopping systems, users are limited to visual information about products, which is far from the shopping experience of a physical store. Furthermore, when users express interest in a product or have a particular emotion, they are unable to provide dynamic information that responds to that interest. Furthermore, real-time communication between users is not adequately supported.

[1362] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for acquiring real-time video from multiple image capture devices, means for transferring the acquired video to a streaming server, means for a user to select a viewpoint using a virtual reality device and receive video from that viewpoint, means for analyzing video data using artificial intelligence and generating supplemental information, means for providing the supplemental information to the user, means for supporting communication between users, means for recognizing user emotions using an emotion engine and generating additional information and effects based on the recognition results, and means for feeding back emotion data to the artificial intelligence analysis server to adjust the user experience. This allows users to enjoy a realistic shopping experience from the comfort of their own home, as if they were in a physical store, and makes it possible to provide dynamic information about interesting products and add emotion-based effects.

[1363] A "camera" is a device that captures images of the real world and records them as digital data.

[1364] A "streaming server" is a server that distributes acquired video data in real time.

[1365] A "virtual reality device" is a device that allows a user to enjoy an experience in a 3D space, and generally includes a head-mounted display and smart glasses.

[1366] "Artificial intelligence" is a technology that analyzes large amounts of data and finds patterns and trends to make decisions and make predictions.

[1367] "Supplementary information" refers to additional information, commentary, promotions, etc. related to the video the user is watching.

[1368] "Communication support means" is a function that supports chat and voice chat between users, enabling smooth information exchange.

[1369] The "emotion engine" is a system that analyzes data such as the user's facial expressions and voice, and recognizes emotions in real time.

[1370] "Emotion data" is data that indicates the user's emotional state as analyzed by the emotion engine.

[1371] The "artificial intelligence analysis server" is a server that analyzes emotional data and video data and makes decisions such as providing information and generating effects.

[1372] The present invention is a system that provides a shopping experience from home that makes you feel as if you are actually in a store, and to achieve this, it combines multiple technologies such as a camera device, a streaming server, a virtual reality device, artificial intelligence, and an emotion engine. Specific embodiments of the system are described below.

[1373] System Configuration

[1374] Camera: Installed in multiple areas of the store, it captures high-resolution video in real time. Specific hardware used includes high-resolution cameras (e.g., Sony α7 series).

[1375] Streaming server: Used to distribute acquired video data to user devices in real time. The specific software used is AWS Elemental MediaLive.

[1376] Virtual reality devices: Devices that allow users to immerse themselves in virtual reality spaces. Specific devices include head-mounted displays (Oculus Rift) and smart glasses (Google Glass).

[1377] Artificial intelligence analysis server: Used to analyze video data and generate supplementary information and effects. Specific software includes Google Cloud Vision API and TensorFlow.

[1378] Emotion engine: Analyzes the user's facial expressions and voice to recognize emotions in real time. Specific software includes the Affectiva SDK.

[1379] System Operation

[1380] Video Acquisition and Streaming

[1381] The server captures real-time video from high-resolution cameras installed in the physical store and streams it to user devices using AWS Elemental MediaLive. Users can then use virtual reality devices to view the store's interior from a specified viewpoint in real time.

[1382] Supplementary Information and Emotion Recognition

[1383] The AI ​​analysis server uses Google Cloud Vision API and TensorFlow to analyze the captured video data and generate supplementary information useful to the user. Meanwhile, the emotion engine analyzes the user's facial expressions and voice to collect emotional data. For example, if the user smiles, a special effect corresponding to that emotion can be added.

[1384] Emotional Data Feedback

[1385] The emotion data collected by the emotion engine is fed back to the AI ​​analysis server, which can then tailor the user experience to be more personalized. For example, if a user shows a strong interest in a particular product, the server can display additional promotional information related to that product.

[1386] User Interactions

[1387] Users can select their viewpoint within the virtual reality space and enjoy shopping while viewing the store in real time. They can also use the chat function to communicate with other users and store staff in real time, allowing them to have an experience equivalent to shopping in a physical store, even from the comfort of their own home.

[1388] Specific examples

[1389] Example 1

[1390] If a user makes an inquisitive facial expression while browsing the fragrance section of a physical store, the emotion engine will recognize this and pop up a special promotion for related products.

[1391] Example 2

[1392] When a user is viewing the clothing section, the AI ​​analysis server recognizes tags within the products and provides detailed information and stock information for the relevant products in real time.

[1393] Prompt Sentence Examples

[1394] "View product information for the fragrance section"

[1395] "Please let me know about promotions for this product."

[1396] Add recommended products to your cart

[1397] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1398] Step 1:

[1399] The server acquires real-time video from multiple high-resolution camera devices installed in the physical store. Specifically, it collects video feeds from each camera and manages them centrally. The input is video data from the camera devices, and the output is video data for transfer to the streaming server.

[1400] Step 2:

[1401] The server transfers the acquired video data to the streaming server. AWS Elemental MediaLive is used to transfer the video data in real time. The input is the video data acquired in step 1, and the output is a streaming data stream. This allows the user's device to receive the video in real time.

[1402] Step 3:

[1403] The user's device uses a virtual reality device to view the video obtained from the streaming server. The user selects a viewpoint from the application and receives video from a specific camera based on that selection. The input is the user's viewpoint selection and streaming data, and the output is the video provided to the user.

[1404] Step 4:

[1405] The AI ​​analysis server uses Google Cloud Vision API and TensorFlow to analyze video data and generate supplemental information. Specifically, it performs object recognition and motion analysis within the video to generate related product information and promotions. The input is streaming data, and the output is supplemental information.

[1406] Step 5:

[1407] The user's device displays the generated supplementary information in real time. Specifically, the information is overlaid on the video being viewed. The input is supplementary information, and the output is a video with additional information provided to the user.

[1408] Step 6:

[1409] The emotion engine analyzes the user's facial expressions and voice to recognize emotions in real time. It uses the Affectiva SDK to analyze data acquired from the user's camera and microphone. The input is the user's facial expressions and voice data, and the output is emotional data.

[1410] Step 7:

[1411] The emotional data obtained by the emotion engine is fed back to the AI ​​analysis server, which uses this data to generate additional information and effects to personalize the user experience. The input is emotional data, and the output is personalized supplementary information and effects.

[1412] Step 8:

[1413] The user's device displays the generated personalized information, allowing the user to enjoy a personalized shopping experience. The input is personalized information, and the output is a customized video provided to the user.

[1414] Step 9:

[1415] The communication support means allows users to chat and voice chat with other users and store staff in real time. Specifically, it sends and receives text messages and voice data. The input is the user's message or voice data, and the output is the delivery of the message or voice to the other party.

[1416] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1417] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1418] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1419] [Fourth embodiment]

[1420] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1421] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1422] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1423] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1424] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1425] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1426] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1427] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1428] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1429] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1430] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1431] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1432] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1433] The system of the present invention is designed to enable users to enjoy events such as sports and live music in real time from the comfort of their own homes using a VR device. Specific means for implementing this system are described below.

[1434] System Configuration

[1435] The system mainly consists of the following components:

[1436] Multiple Cameras

[1437] Streaming Server

[1438] User's device

[1439] VR device

[1440] AI analysis server

[1441] System Operation

[1442] Camera installation and video acquisition

[1443] The server sets up multiple cameras at the event venue, each configured to capture footage from a different perspective. For example, in the case of a sporting event, cameras may be placed around the goal, from the athlete's perspective, or in the spectator seats.

[1444] The cameras capture video in real time and send it to a streaming server, which allows the server to simultaneously obtain video data from multiple cameras.

[1445] Streaming video

[1446] The server processes the acquired video data in real time and transfers it to the user's device via a streaming server, where the user can view the streaming video through a VR device.

[1447] User Interface

[1448] The user's device installs an application, through which viewpoint selection and other operations are performed. The user can select the camera viewpoint they want to see from the viewpoint selection screen displayed within the application. For example, in the case of a music concert, they can select the viewpoint of a specific instrument performer.

[1449] AI-based analysis and information provision

[1450] The AI ​​analysis server analyzes the video data being streamed in real time. For example, in the case of a sporting event, the AI ​​analyzes the movements of players and the position of the ball and generates supplementary information, which can provide information such as highlights of plays and tactical analysis.

[1451] This supplementary information is displayed in real time on the user's device, allowing users to understand and enjoy the event more in-depth, including the progress of the match, set lists of the songs being performed, and lyric information.

[1452] Communication Features

[1453] Users can also communicate with other users who are watching the same event. The server provides chat message and voice chat functions, allowing users to interact with each other. This communication function allows users to enjoy a more interactive experience.

[1454] Specific examples

[1455] Example 1: Watching a soccer game

[1456] 1. The server sets up multiple cameras in the stadium, such as from the goalkeeper's perspective, the spectator's perspective, and the player's perspective.

[1457] 2. The user puts on the VR device at home and launches the application.

[1458] 3. The user selects the goalkeeper perspective within the app and receives streaming video from that perspective.

[1459] 4. The AI ​​analysis server analyzes player movements and generates important play and tactical information.

[1460] 5. The analysis results are displayed in real time on the user's device, allowing them to chat with other viewers.

[1461] Example 2: Watching a live music concert

[1462] 1. The server sets up cameras on multiple musicians on stage and in the audience.

[1463] 2. Through the application, users select the perspective of their favorite instrumentalist.

[1464] 3. The server streams the selected viewpoint to the user's device.

[1465] 4. The AI ​​analysis server provides set lists and lyric information for the songs being performed in real time.

[1466] 5. Users can enjoy the live stream while chatting with other fans.

[1467] In this way, the system of the present invention allows users to enjoy a real event experience from the comfort of their own home. Furthermore, AI analysis and community functions provide new added value that differs from traditional viewing and appreciation methods.

[1468] The processing flow will be explained below.

[1469] Step 1:

[1470] The server performs an initialization procedure for multiple cameras, specifically obtaining the ID of each camera and configuring it to establish a streaming connection, allowing real-time video capture from each camera.

[1471] Step 2:

[1472] The camera acquires video in real time from the location where it is installed. Specifically, the camera captures frames (image data) at regular intervals and stores the video data in a buffer.

[1473] Step 3:

[1474] The camera transmits the captured video data to a server. Specifically, the camera uploads frame data to the server in real time via a network using a streaming protocol.

[1475] Step 4:

[1476] The server processes the received video data to relay it to the streaming server. Specifically, it arranges each received frame in order and sends it to the streaming server. The streaming server manages the video data and prepares it to be provided in response to user requests.

[1477] Step 5:

[1478] The user puts on the VR device and starts the application. Specifically, the application displays the user interface and opens a viewpoint selection screen.

[1479] Step 6:

[1480] The user operates the application interface to select the desired viewpoint. Specifically, the user selects, for example, the "goalkeeper's viewpoint" from the list displayed on the viewpoint selection screen.

[1481] Step 7:

[1482] The device receives the user's selection and requests the video data of that viewpoint from the server. Specifically, the device sends the ID of the selected viewpoint to the server and requests transmission of the corresponding stream.

[1483] Step 8:

[1484] The server acquires the video stream from the requested viewpoint and delivers it to the terminal. Specifically, it selects the corresponding camera stream and starts the process of transferring it to the user's terminal.

[1485] Step 9:

[1486] The AI ​​analysis server analyzes the video data received in real time, specifically recognizing important events in the video (such as player movements and ball position) and generating analysis results.

[1487] Step 10:

[1488] The server sends the AI ​​analysis results to the user's device, specifically, sending data including generated supplementary information to the device in real time.

[1489] Step 11:

[1490] The device displays the analysis results to the user, overlaying the analyzed information on the screen so that the user can instantly access important information.

[1491] Step 12:

[1492] Users can communicate with other users while watching an event by using the chat and voice chat functions within the application to communicate via text messages and voice.

[1493] Step 13:

[1494] The server manages and relays chat and voice chat messages between users, specifically, it delivers received messages to the correct destination and establishes communication between users.

[1495] Example 1

[1496] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1497] Modern event experiences are limited by physical constraints, making it difficult for users in distant locations to participate in events in real time. Furthermore, there are few ways to quickly grasp important moments and detailed information about an event, and there are limited ways to share it with other users through communication. This leads to a decline in overall user satisfaction.

[1498] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1499] In this invention, the server includes means for acquiring real-time video from a plurality of image acquisition devices, means for transferring the acquired video to a data processing device, means for a user to select a viewpoint using a virtual reality device and receive video from that viewpoint, means for analyzing the video data using artificial intelligence and generating supplemental information, and means for providing the supplemental information to users. This allows users who are far away to participate in an event in real time, quickly grasp important scenes and detailed information, and share the event experience through communication with other users.

[1500] An "image capture device" is a device that is installed at an event venue and captures images in real time from different viewpoints.

[1501] "Data processing device" means a device that receives video data transmitted from the image acquisition device and processes and encodes it in real time.

[1502] "User" refers to a person who uses the system to view and interact with events in real time.

[1503] A "virtual reality device" is a device such as a headset or goggles that a user wears to experience virtual reality.

[1504] "Artificial intelligence" refers to advanced algorithms and software used to analyze video data and generate supplemental information.

[1505] "Supplementary information" is information generated from the results of analysis by artificial intelligence, and is intended to provide users with additional understanding and interest.

[1506] The "viewpoint selection interface" is a function of the application that allows the user to select and manipulate different camera viewpoints.

[1507] The "communication function" is a function that allows users to communicate with other users in real time via messages and voice.

[1508] A "streaming server" is a server that transfers processed video data to a user's terminal in real time.

[1509] A "user terminal" is a device such as a PC, smartphone, or tablet operated by the user, which functions in conjunction with the virtual reality device.

[1510] The system of this invention is designed to enable users to watch events such as sports and live music in real time from the comfort of their own homes using a virtual reality device (hereinafter referred to as a VR device). Specific means for implementing this system are described below.

[1511] System Configuration

[1512] The system mainly consists of the following components:

[1513] Multiple image capture devices (cameras)

[1514] Data processing device (streaming server and AI analysis server)

[1515] User's device

[1516] VR device

[1517] System Operation

[1518] Camera installation and video acquisition

[1519] The server installs multiple image capture devices at the event venue. These image capture devices are positioned so that they capture video from different viewpoints. For example, in the case of a sporting event, image capture devices are placed around the goal area, from the players' viewpoints, and in the spectator seats. The image capture devices capture video in real time and send the video to a data processing device. This allows the server to simultaneously acquire video data from multiple image capture devices.

[1520] Streaming video

[1521] The server processes the acquired video data in real time and transfers it to the user's device via a streaming server. The user can then view this streaming video through a VR device.

[1522] User Interface

[1523] A dedicated application is installed on the user's device, and this application allows for viewpoint selection and other operations. The user can select the image capture device viewpoint they want to see from the viewpoint selection interface. For example, in the case of a live music performance, they can select the viewpoint of a specific instrument performer.

[1524] AI-based analysis and information provision

[1525] The AI ​​analysis server analyzes the video data being streamed in real time. For example, in the case of a sporting event, it analyzes the movements of the players and the position of the ball and generates supplementary information. This supplementary information is displayed in real time on the user's device, allowing the user to visually obtain information such as the progress of the game, highlights, and tactical analysis.

[1526] Communication Features

[1527] Users can communicate with other users who are watching the same event. The server provides chat message and voice chat functions, supporting users to interact with each other through these. This communication function allows users to enjoy a more interactive experience.

[1528] Specific examples

[1529] Example 1: Watching a soccer game

[1530] 1. The server installs multiple image capture devices at the stadium's goalkeeper's viewpoint, the spectator's viewpoint, and the player's viewpoint.

[1531] 2. The user puts on the VR device at home and launches the application.

[1532] 3. The user selects the goalkeeper's perspective within the app and receives streaming video from that perspective.

[1533] 4. The AI ​​analysis server analyzes player movements and generates important play and tactical information.

[1534] 5. The analysis results are displayed in real time on the user's device, allowing them to chat with other viewers.

[1535] Example 2: Watching a live music concert

[1536] 1. The server installs image capture devices on multiple instrument performers on stage and in the audience seats.

[1537] 2. Through the application, users can select the perspective of their favorite instrumentalist.

[1538] 3. The server streams the video from the selected viewpoint to the user's device.

[1539] 4. The AI ​​analysis server provides set lists and lyric information for the songs being performed in real time.

[1540] 5. Users can enjoy the live performance while chatting with other fans.

[1541] Example prompts for generative AI models

[1542] While watching a soccer match from the goalkeeper's point of view, generate sentences that provide analysis information on player movements and tactics.

[1543] To enjoy live music, generate text that provides perspective footage of specific instrumentalists and set list information.

[1544] In this way, the system of the present invention allows users to enjoy a real event experience from the comfort of their own home. Furthermore, AI analysis and community functions provide new added value that differs from traditional viewing and appreciation methods.

[1545] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1546] System program processing flow

[1547] Step 1: Installing the image capture device and acquiring images

[1548] The server installs multiple image capture devices (cameras) at the event venue. These cameras are positioned to capture images in real time from different viewpoints. Specifically, in a sporting event, cameras are placed around the goal, from the players' viewpoint, and in the spectator seats. The input from the cameras is real-time images, and the output is captured image data.

[1549] Step 2: Transferring video data

[1550] The camera transmits the captured real-time video to a data processing device (streaming server). In this step, the video data input from the camera is transferred to the streaming server via the network. The output is multiple video data stored in the streaming server.

[1551] Step 3: Processing the video data

[1552] The server processes the received video data in real time. Specifically, it compresses and encodes the video data and converts it into a format suitable for streaming. The input is the video data sent from the camera, and the output is the encoded streaming video data.

[1553] Step 4: Streaming the video data

[1554] The server transfers the processed video data to the user's device via a streaming server. The input is encoded video data, and the output is a video stream that can be played back in real time on the user's device.

[1555] Step 5: Launch the application and select a viewpoint

[1556] The user launches a dedicated application on the device. After the user completes login authentication, a viewpoint selection screen is displayed. Here, the user can select the viewpoint they want to view. The input is the viewpoint selected by the user, and the output is video data based on the selected viewpoint.

[1557] Step 6: Receiving and displaying viewpoint images

[1558] The user's device receives the video from the selected viewpoint and displays it on the VR device. The input is video data sent from the streaming server, and the output is real-time video displayed on the user's VR device.

[1559] Step 7: Analyze the video data

[1560] The AI ​​analytics server analyzes the video data received in real time. Specifically, it identifies player movements and ball position and generates important highlights and tactical information. The input is streaming video data, and the output is supplemental information based on the analysis.

[1561] Step 8: Provide supporting information

[1562] The AI ​​analysis server sends the generated supplementary information to the user's device, which displays it in real time. The input is the analyzed supplementary information, and the output is the information displayed on the user's device.

[1563] Step 9: Providing communication features

[1564] Users communicate with other users through chat and voice chat functions within the application. This function is provided by a server. The input is the user's message or voice data, and the output is the communication data transmitted to other users.

[1565] Through this step-by-step process, users can watch events in real time from the comfort of their own homes and enjoy a variety of interactive experiences.

[1566] (Application example 1)

[1567] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1568] Conventional live event viewing systems make it difficult for users to enjoy the event from multiple angles in real time or to interact with other viewers. They also often fail to obtain important scenes or specific information in real time while watching. This limits the user experience and prevents the full appeal of the event.

[1569] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1570] In this invention, the server includes means for acquiring real-time video from multiple cameras, means for transferring the acquired video to a streaming server, means for a user to select a viewpoint using a VR device and receive video from that viewpoint, means for analyzing video data using AI and generating supplemental information, means for providing the supplemental information to the user, means for supporting communication between users, means for interacting in a virtual space using smart glasses or a head-mounted display, means for providing real-time chat and voice calls with other users, means for automatically detecting highlights of an event from video in real time using a generative AI model, and means for providing the generated highlights to the user based on prompt text. This allows users to enjoy an event from multiple viewpoints in an immersive way, and enables interactive interaction with other viewers and the acquisition of important scenes and supplemental information through real-time analysis using AI.

[1571] "Multiple cameras" refers to multiple imaging devices installed at an event venue to capture footage in real time from different perspectives.

[1572] A "streaming server" refers to a server system that processes acquired video data in real time and distributes it to the user's device.

[1573] "User terminal" refers to a computer device that allows a user to receive event footage, select viewpoints, display supplementary information, and communicate.

[1574] "VR device" refers to devices such as head-mounted displays and smart glasses that users wear to experience virtual reality.

[1575] "AI analysis server" refers to a server system that runs artificial intelligence algorithms to analyze acquired video data and generate supplemental information.

[1576] "Supplementary information" refers to additional information such as highlights of the performance, tactical information, and set lists of songs performed, which is generated by AI analysis based on streaming video.

[1577] "Communication function" refers to the function that allows users to interact with other viewers through chat messages and voice calls while watching an event.

[1578] "Smart glasses" refers to a glasses-type device that users can wear to view images in a virtual space.

[1579] A "head-mounted display" refers to a display device worn on the head that allows users to intuitively experience a virtual space.

[1580] "Virtual space" refers to a digitally constructed simulated environment that users experience visually and aurally through VR devices.

[1581] "Real-time chat" refers to a means of communication that allows users to exchange text messages in real time.

[1582] "Voice Call" means a means by which Users can engage in real-time voice communication.

[1583] "Generative AI model" refers to a machine learning algorithm that automatically generates highlights and supplemental information based on input data.

[1584] "Prompt sentence" refers to text input provided to a generative AI model to generate a specific output.

[1585] This invention is a system that allows users to enjoy events such as sports and live music in real time from home using a VR device. The system consists of multiple cameras, a streaming server, a user's device, a VR device, and an AI analysis server.

[1586] System Configuration

[1587] Hardware Configuration

[1588] Multiple cameras: Installed at the event venue, capturing footage in real time from different perspectives.

[1589] Streaming server: A server system that processes acquired video data in real time and distributes it to user devices. Protocols used include WebRTC and RTMP.

[1590] User's device: A computer device used to receive the video, select viewpoints, display supplementary information, and communicate. Smartphones and PCs are often used.

[1591] VR equipment: A device such as a head-mounted display or smart glasses worn by a user to create a virtual reality experience.

[1592] AI analysis server: A server system that runs artificial intelligence algorithms to analyze video data in real time and generate supplementary information. It uses TensorFlow and OpenCV.

[1593] Software Configuration

[1594] Viewpoint switching function: A function that receives a user's viewpoint switching request and sends the corresponding camera video stream to the user's device.

[1595] Real-time AI analysis: The AI ​​analysis server receives the acquired video data and generates important scenes and supplementary information in real time. Analysis is performed using TensorFlow and OpenCV.

[1596] Communication feature: Supports real-time chat and voice calls between users. Firebase and Agora.io SDK are used.

[1597] Generative AI model: Applies machine learning algorithms to generate key scenes and supplementary information from video footage. Generates appropriate output based on prompts.

[1598] Operation explanation

[1599] Viewpoint switching

[1600] When a user selects a viewpoint, the server uses a WebRTC client to acquire the video stream from the selected camera and transmit it to the user's device. For example, in a live music concert, it is possible to select the viewpoint of a specific instrumentalist.

[1601] Real-time AI analysis

[1602] The video data sent to the streaming server is analyzed by an AI analysis server, which uses TensorFlow models to detect important plays during sporting events or lyrics of live music concerts in real time and notify users.

[1603] Communication Features

[1604] The communication feature uses the Firebase real-time database to manage text messages and the Agora.io SDK to provide voice calls, allowing users to interact with other viewers in real time within the VR space.

[1605] Specific examples

[1606] Example 1: Watching a sports game

[1607] 1. The server sets up cameras at multiple viewpoints in the stadium.

[1608] 2. The user puts on the VR device at home and launches the application.

[1609] 3. The user selects the viewpoint of their choice and receives streaming video from that viewpoint.

[1610] 4. The AI ​​analysis server analyzes players' movements and generates important play and tactical information in real time, which is then provided to users.

[1611] Example 2: Watching a live music concert

[1612] 1. The server sets up cameras on multiple musicians on stage and in the audience.

[1613] 2. The user selects the perspective of their favorite instrument performer and receives streaming footage from that perspective.

[1614] 3. The AI ​​analysis server provides set lists and lyric information for the songs being performed in real time.

[1615] Prompt Sentence Examples

[1616] Here is an example prompt for automatic event highlight detection using an AI model:

[1617] Input: "Auto-detect key plays from this footage and create highlights."

[1618] Output: Automatically highlighted video clip

[1619] This invention allows users to enjoy a more immersive multi-perspective experience, interact with other users, and experience the event in greater depth thanks to real-time analysis by AI.

[1620] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1621] Step 1:

[1622] Multiple cameras are installed at the event venue, each capturing video in real time from a different viewpoint. The server acquires the video data from the cameras and transfers it to the streaming server.

[1623] Input: Live video from the event venue

[1624] Data processing: Video capture and transfer to streaming server

[1625] Output: Raw video data stored on a streaming server

[1626] Step 2:

[1627] The streaming server receives the acquired video data and delivers the video data to the user's terminal in real time.

[1628] Input: Raw video data stored on a streaming server

[1629] Data processing: video data compression and transfer

[1630] Output: Real-time video sent to user device

[1631] Step 3:

[1632] The user selects a viewpoint through the terminal, which then sends a viewpoint switching request to the server, which then delivers the video stream of the selected viewpoint to the user terminal.

[1633] Input: User's request to switch viewpoint

[1634] Data processing: Selecting the video stream corresponding to the viewpoint

[1635] Output: Video of the selected viewpoint sent to the user's device

[1636] Step 4:

[1637] The video data sent to the streaming server is analyzed in real time by an AI analysis server. A generative AI model is used for the analysis to extract important scenes and supplementary information. For example, important plays during a sporting event or lyrics from a live music concert can be automatically extracted.

[1638] Input: Raw video data sent from the streaming server

[1639] Data Computing: Real-time analytics with TensorFlow and OpenCV

[1640] Output: Analyzed important scenes and supplementary information

[1641] Step 5:

[1642] The user's device receives supplemental information sent from the AI ​​analysis server and displays it in real time, such as highlights of important performances and set lists of songs performed.

[1643] Input: Supplementary information sent from the AI ​​analysis server

[1644] Data processing: Converting supplementary information into a display format

[1645] Output: Supplementary information displayed on the user's terminal

[1646] Step 6:

[1647] Users can chat and talk to other users watching the same event in real time through their devices, and the server uses the Firebase real-time database and Agora.io SDK to support this functionality.

[1648] Input: User text messages and voice data

[1649] Data processing: real-time transmission of text messages, compression and transmission of voice data

[1650] Output: Messages and audio delivered to other users

[1651] Step 7:

[1652] When users use the generative AI model as part of real-time AI analysis, they can input prompts to automatically detect event highlights and specific scenes from video.

[1653] Input: Prompt text (e.g. "Automatically detect important plays from this footage and create highlights")

[1654] Data calculation: Execute a generative AI model based on the input prompt.

[1655] Output: Automatically generated highlight reels and clips of specific scenes

[1656] This allows users to enjoy an immersive multi-perspective experience, real-time interactive interactions, and detailed supplemental information about the event through AI analysis.

[1657] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1658] This invention provides a new user experience by combining a system that allows users to experience events such as sports and live music concerts in real time using a VR device from the comfort of their own home with an emotion engine that recognizes the user's emotions. Specific means for implementing this system are described below.

[1659] System Configuration

[1660] The system mainly consists of the following components:

[1661] Multiple Cameras

[1662] Streaming Server

[1663] User's device

[1664] VR device

[1665] AI analysis server

[1666] Emotion Engine

[1667] System Operation

[1668] Camera installation and video acquisition

[1669] The server installs multiple cameras at the event venue and captures video in real time from each camera. This video data is then transferred to the streaming server.

[1670] Streaming video

[1671] The server relays the acquired video to the streaming server in real time, allowing the user's device to receive the video data directly from the streaming server.

[1672] User Interface

[1673] The user's device performs viewpoint selection and other operations via the application. The user can select each camera viewpoint from the viewpoint selection screen displayed within the application.

[1674] AI-based analysis and information provision

[1675] The AI ​​analysis server analyzes the video data being streamed in real time, and based on the results of this analysis, users can receive supplemental information such as highlights of plays and tactical analysis.

[1676] Emotion recognition by emotion engine

[1677] The emotion engine recognizes the user's emotions in real time by analyzing their facial expressions and voice data. Based on the analysis results, the emotion engine generates supplementary information and effects according to the user's emotions. For example, if the user shows a surprised expression, it can respond by displaying a replay of the play at that moment.

[1678] Emotional Data Feedback

[1679] The emotional data recognized by the emotion engine is fed back to the AI ​​analysis server. Based on this data, the system adjusts the user experience across the entire system. For example, if data is obtained showing multiple users expressing excitement, the system can respond by providing more detailed explanatory information.

[1680] Communication Features

[1681] Users can chat and voice chat with other users while watching the event, and the system manages and relays these communications, facilitating interaction between users.

[1682] Specific examples

[1683] Example 1: Watching a soccer game

[1684] 1. The server sets up cameras at multiple viewpoints in the stadium, such as the goalkeeper's viewpoint, the spectator's viewpoint, and the player's viewpoint.

[1685] 2. The user puts on the VR device at home and launches the application.

[1686] 3. The user selects "Goalkeeper View" on the viewpoint selection screen and receives video from that viewpoint.

[1687] 4. The AI ​​analysis server analyzes player movements and ball trajectory to generate important play and tactical information.

[1688] 5. The emotion engine analyzes the user's facial expressions and voice, and when it recognizes moments of surprise or joy, it instantly displays replay footage and detailed commentary.

[1689] 6. Users can enjoy communicating with other viewers using the chat function.

[1690] Example 2: Watching a live music concert

[1691] 1. The server sets up cameras on the stage for multiple musicians and in the audience.

[1692] 2. Through the application, the user selects the perspective of, for example, a drummer.

[1693] 3. The server streams the selected viewpoint to the user's device.

[1694] 4. The AI ​​analysis server provides performance set lists and lyric information in real time.

[1695] 5. The emotion engine recognizes the user's emotions and responds by adding special effects when the user is emotional.

[1696] 6. Users can enjoy the live show while chatting with other fans.

[1697] This system allows users to enjoy live sports and music events from the comfort of their own home through an advanced VR experience, and the introduction of an emotion engine provides a more personalized experience.

[1698] The processing flow will be explained below.

[1699] Step 1:

[1700] The server will set up multiple cameras at the event venue, and will configure the initial settings of these cameras so that they can capture images from each viewpoint in real time.

[1701] Step 2:

[1702] The camera captures video in real time from its installed location, acquiring frames at regular intervals and generating video data.

[1703] Step 3:

[1704] The camera transmits the captured video data to a server. Specifically, the video data is transferred to a streaming server via a network in real time.

[1705] Step 4:

[1706] The server receives video data from multiple cameras and manages it in a streaming server, which prepares to provide the corresponding video stream in response to a viewpoint request from a user.

[1707] Step 5:

[1708] The user puts on the VR device and launches the application, which displays the main screen and offers viewpoint selection options.

[1709] Step 6:

[1710] The user selects the desired viewpoint on the viewpoint selection screen within the application. For example, the user selects the "goalkeeper viewpoint."

[1711] Step 7:

[1712] Based on the user's viewpoint selection, the terminal requests the video data of the viewpoint from the streaming server. Specifically, the terminal transmits the selected viewpoint ID to the streaming server and requests the corresponding stream.

[1713] Step 8:

[1714] The streaming server acquires the video stream of the requested viewpoint and transmits it to the terminal in real time, allowing the user to view the video from the selected viewpoint.

[1715] Step 9:

[1716] The AI ​​analysis server analyzes the streaming video data in real time to detect player movements and important events (e.g., goals, shots).

[1717] Step 10:

[1718] The AI ​​analysis server generates supplemental information based on the analysis results and sends it to the device, allowing users to receive the supplemental information in real time.

[1719] Step 11:

[1720] The device displays the AI ​​analysis results to the user, overlaying supplemental information on the screen in real time.

[1721] Step 12:

[1722] The emotion engine analyzes the user's facial expressions and voice data to recognize their emotions. Specifically, it captures the user's reactions through a camera and microphone and generates emotion data.

[1723] Step 13:

[1724] The emotion engine feeds the recognized emotion data back to the AI ​​analysis server, which then adjusts the overall system experience based on this data, for example, providing special effects or additional information to excited users.

[1725] Step 14:

[1726] Users can communicate with other users while watching an event, using chat and voice chat functions within the application to interact with other viewers.

[1727] Step 15:

[1728] The server manages and relays chat and voice chat messages between users, thereby supporting real-time communication between users.

[1729] This detailed processing flow allows users to enjoy events in real time from home and enjoy a personalized experience through the emotion engine.

[1730] Example 2

[1731] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1732] In conventional virtual reality event viewing systems, users are limited to selecting a viewpoint and viewing experience, and real-time emotion analysis and personalized experiences based on that analysis are not performed. Furthermore, communication functions between users are insufficient, making it insufficient to provide a richer user experience.

[1733] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1734] In this invention, the server includes means for acquiring real-time video from multiple image capture devices, means for transferring the acquired video to a data distribution server, means for allowing a user to select a viewpoint using a virtual reality device and receiving video from that viewpoint, means for analyzing video data using artificial intelligence and generating supplemental information, means for providing the supplemental information to the user, means for supporting communication between users, means for analyzing user emotions and generating supplemental information and effects based on the emotions, and means for feeding back the emotion data to the analysis server and adjusting the overall user experience. This enables the provision of personalized information and the addition of effects according to the user's emotions, further enhancing communication between users.

[1735] An "imaging device" is a device for capturing images in real time.

[1736] The "data distribution server" is a server that distributes acquired video data to user terminals in real time.

[1737] A "virtual reality device" is a device that allows users to immerse themselves in a virtual environment and select and control their viewpoint.

[1738] "Video data" refers to real-time video captured by an imaging device.

[1739] "Artificial intelligence" is a technology that analyzes video data and user emotional data to generate supplementary information and effects.

[1740] "Supplemental Information" is additional information provided to enhance the viewing or viewing experience.

[1741] "Effects" are visual and sound effects added to the video.

[1742] "User emotion" refers to the emotional state analyzed from the user's facial expressions and voice.

[1743] "Feedback" is the process of adjusting the overall system experience based on acquired data.

[1744] "Communication means" refers to a function that supports real-time communication between users.

[1745] This invention provides a new user experience by combining a system that allows users to experience events such as sports and live music concerts in real time using a VR device from the comfort of their own home with an emotion engine that recognizes the user's emotions. Specific means for implementing this system are described below.

[1746] System Configuration

[1747] The system mainly consists of the following components:

[1748] Multiple imaging devices (cameras)

[1749] Data distribution server

[1750] User's device

[1751] Virtual reality device (VR device)

[1752] AI analysis server

[1753] Emotion Engine

[1754] System Operation

[1755] Camera installation and video acquisition

[1756] The server installs multiple camera devices at the event venue, captures video in real time from each camera device, and transfers the video data to a data distribution server.

[1757] Streaming video

[1758] The server relays the acquired video to the data distribution server in real time, allowing the user's device to receive the video data directly from the data distribution server.

[1759] User Interface

[1760] The user's device performs viewpoint selection and other operations via the application. The user can select each camera viewpoint from the viewpoint selection screen displayed within the application.

[1761] AI-based analysis and information provision

[1762] The AI ​​analysis server analyzes the video data being streamed in real time, and based on the results of this analysis, users can receive supplemental information such as highlights of plays and tactical analysis.

[1763] Emotion recognition by emotion engine

[1764] The emotion engine recognizes the user's emotions in real time by analyzing their facial expressions and voice data. Based on the analysis results, the emotion engine generates supplementary information and effects according to the user's emotions. For example, if the user shows a surprised expression, it can respond by displaying a replay of the play at that moment.

[1765] Emotional Data Feedback

[1766] The emotional data recognized by the emotion engine is fed back to the AI ​​analysis server. Based on this data, the system adjusts the user experience across the entire system. For example, if data is obtained showing multiple users expressing excitement, the system can respond by providing more detailed explanatory information.

[1767] Communication Features

[1768] Users can chat and voice chat with other users while watching the event, and the system manages and relays these communications, facilitating interaction between users.

[1769] Specific examples

[1770] Example 1: Watching a soccer game

[1771] 1. The server sets up cameras at multiple viewpoints in the stadium, such as the goalkeeper's viewpoint, the spectator's viewpoint, and the player's viewpoint.

[1772] 2. The user puts on the VR device at home and launches the application.

[1773] 3. The user selects "Goalkeeper View" on the viewpoint selection screen and receives video from that viewpoint.

[1774] 4. The AI ​​analysis server analyzes player movements and ball trajectory to generate important play and tactical information.

[1775] 5. The emotion engine analyzes the user's facial expressions and voice, and when it recognizes moments of surprise or joy, it instantly displays replay footage and detailed commentary.

[1776] 6. Users can enjoy communicating with other viewers using the chat function.

[1777] Example 2: Watching a live music concert

[1778] 1. The server sets up cameras on the stage for multiple musicians and in the audience.

[1779] 2. Through the application, the user selects the perspective of, for example, a drummer.

[1780] 3. The server streams the selected viewpoint to the user's device.

[1781] 4. The AI ​​analysis server provides performance set lists and lyric information in real time.

[1782] 5. The emotion engine recognizes the user's emotions and responds by adding special effects when the user is emotional.

[1783] 6. Users can enjoy the live show while chatting with other fans.

[1784] Prompt Sentence Examples

[1785] Here are some examples of specific prompts for this system:

[1786] 1. Emotion Recognition in Sporting Events: Describe a process for displaying a replay and detailed commentary the moment the user expresses surprise.

[1787] 2. Personalizing Live Music Experience: What is your process for adding effects in real time based on the user's emotions?

[1788] This system allows users to enjoy live sports and music events from the comfort of their own home through an advanced VR experience, and the introduction of an emotion engine provides a more personalized experience.

[1789] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1790] Step 1: Camera installation and video acquisition

[1791] The server will set up multiple camera devices at the event venue. For example, at a sporting event, cameras will be installed to view the goalkeeper, the spectators, and the players, while at a music concert, cameras will be installed to view each musician on stage and in the spectators' seats. The server will acquire images from these cameras in real time.

[1792] Input: Video data from multiple cameras installed at the event venue.

[1793] Data processing: Converts the analog video signal from the camera into a digital signal and formats it for transmission to the streaming server.

[1794] Output: The formatted digital video data is transferred to a streaming server.

[1795] Step 2: Streaming the video

[1796] The server relays the captured video in real time to the data distribution server, which then distributes the video to the user's device.

[1797] Input: Digital video data transferred from the server.

[1798] Data processing: Packetization and streaming optimization for relaying and real-time delivery of digital video data.

[1799] Output: Streaming video data is sent to the user's device.

[1800] Step 3: User Interface

[1801] The user wears the VR device at home and starts the application. The user selects the desired camera viewpoint from the viewpoint selection screen within the application.

[1802] Input: Viewpoint selection information made by the user to the application.

[1803] Data processing: Based on the user's viewpoint selection, the video data of the corresponding camera viewpoint is extracted.

[1804] Output: Real-time video from the selected viewpoint is displayed in the user's VR device.

[1805] Step 4: AI-based video analysis and information provision

[1806] The AI ​​analysis server analyzes the video data being streamed in real time, for example, analyzing the movements of players and the trajectory of the ball, and then generates highlights of the play and tactical information based on that analysis.

[1807] Input: Real-time video data.

[1808] Data Processing: Analyze video data using generative AI models to extract and generate key events and tactical information.

[1809] Output: Analyzed highlights and tactical information are provided to the user.

[1810] Step 5: Emotion Recognition with the Emotion Engine

[1811] The emotion engine analyzes the user's facial expressions and voice data in real time to recognize their emotions. For example, if the user shows a surprised expression, it will display a replay of the play at that moment.

[1812] Input: User's facial and voice data.

[1813] Data processing: Using an AI model, the system analyzes the user's emotions and generates corresponding effects and supplementary information.

[1814] Output: Real-time effects and replay footage are displayed according to the user's emotions.

[1815] Step 6: Feedback of emotional data

[1816] The emotional data recognized by the emotion engine is fed back to the AI ​​analysis server. The system adjusts the user experience based on this data. For example, if multiple users show signs of excitement, more detailed commentary will be provided.

[1817] Input: Recognized user emotion data.

[1818] Data processing: Emotional data is sent to the analysis server and reflected in the user experience of the entire system.

[1819] Output: A tailored user experience based on sentiment data.

[1820] Step 7: Communication Functions

[1821] Users can chat and voice chat with other users while watching the event, and the system manages and relays these communications, facilitating interaction between users.

[1822] Input: Chat messages and voice data between users.

[1823] Data processing: Relaying and logging chat messages and voice data.

[1824] Output: A real-time communication experience.

[1825] (Application example 2)

[1826] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1827] In conventional online shopping systems, users are limited to visual information about products, which is far from the shopping experience of a physical store. Furthermore, when users express interest in a product or have a particular emotion, they are unable to provide dynamic information that responds to that interest. Furthermore, real-time communication between users is not adequately supported.

[1828] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for acquiring real-time video from multiple image capture devices, means for transferring the acquired video to a streaming server, means for a user to select a viewpoint using a virtual reality device and receive video from that viewpoint, means for analyzing video data using artificial intelligence and generating supplemental information, means for providing the supplemental information to the user, means for supporting communication between users, means for recognizing user emotions using an emotion engine and generating additional information and effects based on the recognition results, and means for feeding back emotion data to the artificial intelligence analysis server to adjust the user experience. This allows users to enjoy a realistic shopping experience from the comfort of their own home, as if they were in a physical store, and makes it possible to provide dynamic information about interesting products and add emotion-based effects.

[1829] A "camera" is a device that captures images of the real world and records them as digital data.

[1830] A "streaming server" is a server that distributes acquired video data in real time.

[1831] A "virtual reality device" is a device that allows a user to enjoy an experience in a 3D space, and generally includes a head-mounted display and smart glasses.

[1832] "Artificial intelligence" is a technology that analyzes large amounts of data and finds patterns and trends to make decisions and make predictions.

[1833] "Supplementary information" refers to additional information, commentary, promotions, etc. related to the video the user is watching.

[1834] "Communication support means" is a function that supports chat and voice chat between users, enabling smooth information exchange.

[1835] The "emotion engine" is a system that analyzes data such as the user's facial expressions and voice, and recognizes emotions in real time.

[1836] "Emotion data" is data that indicates the user's emotional state as analyzed by the emotion engine.

[1837] The "artificial intelligence analysis server" is a server that analyzes emotional data and video data and makes decisions such as providing information and generating effects.

[1838] The present invention is a system that provides a shopping experience from home that makes you feel as if you are actually in a store, and to achieve this, it combines multiple technologies such as a camera device, a streaming server, a virtual reality device, artificial intelligence, and an emotion engine. Specific embodiments of the system are described below.

[1839] System Configuration

[1840] Camera: Installed in multiple areas of the store, it captures high-resolution video in real time. Specific hardware used includes high-resolution cameras (e.g., Sony α7 series).

[1841] Streaming server: Used to distribute acquired video data to user devices in real time. The specific software used is AWS Elemental MediaLive.

[1842] Virtual reality devices: Devices that allow users to immerse themselves in virtual reality spaces. Specific devices include head-mounted displays (Oculus Rift) and smart glasses (Google Glass).

[1843] Artificial intelligence analysis server: Used to analyze video data and generate supplementary information and effects. Specific software includes Google Cloud Vision API and TensorFlow.

[1844] Emotion engine: Analyzes the user's facial expressions and voice to recognize emotions in real time. Specific software includes the Affectiva SDK.

[1845] System Operation

[1846] Video Acquisition and Streaming

[1847] The server captures real-time video from high-resolution cameras installed in the physical store and streams it to user devices using AWS Elemental MediaLive. Users can then use virtual reality devices to view the store's interior from a specified viewpoint in real time.

[1848] Supplementary Information and Emotion Recognition

[1849] The AI ​​analysis server uses Google Cloud Vision API and TensorFlow to analyze the captured video data and generate supplementary information useful to the user. Meanwhile, the emotion engine analyzes the user's facial expressions and voice to collect emotional data. For example, if the user smiles, a special effect corresponding to that emotion can be added.

[1850] Emotional Data Feedback

[1851] The emotion data collected by the emotion engine is fed back to the AI ​​analysis server, which can then tailor the user experience to be more personalized. For example, if a user shows a strong interest in a particular product, the server can display additional promotional information related to that product.

[1852] User Interactions

[1853] Users can select their viewpoint within the virtual reality space and enjoy shopping while viewing the store in real time. They can also use the chat function to communicate with other users and store staff in real time, allowing them to have an experience equivalent to shopping in a physical store, even from the comfort of their own home.

[1854] Specific examples

[1855] Example 1

[1856] If a user makes an inquisitive facial expression while browsing the fragrance section of a physical store, the emotion engine will recognize this and pop up a special promotion for related products.

[1857] Example 2

[1858] When a user is viewing the clothing section, the AI ​​analysis server recognizes tags within the products and provides detailed information and stock information for the relevant products in real time.

[1859] Prompt Sentence Examples

[1860] "View product information for the fragrance section"

[1861] "Please let me know about promotions for this product."

[1862] Add recommended products to your cart

[1863] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1864] Step 1:

[1865] The server acquires real-time video from multiple high-resolution camera devices installed in the physical store. Specifically, it collects video feeds from each camera and manages them centrally. The input is video data from the camera devices, and the output is video data for transfer to the streaming server.

[1866] Step 2:

[1867] The server transfers the acquired video data to the streaming server. AWS Elemental MediaLive is used to transfer the video data in real time. The input is the video data acquired in step 1, and the output is a streaming data stream. This allows the user's device to receive the video in real time.

[1868] Step 3:

[1869] The user's device uses a virtual reality device to view the video obtained from the streaming server. The user selects a viewpoint from the application and receives video from a specific camera based on that selection. The input is the user's viewpoint selection and streaming data, and the output is the video provided to the user.

[1870] Step 4:

[1871] The AI ​​analysis server uses Google Cloud Vision API and TensorFlow to analyze video data and generate supplemental information. Specifically, it performs object recognition and motion analysis within the video to generate related product information and promotions. The input is streaming data, and the output is supplemental information.

[1872] Step 5:

[1873] The user's device displays the generated supplementary information in real time. Specifically, the information is overlaid on the video being viewed. The input is supplementary information, and the output is a video with additional information provided to the user.

[1874] Step 6:

[1875] The emotion engine analyzes the user's facial expressions and voice to recognize emotions in real time. It uses the Affectiva SDK to analyze data acquired from the user's camera and microphone. The input is the user's facial expressions and voice data, and the output is emotional data.

[1876] Step 7:

[1877] The emotional data obtained by the emotion engine is fed back to the AI ​​analysis server, which uses this data to generate additional information and effects to personalize the user experience. The input is emotional data, and the output is personalized supplementary information and effects.

[1878] Step 8:

[1879] The user's device displays the generated personalized information, allowing the user to enjoy a personalized shopping experience. The input is personalized information, and the output is a customized video provided to the user.

[1880] Step 9:

[1881] The communication support means allows users to chat and voice chat with other users and store staff in real time. Specifically, it sends and receives text messages and voice data. The input is the user's message or voice data, and the output is the delivery of the message or voice to the other party.

[1882] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1883] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1884] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1885] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1886] FIG. 9 illustrates an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and behaviors arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1887] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1888] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1889] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1890] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1891] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1892] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1893] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1894] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1895] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1896] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1897] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1898] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1899] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1900] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1901] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1902] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1903] The following is further disclosed regarding the above embodiment.

[1904] (Claim 1)

[1905] a means for acquiring real-time video from multiple cameras;

[1906] A means for transferring the acquired video to a streaming server;

[1907] A means for a user to select a viewpoint using a VR device and receive an image from that viewpoint;

[1908] A means of analyzing video data using AI and generating supplemental information;

[1909] a means of providing supplemental information to the user;

[1910] A means to support communication between users;

[1911] A system including:

[1912] (Claim 2)

[1913] 2. The system according to claim 1, wherein the plurality of cameras are installed at different viewpoints, and the user can select different viewpoints.

[1914] (Claim 3)

[1915] The system of claim 1, wherein the AI ​​analyzes video data based on the viewpoint selected by the user and displays information in real time.

[1916] (Claim 4)

[1917] 10. The system of claim 1, further comprising means for enabling a user to chat or voice chat with other users while wearing the VR device.

[1918] "Example 1"

[1919] (Claim 1)

[1920] means for acquiring real-time video from a plurality of image acquisition devices;

[1921] means for transferring the acquired image to a data processing device;

[1922] a means for allowing a user to select a viewpoint using a virtual reality device and receive an image from that viewpoint;

[1923] means for analyzing the video data using artificial intelligence to generate supplemental information;

[1924] a means for providing supplemental information to the user;

[1925] A means of supporting communication between users;

[1926] A means for processing the acquired multiple pieces of video data in real time;

[1927] means for transferring the processed video data to a user terminal in real time;

[1928] software means for performing viewpoint selection and other operations at a user terminal;

[1929] means for providing an interface that enables viewpoint selection operation;

[1930] A means for generating analysis results in real time and providing them to the terminal as supplemental information;

[1931] a means for displaying supplemental information in real time;

[1932] ...

[1933] A system including:

[1934] (Claim 2)

[1935] 2. The system according to claim 1, wherein a plurality of image capture devices are installed at different viewpoints, and a user can select a different viewpoint.

[1936] (Claim 3)

[1937] The system of claim 1, wherein the artificial intelligence analyzes the video data based on the viewpoint selected by the user and displays the information in real time.

[1938] "Application Example 1"

[1939] (Claim 1)

[1940] a means for acquiring real-time video from multiple cameras;

[1941] A means for transferring the acquired video to a streaming server;

[1942] A means for a user to select a viewpoint using a VR device and receive an image from that viewpoint;

[1943] A means of analyzing video data using AI and generating supplemental information;

[1944] a means of providing supplemental information to the user;

[1945] A means to support communication between users;

[1946] A means for interacting in a virtual space using smart glasses or a head-mounted display;

[1947] means for providing real-time chat and voice calls with other users;

[1948] A system including:

[1949] (Claim 2)

[1950] 2. The system according to claim 1, wherein the plurality of cameras are installed at different viewpoints, and the user can select different viewpoints.

[1951] (Claim 3)

[1952] The system of claim 1, wherein the AI ​​analyzes video data based on the viewpoint selected by the user and displays information in real time.

[1953] (Claim 4)

[1954] A means to automatically detect event highlights from video using a generative AI model in real time;

[1955] a means for providing the generated highlights to the user based on a prompt;

[1956] 10. The system of claim 1, comprising:

[1957] "Example 2: Combining Emotion Engines"

[1958] (Claim 1)

[1959] means for acquiring real-time images from a plurality of image capture devices;

[1960] means for transferring the acquired video to a data distribution server;

[1961] a means for a user to select a viewpoint using the virtual reality device and receive an image from that viewpoint;

[1962] means for analyzing the video data using artificial intelligence to generate supplemental information;

[1963] a means of providing supplemental information to the user;

[1964] A means to support communication between users;

[1965] A means for analyzing user emotions and generating supplementary information and effects based on the emotions;

[1966] A means of feeding back emotional data to an analysis server to adjust the overall user experience;

[1967] A system including:

[1968] (Claim 2)

[1969] 2. The system according to claim 1, wherein a plurality of image capturing devices are installed at different viewpoints, and a user can select a different viewpoint.

[1970] (Claim 3)

[1971] The system of claim 1, wherein the artificial intelligence analyzes the video data based on the viewpoint selected by the user and displays the information in real time.

[1972] "Application example 2 when combining emotion engines"

[1973] (Claim 1)

[1974] means for acquiring real-time images from a plurality of image capture devices;

[1975] A means for transferring the acquired video to a streaming server;

[1976] a means for a user to select a viewpoint using the virtual reality device and receive an image from that viewpoint;

[1977] means for analyzing the video data using artificial intelligence to generate supplemental information;

[1978] a means of providing supplemental information to the user;

[1979] A means to support communication between users;

[1980] A means for recognizing a user's emotion using an emotion engine and generating additional information or effects based on the recognition result;

[1981] A means of feeding back emotional data to an AI analysis server to adjust the user experience;

[1982] A system including:

[1983] (Claim 2)

[1984] The system of claim 1, wherein multiple image capturing devices are installed at different viewpoints, allowing the user to select different viewpoints, and further allowing the user to select specific areas of the physical store using a virtual reality device.

[1985] (Claim 3)

[1986] The system of claim 1, wherein the artificial intelligence analyzes the video data based on the user's selected viewpoint and the user's emotions, displays information in real time, and further adds effects according to the emotional data. [Explanation of symbols]

[1987] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. a means for acquiring real-time video from multiple cameras; A means for transferring the acquired video to a streaming server; A means for a user to select a viewpoint using a VR device and receive an image from that viewpoint; A means of analyzing video data using AI and generating supplemental information; a means of providing supplemental information to the user; A means to support communication between users; A system including:

2. 2. The system according to claim 1, wherein a plurality of cameras are installed at different viewpoints, and a user can select a different viewpoint.

3. The system of claim 1, wherein the AI ​​analyzes video data based on a viewpoint selected by the user and displays information in real time.

4. 2. The system according to claim 1, further comprising means for enabling a user to chat or voice chat with other users while wearing the VR device.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A