system

The system enhances game viewing realism by integrating 360-degree panoramic video data from multiple cameras and adjusting effects based on user movements, offering an immersive experience.

JP2026035347APending Publication Date: 2026-03-04SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024138190
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-19
Publication Date
2026-03-04

AI Technical Summary

Technical Problem

Conventional game viewing methods lack realism, particularly for remote viewers, as they fail to provide a sense of speed and size of players' movements in real time, limiting the immersive experience.

Method used

A system that collects video data from multiple cameras within a stadium, integrates it to create 360-degree panoramic data, transmits it in real time, and adjusts video and sound effects based on user movements using a built-in gyro sensor, providing a realistic experience through VR technology.

Benefits of technology

Enables users to experience a game as if they were at the venue, with seamless integration of video data and real-time adjustments for an immersive atmosphere.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026035347000001_ABST
    Figure 2026035347000001_ABST
Patent Text Reader

Abstract

Provide a system. The present invention comprises: a means for collecting video data from a plurality of image capturing devices within a game venue; A means for integrating the collected video data and generating 360-degree video data; A means for transmitting the generated 360-degree omnidirectional video data to a terminal in real time; a means for receiving the video data from the terminal to provide a 360-degree field of view and adjust video and sound effects according to the user's movements; A system including:
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] Conventional methods of watching games have a limited sense of realism through video and audio, making it difficult to provide an experience similar to being at the actual stadium. For fans who cannot attend the stadium, it is particularly difficult to feel the speed and size of players' movements in real time. Therefore, there is a demand for a more realistic and immersive game experience. [Means for solving the problem]

[0005] The present invention solves the above-mentioned problems with a system that includes: means for collecting video data from multiple camera devices within a stadium; means for integrating the collected video data to generate 360-degree panoramic video data; means for transmitting the generated 360-degree panoramic video data to a device in real time; and means for the device to receive the video data, provide a 360-degree panoramic view, and adjust the video and sound effects according to the user's movements. This allows users to experience the realism of the stadium from the comfort of their own homes. The system also includes means for automatically recognizing overlapping portions of the video data from the camera devices and seamlessly combining them. Furthermore, the device includes means for detecting the user's head movement using a built-in gyro sensor and adjusting the video and sound effects in real time based on that movement, thereby providing a more realistic sense of realism.

[0006] "Filming equipment" refers to equipment installed within the venue for capturing video data.

[0007] "Video data" refers to digital video information acquired from a photographing device.

[0008] "Merge" refers to the process of combining image data collected from multiple imaging devices and editing it into one continuous image data.

[0009] "360-degree omnidirectional video data" refers to video data that captures the entire surrounding environment of the match venue, and this data is provided so that users can view in all directions.

[0010] "Terminal" refers to an electronic device worn by a user to receive and display video data, specifically equipment such as VR goggles.

[0011] "Real-time" refers to the transmission and processing of video and audio with almost no delay.

[0012] "Field of view" refers to the range that a user can see from the video data displayed on the terminal.

[0013] "Sound effects" are audio data that are played in sync with the video and are used to enhance the sense of realism.

[0014] "Built-in gyro sensor" refers to a sensor built into the device that detects the user's head movements.

[0015] "User movement" refers to the head and body movements made by the user while wearing the VR goggles.

[0016] The "overlapped portion" refers to the overlapping portion of video data captured by multiple image capture devices.

[0017] "Seamless combining" refers to naturally combining overlapping portions of video data to generate a consistent video. [Brief explanation of the drawings]

[0018] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0019] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0020] First, the terms used in the following description will be explained.

[0021] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0022] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0023] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0024] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0025] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0026] [First embodiment]

[0027] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0028] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0029] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0030] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0031] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0032] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0033] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0034] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0035] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0036] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0037] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0038] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0039] The present invention relates to a system that uses VR technology to enhance the sense of realism when watching a game. Hereinafter, an embodiment of the present invention will be described in detail.

[0040] Server Processing

[0041] The server collects video data from multiple camera systems within the venue, including fixed cameras, mobile cameras, and even drone cameras. The collected video data is buffered and the footage from each camera is sorted chronologically. A process then automatically recognizes overlapping areas of the footage and seamlessly combines them to generate 360-degree panoramic video data.

[0042] The generated 360-degree video data is encoded in real time, divided into packets of a certain size, and transmitted to the device with low latency, allowing users to experience the video smoothly.

[0043] Terminal processing (VR goggle processing)

[0044] The device (VR goggles) receives real-time video data sent from the server. The device reassembles the received data packets and stores them in a temporary buffer as video data. Next, the device decodes the video data and renders the video in real time according to the user's viewpoint.

[0045] The device uses a built-in gyro sensor to capture the user's head movement data and adjusts the visual and audio effects in real time based on that movement, providing users with a visual and audio experience that makes them feel like they're at the stadium.

[0046] User operations

[0047] First, users put on VR goggles and log in to the system. Users can freely change their viewpoint by moving their head, and enjoy a 360-degree panoramic view of their favorite scenes in real time. For example, they can stand in the perspective of their favorite player and experience the moment they make a shot up close.

[0048] Specific examples

[0049] For example, if you were streaming a basketball game, it would look like this:

[0050] Server processing: Images are collected from multiple cameras installed within the venue, and overlapping areas of the images are automatically recognized and seamlessly combined. This combined video data is encoded in real time and sent to the terminal.

[0051] Device processing: The device receives the video data sent from the server and stores it in a temporary buffer. It then decodes it and renders the video in real time according to the user's head movements. It uses the built-in gyro sensor to adjust the user's viewpoint and adjusts the sound effects in real time.

[0052] User operation: The user puts on the VR goggles and logs in to the system. The user can change the viewpoint by moving their head and enjoy 360-degree omnidirectional images. For example, they can switch to a courtside viewpoint and experience the players' movements and plays in real time.

[0053] This allows users to experience the immersive atmosphere of a live game venue from the comfort of their own home.

[0054] The processing flow will be explained below.

[0055] Server Processing

[0056] Step 1: Collect video data

[0057] The server receives video data from multiple camera devices installed within the venue, including fixed, mobile and drone cameras, each capturing the match from a different perspective or angle.

[0058] Step 2: Buffering video data

[0059] The server temporarily stores the received video data in a buffer, whereby the video data from each camera is temporarily stored.

[0060] Step 3: Synchronize and sort data

[0061] The server rearranges the video data in the temporary buffer in chronological order and synchronizes them based on the timestamps.

[0062] Step 4: Video data integration

[0063] The server then combines the sorted video data from multiple cameras, automatically recognizing overlapping areas and seamlessly joining them together.

[0064] Step 5: Encode the video data

[0065] The server encodes the integrated 360-degree panoramic video data in real time.

[0066] Step 6: Packetize the data

[0067] The server divides the encoded video data into packets of a fixed size and prepares them for transfer.

[0068] Step 7: Sending data

[0069] The server then transmits the divided data packets to the terminal via a high-speed network, aiming for low latency.

[0070] Terminal (VR goggles) processing

[0071] Step 1: Receiving a data packet

[0072] The terminal receives the data packets sent from the server through a high-speed network.

[0073] Step 2: Reassembling the data packets

[0074] The terminal reassembles the received data packets into the original video data.

[0075] Step 3: Decoding the video data

[0076] The terminal decodes the reassembled video data and prepares it for display.

[0077] Step 4: Rendering the Perspective

[0078] Based on the decoded video data, the device renders a 360-degree panoramic image tailored to the user's viewpoint in real time.

[0079] Step 5: Acquire motion data

[0080] The device uses a built-in gyro sensor to detect the user's head movements.

[0081] Step 6: Adjusting the image and sound

[0082] The device adjusts the video and audio effects in real time based on the acquired motion data, providing a sense of realism in both visual and audio.

[0083] User operations

[0084] Step 1: Put on the VR goggles and log in

[0085] The user puts on the VR goggles and logs into the system. To log in, the user must enter their account information.

[0086] Step 2: Change your perspective

[0087] Users can freely change their viewpoint by moving their head, enjoying a 360-degree view, and can also focus on a specific player or play.

[0088] Step 3: Interactive Experience

[0089] Users receive real-time visual and auditory feedback, making them feel as if they are at the stadium, with sounds such as the cheers of the crowd and the footsteps of the players reproduced in real time.

[0090] Example: A basketball game

[0091] Server Action:

[0092] The server collects video from multiple cameras in the stadium, stores it in a temporary buffer, and then sorts it chronologically. It then automatically recognizes overlapping areas of the video and seamlessly combines them to generate a 360-degree panoramic video. The video is then encoded, divided into packets, and sent to the terminal.

[0093] Terminal handling:

[0094] The device receives and reassembles data packets, decodes the video data, renders the video in real time according to the user's viewpoint, detects head movement using the built-in gyro sensor, and adjusts the video and sound effects in real time to create a sense of realism.

[0095] User Action:

[0096] Users put on VR goggles and log in. They can freely change the viewpoint by moving their head and enjoy an interactive experience of watching specific scenes. For example, they can switch to a courtside viewpoint and experience the players' movements and plays in real time.

[0097] Example 1

[0098] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0099] With conventional game viewing systems, it was difficult for users to experience the immersive atmosphere of a game in real time from a remote location. Furthermore, there was a lack of technology to seamlessly integrate video data from multiple cameras and provide 360-degree panoramic video. As a result, it was difficult for users to enjoy a large-scale game as if they were actually there.

[0100] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0101] In this invention, the server includes means for collecting video data from multiple camera devices within the venue, means for storing the collected video data in a buffer and sorting it in chronological order based on timestamps, means for automatically recognizing overlapping portions of the video data and combining them seamlessly, means for encoding the combined 360-degree omnidirectional video data in real time, dividing it into packets, and transmitting them to the terminal, means for the terminal to reconstruct the received data packets and store them in a temporary buffer as video data, means for the terminal to decode the reconstructed video data and render an image according to the user's viewpoint in real time, and means for the terminal to acquire the user's head movement using a built-in gyro sensor and adjust video and sound effects in real time based on the movement, thereby enabling users to experience the same sense of realism as if they were at the venue, even from a remote location.

[0102] A "venue" refers to a place where a sport or event is held, and is a facility equipped to allow spectators to watch the event.

[0103] "Photography device" refers to a device for acquiring video data, and includes fixed cameras, movable cameras, drone cameras, etc.

[0104] "Video data" refers to digital video information captured by a camera, and may include timestamps and coordinate information.

[0105] A "buffer" is a temporary data storage area used to temporarily store collected video data and reconstructed data.

[0106] A "timestamp" is digital information that indicates the time at which video data was collected, and is used to accurately manage the time series of data.

[0107] The "overlap" refers to the area where images captured by multiple cameras overlap, and is an important element for seamless image stitching.

[0108] "Seamless combining" refers to the process of integrating multiple pieces of video data into one continuous image without discontinuities or boundaries.

[0109] "360-degree omnidirectional video data" is video data that covers the field of view in all directions, and is in a format that allows the user to freely change the viewpoint in any direction.

[0110] "Encoding" is the process of compressing video data using a certain format or codec and converting it into a transmittable form.

[0111] A "packet" refers to each unit of data divided when digital information is transmitted over a network, and each packet is accompanied by transmission route information and error check information.

[0112] A "terminal" is a device that the user directly operates, which in this case refers to VR goggles.

[0113] "Decoding" is the process of restoring encoded video data to its original format so that the device can play the video.

[0114] A "gyro sensor" is a sensor that detects the angular velocity and rotation of a device and is used to accurately capture the user's head movements.

[0115] "Rendering in real time" refers to the process of generating and displaying images in real time in response to the user's viewpoint and movements.

[0116] "Sound effects" refers to the audio and sound effects that accompany the video, and enhance the sense of realism by adjusting the sense of direction and distance according to the user's movements.

[0117] The present invention relates to a system that uses VR technology to enhance the sense of realism when watching a game. Hereinafter, an embodiment of the present invention will be described in detail.

[0118] Server Processing

[0119] The server collects video data from multiple camera devices located within the venue. These include fixed cameras, mobile cameras, and drone cameras. The collected video data is first stored in a buffer. Each piece of video data is assigned a timestamp, which is used to sort the data in chronological order. The server then automatically recognizes overlapping areas of the video data and seamlessly combines them. This process uses an image processing library (e.g., OpenCV).

[0120] Specifically, the server uses an edge detection algorithm (e.g., SIFT) to identify adjacent overlapping areas and integrate them into one continuous video. This integrated 360-degree panoramic video data is then encoded in real time using the H.264 or H.265 codec. The encoded data is then divided into packets of a fixed size and sent to the device using the UDP protocol. The transmission process uses the ffmpeg library.

[0121] Terminal processing (VR goggle processing)

[0122] The terminal receives real-time video data packets sent from the server. The received packets are stored in a temporary buffer and reconstructed based on the sequence numbers. The terminal then decodes the video data using a decoding library (e.g., libavcodec). The decoded video data is stored in the temporary buffer again.

[0123] The device then renders this video data in real time according to the user's viewpoint. A game engine (e.g., Unity or Unreal Engine) is used for rendering. The device uses a built-in gyro sensor to detect the user's head movements and adjusts the video viewpoint and sound effects in real time based on those movements. An audio engine (e.g., OpenAL or FMOD) is used to adjust the sound effects.

[0124] User operations

[0125] First, users put on VR goggles and log in to the system. This process is done using facial recognition, ID, and password. Users can freely change their viewpoint by moving their head. They can view 360-degree images and select and watch their favorite scenes in real time.

[0126] Specific examples

[0127] For example, when streaming a basketball game, the server collects footage from multiple cameras installed around the stadium. The collected footage is then sorted chronologically based on timestamps, overlapping parts are recognized and seamlessly joined, and then encoded and sent to the device.

[0128] The device decodes the received video data and renders it in real time according to the user's viewpoint. Users can change the viewpoint by moving their head, for example, to a courtside perspective, and experience the players' movements and plays in real time. Sound effects are also adjusted in real time, making users feel as if they are at the game venue.

[0129] Prompt Sentence Examples

[0130] The following are examples of prompt sentences to input into a generative AI model:

[0131] "Please explain in detail the terminal processing of the game viewing system using VR technology. In particular, please provide details on the reception and reassembly of data packets, decoding and buffering, real-time rendering based on the user's viewpoint, and adjustment of sound effects."

[0132] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0133] Server Processing

[0134] Step 1: Collect video data

[0135] The server collects video data in real time from multiple camera devices (fixed cameras, mobile cameras, drone cameras, etc.) installed within the venue. These cameras assign a timestamp to each video frame and send it to the server. The server temporarily stores the received video data in a buffer. The input is real-time video data from the camera devices, and the output is the video data stored in the buffer.

[0136] Step 2: Adjust the timeline of the video data

[0137] The server sorts the video data stored in the buffer into chronological order based on the timestamps, ensuring accurate synchronization of the video from each camera. The input is time-stamped video data, and the output is video data sorted in chronological order.

[0138] Step 3: Seamlessly stitch together footage

[0139] The server automatically recognizes overlapping areas of the video data from each camera and seamlessly combines them. This process uses an image processing library (e.g., OpenCV) and utilizes an edge detection algorithm (e.g., SIFT). The input is multiple video data sorted in chronological order, and the output is seamlessly combined 360-degree omnidirectional video data.

[0140] Step 4: Encode and transmit video data

[0141] The server encodes the combined 360-degree panoramic video data in real time using the H.264 or H.265 codec. The encoded data is then divided into packets of a fixed size and sent to the device using the UDP protocol. The input is the seamlessly combined 360-degree panoramic video data, and the output is the encoded data packets sent to the device.

[0142] Terminal processing (VR goggle processing)

[0143] Step 1: Receiving and reassembling data packets

[0144] The terminal receives data packets sent from the server. The received packets are stored in a temporary buffer and reassembled into the correct order based on the sequence numbers. The input is the data packets from the server, and the output is the assembled video data.

[0145] Step 2: Decoding and Buffering

[0146] The device decodes the reconstructed video data. This process uses a decoding library (e.g., libavcodec) and may utilize hardware acceleration. The decoded video is again stored in a temporary buffer. The input is the reconstructed video data, and the output is the decoded video data.

[0147] Step 3: Real-time rendering based on user perspective

[0148] The device uses a built-in gyro sensor to capture the user's head movements and renders video data in real time based on those movements. A 360-degree panoramic image is displayed according to the user's viewpoint. This process uses a game engine (e.g., Unity or Unreal Engine). The input is the decoded video data and gyro sensor movement data, and the output is a rendered image according to the user's viewpoint.

[0149] Step 4: Real-time sound effect adjustment

[0150] The device adjusts the sound effects in real time according to the user's head movements, so that the sound has a sense of direction according to the user's movements. This process uses an acoustic engine (e.g., OpenAL or FMOD). The input is the user's movement data, and the output is the adjusted sound effects.

[0151] User operations

[0152] Step 1: Put on the VR goggles and log in to the system

[0153] The user puts on the VR goggles and logs in to the system. The login process uses facial recognition, ID, and password. The input is the user's authentication information, and the output is the system login status.

[0154] Step 2: Freely change the viewpoint

[0155] Users can freely change their viewpoint by moving their head, allowing them to view a 360-degree panoramic view and select and watch their favorite scenes in real time. The input is the user's head movement data, and the output is a rendered image based on the user's viewpoint.

[0156] (Application example 1)

[0157] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0158] Conventional game viewing systems have limited means for enhancing the sense of realism, limiting the ability to experience the game from specific viewpoints and angles. It has also been difficult for users to freely change viewpoints or adjust video and audio effects in real time. This has prevented users from experiencing the game as if they were actually at the venue.

[0159] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0160] In this invention, the server includes means for collecting video data from multiple camera devices within the venue, means for integrating the collected video data to generate 360-degree omnidirectional video data, means for transmitting the generated 360-degree omnidirectional video data to the terminal in real time, means for the terminal to receive the video data and provide a 360-degree omnidirectional field of view and adjust the video and sound effects according to the user's movements, means for the terminal to detect the user's movements using a built-in gyro sensor and adjust the video and sound effects based on the movements in real time, and means for temporarily storing the collected video data in a buffer and realizing seamless playback. This allows users to experience a realistic game viewing experience in real time and to freely change the viewpoint to watch from different viewpoints within the venue.

[0161] "Game venue" means a particular location where a sport or event is held.

[0162] "Filming equipment" refers to equipment used to collect video data, including fixed cameras, mobile cameras, and drone cameras.

[0163] "Video data" refers to visual information captured by a camera within the venue.

[0164] "Integration" refers to the process of combining multiple pieces of video data and reconstructing them into a single, continuous image.

[0165] "360-degree panoramic video" refers to video data that covers all directions within the venue, and has the ability to freely change the viewpoint.

[0166] "Real-time" refers to a situation in which information is processed almost instantly, with little time delay.

[0167] "Terminal" refers to a device that a user uses to watch video, including smartphones and head-mounted displays.

[0168] "Providing" refers to the act of making a particular service or feature available to a user.

[0169] "Movement" refers to the user's actions of moving their head or body.

[0170] "Sound effects" refers to audio elements such as voice and music that are provided when watching a video.

[0171] "Adjustment" refers to the act of changing or modifying functions or effects to suit specific conditions or environments.

[0172] A "gyro sensor" is a sensor for measuring angular velocity and is used to detect the movement of the user's head.

[0173] "Decoding" refers to the process of restoring encoded data to its original form.

[0174] "Rendering" refers to the process of visually displaying video data.

[0175] A "temporary buffer" refers to a memory area for temporarily storing data.

[0176] "Seamless" refers to playback or joining that is performed without interruption, with continuity maintained.

[0177] This invention relates to a system that uses VR technology to enhance the sense of realism when watching a match. This system collects video data from multiple camera devices within the match venue and provides it to users in real time, allowing them to enjoy the match from a 360-degree, all-around perspective.

[0178] Server Processing

[0179] The server collects video data from multiple camera devices installed within the venue, including fixed cameras, mobile cameras, and drone cameras. The collected video data is first stored in a buffer, and the footage from each camera is sorted chronologically. Next, overlapping areas of the footage are automatically recognized and seamlessly combined, generating 360-degree video data. The generated 360-degree video data is then encoded in real time, divided into packets of a fixed size, and sent to the device. This transmission is performed with low latency, allowing users to experience the video smoothly.

[0180] Terminal handling

[0181] The device (e.g., a smartphone or head-mounted display) receives real-time video data sent from the server. The device reassembles the received data packets and stores them in a temporary buffer. It then decodes the video data and renders the video in real time according to the user's movements. The device uses its built-in gyro sensor to acquire data on the user's head movement and adjusts the video and sound effects in real time based on that movement. This allows the user to enjoy a visual and auditory experience that makes them feel as if they are at the stadium.

[0182] User operations

[0183] First, users put on the device and log in to the system. They can freely change their viewpoint by moving their head, looking around in all 360 degrees to watch any scene they like in real time. For example, users can stand in the perspective of a specific player and experience important moments of the game up close. Users can also select different viewpoints to enjoy views from different locations within the stadium.

[0184] Specific examples

[0185] For example, when streaming a basketball game, the process goes like this: The server collects footage from multiple cameras installed within the venue, automatically recognizes overlapping areas of the footage, and seamlessly combines them. This combined video data is encoded in real time and sent to the device. The device receives the video data sent from the server and stores it in a temporary buffer. It then decodes it and renders the video in real time according to the user's head movements. The built-in gyro sensor adjusts the user's viewpoint, and sound effects are also adjusted in real time. For example, you can switch to a courtside view to experience the players' movements and plays in real time.

[0186] Prompt Sentence Examples

[0187] "I want to develop a VR application that allows users to freely watch soccer matches from a 360-degree perspective. Please tell me specifically how to collect video data from a server in real time and render the video on an HMD according to the user's viewpoint."

[0188] Thus, according to the present invention, the user can experience watching a game with a sense of real-time presence, and by freely changing the viewpoint, the user can watch the game from different viewpoints within the venue.

[0189] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0190] Step 1:

[0191] The server collects video data from multiple camera devices installed within the venue.

[0192] Input: Video data from fixed, mobile, and drone cameras.

[0193] Output: Collected multi-camera video data.

[0194] Specific operation: The server continuously receives video data from each imaging device and temporarily stores it in a buffer.

[0195] Step 2:

[0196] The server integrates the collected video data and generates 360-degree panoramic video data.

[0197] Input: Single camera video data (multiple).

[0198] Output: Integrated 360° video data.

[0199] Specific operation: The time series of video data is organized and overlapping parts are seamlessly combined using an automatic recognition system (computer vision algorithm).

[0200] Step 3:

[0201] The server encodes the generated 360-degree omnidirectional video data in real time, divides it into packets of a certain size, and transmits it to the terminal.

[0202] Input: Integrated 360° video data.

[0203] Output: The encoded data packet.

[0204] Specific operation: Video data is encoded and compressed, divided into packets, and sent to the terminal via a low-latency network.

[0205] Step 4:

[0206] The terminal receives real-time video data transmitted from the server and stores it in a temporary buffer.

[0207] Input: The encoded data packet.

[0208] Output: Data packets stored in a temporary buffer.

[0209] Specific operation: Starts the process of receiving data from the network and storing it in a temporary buffer.

[0210] Step 5:

[0211] The terminal decodes the data packets stored in the temporary buffer and assembles them into a video.

[0212] Input: Data packets stored in a temporary buffer.

[0213] Output: Decoded video data.

[0214] Specific operation: Using a decoding algorithm, the packets are converted into the original video data and a continuous video is reassembled.

[0215] Step 6:

[0216] The device uses a built-in gyro sensor to detect the user's head movements and adjusts visual and sound effects in real time based on those movements.

[0217] Input: Motion data from gyro sensor, decoded video data.

[0218] Output: Visual and sound effects adjusted to the user's point of view.

[0219] How it works: Using gyro sensor data, it tracks the user's head movements and dynamically adjusts the viewpoint and sound in real time.

[0220] Step 7:

[0221] The user wears the device and logs in to the system. While watching the captured video, the user can freely change the viewpoint by moving their head.

[0222] Input: Input data from the user interface.

[0223] Output: The visual and audio data that the user sees.

[0224] Specific operation: After the user wears the device and logs in to the system, they can watch the game while the video and audio are changed in real time based on their head movements.

[0225] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0226] The present invention relates to a system that uses VR technology and emotion recognition technology to enhance the sense of realism when watching a game. Hereinafter, an embodiment of the present invention will be described in detail.

[0227] Server Processing

[0228] The server collects video data from multiple camera systems within the venue, including fixed cameras, mobile cameras, and even drone cameras. The collected video data is buffered and the video data from each camera is sorted chronologically. A process then automatically recognizes overlapping areas of the video and seamlessly combines them to generate 360-degree video data.

[0229] The generated 360-degree video data is encoded in real time, divided into packets of a certain size, and transmitted to the device with low latency, allowing users to enjoy a smooth experience.

[0230] Terminal processing (VR goggle processing)

[0231] The device (VR goggles) receives real-time video data sent from the server. The device reassembles the received data packets and stores them in a temporary buffer as video data. Next, the device decodes the video data and renders the video in real time according to the user's viewpoint.

[0232] The device uses a built-in gyro sensor to capture the user's head movement data and adjusts the visual and audio effects in real time based on that movement, providing users with a visual and audio experience that makes them feel like they're at the stadium.

[0233] Furthermore, the device includes an emotion engine that uses a camera to detect the user's facial expressions and biometric signals (e.g., heart rate, brain waves, etc.) and recognizes the user's emotions in real time based on the detected emotions. The recognized emotion data is then sent to a server.

[0234] Emotion data processing on the server

[0235] The server receives the user's emotional data sent from the device. Based on this emotional data, the server can automatically adjust the video and sound effects. For example, if the user is excited, it can emphasize the display of important scenes and enhance the sound effects.

[0236] User operations

[0237] First, users put on VR goggles and log in to the system. Users can freely change their viewpoint by moving their head, looking around in all 360 degrees and watching their favorite scenes in real time. Furthermore, camera angles and content are automatically selected according to the user's emotions, giving users an even more immersive experience.

[0238] Specific examples

[0239] For example, if you were streaming a basketball game, it would look like this:

[0240] Server processing: Images are collected from multiple cameras installed within the venue, temporarily stored in a buffer, and then sorted chronologically. Next, overlapping areas of the images are automatically recognized and seamlessly combined to generate a 360-degree panoramic image. The image is then encoded, divided into packets, and sent to the terminal.

[0241] Device processing: The device receives and reassembles data packets and decodes the video data. It renders the video in real time according to the user's viewpoint and detects head movements using the built-in gyro sensor. It adjusts the video and sound effects in real time. The device's camera also detects the user's facial expressions and biometric signals, analyzes the data using the emotion engine, and sends it to the server as emotion data.

[0242] Server emotion data processing: The server receives the emotion data and adjusts the video and sound effects in real time according to the user's emotion. For example, if the user is excited, the video will be adjusted to emphasize particularly dynamic scenes or important plays.

[0243] User operation: The user puts on the VR goggles and logs in. They can freely change the viewpoint by moving their head, enjoying an interactive experience of watching specific scenes. Furthermore, camera angles and specific content are automatically selected according to the user's emotions, enhancing the sense of realism. For example, if the user is feeling nervous, they can receive feedback such as scenes and angles that will help them relax.

[0244] This allows users to experience the realistic sensation of being at a game venue from the comfort of their own home, and furthermore, to enjoy a visual and auditory experience that is customized to suit their individual emotions.

[0245] The processing flow will be explained below.

[0246] Server Processing

[0247] Step 1: Collect video data

[0248] The server receives video data from multiple camera devices installed within the venue, including fixed, mobile and drone cameras, each capturing the match from a different perspective or angle.

[0249] Step 2: Buffering video data

[0250] The server temporarily stores the received video data in a buffer, whereby the video data from each camera is temporarily stored.

[0251] Step 3: Synchronize and sort data

[0252] The server rearranges the video data in the temporary buffer in chronological order and synchronizes them based on the timestamps.

[0253] Step 4: Video data integration

[0254] The server then combines the sorted video data from multiple cameras, automatically recognizing overlapping areas and seamlessly joining them together.

[0255] Step 5: Encode the video data

[0256] The server encodes the integrated 360-degree panoramic video data in real time.

[0257] Step 6: Packetize the data

[0258] The server divides the encoded video data into packets of a fixed size and prepares them for transfer.

[0259] Step 7: Sending data

[0260] The server then transmits the divided data packets to the terminal via a high-speed network, aiming for low latency.

[0261] Step 8: Receiving emotion data

[0262] The server receives the user's emotional data transmitted from the device, including emotional information analyzed from the user's facial expressions and biometric signals.

[0263] Step 9: Adjust based on sentiment data

[0264] The server adjusts the visual and audio effects in real time based on the received emotion data, selecting specific camera angles and content to display according to the user's emotion.

[0265] Terminal (VR goggles) processing

[0266] Step 1: Receiving a data packet

[0267] The terminal receives the data packets sent from the server through a high-speed network.

[0268] Step 2: Reassembling the data packets

[0269] The terminal reassembles the received data packets into the original video data.

[0270] Step 3: Decoding the video data

[0271] The terminal decodes the reassembled video data and prepares it for display.

[0272] Step 4: Rendering the Perspective

[0273] Based on the decoded video data, the device renders a 360-degree panoramic image tailored to the user's viewpoint in real time.

[0274] Step 5: Acquire motion data

[0275] The device uses a built-in gyro sensor to detect the user's head movements.

[0276] Step 6: Adjusting the image and sound

[0277] The device adjusts the video and audio effects in real time based on the acquired motion data, providing a sense of realism in both visual and audio.

[0278] Step 7: Obtaining emotion data

[0279] The device uses cameras and sensors to detect the user's facial expressions and biometric signals, recognizing the user's emotions in real time.

[0280] Step 8: Sending Emotion Data

[0281] The terminal transmits the recognized emotion data of the user to the server.

[0282] User operations

[0283] Step 1: Put on the VR goggles and log in

[0284] The user puts on the VR goggles and logs into the system. To log in, the user must enter their account information.

[0285] Step 2: Change your perspective

[0286] Users can freely change their viewpoint by moving their head, enjoying a 360-degree view, and can also focus on a specific player or play.

[0287] Step 3: Interactive Experience

[0288] Users receive real-time visual and auditory feedback, making them feel as if they are at the stadium, with sounds such as the cheers of the crowd and the footsteps of the players reproduced in real time.

[0289] Step 4: Displaying emotions

[0290] Depending on the user's emotions, specific camera angles and content can be automatically selected and displayed. For example, if the user is excited, important scenes can be displayed more prominently and sound effects can be enhanced.

[0291] Example 2

[0292] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0293] Conventional game viewing systems have limited visual and auditory experiences, making it difficult to fully recreate the sense of presence of a game. Furthermore, it is not possible to optimize the content played based on the user's real-time emotions and movements, making it difficult to provide a personalized experience for each user.

[0294] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0295] In this invention, the server includes means for collecting video data from multiple camera devices within the venue, means for integrating the collected video data to generate 360-degree omnidirectional video data, and means for transmitting the generated 360-degree omnidirectional video data to the terminal in real time, thereby enabling the user to experience 360-degree omnidirectional real-time video, and for the terminal to receive the video data and provide a 360-degree omnidirectional field of view, and to adjust the video and sound effects according to the user's movements.

[0296] Furthermore, by including a means for the terminal to decode the video data sent from the server and render the video in real time according to the user's viewpoint, and a means for acquiring the user's emotional data and adjusting the video and sound effects based on that, it is possible to provide a more immersive and individually optimized visual and auditory experience.

[0297] A "game venue" is a specific location where a sport or event is played.

[0298] "Filming equipment" refers to equipment such as cameras and drone cameras used to collect video data.

[0299] "Video data" refers to video information collected from a camera and includes visual content.

[0300] "Collection" refers to the act of collecting video data from multiple imaging devices.

[0301] "Integration" is the process of bringing together collected video data.

[0302] "360-degree omnidirectional video data" is video data that provides visual information from all directions.

[0303] "Generation" is the act of creating new video data.

[0304] "Real-time" refers to almost instantaneous processing, with very little delay.

[0305] A "terminal" is a device that receives and displays video data, such as VR goggles.

[0306] "Transmitting" is the act of sending data to another device or system.

[0307] "Rendering" is the process of generating images or videos from digital data.

[0308] A "gyro sensor" is a sensor that detects the movement of the user's head.

[0309] "Emotion data" is data that indicates the emotional state of the user.

[0310] "Decoding" is the process of converting encoded data into a playable format.

[0311] "Adjustment" is the act of changing visual and sound effects based on the user's movements and emotions.

[0312] The present invention relates to a system that uses VR technology and emotion recognition technology to enhance the sense of realism when watching a game. Specific embodiments of the present invention will be described below.

[0313] The server collects video data from multiple camera devices within the venue. This includes fixed cameras, mobile cameras, and drone cameras. The video data from each camera is temporarily stored in a buffer and sorted chronologically. The server then automatically recognizes overlapping areas of the footage and seamlessly combines them to generate 360-degree panoramic video data.

[0314] The generated 360-degree video data is encoded in real time using codecs such as H.264 or H.265. The encoded data is divided into packets of a fixed size and transmitted to the device with low latency.

[0315] The device (VR goggles) receives video data packets sent from the server. The received data is stored in a buffer and decoded. The decoded video data is rendered in real time according to the user's viewpoint. The device uses a built-in gyro sensor to detect the user's head movement and adjusts the video and sound effects based on that movement.

[0316] Furthermore, the device is equipped with a camera and biometric sensors to detect the user's facial expressions and biometric signals (e.g., heart rate, brain waves). This data is analyzed by the emotion engine and sent to the server as user emotion data.

[0317] The server receives the emotion data sent from the terminal and adjusts the video and sound effects based on the user's emotional state. For example, if the user is excited, the server adjusts the video to emphasize important scenes or dynamic play.

[0318] Users can use the above services by wearing VR goggles and logging in to the system. Users can freely change their viewpoint by moving their head, enjoying a 360-degree panoramic view. In addition, camera angles and content are automatically selected based on the user's emotions, creating an even more immersive experience.

[0319] Specific examples

[0320] For example, when streaming a basketball game, the following steps are taken: Video data is collected from multiple cameras installed in the stadium and stored in a temporary buffer. The images from each camera are rearranged in chronological order and seamlessly combined to generate a 360-degree panoramic video. The video is then encoded, divided into packets, and sent to the device.

[0321] The device reassembles the received data packets and renders the decoded video data in real time. The built-in gyro sensor detects the user's head movements and adjusts the video and audio in real time. The device's camera and biometric sensors also detect the user's facial expressions and biometric signals, and the emotion engine analyzes the emotional data and transmits it to the server.

[0322] The server receives the emotion data and adjusts the video and sound effects according to the user's emotions. For example, if the user is excited, the video can be adjusted to emphasize dynamic scenes and important plays, allowing the user to enjoy the game more realistically.

[0323] Prompt Sentence Examples

[0324] "I would like to develop a system that uses VR technology to make watching basketball games more immersive, and also utilizes emotion recognition technology to adjust video and sound effects. This system will incorporate a mechanism to optimize video in real time according to the user's movements and emotions. Please tell me specifically how you will design the system and how you will use emotion data."

[0325] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0326] Step 1:

[0327] "Video data collection"

[0328] The server collects video data from fixed cameras, mobile cameras, drone cameras, etc. within the venue. The video data from each camera is temporarily stored in the server's buffer.

[0329] Input: Video streams from each imaging modality.

[0330] Output: Multiple video data stored in a buffer.

[0331] Specific operation: The server uses the stream URL of the camera registered in advance to pull in video data using a specified protocol (such as RTSP).

[0332] Step 2:

[0333] "Sorting time series data"

[0334] The server then sorts the video data stored in the buffer into chronological order based on the timestamps of each camera, thereby synchronizing the video data.

[0335] Input: Multiple video data stored in a buffer.

[0336] Output: Video data sorted in chronological order.

[0337] What it does: The server analyzes the timestamp information and runs an algorithm to arrange the video frames from each camera in the proper order.

[0338] Step 3:

[0339] "Seamless joining process"

[0340] The server automatically recognizes overlapping areas of the images and seamlessly combines them to generate 360-degree panoramic video data.

[0341] Input: Multiple video data sorted in chronological order.

[0342] Output: 360-degree panoramic video data.

[0343] How it works: The server detects overlapping areas and uses image processing algorithms to merge them together, creating a continuous omnidirectional image.

[0344] Step 4:

[0345] "Encoding 360-degree video"

[0346] The server encodes the generated 360-degree panoramic video data in real time using the H.264 or H.265 codec.

[0347] Input: 360-degree omnidirectional video data.

[0348] Output: Encoded video data.

[0349] Specific operation: The server inputs the video data into the codec, which performs compression and encoding processes, thereby reducing the data size appropriately and making transmission more efficient.

[0350] Step 5:

[0351] "Sending video packets"

[0352] The server divides the encoded video data into packets of a fixed size and transmits them to the terminal with low latency.

[0353] Input: Encoded video data.

[0354] Output: Split data packets.

[0355] Specific operation: The server packetizes the video data and transmits it to the terminal using the UDP or TCP protocol.

[0356] Step 6:

[0357] "Receiving Data Packets"

[0358] The terminal receives the video data packets sent from the server, checking for packet loss and duplication to ensure accurate reception.

[0359] Input: Data packet from the server.

[0360] Output: Data packets that have been received and processed.

[0361] Specific operation: The terminal performs packet reception processing at the network layer, and processes error checks and retransmission requests.

[0362] Step 7:

[0363] "Assembling Data Packets"

[0364] The received data packets are reassembled into the original video data and stored in a temporary buffer.

[0365] Input: Received data packets.

[0366] Output: Video data stored in a temporary buffer.

[0367] Specific operation: The terminal rearranges, reconstructs, and temporarily stores the data packets, referring to the timestamp information.

[0368] Step 8:

[0369] "Video data decoding"

[0370] The terminal decodes the video data in the buffer using an H.264 or H.265 decoder and acquires it as successive video frames.

[0371] Input: Video data stored in a temporary buffer.

[0372] Output: Decoded video frames.

[0373] Specific operation: The terminal decodes the video frame using dedicated decoding hardware or software.

[0374] Step 9:

[0375] "Video rendering"

[0376] The decoded video is rendered in real time according to the user's viewpoint and displayed in the VR goggles.

[0377] Input: Decoded video frames.

[0378] Output: Rendered video frames.

[0379] Specific operation: The terminal uses a rendering engine to generate video frames according to the user's viewpoint and displays them on the display.

[0380] Step 10:

[0381] "Acquiring head movement data"

[0382] The device's built-in gyro sensor detects the user's head movements and uses that data to adjust visual and sound effects in real time.

[0383] Input: User's head movement.

[0384] Output: Coordinated video and audio.

[0385] Specific operation: The device obtains data in real time from the gyro sensor and dynamically adjusts image rendering and audio filtering.

[0386] Step 11:

[0387] "Acquiring emotion data"

[0388] The device's camera and biometric sensors capture biometric data such as the user's facial expressions and heart rate, which is then analyzed by the emotion engine.

[0389] Input: User's facial expression data and biometric signals.

[0390] Output: Parsed emotion data.

[0391] Specific operation: The device's camera extracts feature points using a facial expression recognition algorithm, and biometric sensors measure heart rate and skin galvanic response, which are then comprehensively analyzed by the emotion engine.

[0392] Step 12:

[0393] "Sending emotional data"

[0394] The analyzed emotional data is sent to the server in real time.

[0395] Input: Parsed emotion data.

[0396] Output: Emotion data sent to the server.

[0397] Specific operation: The device encodes the emotion data into an appropriate format and makes a transmission request to the server.

[0398] Step 13:

[0399] "Receiving and adjusting emotional data"

[0400] The server receives the emotional data sent from the terminal and adjusts the visual and sound effects based on the user's emotional state.

[0401] Input: Emotion data sent from the device.

[0402] Output: Coordinated video and audio.

[0403] Specific operation: The server analyzes the emotional data and dynamically generates content optimized for the user's emotional state, emphasizing specific visual scenes and sound effects.

[0404] Through the above processing steps, the system can provide the user with a realistic sense of being at the game, and can also provide a customized visual and auditory experience according to the user's emotional state.

[0405] (Application example 2)

[0406] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0407] While conventional VR-based game viewing systems provide users with a high sense of realism, they have difficulty providing a personalized experience that reflects the emotional and physical state of each user. Therefore, a new system is needed that can provide appropriate visual and audio effects in real time according to the emotional state and preferences of various users.

[0408] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for collecting video data from multiple camera devices within the venue, means for integrating the collected video data to generate 360-degree omnidirectional video data, and means for transmitting the generated 360-degree omnidirectional video data to the terminal in real time. This allows the user to experience the immersive atmosphere of the venue in real time. The terminal also includes means for adjusting video and audio effects according to the user's movements, means for sensing the user's facial expressions and biometric signals and recognizing the user's emotions in real time based on the detected facial expressions and biometric signals, and means for automatically adjusting video and audio effects based on the recognized emotional data. This allows the user to enjoy a personalized visual and audio experience that is tailored to their emotions and physical state.

[0409] "Game venue" refers to a location where various sports and entertainment matches are played.

[0410] "Filming equipment" refers to equipment such as cameras and drones used to collect video data.

[0411] "Video data" refers to video data collected by an imaging device.

[0412] "360-degree omnidirectional video data" refers to video that can be viewed in all directions and is generated by integrating video data collected from multiple imaging devices.

[0413] "Terminal" refers to a device (e.g., VR goggles) that receives video data and provides a visual and auditory experience to the user.

[0414] "User movement" refers to the user's body movement, particularly head movement, detected by the terminal.

[0415] "Sound effects" refers to the audio data corresponding to the video and the method of reproducing it.

[0416] "Expression" refers to the facial expression of the user.

[0417] "Biological signals" refer to physiological data such as a user's heart rate and brain waves.

[0418] "Emotion data" refers to the emotional state recognized based on the user's facial expressions and biometric signals.

[0419] "Real-time" refers to near-instant processing and transmission.

[0420] A "gyro sensor" refers to an inertial sensor used to detect the movement of a device.

[0421] "Synthesis" refers to the process of combining collected data into one linked data set.

[0422] "Seamlessly combining" refers to combining different video data continuously without interruption.

[0423] "Automatic adjustment" refers to the system automatically adjusting the video and audio based on the user's emotional data.

[0424] This invention relates to a system that enhances the sense of realism when a user watches a game using VR goggles and automatically adjusts video and sound effects according to the emotions of each individual user. The specific configuration and operation of the system are described below.

[0425] Server Processing

[0426] The server collects video data from multiple camera devices within the venue. The camera devices include fixed cameras, mobile cameras, and drone cameras. The collected video data is first stored in a buffer, and the video data from each camera is sorted chronologically. The server then automatically recognizes overlapping areas of the video and seamlessly combines them to generate 360-degree panoramic video data. The generated video data is then encoded in real time, divided into packets of a fixed size, and transmitted to the device. This transmission is performed with low latency, ensuring a smooth user experience.

[0427] Terminal handling

[0428] The device (VR goggles) receives real-time video data sent from the server. It reassembles the received data packets and stores them in a temporary buffer. It then decodes the video data and renders the video in real time according to the user's viewpoint. The device uses a built-in gyro sensor to acquire data on the user's head movement and adjusts the video and sound effects in real time based on that movement. The device also uses a camera to detect the user's facial expressions and biometric signals (e.g., heart rate, brain waves, etc.), and includes an emotion engine that recognizes the user's emotions in real time based on this data. The recognized emotion data is sent to the server in real time.

[0429] Emotion data processing on the server

[0430] The server receives the user's emotional data sent from the device. Based on this emotional data, it can automatically adjust the visual and sound effects. For example, if the user is excited, it can highlight important scenes and thrilling plays and enhance the sound effects.

[0431] User operations

[0432] First, users put on VR goggles and log in to the system. Users can freely change their viewpoint by moving their head, and enjoy a 360-degree panoramic view of their favorite scenes in real time. Furthermore, camera angles and content are automatically selected based on the user's emotions, creating an even more immersive experience.

[0433] Hardware and software used

[0434] The system is implemented using the following hardware and software:

[0435] Hardware: VR goggles (e.g., Oculus Rift, HTC Vive), fixed cameras, movable cameras, drone cameras, heart rate monitors, electroencephalographs

[0436] Software: OpenCV (used for image acquisition and processing), Python socket library (used for sending and receiving image data), Unity or Unreal Engine (for VR rendering and user interface construction)

[0437] Specific examples

[0438] For example, when live streaming a soccer match, the following prompt sentence can be input into the generative AI model to generate video that highlights important scenes from the match.

[0439] "Analyze the user's emotion recognition data in real time during a soccer match. If the user is excited, highlight the goal scene and the play before the goal. If the user is nervous, switch to a more relaxing video."

[0440] This invention allows users to enjoy a highly realistic and personalized viewing experience.

[0441] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0442] Step 1:

[0443] The server collects video data from multiple camera devices (fixed cameras, mobile cameras, drone cameras) within the venue. The collected video data is temporarily stored in a buffer, and the video data from each camera is sorted chronologically. This ensures that the video data is arranged in the appropriate order.

[0444] Step 2:

[0445] The server automatically recognizes overlapping areas in the collected video data and seamlessly combines them. It uses OpenCV to detect overlapping areas in each frame and combines them continuously. The combined video data becomes 360-degree omnidirectional video data.

[0446] Step 3:

[0447] The server encodes the generated 360-degree omnidirectional video data in real time, divides it into packets of a fixed size, and transmits them to the device. This transmission is performed with low latency, and the data is sent to the device via the network. Video codecs such as H.264 and HEVC are used for encoding.

[0448] Step 4:

[0449] The device (VR goggles) receives real-time video data sent from the server, reassembles the received data packets, and stores them in a temporary buffer. Then, it decodes the video data. For decoding, it uses FFmpeg or other decoding libraries.

[0450] Step 5:

[0451] The device uses a built-in gyro sensor to capture data on the user's head movements, and uses this data to render video and sound effects in real time. Rendering is done using Unity or Unreal Engine, and the video is displayed according to the user's viewpoint.

[0452] Step 6:

[0453] The device uses a camera and biosensors to detect the user's facial expressions and biometric signals (heart rate, brain waves, etc.). It then uses an emotion engine to analyze this data and extract the user's emotional data. For example, it uses a machine learning model to infer emotions from facial expressions and biometric signals.

[0454] Step 7:

[0455] The device transmits the recognized emotion data to the server in real time with low latency, so the user's emotional state is immediately conveyed to the server.

[0456] Step 8:

[0457] The server automatically adjusts visual and audio effects based on the received user emotion data. For example, if the user is excited, it will highlight important scenes or thrilling action scenes and enhance audio effects. It uses a generative AI model to calculate the appropriate visual and audio settings in real time.

[0458] Step 9:

[0459] Users put on VR goggles, log in, and begin watching the game. They can freely change their viewpoint by moving their head, and enjoy a 360-degree panoramic view of the game to watch any scene they like. Furthermore, camera angles and content are automatically selected based on the user's emotions, creating an even more immersive experience.

[0460] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0461] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (registered trademark) (Internet search engine).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0462] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0463] [Second embodiment]

[0464] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0465] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0466] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0467] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0468] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0469] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0470] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0471] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0472] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0473] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0474] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0475] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0476] The present invention relates to a system that uses VR technology to enhance the sense of realism when watching a game. Hereinafter, an embodiment of the present invention will be described in detail.

[0477] Server Processing

[0478] The server collects video data from multiple camera systems within the venue, including fixed cameras, mobile cameras, and even drone cameras. The collected video data is buffered and the footage from each camera is sorted chronologically. A process then automatically recognizes overlapping areas of the footage and seamlessly combines them to generate 360-degree panoramic video data.

[0479] The generated 360-degree video data is encoded in real time, divided into packets of a certain size, and transmitted to the device with low latency, allowing users to experience the video smoothly.

[0480] Terminal processing (VR goggle processing)

[0481] The device (VR goggles) receives real-time video data sent from the server. The device reassembles the received data packets and stores them in a temporary buffer as video data. Next, the device decodes the video data and renders the video in real time according to the user's viewpoint.

[0482] The device uses a built-in gyro sensor to capture the user's head movement data and adjusts the visual and audio effects in real time based on that movement, providing users with a visual and audio experience that makes them feel like they're at the stadium.

[0483] User operations

[0484] First, users put on VR goggles and log in to the system. Users can freely change their viewpoint by moving their head, and enjoy a 360-degree panoramic view of their favorite scenes in real time. For example, they can stand in the perspective of their favorite player and experience the moment they make a shot up close.

[0485] Specific examples

[0486] For example, if you were streaming a basketball game, it would look like this:

[0487] Server processing: Images are collected from multiple cameras installed within the venue, and overlapping areas of the images are automatically recognized and seamlessly combined. This combined video data is encoded in real time and sent to the terminal.

[0488] Device processing: The device receives the video data sent from the server and stores it in a temporary buffer. It then decodes it and renders the video in real time according to the user's head movements. It uses the built-in gyro sensor to adjust the user's viewpoint and adjusts the sound effects in real time.

[0489] User operation: The user puts on the VR goggles and logs in to the system. The user can change the viewpoint by moving their head and enjoy 360-degree omnidirectional images. For example, they can switch to a courtside viewpoint and experience the players' movements and plays in real time.

[0490] This allows users to experience the immersive atmosphere of a live game venue from the comfort of their own home.

[0491] The processing flow will be explained below.

[0492] Server Processing

[0493] Step 1: Collect video data

[0494] The server receives video data from multiple camera devices installed within the venue, including fixed, mobile and drone cameras, each capturing the match from a different perspective or angle.

[0495] Step 2: Buffering video data

[0496] The server temporarily stores the received video data in a buffer, whereby the video data from each camera is temporarily stored.

[0497] Step 3: Synchronize and sort data

[0498] The server rearranges the video data in the temporary buffer in chronological order and synchronizes them based on the timestamps.

[0499] Step 4: Video data integration

[0500] The server then combines the sorted video data from multiple cameras, automatically recognizing overlapping areas and seamlessly joining them together.

[0501] Step 5: Encode the video data

[0502] The server encodes the integrated 360-degree panoramic video data in real time.

[0503] Step 6: Packetize the data

[0504] The server divides the encoded video data into packets of a fixed size and prepares them for transfer.

[0505] Step 7: Sending data

[0506] The server then transmits the divided data packets to the terminal via a high-speed network, aiming for low latency.

[0507] Terminal (VR goggles) processing

[0508] Step 1: Receiving a data packet

[0509] The terminal receives the data packets sent from the server through a high-speed network.

[0510] Step 2: Reassembling the data packets

[0511] The terminal reassembles the received data packets into the original video data.

[0512] Step 3: Decoding the video data

[0513] The terminal decodes the reassembled video data and prepares it for display.

[0514] Step 4: Rendering the Perspective

[0515] Based on the decoded video data, the device renders a 360-degree panoramic image tailored to the user's viewpoint in real time.

[0516] Step 5: Acquire motion data

[0517] The device uses a built-in gyro sensor to detect the user's head movements.

[0518] Step 6: Adjusting the image and sound

[0519] The device adjusts the video and audio effects in real time based on the acquired motion data, providing a sense of realism in both visual and audio.

[0520] User operations

[0521] Step 1: Put on the VR goggles and log in

[0522] The user puts on the VR goggles and logs into the system. To log in, the user must enter their account information.

[0523] Step 2: Change your perspective

[0524] Users can freely change their viewpoint by moving their head, enjoying a 360-degree view, and can also focus on a specific player or play.

[0525] Step 3: Interactive Experience

[0526] Users receive real-time visual and auditory feedback, making them feel as if they are at the stadium, with sounds such as the cheers of the crowd and the footsteps of the players reproduced in real time.

[0527] Example: A basketball game

[0528] Server Action:

[0529] The server collects video from multiple cameras in the stadium, stores it in a temporary buffer, and then sorts it chronologically. It then automatically recognizes overlapping areas of the video and seamlessly combines them to generate a 360-degree panoramic video. The video is then encoded, divided into packets, and sent to the terminal.

[0530] Terminal handling:

[0531] The device receives and reassembles data packets, decodes the video data, renders the video in real time according to the user's viewpoint, detects head movement using the built-in gyro sensor, and adjusts the video and sound effects in real time to create a sense of realism.

[0532] User Action:

[0533] Users put on VR goggles and log in. They can freely change the viewpoint by moving their head and enjoy an interactive experience of watching specific scenes. For example, they can switch to a courtside viewpoint and experience the players' movements and plays in real time.

[0534] Example 1

[0535] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0536] With conventional game viewing systems, it was difficult for users to experience the immersive atmosphere of a game in real time from a remote location. Furthermore, there was a lack of technology to seamlessly integrate video data from multiple cameras and provide 360-degree panoramic video. As a result, it was difficult for users to enjoy a large-scale game as if they were actually there.

[0537] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0538] In this invention, the server includes means for collecting video data from multiple camera devices within the venue, means for storing the collected video data in a buffer and sorting it in chronological order based on timestamps, means for automatically recognizing overlapping portions of the video data and combining them seamlessly, means for encoding the combined 360-degree omnidirectional video data in real time, dividing it into packets, and transmitting them to the terminal, means for the terminal to reconstruct the received data packets and store them in a temporary buffer as video data, means for the terminal to decode the reconstructed video data and render an image according to the user's viewpoint in real time, and means for the terminal to acquire the user's head movement using a built-in gyro sensor and adjust video and sound effects in real time based on the movement, thereby enabling users to experience the same sense of realism as if they were at the venue, even from a remote location.

[0539] A "venue" refers to a place where a sport or event is held, and is a facility equipped to allow spectators to watch the event.

[0540] "Photography device" refers to a device for acquiring video data, and includes fixed cameras, movable cameras, drone cameras, etc.

[0541] "Video data" refers to digital video information captured by a camera, and may include timestamps and coordinate information.

[0542] A "buffer" is a temporary data storage area used to temporarily store collected video data and reconstructed data.

[0543] A "timestamp" is digital information that indicates the time at which video data was collected, and is used to accurately manage the time series of data.

[0544] The "overlap" refers to the area where images captured by multiple cameras overlap, and is an important element for seamless image stitching.

[0545] "Seamless combining" refers to the process of integrating multiple pieces of video data into one continuous image without discontinuities or boundaries.

[0546] "360-degree omnidirectional video data" is video data that covers the field of view in all directions, and is in a format that allows the user to freely change the viewpoint in any direction.

[0547] "Encoding" is the process of compressing video data using a certain format or codec and converting it into a transmittable form.

[0548] A "packet" refers to each unit of data divided when digital information is transmitted over a network, and each packet is accompanied by transmission route information and error check information.

[0549] A "terminal" is a device that the user directly operates, which in this case refers to VR goggles.

[0550] "Decoding" is the process of restoring encoded video data to its original format so that the device can play the video.

[0551] A "gyro sensor" is a sensor that detects the angular velocity and rotation of a device and is used to accurately capture the user's head movements.

[0552] "Rendering in real time" refers to the process of generating and displaying images in real time in response to the user's viewpoint and movements.

[0553] "Sound effects" refers to the audio and sound effects that accompany the video, and enhance the sense of realism by adjusting the sense of direction and distance according to the user's movements.

[0554] The present invention relates to a system that uses VR technology to enhance the sense of realism when watching a game. Hereinafter, an embodiment of the present invention will be described in detail.

[0555] Server Processing

[0556] The server collects video data from multiple camera devices located within the venue. These include fixed cameras, mobile cameras, and drone cameras. The collected video data is first stored in a buffer. Each piece of video data is assigned a timestamp, which is used to sort the data in chronological order. The server then automatically recognizes overlapping areas of the video data and seamlessly combines them. This process uses an image processing library (e.g., OpenCV).

[0557] Specifically, the server uses an edge detection algorithm (e.g., SIFT) to identify adjacent overlapping areas and integrate them into one continuous video. This integrated 360-degree panoramic video data is then encoded in real time using the H.264 or H.265 codec. The encoded data is then divided into packets of a fixed size and sent to the device using the UDP protocol. The transmission process uses the ffmpeg library.

[0558] Terminal processing (VR goggle processing)

[0559] The terminal receives real-time video data packets sent from the server. The received packets are stored in a temporary buffer and reconstructed based on the sequence numbers. The terminal then decodes the video data using a decoding library (e.g., libavcodec). The decoded video data is stored in the temporary buffer again.

[0560] The device then renders this video data in real time according to the user's viewpoint. A game engine (e.g., Unity or Unreal Engine) is used for rendering. The device uses a built-in gyro sensor to detect the user's head movements and adjusts the video viewpoint and sound effects in real time based on those movements. An audio engine (e.g., OpenAL or FMOD) is used to adjust the sound effects.

[0561] User operations

[0562] First, users put on VR goggles and log in to the system. This process is done using facial recognition, ID, and password. Users can freely change their viewpoint by moving their head. They can view 360-degree images and select and watch their favorite scenes in real time.

[0563] Specific examples

[0564] For example, when streaming a basketball game, the server collects footage from multiple cameras installed around the stadium. The collected footage is then sorted chronologically based on timestamps, overlapping parts are recognized and seamlessly joined, and then encoded and sent to the device.

[0565] The device decodes the received video data and renders it in real time according to the user's viewpoint. Users can change the viewpoint by moving their head, for example, to a courtside perspective, and experience the players' movements and plays in real time. Sound effects are also adjusted in real time, making users feel as if they are at the game venue.

[0566] Prompt Sentence Examples

[0567] The following are examples of prompt sentences to input into a generative AI model:

[0568] "Please explain in detail the terminal processing of the game viewing system using VR technology. In particular, please provide details on the reception and reassembly of data packets, decoding and buffering, real-time rendering based on the user's viewpoint, and adjustment of sound effects."

[0569] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0570] Server Processing

[0571] Step 1: Collect video data

[0572] The server collects video data in real time from multiple camera devices (fixed cameras, mobile cameras, drone cameras, etc.) installed within the venue. These cameras assign a timestamp to each video frame and send it to the server. The server temporarily stores the received video data in a buffer. The input is real-time video data from the camera devices, and the output is the video data stored in the buffer.

[0573] Step 2: Adjust the timeline of the video data

[0574] The server sorts the video data stored in the buffer into chronological order based on the timestamps, ensuring accurate synchronization of the video from each camera. The input is time-stamped video data, and the output is video data sorted in chronological order.

[0575] Step 3: Seamlessly stitch together footage

[0576] The server automatically recognizes overlapping areas of the video data from each camera and seamlessly combines them. This process uses an image processing library (e.g., OpenCV) and utilizes an edge detection algorithm (e.g., SIFT). The input is multiple video data sorted in chronological order, and the output is seamlessly combined 360-degree omnidirectional video data.

[0577] Step 4: Encode and transmit video data

[0578] The server encodes the combined 360-degree panoramic video data in real time using the H.264 or H.265 codec. The encoded data is then divided into packets of a fixed size and sent to the device using the UDP protocol. The input is the seamlessly combined 360-degree panoramic video data, and the output is the encoded data packets sent to the device.

[0579] Terminal processing (VR goggle processing)

[0580] Step 1: Receiving and reassembling data packets

[0581] The terminal receives data packets sent from the server. The received packets are stored in a temporary buffer and reassembled into the correct order based on the sequence numbers. The input is the data packets from the server, and the output is the assembled video data.

[0582] Step 2: Decoding and Buffering

[0583] The device decodes the reconstructed video data. This process uses a decoding library (e.g., libavcodec) and may utilize hardware acceleration. The decoded video is again stored in a temporary buffer. The input is the reconstructed video data, and the output is the decoded video data.

[0584] Step 3: Real-time rendering based on user perspective

[0585] The device uses a built-in gyro sensor to capture the user's head movements and renders video data in real time based on those movements. A 360-degree panoramic image is displayed according to the user's viewpoint. This process uses a game engine (e.g., Unity or Unreal Engine). The input is the decoded video data and gyro sensor movement data, and the output is a rendered image according to the user's viewpoint.

[0586] Step 4: Real-time sound effect adjustment

[0587] The device adjusts the sound effects in real time according to the user's head movements, so that the sound has a sense of direction according to the user's movements. This process uses an acoustic engine (e.g., OpenAL or FMOD). The input is the user's movement data, and the output is the adjusted sound effects.

[0588] User operations

[0589] Step 1: Put on the VR goggles and log in to the system

[0590] The user puts on the VR goggles and logs in to the system. The login process uses facial recognition, ID, and password. The input is the user's authentication information, and the output is the system login status.

[0591] Step 2: Freely change the viewpoint

[0592] Users can freely change their viewpoint by moving their head, allowing them to view a 360-degree panoramic view and select and watch their favorite scenes in real time. The input is the user's head movement data, and the output is a rendered image based on the user's viewpoint.

[0593] (Application example 1)

[0594] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0595] Conventional game viewing systems have limited means for enhancing the sense of realism, limiting the ability to experience the game from specific viewpoints and angles. It has also been difficult for users to freely change viewpoints or adjust video and audio effects in real time. This has prevented users from experiencing the game as if they were actually at the venue.

[0596] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0597] In this invention, the server includes means for collecting video data from multiple camera devices within the venue, means for integrating the collected video data to generate 360-degree omnidirectional video data, means for transmitting the generated 360-degree omnidirectional video data to the terminal in real time, means for the terminal to receive the video data and provide a 360-degree omnidirectional field of view and adjust the video and sound effects according to the user's movements, means for the terminal to detect the user's movements using a built-in gyro sensor and adjust the video and sound effects based on the movements in real time, and means for temporarily storing the collected video data in a buffer and realizing seamless playback. This allows users to experience a realistic game viewing experience in real time and to freely change the viewpoint to watch from different viewpoints within the venue.

[0598] "Game venue" means a particular location where a sport or event is held.

[0599] "Filming equipment" refers to equipment used to collect video data, including fixed cameras, mobile cameras, and drone cameras.

[0600] "Video data" refers to visual information captured by a camera within the venue.

[0601] "Integration" refers to the process of combining multiple pieces of video data and reconstructing them into a single, continuous image.

[0602] "360-degree panoramic video" refers to video data that covers all directions within the venue, and has the ability to freely change the viewpoint.

[0603] "Real-time" refers to a situation in which information is processed almost instantly, with little time delay.

[0604] "Terminal" refers to a device that a user uses to watch video, including smartphones and head-mounted displays.

[0605] "Providing" refers to the act of making a particular service or feature available to a user.

[0606] "Movement" refers to the user's actions of moving their head or body.

[0607] "Sound effects" refers to audio elements such as voice and music that are provided when watching a video.

[0608] "Adjustment" refers to the act of changing or modifying functions or effects to suit specific conditions or environments.

[0609] A "gyro sensor" is a sensor for measuring angular velocity and is used to detect the movement of the user's head.

[0610] "Decoding" refers to the process of restoring encoded data to its original form.

[0611] "Rendering" refers to the process of visually displaying video data.

[0612] A "temporary buffer" refers to a memory area for temporarily storing data.

[0613] "Seamless" refers to playback or joining that is performed without interruption, with continuity maintained.

[0614] This invention relates to a system that uses VR technology to enhance the sense of realism when watching a match. This system collects video data from multiple camera devices within the match venue and provides it to users in real time, allowing them to enjoy the match from a 360-degree, all-around perspective.

[0615] Server Processing

[0616] The server collects video data from multiple camera devices installed within the venue, including fixed cameras, mobile cameras, and drone cameras. The collected video data is first stored in a buffer, and the footage from each camera is sorted chronologically. Next, overlapping areas of the footage are automatically recognized and seamlessly combined, generating 360-degree video data. The generated 360-degree video data is then encoded in real time, divided into packets of a fixed size, and sent to the device. This transmission is performed with low latency, allowing users to experience the video smoothly.

[0617] Terminal handling

[0618] The device (e.g., a smartphone or head-mounted display) receives real-time video data sent from the server. The device reassembles the received data packets and stores them in a temporary buffer. It then decodes the video data and renders the video in real time according to the user's movements. The device uses its built-in gyro sensor to acquire data on the user's head movement and adjusts the video and sound effects in real time based on that movement. This allows the user to enjoy a visual and auditory experience that makes them feel as if they are at the stadium.

[0619] User operations

[0620] First, users put on the device and log in to the system. They can freely change their viewpoint by moving their head, looking around in all 360 degrees to watch any scene they like in real time. For example, users can stand in the perspective of a specific player and experience important moments of the game up close. Users can also select different viewpoints to enjoy views from different locations within the stadium.

[0621] Specific examples

[0622] For example, when streaming a basketball game, the process goes like this: The server collects footage from multiple cameras installed within the venue, automatically recognizes overlapping areas of the footage, and seamlessly combines them. This combined video data is encoded in real time and sent to the device. The device receives the video data sent from the server and stores it in a temporary buffer. It then decodes it and renders the video in real time according to the user's head movements. The built-in gyro sensor adjusts the user's viewpoint, and sound effects are also adjusted in real time. For example, you can switch to a courtside view to experience the players' movements and plays in real time.

[0623] Prompt Sentence Examples

[0624] "I want to develop a VR application that allows users to freely watch soccer matches from a 360-degree perspective. Please tell me specifically how to collect video data from a server in real time and render the video on an HMD according to the user's viewpoint."

[0625] Thus, according to the present invention, the user can experience watching a game with a sense of real-time presence, and by freely changing the viewpoint, the user can watch the game from different viewpoints within the venue.

[0626] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0627] Step 1:

[0628] The server collects video data from multiple camera devices installed within the venue.

[0629] Input: Video data from fixed, mobile, and drone cameras.

[0630] Output: Collected multi-camera video data.

[0631] Specific operation: The server continuously receives video data from each imaging device and temporarily stores it in a buffer.

[0632] Step 2:

[0633] The server integrates the collected video data and generates 360-degree panoramic video data.

[0634] Input: Single camera video data (multiple).

[0635] Output: Integrated 360° video data.

[0636] Specific operation: The time series of video data is organized and overlapping parts are seamlessly combined using an automatic recognition system (computer vision algorithm).

[0637] Step 3:

[0638] The server encodes the generated 360-degree omnidirectional video data in real time, divides it into packets of a certain size, and transmits it to the terminal.

[0639] Input: Integrated 360° video data.

[0640] Output: The encoded data packet.

[0641] Specific operation: Video data is encoded and compressed, divided into packets, and sent to the terminal via a low-latency network.

[0642] Step 4:

[0643] The terminal receives real-time video data transmitted from the server and stores it in a temporary buffer.

[0644] Input: The encoded data packet.

[0645] Output: Data packets stored in a temporary buffer.

[0646] Specific operation: Starts the process of receiving data from the network and storing it in a temporary buffer.

[0647] Step 5:

[0648] The terminal decodes the data packets stored in the temporary buffer and assembles them into a video.

[0649] Input: Data packets stored in a temporary buffer.

[0650] Output: Decoded video data.

[0651] Specific operation: Using a decoding algorithm, the packets are converted into the original video data and a continuous video is reassembled.

[0652] Step 6:

[0653] The device uses a built-in gyro sensor to detect the user's head movements and adjusts visual and sound effects in real time based on those movements.

[0654] Input: Motion data from gyro sensor, decoded video data.

[0655] Output: Visual and sound effects adjusted to the user's point of view.

[0656] How it works: Using gyro sensor data, it tracks the user's head movements and dynamically adjusts the viewpoint and sound in real time.

[0657] Step 7:

[0658] The user wears the device and logs in to the system. While watching the captured video, the user can freely change the viewpoint by moving their head.

[0659] Input: Input data from the user interface.

[0660] Output: The visual and audio data that the user sees.

[0661] Specific operation: After the user wears the device and logs in to the system, they can watch the game while the video and audio are changed in real time based on their head movements.

[0662] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0663] The present invention relates to a system that uses VR technology and emotion recognition technology to enhance the sense of realism when watching a game. Hereinafter, an embodiment of the present invention will be described in detail.

[0664] Server Processing

[0665] The server collects video data from multiple camera systems within the venue, including fixed cameras, mobile cameras, and even drone cameras. The collected video data is buffered and the video data from each camera is sorted chronologically. A process then automatically recognizes overlapping areas of the video and seamlessly combines them to generate 360-degree video data.

[0666] The generated 360-degree video data is encoded in real time, divided into packets of a certain size, and transmitted to the device with low latency, allowing users to enjoy a smooth experience.

[0667] Terminal processing (VR goggle processing)

[0668] The device (VR goggles) receives real-time video data sent from the server. The device reassembles the received data packets and stores them in a temporary buffer as video data. Next, the device decodes the video data and renders the video in real time according to the user's viewpoint.

[0669] The device uses a built-in gyro sensor to capture the user's head movement data and adjusts the visual and audio effects in real time based on that movement, providing users with a visual and audio experience that makes them feel like they're at the stadium.

[0670] Furthermore, the device includes an emotion engine that uses a camera to detect the user's facial expressions and biometric signals (e.g., heart rate, brain waves, etc.) and recognizes the user's emotions in real time based on the detected emotions. The recognized emotion data is then sent to a server.

[0671] Emotion data processing on the server

[0672] The server receives the user's emotional data sent from the device. Based on this emotional data, the server can automatically adjust the video and sound effects. For example, if the user is excited, it can emphasize the display of important scenes and enhance the sound effects.

[0673] User operations

[0674] First, users put on VR goggles and log in to the system. Users can freely change their viewpoint by moving their head, looking around in all 360 degrees and watching their favorite scenes in real time. Furthermore, camera angles and content are automatically selected according to the user's emotions, giving users an even more immersive experience.

[0675] Specific examples

[0676] For example, if you were streaming a basketball game, it would look like this:

[0677] Server processing: Images are collected from multiple cameras installed within the venue, temporarily stored in a buffer, and then sorted chronologically. Next, overlapping areas of the images are automatically recognized and seamlessly combined to generate a 360-degree panoramic image. The image is then encoded, divided into packets, and sent to the terminal.

[0678] Device processing: The device receives and reassembles data packets and decodes the video data. It renders the video in real time according to the user's viewpoint and detects head movements using the built-in gyro sensor. It adjusts the video and sound effects in real time. The device's camera also detects the user's facial expressions and biometric signals, analyzes the data using the emotion engine, and sends it to the server as emotion data.

[0679] Server emotion data processing: The server receives the emotion data and adjusts the video and sound effects in real time according to the user's emotion. For example, if the user is excited, the video will be adjusted to emphasize particularly dynamic scenes or important plays.

[0680] User operation: The user puts on the VR goggles and logs in. They can freely change the viewpoint by moving their head, enjoying an interactive experience of watching specific scenes. Furthermore, camera angles and specific content are automatically selected according to the user's emotions, enhancing the sense of realism. For example, if the user is feeling nervous, they can receive feedback such as scenes and angles that will help them relax.

[0681] This allows users to experience the realistic sensation of being at a game venue from the comfort of their own home, and furthermore, to enjoy a visual and auditory experience that is customized to suit their individual emotions.

[0682] The processing flow will be explained below.

[0683] Server Processing

[0684] Step 1: Collect video data

[0685] The server receives video data from multiple camera devices installed within the venue, including fixed, mobile and drone cameras, each capturing the match from a different perspective or angle.

[0686] Step 2: Buffering video data

[0687] The server temporarily stores the received video data in a buffer, whereby the video data from each camera is temporarily stored.

[0688] Step 3: Synchronize and sort data

[0689] The server rearranges the video data in the temporary buffer in chronological order and synchronizes them based on the timestamps.

[0690] Step 4: Video data integration

[0691] The server then combines the sorted video data from multiple cameras, automatically recognizing overlapping areas and seamlessly joining them together.

[0692] Step 5: Encode the video data

[0693] The server encodes the integrated 360-degree panoramic video data in real time.

[0694] Step 6: Packetize the data

[0695] The server divides the encoded video data into packets of a fixed size and prepares them for transfer.

[0696] Step 7: Sending data

[0697] The server then transmits the divided data packets to the terminal via a high-speed network, aiming for low latency.

[0698] Step 8: Receiving emotion data

[0699] The server receives the user's emotional data transmitted from the device, including emotional information analyzed from the user's facial expressions and biometric signals.

[0700] Step 9: Adjust based on sentiment data

[0701] The server adjusts the visual and audio effects in real time based on the received emotion data, selecting specific camera angles and content to display according to the user's emotion.

[0702] Terminal (VR goggles) processing

[0703] Step 1: Receiving a data packet

[0704] The terminal receives the data packets sent from the server through a high-speed network.

[0705] Step 2: Reassembling the data packets

[0706] The terminal reassembles the received data packets into the original video data.

[0707] Step 3: Decoding the video data

[0708] The terminal decodes the reassembled video data and prepares it for display.

[0709] Step 4: Rendering the Perspective

[0710] Based on the decoded video data, the device renders a 360-degree panoramic image tailored to the user's viewpoint in real time.

[0711] Step 5: Acquire motion data

[0712] The device uses a built-in gyro sensor to detect the user's head movements.

[0713] Step 6: Adjusting the image and sound

[0714] The device adjusts the video and audio effects in real time based on the acquired motion data, providing a sense of realism in both visual and audio.

[0715] Step 7: Obtaining emotion data

[0716] The device uses cameras and sensors to detect the user's facial expressions and biometric signals, recognizing the user's emotions in real time.

[0717] Step 8: Sending Emotion Data

[0718] The terminal transmits the recognized emotion data of the user to the server.

[0719] User operations

[0720] Step 1: Put on the VR goggles and log in

[0721] The user puts on the VR goggles and logs into the system. To log in, the user must enter their account information.

[0722] Step 2: Change your perspective

[0723] Users can freely change their viewpoint by moving their head, enjoying a 360-degree view, and can also focus on a specific player or play.

[0724] Step 3: Interactive Experience

[0725] Users receive real-time visual and auditory feedback, making them feel as if they are at the stadium, with sounds such as the cheers of the crowd and the footsteps of the players reproduced in real time.

[0726] Step 4: Displaying emotions

[0727] Depending on the user's emotions, specific camera angles and content can be automatically selected and displayed. For example, if the user is excited, important scenes can be displayed more prominently and sound effects can be enhanced.

[0728] Example 2

[0729] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0730] Conventional game viewing systems have limited visual and auditory experiences, making it difficult to fully recreate the sense of presence of a game. Furthermore, it is not possible to optimize the content played based on the user's real-time emotions and movements, making it difficult to provide a personalized experience for each user.

[0731] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0732] In this invention, the server includes means for collecting video data from multiple camera devices within the venue, means for integrating the collected video data to generate 360-degree omnidirectional video data, and means for transmitting the generated 360-degree omnidirectional video data to the terminal in real time, thereby enabling the user to experience 360-degree omnidirectional real-time video, and for the terminal to receive the video data and provide a 360-degree omnidirectional field of view, and to adjust the video and sound effects according to the user's movements.

[0733] Furthermore, by including a means for the terminal to decode the video data sent from the server and render the video in real time according to the user's viewpoint, and a means for acquiring the user's emotional data and adjusting the video and sound effects based on that, it is possible to provide a more immersive and individually optimized visual and auditory experience.

[0734] A "game venue" is a specific location where a sport or event is played.

[0735] "Filming equipment" refers to equipment such as cameras and drone cameras used to collect video data.

[0736] "Video data" refers to video information collected from a camera and includes visual content.

[0737] "Collection" refers to the act of collecting video data from multiple imaging devices.

[0738] "Integration" is the process of bringing together collected video data.

[0739] "360-degree omnidirectional video data" is video data that provides visual information from all directions.

[0740] "Generation" is the act of creating new video data.

[0741] "Real-time" refers to almost instantaneous processing, with very little delay.

[0742] A "terminal" is a device that receives and displays video data, such as VR goggles.

[0743] "Transmitting" is the act of sending data to another device or system.

[0744] "Rendering" is the process of generating images or videos from digital data.

[0745] A "gyro sensor" is a sensor that detects the movement of the user's head.

[0746] "Emotion data" is data that indicates the emotional state of the user.

[0747] "Decoding" is the process of converting encoded data into a playable format.

[0748] "Adjustment" is the act of changing visual and sound effects based on the user's movements and emotions.

[0749] The present invention relates to a system that uses VR technology and emotion recognition technology to enhance the sense of realism when watching a game. Specific embodiments of the present invention will be described below.

[0750] The server collects video data from multiple camera devices within the venue. This includes fixed cameras, mobile cameras, and drone cameras. The video data from each camera is temporarily stored in a buffer and sorted chronologically. The server then automatically recognizes overlapping areas of the footage and seamlessly combines them to generate 360-degree panoramic video data.

[0751] The generated 360-degree video data is encoded in real time using codecs such as H.264 or H.265. The encoded data is divided into packets of a fixed size and transmitted to the device with low latency.

[0752] The device (VR goggles) receives video data packets sent from the server. The received data is stored in a buffer and decoded. The decoded video data is rendered in real time according to the user's viewpoint. The device uses a built-in gyro sensor to detect the user's head movement and adjusts the video and sound effects based on that movement.

[0753] Furthermore, the device is equipped with a camera and biometric sensors to detect the user's facial expressions and biometric signals (e.g., heart rate, brain waves). This data is analyzed by the emotion engine and sent to the server as user emotion data.

[0754] The server receives the emotion data sent from the terminal and adjusts the video and sound effects based on the user's emotional state. For example, if the user is excited, the server adjusts the video to emphasize important scenes or dynamic play.

[0755] Users can use the above services by wearing VR goggles and logging in to the system. Users can freely change their viewpoint by moving their head, enjoying a 360-degree panoramic view. In addition, camera angles and content are automatically selected based on the user's emotions, creating an even more immersive experience.

[0756] Specific examples

[0757] For example, when streaming a basketball game, the following steps are taken: Video data is collected from multiple cameras installed in the stadium and stored in a temporary buffer. The images from each camera are rearranged in chronological order and seamlessly combined to generate a 360-degree panoramic video. The video is then encoded, divided into packets, and sent to the device.

[0758] The device reassembles the received data packets and renders the decoded video data in real time. The built-in gyro sensor detects the user's head movements and adjusts the video and audio in real time. The device's camera and biometric sensors also detect the user's facial expressions and biometric signals, and the emotion engine analyzes the emotional data and transmits it to the server.

[0759] The server receives the emotion data and adjusts the video and sound effects according to the user's emotions. For example, if the user is excited, the video can be adjusted to emphasize dynamic scenes and important plays, allowing the user to enjoy the game more realistically.

[0760] Prompt Sentence Examples

[0761] "I would like to develop a system that uses VR technology to make watching basketball games more immersive, and also utilizes emotion recognition technology to adjust video and sound effects. This system will incorporate a mechanism to optimize video in real time according to the user's movements and emotions. Please tell me specifically how you will design the system and how you will use emotion data."

[0762] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0763] Step 1:

[0764] "Video data collection"

[0765] The server collects video data from fixed cameras, mobile cameras, drone cameras, etc. within the venue. The video data from each camera is temporarily stored in the server's buffer.

[0766] Input: Video streams from each imaging modality.

[0767] Output: Multiple video data stored in a buffer.

[0768] Specific operation: The server uses the stream URL of the camera registered in advance to pull in video data using a specified protocol (such as RTSP).

[0769] Step 2:

[0770] "Sorting time series data"

[0771] The server then sorts the video data stored in the buffer into chronological order based on the timestamps of each camera, thereby synchronizing the video data.

[0772] Input: Multiple video data stored in a buffer.

[0773] Output: Video data sorted in chronological order.

[0774] What it does: The server analyzes the timestamp information and runs an algorithm to arrange the video frames from each camera in the proper order.

[0775] Step 3:

[0776] "Seamless joining process"

[0777] The server automatically recognizes overlapping areas of the images and seamlessly combines them to generate 360-degree panoramic video data.

[0778] Input: Multiple video data sorted in chronological order.

[0779] Output: 360-degree panoramic video data.

[0780] How it works: The server detects overlapping areas and uses image processing algorithms to merge them together, creating a continuous omnidirectional image.

[0781] Step 4:

[0782] "Encoding 360-degree video"

[0783] The server encodes the generated 360-degree panoramic video data in real time using the H.264 or H.265 codec.

[0784] Input: 360-degree omnidirectional video data.

[0785] Output: Encoded video data.

[0786] Specific operation: The server inputs the video data into the codec, which performs compression and encoding processes, thereby reducing the data size appropriately and making transmission more efficient.

[0787] Step 5:

[0788] "Sending video packets"

[0789] The server divides the encoded video data into packets of a fixed size and transmits them to the terminal with low latency.

[0790] Input: Encoded video data.

[0791] Output: Split data packets.

[0792] Specific operation: The server packetizes the video data and transmits it to the terminal using the UDP or TCP protocol.

[0793] Step 6:

[0794] "Receiving Data Packets"

[0795] The terminal receives the video data packets sent from the server, checking for packet loss and duplication to ensure accurate reception.

[0796] Input: Data packet from the server.

[0797] Output: Data packets that have been received and processed.

[0798] Specific operation: The terminal performs packet reception processing at the network layer, and processes error checks and retransmission requests.

[0799] Step 7:

[0800] "Assembling Data Packets"

[0801] The received data packets are reassembled into the original video data and stored in a temporary buffer.

[0802] Input: Received data packets.

[0803] Output: Video data stored in a temporary buffer.

[0804] Specific operation: The terminal rearranges, reconstructs, and temporarily stores the data packets, referring to the timestamp information.

[0805] Step 8:

[0806] "Video data decoding"

[0807] The terminal decodes the video data in the buffer using an H.264 or H.265 decoder and acquires it as successive video frames.

[0808] Input: Video data stored in a temporary buffer.

[0809] Output: Decoded video frames.

[0810] Specific operation: The terminal decodes the video frame using dedicated decoding hardware or software.

[0811] Step 9:

[0812] "Video rendering"

[0813] The decoded video is rendered in real time according to the user's viewpoint and displayed in the VR goggles.

[0814] Input: Decoded video frames.

[0815] Output: Rendered video frames.

[0816] Specific operation: The terminal uses a rendering engine to generate video frames according to the user's viewpoint and displays them on the display.

[0817] Step 10:

[0818] "Acquiring head movement data"

[0819] The device's built-in gyro sensor detects the user's head movements and uses that data to adjust visual and sound effects in real time.

[0820] Input: User's head movement.

[0821] Output: Coordinated video and audio.

[0822] Specific operation: The device obtains data in real time from the gyro sensor and dynamically adjusts image rendering and audio filtering.

[0823] Step 11:

[0824] "Acquiring emotion data"

[0825] The device's camera and biometric sensors capture biometric data such as the user's facial expressions and heart rate, which is then analyzed by the emotion engine.

[0826] Input: User's facial expression data and biometric signals.

[0827] Output: Parsed emotion data.

[0828] Specific operation: The device's camera extracts feature points using a facial expression recognition algorithm, and biometric sensors measure heart rate and skin galvanic response, which are then comprehensively analyzed by the emotion engine.

[0829] Step 12:

[0830] "Sending emotional data"

[0831] The analyzed emotional data is sent to the server in real time.

[0832] Input: Parsed emotion data.

[0833] Output: Emotion data sent to the server.

[0834] Specific operation: The device encodes the emotion data into an appropriate format and makes a transmission request to the server.

[0835] Step 13:

[0836] "Receiving and adjusting emotional data"

[0837] The server receives the emotional data sent from the terminal and adjusts the visual and sound effects based on the user's emotional state.

[0838] Input: Emotion data sent from the device.

[0839] Output: Coordinated video and audio.

[0840] Specific operation: The server analyzes the emotional data and dynamically generates content optimized for the user's emotional state, emphasizing specific visual scenes and sound effects.

[0841] Through the above processing steps, the system can provide the user with a realistic sense of being at the game, and can also provide a customized visual and auditory experience according to the user's emotional state.

[0842] (Application example 2)

[0843] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0844] While conventional VR-based game viewing systems provide users with a high sense of realism, they have difficulty providing a personalized experience that reflects the emotional and physical state of each user. Therefore, a new system is needed that can provide appropriate visual and audio effects in real time according to the emotional state and preferences of various users.

[0845] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for collecting video data from multiple camera devices within the venue, means for integrating the collected video data to generate 360-degree omnidirectional video data, and means for transmitting the generated 360-degree omnidirectional video data to the terminal in real time. This allows the user to experience the immersive atmosphere of the venue in real time. The terminal also includes means for adjusting video and audio effects according to the user's movements, means for sensing the user's facial expressions and biometric signals and recognizing the user's emotions in real time based on the detected facial expressions and biometric signals, and means for automatically adjusting video and audio effects based on the recognized emotional data. This allows the user to enjoy a personalized visual and audio experience that is tailored to their emotions and physical state.

[0846] "Game venue" refers to a location where various sports and entertainment matches are played.

[0847] "Filming equipment" refers to equipment such as cameras and drones used to collect video data.

[0848] "Video data" refers to video data collected by an imaging device.

[0849] "360-degree omnidirectional video data" refers to video that can be viewed in all directions and is generated by integrating video data collected from multiple imaging devices.

[0850] "Terminal" refers to a device (e.g., VR goggles) that receives video data and provides a visual and auditory experience to the user.

[0851] "User movement" refers to the user's body movement, particularly head movement, detected by the terminal.

[0852] "Sound effects" refers to the audio data corresponding to the video and the method of reproducing it.

[0853] "Expression" refers to the facial expression of the user.

[0854] "Biological signals" refer to physiological data such as a user's heart rate and brain waves.

[0855] "Emotion data" refers to the emotional state recognized based on the user's facial expressions and biometric signals.

[0856] "Real-time" refers to near-instant processing and transmission.

[0857] A "gyro sensor" refers to an inertial sensor used to detect the movement of a device.

[0858] "Synthesis" refers to the process of combining collected data into one linked data set.

[0859] "Seamlessly combining" refers to combining different video data continuously without interruption.

[0860] "Automatic adjustment" refers to the system automatically adjusting the video and audio based on the user's emotional data.

[0861] This invention relates to a system that enhances the sense of realism when a user watches a game using VR goggles and automatically adjusts video and sound effects according to the emotions of each individual user. The specific configuration and operation of the system are described below.

[0862] Server Processing

[0863] The server collects video data from multiple camera devices within the venue. The camera devices include fixed cameras, mobile cameras, and drone cameras. The collected video data is first stored in a buffer, and the video data from each camera is sorted chronologically. The server then automatically recognizes overlapping areas of the video and seamlessly combines them to generate 360-degree panoramic video data. The generated video data is then encoded in real time, divided into packets of a fixed size, and transmitted to the device. This transmission is performed with low latency, ensuring a smooth user experience.

[0864] Terminal handling

[0865] The device (VR goggles) receives real-time video data sent from the server. It reassembles the received data packets and stores them in a temporary buffer. It then decodes the video data and renders the video in real time according to the user's viewpoint. The device uses a built-in gyro sensor to acquire data on the user's head movement and adjusts the video and sound effects in real time based on that movement. The device also uses a camera to detect the user's facial expressions and biometric signals (e.g., heart rate, brain waves, etc.), and includes an emotion engine that recognizes the user's emotions in real time based on this data. The recognized emotion data is sent to the server in real time.

[0866] Emotion data processing on the server

[0867] The server receives the user's emotional data sent from the device. Based on this emotional data, it can automatically adjust the visual and sound effects. For example, if the user is excited, it can highlight important scenes and thrilling plays and enhance the sound effects.

[0868] User operations

[0869] First, users put on VR goggles and log in to the system. Users can freely change their viewpoint by moving their head, and enjoy a 360-degree panoramic view of their favorite scenes in real time. Furthermore, camera angles and content are automatically selected based on the user's emotions, creating an even more immersive experience.

[0870] Hardware and software used

[0871] The system is implemented using the following hardware and software:

[0872] Hardware: VR goggles (e.g., Oculus Rift, HTC Vive), fixed cameras, movable cameras, drone cameras, heart rate monitors, electroencephalographs

[0873] Software: OpenCV (used for image acquisition and processing), Python socket library (used for sending and receiving image data), Unity or Unreal Engine (for VR rendering and user interface construction)

[0874] Specific examples

[0875] For example, when live streaming a soccer match, the following prompt sentence can be input into the generative AI model to generate video that highlights important scenes from the match.

[0876] "Analyze the user's emotion recognition data in real time during a soccer match. If the user is excited, highlight the goal scene and the play before the goal. If the user is nervous, switch to a more relaxing video."

[0877] This invention allows users to enjoy a highly realistic and personalized viewing experience.

[0878] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0879] Step 1:

[0880] The server collects video data from multiple camera devices (fixed cameras, mobile cameras, drone cameras) within the venue. The collected video data is temporarily stored in a buffer, and the video data from each camera is sorted chronologically. This ensures that the video data is arranged in the appropriate order.

[0881] Step 2:

[0882] The server automatically recognizes overlapping areas in the collected video data and seamlessly combines them. It uses OpenCV to detect overlapping areas in each frame and combines them continuously. The combined video data becomes 360-degree omnidirectional video data.

[0883] Step 3:

[0884] The server encodes the generated 360-degree omnidirectional video data in real time, divides it into packets of a fixed size, and transmits them to the device. This transmission is performed with low latency, and the data is sent to the device via the network. Video codecs such as H.264 and HEVC are used for encoding.

[0885] Step 4:

[0886] The device (VR goggles) receives real-time video data sent from the server, reassembles the received data packets, and stores them in a temporary buffer. Then, it decodes the video data. For decoding, it uses FFmpeg or other decoding libraries.

[0887] Step 5:

[0888] The device uses a built-in gyro sensor to capture data on the user's head movements, and uses this data to render video and sound effects in real time. Rendering is done using Unity or Unreal Engine, and the video is displayed according to the user's viewpoint.

[0889] Step 6:

[0890] The device uses a camera and biosensors to detect the user's facial expressions and biometric signals (heart rate, brain waves, etc.). It then uses an emotion engine to analyze this data and extract the user's emotional data. For example, it uses a machine learning model to infer emotions from facial expressions and biometric signals.

[0891] Step 7:

[0892] The device transmits the recognized emotion data to the server in real time with low latency, so the user's emotional state is immediately conveyed to the server.

[0893] Step 8:

[0894] The server automatically adjusts visual and audio effects based on the received user emotion data. For example, if the user is excited, it will highlight important scenes or thrilling action scenes and enhance audio effects. It uses a generative AI model to calculate the appropriate visual and audio settings in real time.

[0895] Step 9:

[0896] Users put on VR goggles, log in, and begin watching the game. They can freely change their viewpoint by moving their head, and enjoy a 360-degree panoramic view of the game to watch any scene they like. Furthermore, camera angles and content are automatically selected based on the user's emotions, creating an even more immersive experience.

[0897] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0898] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0899] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0900] [Third embodiment]

[0901] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0902] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0903] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0904] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0905] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0906] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0907] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0908] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0909] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0910] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0911] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0912] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0913] The present invention relates to a system that uses VR technology to enhance the sense of realism when watching a game. Hereinafter, an embodiment of the present invention will be described in detail.

[0914] Server Processing

[0915] The server collects video data from multiple camera systems within the venue, including fixed cameras, mobile cameras, and even drone cameras. The collected video data is buffered and the footage from each camera is sorted chronologically. A process then automatically recognizes overlapping areas of the footage and seamlessly combines them to generate 360-degree panoramic video data.

[0916] The generated 360-degree video data is encoded in real time, divided into packets of a certain size, and transmitted to the device with low latency, allowing users to experience the video smoothly.

[0917] Terminal processing (VR goggle processing)

[0918] The device (VR goggles) receives real-time video data sent from the server. The device reassembles the received data packets and stores them in a temporary buffer as video data. Next, the device decodes the video data and renders the video in real time according to the user's viewpoint.

[0919] The device uses a built-in gyro sensor to capture the user's head movement data and adjusts the visual and audio effects in real time based on that movement, providing users with a visual and audio experience that makes them feel like they're at the stadium.

[0920] User operations

[0921] First, users put on VR goggles and log in to the system. Users can freely change their viewpoint by moving their head, and enjoy a 360-degree panoramic view of their favorite scenes in real time. For example, they can stand in the perspective of their favorite player and experience the moment they make a shot up close.

[0922] Specific examples

[0923] For example, if you were streaming a basketball game, it would look like this:

[0924] Server processing: Images are collected from multiple cameras installed within the venue, and overlapping areas of the images are automatically recognized and seamlessly combined. This combined video data is encoded in real time and sent to the terminal.

[0925] Device processing: The device receives the video data sent from the server and stores it in a temporary buffer. It then decodes it and renders the video in real time according to the user's head movements. It uses the built-in gyro sensor to adjust the user's viewpoint and adjusts the sound effects in real time.

[0926] User operation: The user puts on the VR goggles and logs in to the system. The user can change the viewpoint by moving their head and enjoy 360-degree omnidirectional images. For example, they can switch to a courtside viewpoint and experience the players' movements and plays in real time.

[0927] This allows users to experience the immersive atmosphere of a live game venue from the comfort of their own home.

[0928] The processing flow will be explained below.

[0929] Server Processing

[0930] Step 1: Collect video data

[0931] The server receives video data from multiple camera devices installed within the venue, including fixed, mobile and drone cameras, each capturing the match from a different perspective or angle.

[0932] Step 2: Buffering video data

[0933] The server temporarily stores the received video data in a buffer, whereby the video data from each camera is temporarily stored.

[0934] Step 3: Synchronize and sort data

[0935] The server rearranges the video data in the temporary buffer in chronological order and synchronizes them based on the timestamps.

[0936] Step 4: Video data integration

[0937] The server then combines the sorted video data from multiple cameras, automatically recognizing overlapping areas and seamlessly joining them together.

[0938] Step 5: Encode the video data

[0939] The server encodes the integrated 360-degree panoramic video data in real time.

[0940] Step 6: Packetize the data

[0941] The server divides the encoded video data into packets of a fixed size and prepares them for transfer.

[0942] Step 7: Sending data

[0943] The server then transmits the divided data packets to the terminal via a high-speed network, aiming for low latency.

[0944] Terminal (VR goggles) processing

[0945] Step 1: Receiving a data packet

[0946] The terminal receives the data packets sent from the server through a high-speed network.

[0947] Step 2: Reassembling the data packets

[0948] The terminal reassembles the received data packets into the original video data.

[0949] Step 3: Decoding the video data

[0950] The terminal decodes the reassembled video data and prepares it for display.

[0951] Step 4: Rendering the Perspective

[0952] Based on the decoded video data, the device renders a 360-degree panoramic image tailored to the user's viewpoint in real time.

[0953] Step 5: Acquire motion data

[0954] The device uses a built-in gyro sensor to detect the user's head movements.

[0955] Step 6: Adjusting the image and sound

[0956] The device adjusts the video and audio effects in real time based on the acquired motion data, providing a sense of realism in both visual and audio.

[0957] User operations

[0958] Step 1: Put on the VR goggles and log in

[0959] The user puts on the VR goggles and logs into the system. To log in, the user must enter their account information.

[0960] Step 2: Change your perspective

[0961] Users can freely change their viewpoint by moving their head, enjoying a 360-degree view, and can also focus on a specific player or play.

[0962] Step 3: Interactive Experience

[0963] Users receive real-time visual and auditory feedback, making them feel as if they are at the stadium, with sounds such as the cheers of the crowd and the footsteps of the players reproduced in real time.

[0964] Example: A basketball game

[0965] Server Action:

[0966] The server collects video from multiple cameras in the stadium, stores it in a temporary buffer, and then sorts it chronologically. It then automatically recognizes overlapping areas of the video and seamlessly combines them to generate a 360-degree panoramic video. The video is then encoded, divided into packets, and sent to the terminal.

[0967] Terminal handling:

[0968] The device receives and reassembles data packets, decodes the video data, renders the video in real time according to the user's viewpoint, detects head movement using the built-in gyro sensor, and adjusts the video and sound effects in real time to create a sense of realism.

[0969] User Action:

[0970] Users put on VR goggles and log in. They can freely change the viewpoint by moving their head and enjoy an interactive experience of watching specific scenes. For example, they can switch to a courtside viewpoint and experience the players' movements and plays in real time.

[0971] Example 1

[0972] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0973] With conventional game viewing systems, it was difficult for users to experience the immersive atmosphere of a game in real time from a remote location. Furthermore, there was a lack of technology to seamlessly integrate video data from multiple cameras and provide 360-degree panoramic video. As a result, it was difficult for users to enjoy a large-scale game as if they were actually there.

[0974] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0975] In this invention, the server includes means for collecting video data from multiple camera devices within the venue, means for storing the collected video data in a buffer and sorting it in chronological order based on timestamps, means for automatically recognizing overlapping portions of the video data and combining them seamlessly, means for encoding the combined 360-degree omnidirectional video data in real time, dividing it into packets, and transmitting them to the terminal, means for the terminal to reconstruct the received data packets and store them in a temporary buffer as video data, means for the terminal to decode the reconstructed video data and render an image according to the user's viewpoint in real time, and means for the terminal to acquire the user's head movement using a built-in gyro sensor and adjust video and sound effects in real time based on the movement, thereby enabling users to experience the same sense of realism as if they were at the venue, even from a remote location.

[0976] A "venue" refers to a place where a sport or event is held, and is a facility equipped to allow spectators to watch the event.

[0977] "Photography device" refers to a device for acquiring video data, and includes fixed cameras, movable cameras, drone cameras, etc.

[0978] "Video data" refers to digital video information captured by a camera, and may include timestamps and coordinate information.

[0979] A "buffer" is a temporary data storage area used to temporarily store collected video data and reconstructed data.

[0980] A "timestamp" is digital information that indicates the time at which video data was collected, and is used to accurately manage the time series of data.

[0981] The "overlap" refers to the area where images captured by multiple cameras overlap, and is an important element for seamless image stitching.

[0982] "Seamless combining" refers to the process of integrating multiple pieces of video data into one continuous image without discontinuities or boundaries.

[0983] "360-degree omnidirectional video data" is video data that covers the field of view in all directions, and is in a format that allows the user to freely change the viewpoint in any direction.

[0984] "Encoding" is the process of compressing video data using a certain format or codec and converting it into a transmittable form.

[0985] A "packet" refers to each unit of data divided when digital information is transmitted over a network, and each packet is accompanied by transmission route information and error check information.

[0986] A "terminal" is a device that the user directly operates, which in this case refers to VR goggles.

[0987] "Decoding" is the process of restoring encoded video data to its original format so that the device can play the video.

[0988] A "gyro sensor" is a sensor that detects the angular velocity and rotation of a device and is used to accurately capture the user's head movements.

[0989] "Rendering in real time" refers to the process of generating and displaying images in real time in response to the user's viewpoint and movements.

[0990] "Sound effects" refers to the audio and sound effects that accompany the video, and enhance the sense of realism by adjusting the sense of direction and distance according to the user's movements.

[0991] The present invention relates to a system that uses VR technology to enhance the sense of realism when watching a game. Hereinafter, an embodiment of the present invention will be described in detail.

[0992] Server Processing

[0993] The server collects video data from multiple camera devices located within the venue. These include fixed cameras, mobile cameras, and drone cameras. The collected video data is first stored in a buffer. Each piece of video data is assigned a timestamp, which is used to sort the data in chronological order. The server then automatically recognizes overlapping areas of the video data and seamlessly combines them. This process uses an image processing library (e.g., OpenCV).

[0994] Specifically, the server uses an edge detection algorithm (e.g., SIFT) to identify adjacent overlapping areas and integrate them into one continuous video. This integrated 360-degree panoramic video data is then encoded in real time using the H.264 or H.265 codec. The encoded data is then divided into packets of a fixed size and sent to the device using the UDP protocol. The transmission process uses the ffmpeg library.

[0995] Terminal processing (VR goggle processing)

[0996] The terminal receives real-time video data packets sent from the server. The received packets are stored in a temporary buffer and reconstructed based on the sequence numbers. The terminal then decodes the video data using a decoding library (e.g., libavcodec). The decoded video data is stored in the temporary buffer again.

[0997] The device then renders this video data in real time according to the user's viewpoint. A game engine (e.g., Unity or Unreal Engine) is used for rendering. The device uses a built-in gyro sensor to detect the user's head movements and adjusts the video viewpoint and sound effects in real time based on those movements. An audio engine (e.g., OpenAL or FMOD) is used to adjust the sound effects.

[0998] User operations

[0999] First, users put on VR goggles and log in to the system. This process is done using facial recognition, ID, and password. Users can freely change their viewpoint by moving their head. They can view 360-degree images and select and watch their favorite scenes in real time.

[1000] Specific examples

[1001] For example, when streaming a basketball game, the server collects footage from multiple cameras installed around the stadium. The collected footage is then sorted chronologically based on timestamps, overlapping parts are recognized and seamlessly joined, and then encoded and sent to the device.

[1002] The device decodes the received video data and renders it in real time according to the user's viewpoint. Users can change the viewpoint by moving their head, for example, to a courtside perspective, and experience the players' movements and plays in real time. Sound effects are also adjusted in real time, making users feel as if they are at the game venue.

[1003] Prompt Sentence Examples

[1004] The following are examples of prompt sentences to input into a generative AI model:

[1005] "Please explain in detail the terminal processing of the game viewing system using VR technology. In particular, please provide details on the reception and reassembly of data packets, decoding and buffering, real-time rendering based on the user's viewpoint, and adjustment of sound effects."

[1006] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1007] Server Processing

[1008] Step 1: Collect video data

[1009] The server collects video data in real time from multiple camera devices (fixed cameras, mobile cameras, drone cameras, etc.) installed within the venue. These cameras assign a timestamp to each video frame and send it to the server. The server temporarily stores the received video data in a buffer. The input is real-time video data from the camera devices, and the output is the video data stored in the buffer.

[1010] Step 2: Adjust the timeline of the video data

[1011] The server sorts the video data stored in the buffer into chronological order based on the timestamps, ensuring accurate synchronization of the video from each camera. The input is time-stamped video data, and the output is video data sorted in chronological order.

[1012] Step 3: Seamlessly stitch together footage

[1013] The server automatically recognizes overlapping areas of the video data from each camera and seamlessly combines them. This process uses an image processing library (e.g., OpenCV) and utilizes an edge detection algorithm (e.g., SIFT). The input is multiple video data sorted in chronological order, and the output is seamlessly combined 360-degree omnidirectional video data.

[1014] Step 4: Encode and transmit video data

[1015] The server encodes the combined 360-degree panoramic video data in real time using the H.264 or H.265 codec. The encoded data is then divided into packets of a fixed size and sent to the device using the UDP protocol. The input is the seamlessly combined 360-degree panoramic video data, and the output is the encoded data packets sent to the device.

[1016] Terminal processing (VR goggle processing)

[1017] Step 1: Receiving and reassembling data packets

[1018] The terminal receives data packets sent from the server. The received packets are stored in a temporary buffer and reassembled into the correct order based on the sequence numbers. The input is the data packets from the server, and the output is the assembled video data.

[1019] Step 2: Decoding and Buffering

[1020] The device decodes the reconstructed video data. This process uses a decoding library (e.g., libavcodec) and may utilize hardware acceleration. The decoded video is again stored in a temporary buffer. The input is the reconstructed video data, and the output is the decoded video data.

[1021] Step 3: Real-time rendering based on user perspective

[1022] The device uses a built-in gyro sensor to capture the user's head movements and renders video data in real time based on those movements. A 360-degree panoramic image is displayed according to the user's viewpoint. This process uses a game engine (e.g., Unity or Unreal Engine). The input is the decoded video data and gyro sensor movement data, and the output is a rendered image according to the user's viewpoint.

[1023] Step 4: Real-time sound effect adjustment

[1024] The device adjusts the sound effects in real time according to the user's head movements, so that the sound has a sense of direction according to the user's movements. This process uses an acoustic engine (e.g., OpenAL or FMOD). The input is the user's movement data, and the output is the adjusted sound effects.

[1025] User operations

[1026] Step 1: Put on the VR goggles and log in to the system

[1027] The user puts on the VR goggles and logs in to the system. The login process uses facial recognition, ID, and password. The input is the user's authentication information, and the output is the system login status.

[1028] Step 2: Freely change the viewpoint

[1029] Users can freely change their viewpoint by moving their head, allowing them to view a 360-degree panoramic view and select and watch their favorite scenes in real time. The input is the user's head movement data, and the output is a rendered image based on the user's viewpoint.

[1030] (Application example 1)

[1031] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1032] Conventional game viewing systems have limited means for enhancing the sense of realism, limiting the ability to experience the game from specific viewpoints and angles. It has also been difficult for users to freely change viewpoints or adjust video and audio effects in real time. This has prevented users from experiencing the game as if they were actually at the venue.

[1033] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1034] In this invention, the server includes means for collecting video data from multiple camera devices within the venue, means for integrating the collected video data to generate 360-degree omnidirectional video data, means for transmitting the generated 360-degree omnidirectional video data to the terminal in real time, means for the terminal to receive the video data and provide a 360-degree omnidirectional field of view and adjust the video and sound effects according to the user's movements, means for the terminal to detect the user's movements using a built-in gyro sensor and adjust the video and sound effects based on the movements in real time, and means for temporarily storing the collected video data in a buffer and realizing seamless playback. This allows users to experience a realistic game viewing experience in real time and to freely change the viewpoint to watch from different viewpoints within the venue.

[1035] "Game venue" means a particular location where a sport or event is held.

[1036] "Filming equipment" refers to equipment used to collect video data, including fixed cameras, mobile cameras, and drone cameras.

[1037] "Video data" refers to visual information captured by a camera within the venue.

[1038] "Integration" refers to the process of combining multiple pieces of video data and reconstructing them into a single, continuous image.

[1039] "360-degree panoramic video" refers to video data that covers all directions within the venue, and has the ability to freely change the viewpoint.

[1040] "Real-time" refers to a situation in which information is processed almost instantly, with little time delay.

[1041] "Terminal" refers to a device that a user uses to watch video, including smartphones and head-mounted displays.

[1042] "Providing" refers to the act of making a particular service or feature available to a user.

[1043] "Movement" refers to the user's actions of moving their head or body.

[1044] "Sound effects" refers to audio elements such as voice and music that are provided when watching a video.

[1045] "Adjustment" refers to the act of changing or modifying functions or effects to suit specific conditions or environments.

[1046] A "gyro sensor" is a sensor for measuring angular velocity and is used to detect the movement of the user's head.

[1047] "Decoding" refers to the process of restoring encoded data to its original form.

[1048] "Rendering" refers to the process of visually displaying video data.

[1049] A "temporary buffer" refers to a memory area for temporarily storing data.

[1050] "Seamless" refers to playback or joining that is performed without interruption, with continuity maintained.

[1051] This invention relates to a system that uses VR technology to enhance the sense of realism when watching a match. This system collects video data from multiple camera devices within the match venue and provides it to users in real time, allowing them to enjoy the match from a 360-degree, all-around perspective.

[1052] Server Processing

[1053] The server collects video data from multiple camera devices installed within the venue, including fixed cameras, mobile cameras, and drone cameras. The collected video data is first stored in a buffer, and the footage from each camera is sorted chronologically. Next, overlapping areas of the footage are automatically recognized and seamlessly combined, generating 360-degree video data. The generated 360-degree video data is then encoded in real time, divided into packets of a fixed size, and sent to the device. This transmission is performed with low latency, allowing users to experience the video smoothly.

[1054] Terminal handling

[1055] The device (e.g., a smartphone or head-mounted display) receives real-time video data sent from the server. The device reassembles the received data packets and stores them in a temporary buffer. It then decodes the video data and renders the video in real time according to the user's movements. The device uses its built-in gyro sensor to acquire data on the user's head movement and adjusts the video and sound effects in real time based on that movement. This allows the user to enjoy a visual and auditory experience that makes them feel as if they are at the stadium.

[1056] User operations

[1057] First, users put on the device and log in to the system. They can freely change their viewpoint by moving their head, looking around in all 360 degrees to watch any scene they like in real time. For example, users can stand in the perspective of a specific player and experience important moments of the game up close. Users can also select different viewpoints to enjoy views from different locations within the stadium.

[1058] Specific examples

[1059] For example, when streaming a basketball game, the process goes like this: The server collects footage from multiple cameras installed within the venue, automatically recognizes overlapping areas of the footage, and seamlessly combines them. This combined video data is encoded in real time and sent to the device. The device receives the video data sent from the server and stores it in a temporary buffer. It then decodes it and renders the video in real time according to the user's head movements. The built-in gyro sensor adjusts the user's viewpoint, and sound effects are also adjusted in real time. For example, you can switch to a courtside view to experience the players' movements and plays in real time.

[1060] Prompt Sentence Examples

[1061] "I want to develop a VR application that allows users to freely watch soccer matches from a 360-degree perspective. Please tell me specifically how to collect video data from a server in real time and render the video on an HMD according to the user's viewpoint."

[1062] Thus, according to the present invention, the user can experience watching a game with a sense of real-time presence, and by freely changing the viewpoint, the user can watch the game from different viewpoints within the venue.

[1063] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1064] Step 1:

[1065] The server collects video data from multiple camera devices installed within the venue.

[1066] Input: Video data from fixed, mobile, and drone cameras.

[1067] Output: Collected multi-camera video data.

[1068] Specific operation: The server continuously receives video data from each imaging device and temporarily stores it in a buffer.

[1069] Step 2:

[1070] The server integrates the collected video data and generates 360-degree panoramic video data.

[1071] Input: Single camera video data (multiple).

[1072] Output: Integrated 360° video data.

[1073] Specific operation: The time series of video data is organized and overlapping parts are seamlessly combined using an automatic recognition system (computer vision algorithm).

[1074] Step 3:

[1075] The server encodes the generated 360-degree omnidirectional video data in real time, divides it into packets of a certain size, and transmits it to the terminal.

[1076] Input: Integrated 360° video data.

[1077] Output: The encoded data packet.

[1078] Specific operation: Video data is encoded and compressed, divided into packets, and sent to the terminal via a low-latency network.

[1079] Step 4:

[1080] The terminal receives real-time video data transmitted from the server and stores it in a temporary buffer.

[1081] Input: The encoded data packet.

[1082] Output: Data packets stored in a temporary buffer.

[1083] Specific operation: Starts the process of receiving data from the network and storing it in a temporary buffer.

[1084] Step 5:

[1085] The terminal decodes the data packets stored in the temporary buffer and assembles them into a video.

[1086] Input: Data packets stored in a temporary buffer.

[1087] Output: Decoded video data.

[1088] Specific operation: Using a decoding algorithm, the packets are converted into the original video data and a continuous video is reassembled.

[1089] Step 6:

[1090] The device uses a built-in gyro sensor to detect the user's head movements and adjusts visual and sound effects in real time based on those movements.

[1091] Input: Motion data from gyro sensor, decoded video data.

[1092] Output: Visual and sound effects adjusted to the user's point of view.

[1093] How it works: Using gyro sensor data, it tracks the user's head movements and dynamically adjusts the viewpoint and sound in real time.

[1094] Step 7:

[1095] The user wears the device and logs in to the system. While watching the captured video, the user can freely change the viewpoint by moving their head.

[1096] Input: Input data from the user interface.

[1097] Output: The visual and audio data that the user sees.

[1098] Specific operation: After the user wears the device and logs in to the system, they can watch the game while the video and audio are changed in real time based on their head movements.

[1099] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1100] The present invention relates to a system that uses VR technology and emotion recognition technology to enhance the sense of realism when watching a game. Hereinafter, an embodiment of the present invention will be described in detail.

[1101] Server Processing

[1102] The server collects video data from multiple camera systems within the venue, including fixed cameras, mobile cameras, and even drone cameras. The collected video data is buffered and the video data from each camera is sorted chronologically. A process then automatically recognizes overlapping areas of the video and seamlessly combines them to generate 360-degree video data.

[1103] The generated 360-degree video data is encoded in real time, divided into packets of a certain size, and transmitted to the device with low latency, allowing users to enjoy a smooth experience.

[1104] Terminal processing (VR goggle processing)

[1105] The device (VR goggles) receives real-time video data sent from the server. The device reassembles the received data packets and stores them in a temporary buffer as video data. Next, the device decodes the video data and renders the video in real time according to the user's viewpoint.

[1106] The device uses a built-in gyro sensor to capture the user's head movement data and adjusts the visual and audio effects in real time based on that movement, providing users with a visual and audio experience that makes them feel like they're at the stadium.

[1107] Furthermore, the device includes an emotion engine that uses a camera to detect the user's facial expressions and biometric signals (e.g., heart rate, brain waves, etc.) and recognizes the user's emotions in real time based on the detected emotions. The recognized emotion data is then sent to a server.

[1108] Emotion data processing on the server

[1109] The server receives the user's emotional data sent from the device. Based on this emotional data, the server can automatically adjust the video and sound effects. For example, if the user is excited, it can emphasize the display of important scenes and enhance the sound effects.

[1110] User operations

[1111] First, users put on VR goggles and log in to the system. Users can freely change their viewpoint by moving their head, looking around in all 360 degrees and watching their favorite scenes in real time. Furthermore, camera angles and content are automatically selected according to the user's emotions, giving users an even more immersive experience.

[1112] Specific examples

[1113] For example, if you were streaming a basketball game, it would look like this:

[1114] Server processing: Images are collected from multiple cameras installed within the venue, temporarily stored in a buffer, and then sorted chronologically. Next, overlapping areas of the images are automatically recognized and seamlessly combined to generate a 360-degree panoramic image. The image is then encoded, divided into packets, and sent to the terminal.

[1115] Device processing: The device receives and reassembles data packets and decodes the video data. It renders the video in real time according to the user's viewpoint and detects head movements using the built-in gyro sensor. It adjusts the video and sound effects in real time. The device's camera also detects the user's facial expressions and biometric signals, analyzes the data using the emotion engine, and sends it to the server as emotion data.

[1116] Server emotion data processing: The server receives the emotion data and adjusts the video and sound effects in real time according to the user's emotion. For example, if the user is excited, the video will be adjusted to emphasize particularly dynamic scenes or important plays.

[1117] User operation: The user puts on the VR goggles and logs in. They can freely change the viewpoint by moving their head, enjoying an interactive experience of watching specific scenes. Furthermore, camera angles and specific content are automatically selected according to the user's emotions, enhancing the sense of realism. For example, if the user is feeling nervous, they can receive feedback such as scenes and angles that will help them relax.

[1118] This allows users to experience the realistic sensation of being at a game venue from the comfort of their own home, and furthermore, to enjoy a visual and auditory experience that is customized to suit their individual emotions.

[1119] The processing flow will be explained below.

[1120] Server Processing

[1121] Step 1: Collect video data

[1122] The server receives video data from multiple camera devices installed within the venue, including fixed, mobile and drone cameras, each capturing the match from a different perspective or angle.

[1123] Step 2: Buffering video data

[1124] The server temporarily stores the received video data in a buffer, whereby the video data from each camera is temporarily stored.

[1125] Step 3: Synchronize and sort data

[1126] The server rearranges the video data in the temporary buffer in chronological order and synchronizes them based on the timestamps.

[1127] Step 4: Video data integration

[1128] The server then combines the sorted video data from multiple cameras, automatically recognizing overlapping areas and seamlessly joining them together.

[1129] Step 5: Encode the video data

[1130] The server encodes the integrated 360-degree panoramic video data in real time.

[1131] Step 6: Packetize the data

[1132] The server divides the encoded video data into packets of a fixed size and prepares them for transfer.

[1133] Step 7: Sending data

[1134] The server then transmits the divided data packets to the terminal via a high-speed network, aiming for low latency.

[1135] Step 8: Receiving emotion data

[1136] The server receives the user's emotional data transmitted from the device, including emotional information analyzed from the user's facial expressions and biometric signals.

[1137] Step 9: Adjust based on sentiment data

[1138] The server adjusts the visual and audio effects in real time based on the received emotion data, selecting specific camera angles and content to display according to the user's emotion.

[1139] Terminal (VR goggles) processing

[1140] Step 1: Receiving a data packet

[1141] The terminal receives the data packets sent from the server through a high-speed network.

[1142] Step 2: Reassembling the data packets

[1143] The terminal reassembles the received data packets into the original video data.

[1144] Step 3: Decoding the video data

[1145] The terminal decodes the reassembled video data and prepares it for display.

[1146] Step 4: Rendering the Perspective

[1147] Based on the decoded video data, the device renders a 360-degree panoramic image tailored to the user's viewpoint in real time.

[1148] Step 5: Acquire motion data

[1149] The device uses a built-in gyro sensor to detect the user's head movements.

[1150] Step 6: Adjusting the image and sound

[1151] The device adjusts the video and audio effects in real time based on the acquired motion data, providing a sense of realism in both visual and audio.

[1152] Step 7: Obtaining emotion data

[1153] The device uses cameras and sensors to detect the user's facial expressions and biometric signals, recognizing the user's emotions in real time.

[1154] Step 8: Sending Emotion Data

[1155] The terminal transmits the recognized emotion data of the user to the server.

[1156] User operations

[1157] Step 1: Put on the VR goggles and log in

[1158] The user puts on the VR goggles and logs into the system. To log in, the user must enter their account information.

[1159] Step 2: Change your perspective

[1160] Users can freely change their viewpoint by moving their head, enjoying a 360-degree view, and can also focus on a specific player or play.

[1161] Step 3: Interactive Experience

[1162] Users receive real-time visual and auditory feedback, making them feel as if they are at the stadium, with sounds such as the cheers of the crowd and the footsteps of the players reproduced in real time.

[1163] Step 4: Displaying emotions

[1164] Depending on the user's emotions, specific camera angles and content can be automatically selected and displayed. For example, if the user is excited, important scenes can be displayed more prominently and sound effects can be enhanced.

[1165] Example 2

[1166] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1167] Conventional game viewing systems have limited visual and auditory experiences, making it difficult to fully recreate the sense of presence of a game. Furthermore, it is not possible to optimize the content played based on the user's real-time emotions and movements, making it difficult to provide a personalized experience for each user.

[1168] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1169] In this invention, the server includes means for collecting video data from multiple camera devices within the venue, means for integrating the collected video data to generate 360-degree omnidirectional video data, and means for transmitting the generated 360-degree omnidirectional video data to the terminal in real time, thereby enabling the user to experience 360-degree omnidirectional real-time video, and for the terminal to receive the video data and provide a 360-degree omnidirectional field of view, and to adjust the video and sound effects according to the user's movements.

[1170] Furthermore, by including a means for the terminal to decode the video data sent from the server and render the video in real time according to the user's viewpoint, and a means for acquiring the user's emotional data and adjusting the video and sound effects based on that, it is possible to provide a more immersive and individually optimized visual and auditory experience.

[1171] A "game venue" is a specific location where a sport or event is played.

[1172] "Filming equipment" refers to equipment such as cameras and drone cameras used to collect video data.

[1173] "Video data" refers to video information collected from a camera and includes visual content.

[1174] "Collection" refers to the act of collecting video data from multiple imaging devices.

[1175] "Integration" is the process of bringing together collected video data.

[1176] "360-degree omnidirectional video data" is video data that provides visual information from all directions.

[1177] "Generation" is the act of creating new video data.

[1178] "Real-time" refers to almost instantaneous processing, with very little delay.

[1179] A "terminal" is a device that receives and displays video data, such as VR goggles.

[1180] "Transmitting" is the act of sending data to another device or system.

[1181] "Rendering" is the process of generating images or videos from digital data.

[1182] A "gyro sensor" is a sensor that detects the movement of the user's head.

[1183] "Emotion data" is data that indicates the emotional state of the user.

[1184] "Decoding" is the process of converting encoded data into a playable format.

[1185] "Adjustment" is the act of changing visual and sound effects based on the user's movements and emotions.

[1186] The present invention relates to a system that uses VR technology and emotion recognition technology to enhance the sense of realism when watching a game. Specific embodiments of the present invention will be described below.

[1187] The server collects video data from multiple camera devices within the venue. This includes fixed cameras, mobile cameras, and drone cameras. The video data from each camera is temporarily stored in a buffer and sorted chronologically. The server then automatically recognizes overlapping areas of the footage and seamlessly combines them to generate 360-degree panoramic video data.

[1188] The generated 360-degree video data is encoded in real time using codecs such as H.264 or H.265. The encoded data is divided into packets of a fixed size and transmitted to the device with low latency.

[1189] The device (VR goggles) receives video data packets sent from the server. The received data is stored in a buffer and decoded. The decoded video data is rendered in real time according to the user's viewpoint. The device uses a built-in gyro sensor to detect the user's head movement and adjusts the video and sound effects based on that movement.

[1190] Furthermore, the device is equipped with a camera and biometric sensors to detect the user's facial expressions and biometric signals (e.g., heart rate, brain waves). This data is analyzed by the emotion engine and sent to the server as user emotion data.

[1191] The server receives the emotion data sent from the terminal and adjusts the video and sound effects based on the user's emotional state. For example, if the user is excited, the server adjusts the video to emphasize important scenes or dynamic play.

[1192] Users can use the above services by wearing VR goggles and logging in to the system. Users can freely change their viewpoint by moving their head, enjoying a 360-degree panoramic view. In addition, camera angles and content are automatically selected based on the user's emotions, creating an even more immersive experience.

[1193] Specific examples

[1194] For example, when streaming a basketball game, the following steps are taken: Video data is collected from multiple cameras installed in the stadium and stored in a temporary buffer. The images from each camera are rearranged in chronological order and seamlessly combined to generate a 360-degree panoramic video. The video is then encoded, divided into packets, and sent to the device.

[1195] The device reassembles the received data packets and renders the decoded video data in real time. The built-in gyro sensor detects the user's head movements and adjusts the video and audio in real time. The device's camera and biometric sensors also detect the user's facial expressions and biometric signals, and the emotion engine analyzes the emotional data and transmits it to the server.

[1196] The server receives the emotion data and adjusts the video and sound effects according to the user's emotions. For example, if the user is excited, the video can be adjusted to emphasize dynamic scenes and important plays, allowing the user to enjoy the game more realistically.

[1197] Prompt Sentence Examples

[1198] "I would like to develop a system that uses VR technology to make watching basketball games more immersive, and also utilizes emotion recognition technology to adjust video and sound effects. This system will incorporate a mechanism to optimize video in real time according to the user's movements and emotions. Please tell me specifically how you will design the system and how you will use emotion data."

[1199] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1200] Step 1:

[1201] "Video data collection"

[1202] The server collects video data from fixed cameras, mobile cameras, drone cameras, etc. within the venue. The video data from each camera is temporarily stored in the server's buffer.

[1203] Input: Video streams from each imaging modality.

[1204] Output: Multiple video data stored in a buffer.

[1205] Specific operation: The server uses the stream URL of the camera registered in advance to pull in video data using a specified protocol (such as RTSP).

[1206] Step 2:

[1207] "Sorting time series data"

[1208] The server then sorts the video data stored in the buffer into chronological order based on the timestamps of each camera, thereby synchronizing the video data.

[1209] Input: Multiple video data stored in a buffer.

[1210] Output: Video data sorted in chronological order.

[1211] What it does: The server analyzes the timestamp information and runs an algorithm to arrange the video frames from each camera in the proper order.

[1212] Step 3:

[1213] "Seamless joining process"

[1214] The server automatically recognizes overlapping areas of the images and seamlessly combines them to generate 360-degree panoramic video data.

[1215] Input: Multiple video data sorted in chronological order.

[1216] Output: 360-degree panoramic video data.

[1217] How it works: The server detects overlapping areas and uses image processing algorithms to merge them together, creating a continuous omnidirectional image.

[1218] Step 4:

[1219] "Encoding 360-degree video"

[1220] The server encodes the generated 360-degree panoramic video data in real time using the H.264 or H.265 codec.

[1221] Input: 360-degree omnidirectional video data.

[1222] Output: Encoded video data.

[1223] Specific operation: The server inputs the video data into the codec, which performs compression and encoding processes, thereby reducing the data size appropriately and making transmission more efficient.

[1224] Step 5:

[1225] "Sending video packets"

[1226] The server divides the encoded video data into packets of a fixed size and transmits them to the terminal with low latency.

[1227] Input: Encoded video data.

[1228] Output: Split data packets.

[1229] Specific operation: The server packetizes the video data and transmits it to the terminal using the UDP or TCP protocol.

[1230] Step 6:

[1231] "Receiving Data Packets"

[1232] The terminal receives the video data packets sent from the server, checking for packet loss and duplication to ensure accurate reception.

[1233] Input: Data packet from the server.

[1234] Output: Data packets that have been received and processed.

[1235] Specific operation: The terminal performs packet reception processing at the network layer, and processes error checks and retransmission requests.

[1236] Step 7:

[1237] "Assembling Data Packets"

[1238] The received data packets are reassembled into the original video data and stored in a temporary buffer.

[1239] Input: Received data packets.

[1240] Output: Video data stored in a temporary buffer.

[1241] Specific operation: The terminal rearranges, reconstructs, and temporarily stores the data packets, referring to the timestamp information.

[1242] Step 8:

[1243] "Video data decoding"

[1244] The terminal decodes the video data in the buffer using an H.264 or H.265 decoder and acquires it as successive video frames.

[1245] Input: Video data stored in a temporary buffer.

[1246] Output: Decoded video frames.

[1247] Specific operation: The terminal decodes the video frame using dedicated decoding hardware or software.

[1248] Step 9:

[1249] "Video rendering"

[1250] The decoded video is rendered in real time according to the user's viewpoint and displayed in the VR goggles.

[1251] Input: Decoded video frames.

[1252] Output: Rendered video frames.

[1253] Specific operation: The terminal uses a rendering engine to generate video frames according to the user's viewpoint and displays them on the display.

[1254] Step 10:

[1255] "Acquiring head movement data"

[1256] The device's built-in gyro sensor detects the user's head movements and uses that data to adjust visual and sound effects in real time.

[1257] Input: User's head movement.

[1258] Output: Coordinated video and audio.

[1259] Specific operation: The device obtains data in real time from the gyro sensor and dynamically adjusts image rendering and audio filtering.

[1260] Step 11:

[1261] "Acquiring emotion data"

[1262] The device's camera and biometric sensors capture biometric data such as the user's facial expressions and heart rate, which is then analyzed by the emotion engine.

[1263] Input: User's facial expression data and biometric signals.

[1264] Output: Parsed emotion data.

[1265] Specific operation: The device's camera extracts feature points using a facial expression recognition algorithm, and biometric sensors measure heart rate and skin galvanic response, which are then comprehensively analyzed by the emotion engine.

[1266] Step 12:

[1267] "Sending emotional data"

[1268] The analyzed emotional data is sent to the server in real time.

[1269] Input: Parsed emotion data.

[1270] Output: Emotion data sent to the server.

[1271] Specific operation: The device encodes the emotion data into an appropriate format and makes a transmission request to the server.

[1272] Step 13:

[1273] "Receiving and adjusting emotional data"

[1274] The server receives the emotional data sent from the terminal and adjusts the visual and sound effects based on the user's emotional state.

[1275] Input: Emotion data sent from the device.

[1276] Output: Coordinated video and audio.

[1277] Specific operation: The server analyzes the emotional data and dynamically generates content optimized for the user's emotional state, emphasizing specific visual scenes and sound effects.

[1278] Through the above processing steps, the system can provide the user with a realistic sense of being at the game, and can also provide a customized visual and auditory experience according to the user's emotional state.

[1279] (Application example 2)

[1280] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1281] While conventional VR-based game viewing systems provide users with a high sense of realism, they have difficulty providing a personalized experience that reflects the emotional and physical state of each user. Therefore, a new system is needed that can provide appropriate visual and audio effects in real time according to the emotional state and preferences of various users.

[1282] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for collecting video data from multiple camera devices within the venue, means for integrating the collected video data to generate 360-degree omnidirectional video data, and means for transmitting the generated 360-degree omnidirectional video data to the terminal in real time. This allows the user to experience the immersive atmosphere of the venue in real time. The terminal also includes means for adjusting video and audio effects according to the user's movements, means for sensing the user's facial expressions and biometric signals and recognizing the user's emotions in real time based on the detected facial expressions and biometric signals, and means for automatically adjusting video and audio effects based on the recognized emotional data. This allows the user to enjoy a personalized visual and audio experience that is tailored to their emotions and physical state.

[1283] "Game venue" refers to a location where various sports and entertainment matches are played.

[1284] "Filming equipment" refers to equipment such as cameras and drones used to collect video data.

[1285] "Video data" refers to video data collected by an imaging device.

[1286] "360-degree omnidirectional video data" refers to video that can be viewed in all directions and is generated by integrating video data collected from multiple imaging devices.

[1287] "Terminal" refers to a device (e.g., VR goggles) that receives video data and provides a visual and auditory experience to the user.

[1288] "User movement" refers to the user's body movement, particularly head movement, detected by the terminal.

[1289] "Sound effects" refers to the audio data corresponding to the video and the method of reproducing it.

[1290] "Expression" refers to the facial expression of the user.

[1291] "Biological signals" refer to physiological data such as a user's heart rate and brain waves.

[1292] "Emotion data" refers to the emotional state recognized based on the user's facial expressions and biometric signals.

[1293] "Real-time" refers to near-instant processing and transmission.

[1294] A "gyro sensor" refers to an inertial sensor used to detect the movement of a device.

[1295] "Synthesis" refers to the process of combining collected data into one linked data set.

[1296] "Seamlessly combining" refers to combining different video data continuously without interruption.

[1297] "Automatic adjustment" refers to the system automatically adjusting the video and audio based on the user's emotional data.

[1298] This invention relates to a system that enhances the sense of realism when a user watches a game using VR goggles and automatically adjusts video and sound effects according to the emotions of each individual user. The specific configuration and operation of the system are described below.

[1299] Server Processing

[1300] The server collects video data from multiple camera devices within the venue. The camera devices include fixed cameras, mobile cameras, and drone cameras. The collected video data is first stored in a buffer, and the video data from each camera is sorted chronologically. The server then automatically recognizes overlapping areas of the video and seamlessly combines them to generate 360-degree panoramic video data. The generated video data is then encoded in real time, divided into packets of a fixed size, and transmitted to the device. This transmission is performed with low latency, ensuring a smooth user experience.

[1301] Terminal handling

[1302] The device (VR goggles) receives real-time video data sent from the server. It reassembles the received data packets and stores them in a temporary buffer. It then decodes the video data and renders the video in real time according to the user's viewpoint. The device uses a built-in gyro sensor to acquire data on the user's head movement and adjusts the video and sound effects in real time based on that movement. The device also uses a camera to detect the user's facial expressions and biometric signals (e.g., heart rate, brain waves, etc.), and includes an emotion engine that recognizes the user's emotions in real time based on this data. The recognized emotion data is sent to the server in real time.

[1303] Emotion data processing on the server

[1304] The server receives the user's emotional data sent from the device. Based on this emotional data, it can automatically adjust the visual and sound effects. For example, if the user is excited, it can highlight important scenes and thrilling plays and enhance the sound effects.

[1305] User operations

[1306] First, users put on VR goggles and log in to the system. Users can freely change their viewpoint by moving their head, and enjoy a 360-degree panoramic view of their favorite scenes in real time. Furthermore, camera angles and content are automatically selected based on the user's emotions, creating an even more immersive experience.

[1307] Hardware and software used

[1308] The system is implemented using the following hardware and software:

[1309] Hardware: VR goggles (e.g., Oculus Rift, HTC Vive), fixed cameras, movable cameras, drone cameras, heart rate monitors, electroencephalographs

[1310] Software: OpenCV (used for image acquisition and processing), Python socket library (used for sending and receiving image data), Unity or Unreal Engine (for VR rendering and user interface construction)

[1311] Specific examples

[1312] For example, when live streaming a soccer match, the following prompt sentence can be input into the generative AI model to generate video that highlights important scenes from the match.

[1313] "Analyze the user's emotion recognition data in real time during a soccer match. If the user is excited, highlight the goal scene and the play before the goal. If the user is nervous, switch to a more relaxing video."

[1314] This invention allows users to enjoy a highly realistic and personalized viewing experience.

[1315] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1316] Step 1:

[1317] The server collects video data from multiple camera devices (fixed cameras, mobile cameras, drone cameras) within the venue. The collected video data is temporarily stored in a buffer, and the video data from each camera is sorted chronologically. This ensures that the video data is arranged in the appropriate order.

[1318] Step 2:

[1319] The server automatically recognizes overlapping areas in the collected video data and seamlessly combines them. It uses OpenCV to detect overlapping areas in each frame and combines them continuously. The combined video data becomes 360-degree omnidirectional video data.

[1320] Step 3:

[1321] The server encodes the generated 360-degree omnidirectional video data in real time, divides it into packets of a fixed size, and transmits them to the device. This transmission is performed with low latency, and the data is sent to the device via the network. Video codecs such as H.264 and HEVC are used for encoding.

[1322] Step 4:

[1323] The device (VR goggles) receives real-time video data sent from the server, reassembles the received data packets, and stores them in a temporary buffer. Then, it decodes the video data. For decoding, it uses FFmpeg or other decoding libraries.

[1324] Step 5:

[1325] The device uses a built-in gyro sensor to capture data on the user's head movements, and uses this data to render video and sound effects in real time. Rendering is done using Unity or Unreal Engine, and the video is displayed according to the user's viewpoint.

[1326] Step 6:

[1327] The device uses a camera and biosensors to detect the user's facial expressions and biometric signals (heart rate, brain waves, etc.). It then uses an emotion engine to analyze this data and extract the user's emotional data. For example, it uses a machine learning model to infer emotions from facial expressions and biometric signals.

[1328] Step 7:

[1329] The device transmits the recognized emotion data to the server in real time with low latency, so the user's emotional state is immediately conveyed to the server.

[1330] Step 8:

[1331] The server automatically adjusts visual and audio effects based on the received user emotion data. For example, if the user is excited, it will highlight important scenes or thrilling action scenes and enhance audio effects. It uses a generative AI model to calculate the appropriate visual and audio settings in real time.

[1332] Step 9:

[1333] Users put on VR goggles, log in, and begin watching the game. They can freely change their viewpoint by moving their head, and enjoy a 360-degree panoramic view of the game to watch any scene they like. Furthermore, camera angles and content are automatically selected based on the user's emotions, creating an even more immersive experience.

[1334] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1335] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1336] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1337] [Fourth embodiment]

[1338] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1339] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1340] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1341] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1342] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1343] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1344] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1345] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1346] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1347] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1348] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1349] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1350] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1351] The present invention relates to a system that uses VR technology to enhance the sense of realism when watching a game. Hereinafter, an embodiment of the present invention will be described in detail.

[1352] Server Processing

[1353] The server collects video data from multiple camera systems within the venue, including fixed cameras, mobile cameras, and even drone cameras. The collected video data is buffered and the footage from each camera is sorted chronologically. A process then automatically recognizes overlapping areas of the footage and seamlessly combines them to generate 360-degree panoramic video data.

[1354] The generated 360-degree video data is encoded in real time, divided into packets of a certain size, and transmitted to the device with low latency, allowing users to experience the video smoothly.

[1355] Terminal processing (VR goggle processing)

[1356] The device (VR goggles) receives real-time video data sent from the server. The device reassembles the received data packets and stores them in a temporary buffer as video data. Next, the device decodes the video data and renders the video in real time according to the user's viewpoint.

[1357] The device uses a built-in gyro sensor to capture the user's head movement data and adjusts the visual and audio effects in real time based on that movement, providing users with a visual and audio experience that makes them feel like they're at the stadium.

[1358] User operations

[1359] First, users put on VR goggles and log in to the system. Users can freely change their viewpoint by moving their head, and enjoy a 360-degree panoramic view of their favorite scenes in real time. For example, they can stand in the perspective of their favorite player and experience the moment they make a shot up close.

[1360] Specific examples

[1361] For example, if you were streaming a basketball game, it would look like this:

[1362] Server processing: Images are collected from multiple cameras installed within the venue, and overlapping areas of the images are automatically recognized and seamlessly combined. This combined video data is encoded in real time and sent to the terminal.

[1363] Device processing: The device receives the video data sent from the server and stores it in a temporary buffer. It then decodes it and renders the video in real time according to the user's head movements. It uses the built-in gyro sensor to adjust the user's viewpoint and adjusts the sound effects in real time.

[1364] User operation: The user puts on the VR goggles and logs in to the system. The user can change the viewpoint by moving their head and enjoy 360-degree omnidirectional images. For example, they can switch to a courtside viewpoint and experience the players' movements and plays in real time.

[1365] This allows users to experience the immersive atmosphere of a live game venue from the comfort of their own home.

[1366] The processing flow will be explained below.

[1367] Server Processing

[1368] Step 1: Collect video data

[1369] The server receives video data from multiple camera devices installed within the venue, including fixed, mobile and drone cameras, each capturing the match from a different perspective or angle.

[1370] Step 2: Buffering video data

[1371] The server temporarily stores the received video data in a buffer, whereby the video data from each camera is temporarily stored.

[1372] Step 3: Synchronize and sort data

[1373] The server rearranges the video data in the temporary buffer in chronological order and synchronizes them based on the timestamps.

[1374] Step 4: Video data integration

[1375] The server then combines the sorted video data from multiple cameras, automatically recognizing overlapping areas and seamlessly joining them together.

[1376] Step 5: Encode the video data

[1377] The server encodes the integrated 360-degree panoramic video data in real time.

[1378] Step 6: Packetize the data

[1379] The server divides the encoded video data into packets of a fixed size and prepares them for transfer.

[1380] Step 7: Sending data

[1381] The server then transmits the divided data packets to the terminal via a high-speed network, aiming for low latency.

[1382] Terminal (VR goggles) processing

[1383] Step 1: Receiving a data packet

[1384] The terminal receives the data packets sent from the server through a high-speed network.

[1385] Step 2: Reassembling the data packets

[1386] The terminal reassembles the received data packets into the original video data.

[1387] Step 3: Decoding the video data

[1388] The terminal decodes the reassembled video data and prepares it for display.

[1389] Step 4: Rendering the Perspective

[1390] Based on the decoded video data, the device renders a 360-degree panoramic image tailored to the user's viewpoint in real time.

[1391] Step 5: Acquire motion data

[1392] The device uses a built-in gyro sensor to detect the user's head movements.

[1393] Step 6: Adjusting the image and sound

[1394] The device adjusts the video and audio effects in real time based on the acquired motion data, providing a sense of realism in both visual and audio.

[1395] User operations

[1396] Step 1: Put on the VR goggles and log in

[1397] The user puts on the VR goggles and logs into the system. To log in, the user must enter their account information.

[1398] Step 2: Change your perspective

[1399] Users can freely change their viewpoint by moving their head, enjoying a 360-degree view, and can also focus on a specific player or play.

[1400] Step 3: Interactive Experience

[1401] Users receive real-time visual and auditory feedback, making them feel as if they are at the stadium, with sounds such as the cheers of the crowd and the footsteps of the players reproduced in real time.

[1402] Example: A basketball game

[1403] Server Action:

[1404] The server collects video from multiple cameras in the stadium, stores it in a temporary buffer, and then sorts it chronologically. It then automatically recognizes overlapping areas of the video and seamlessly combines them to generate a 360-degree panoramic video. The video is then encoded, divided into packets, and sent to the terminal.

[1405] Terminal handling:

[1406] The device receives and reassembles data packets, decodes the video data, renders the video in real time according to the user's viewpoint, detects head movement using the built-in gyro sensor, and adjusts the video and sound effects in real time to create a sense of realism.

[1407] User Action:

[1408] Users put on VR goggles and log in. They can freely change the viewpoint by moving their head and enjoy an interactive experience of watching specific scenes. For example, they can switch to a courtside viewpoint and experience the players' movements and plays in real time.

[1409] Example 1

[1410] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1411] With conventional game viewing systems, it was difficult for users to experience the immersive atmosphere of a game in real time from a remote location. Furthermore, there was a lack of technology to seamlessly integrate video data from multiple cameras and provide 360-degree panoramic video. As a result, it was difficult for users to enjoy a large-scale game as if they were actually there.

[1412] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1413] In this invention, the server includes means for collecting video data from multiple camera devices within the venue, means for storing the collected video data in a buffer and sorting it in chronological order based on timestamps, means for automatically recognizing overlapping portions of the video data and combining them seamlessly, means for encoding the combined 360-degree omnidirectional video data in real time, dividing it into packets, and transmitting them to the terminal, means for the terminal to reconstruct the received data packets and store them in a temporary buffer as video data, means for the terminal to decode the reconstructed video data and render an image according to the user's viewpoint in real time, and means for the terminal to acquire the user's head movement using a built-in gyro sensor and adjust video and sound effects in real time based on the movement, thereby enabling users to experience the same sense of realism as if they were at the venue, even from a remote location.

[1414] A "venue" refers to a place where a sport or event is held, and is a facility equipped to allow spectators to watch the event.

[1415] "Photography device" refers to a device for acquiring video data, and includes fixed cameras, movable cameras, drone cameras, etc.

[1416] "Video data" refers to digital video information captured by a camera, and may include timestamps and coordinate information.

[1417] A "buffer" is a temporary data storage area used to temporarily store collected video data and reconstructed data.

[1418] A "timestamp" is digital information that indicates the time at which video data was collected, and is used to accurately manage the time series of data.

[1419] The "overlap" refers to the area where images captured by multiple cameras overlap, and is an important element for seamless image stitching.

[1420] "Seamless combining" refers to the process of integrating multiple pieces of video data into one continuous image without discontinuities or boundaries.

[1421] "360-degree omnidirectional video data" is video data that covers the field of view in all directions, and is in a format that allows the user to freely change the viewpoint in any direction.

[1422] "Encoding" is the process of compressing video data using a certain format or codec and converting it into a transmittable form.

[1423] A "packet" refers to each unit of data divided when digital information is transmitted over a network, and each packet is accompanied by transmission route information and error check information.

[1424] A "terminal" is a device that the user directly operates, which in this case refers to VR goggles.

[1425] "Decoding" is the process of restoring encoded video data to its original format so that the device can play the video.

[1426] A "gyro sensor" is a sensor that detects the angular velocity and rotation of a device and is used to accurately capture the user's head movements.

[1427] "Rendering in real time" refers to the process of generating and displaying images in real time in response to the user's viewpoint and movements.

[1428] "Sound effects" refers to the audio and sound effects that accompany the video, and enhance the sense of realism by adjusting the sense of direction and distance according to the user's movements.

[1429] The present invention relates to a system that uses VR technology to enhance the sense of realism when watching a game. Hereinafter, an embodiment of the present invention will be described in detail.

[1430] Server Processing

[1431] The server collects video data from multiple camera devices located within the venue. These include fixed cameras, mobile cameras, and drone cameras. The collected video data is first stored in a buffer. Each piece of video data is assigned a timestamp, which is used to sort the data in chronological order. The server then automatically recognizes overlapping areas of the video data and seamlessly combines them. This process uses an image processing library (e.g., OpenCV).

[1432] Specifically, the server uses an edge detection algorithm (e.g., SIFT) to identify adjacent overlapping areas and integrate them into one continuous video. This integrated 360-degree panoramic video data is then encoded in real time using the H.264 or H.265 codec. The encoded data is then divided into packets of a fixed size and sent to the device using the UDP protocol. The transmission process uses the ffmpeg library.

[1433] Terminal processing (VR goggle processing)

[1434] The terminal receives real-time video data packets sent from the server. The received packets are stored in a temporary buffer and reconstructed based on the sequence numbers. The terminal then decodes the video data using a decoding library (e.g., libavcodec). The decoded video data is stored in the temporary buffer again.

[1435] The device then renders this video data in real time according to the user's viewpoint. A game engine (e.g., Unity or Unreal Engine) is used for rendering. The device uses a built-in gyro sensor to detect the user's head movements and adjusts the video viewpoint and sound effects in real time based on those movements. An audio engine (e.g., OpenAL or FMOD) is used to adjust the sound effects.

[1436] User operations

[1437] First, users put on VR goggles and log in to the system. This process is done using facial recognition, ID, and password. Users can freely change their viewpoint by moving their head. They can view 360-degree images and select and watch their favorite scenes in real time.

[1438] Specific examples

[1439] For example, when streaming a basketball game, the server collects footage from multiple cameras installed around the stadium. The collected footage is then sorted chronologically based on timestamps, overlapping parts are recognized and seamlessly joined, and then encoded and sent to the device.

[1440] The device decodes the received video data and renders it in real time according to the user's viewpoint. Users can change the viewpoint by moving their head, for example, to a courtside perspective, and experience the players' movements and plays in real time. Sound effects are also adjusted in real time, making users feel as if they are at the game venue.

[1441] Prompt Sentence Examples

[1442] The following are examples of prompt sentences to input into a generative AI model:

[1443] "Please explain in detail the terminal processing of the game viewing system using VR technology. In particular, please provide details on the reception and reassembly of data packets, decoding and buffering, real-time rendering based on the user's viewpoint, and adjustment of sound effects."

[1444] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1445] Server Processing

[1446] Step 1: Collect video data

[1447] The server collects video data in real time from multiple camera devices (fixed cameras, mobile cameras, drone cameras, etc.) installed within the venue. These cameras assign a timestamp to each video frame and send it to the server. The server temporarily stores the received video data in a buffer. The input is real-time video data from the camera devices, and the output is the video data stored in the buffer.

[1448] Step 2: Adjust the timeline of the video data

[1449] The server sorts the video data stored in the buffer into chronological order based on the timestamps, ensuring accurate synchronization of the video from each camera. The input is time-stamped video data, and the output is video data sorted in chronological order.

[1450] Step 3: Seamlessly stitch together footage

[1451] The server automatically recognizes overlapping areas of the video data from each camera and seamlessly combines them. This process uses an image processing library (e.g., OpenCV) and utilizes an edge detection algorithm (e.g., SIFT). The input is multiple video data sorted in chronological order, and the output is seamlessly combined 360-degree omnidirectional video data.

[1452] Step 4: Encode and transmit video data

[1453] The server encodes the combined 360-degree panoramic video data in real time using the H.264 or H.265 codec. The encoded data is then divided into packets of a fixed size and sent to the device using the UDP protocol. The input is the seamlessly combined 360-degree panoramic video data, and the output is the encoded data packets sent to the device.

[1454] Terminal processing (VR goggle processing)

[1455] Step 1: Receiving and reassembling data packets

[1456] The terminal receives data packets sent from the server. The received packets are stored in a temporary buffer and reassembled into the correct order based on the sequence numbers. The input is the data packets from the server, and the output is the assembled video data.

[1457] Step 2: Decoding and Buffering

[1458] The device decodes the reconstructed video data. This process uses a decoding library (e.g., libavcodec) and may utilize hardware acceleration. The decoded video is again stored in a temporary buffer. The input is the reconstructed video data, and the output is the decoded video data.

[1459] Step 3: Real-time rendering based on user perspective

[1460] The device uses a built-in gyro sensor to capture the user's head movements and renders video data in real time based on those movements. A 360-degree panoramic image is displayed according to the user's viewpoint. This process uses a game engine (e.g., Unity or Unreal Engine). The input is the decoded video data and gyro sensor movement data, and the output is a rendered image according to the user's viewpoint.

[1461] Step 4: Real-time sound effect adjustment

[1462] The device adjusts the sound effects in real time according to the user's head movements, so that the sound has a sense of direction according to the user's movements. This process uses an acoustic engine (e.g., OpenAL or FMOD). The input is the user's movement data, and the output is the adjusted sound effects.

[1463] User operations

[1464] Step 1: Put on the VR goggles and log in to the system

[1465] The user puts on the VR goggles and logs in to the system. The login process uses facial recognition, ID, and password. The input is the user's authentication information, and the output is the system login status.

[1466] Step 2: Freely change the viewpoint

[1467] Users can freely change their viewpoint by moving their head, allowing them to view a 360-degree panoramic view and select and watch their favorite scenes in real time. The input is the user's head movement data, and the output is a rendered image based on the user's viewpoint.

[1468] (Application example 1)

[1469] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1470] Conventional game viewing systems have limited means for enhancing the sense of realism, limiting the ability to experience the game from specific viewpoints and angles. It has also been difficult for users to freely change viewpoints or adjust video and audio effects in real time. This has prevented users from experiencing the game as if they were actually at the venue.

[1471] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1472] In this invention, the server includes means for collecting video data from multiple camera devices within the venue, means for integrating the collected video data to generate 360-degree omnidirectional video data, means for transmitting the generated 360-degree omnidirectional video data to the terminal in real time, means for the terminal to receive the video data and provide a 360-degree omnidirectional field of view and adjust the video and sound effects according to the user's movements, means for the terminal to detect the user's movements using a built-in gyro sensor and adjust the video and sound effects based on the movements in real time, and means for temporarily storing the collected video data in a buffer and realizing seamless playback. This allows users to experience a realistic game viewing experience in real time and to freely change the viewpoint to watch from different viewpoints within the venue.

[1473] "Game venue" means a particular location where a sport or event is held.

[1474] "Filming equipment" refers to equipment used to collect video data, including fixed cameras, mobile cameras, and drone cameras.

[1475] "Video data" refers to visual information captured by a camera within the venue.

[1476] "Integration" refers to the process of combining multiple pieces of video data and reconstructing them into a single, continuous image.

[1477] "360-degree panoramic video" refers to video data that covers all directions within the venue, and has the ability to freely change the viewpoint.

[1478] "Real-time" refers to a situation in which information is processed almost instantly, with little time delay.

[1479] "Terminal" refers to a device that a user uses to watch video, including smartphones and head-mounted displays.

[1480] "Providing" refers to the act of making a particular service or feature available to a user.

[1481] "Movement" refers to the user's actions of moving their head or body.

[1482] "Sound effects" refers to audio elements such as voice and music that are provided when watching a video.

[1483] "Adjustment" refers to the act of changing or modifying functions or effects to suit specific conditions or environments.

[1484] A "gyro sensor" is a sensor for measuring angular velocity and is used to detect the movement of the user's head.

[1485] "Decoding" refers to the process of restoring encoded data to its original form.

[1486] "Rendering" refers to the process of visually displaying video data.

[1487] A "temporary buffer" refers to a memory area for temporarily storing data.

[1488] "Seamless" refers to playback or joining that is performed without interruption, with continuity maintained.

[1489] This invention relates to a system that uses VR technology to enhance the sense of realism when watching a match. This system collects video data from multiple camera devices within the match venue and provides it to users in real time, allowing them to enjoy the match from a 360-degree, all-around perspective.

[1490] Server Processing

[1491] The server collects video data from multiple camera devices installed within the venue, including fixed cameras, mobile cameras, and drone cameras. The collected video data is first stored in a buffer, and the footage from each camera is sorted chronologically. Next, overlapping areas of the footage are automatically recognized and seamlessly combined, generating 360-degree video data. The generated 360-degree video data is then encoded in real time, divided into packets of a fixed size, and sent to the device. This transmission is performed with low latency, allowing users to experience the video smoothly.

[1492] Terminal handling

[1493] The device (e.g., a smartphone or head-mounted display) receives real-time video data sent from the server. The device reassembles the received data packets and stores them in a temporary buffer. It then decodes the video data and renders the video in real time according to the user's movements. The device uses its built-in gyro sensor to acquire data on the user's head movement and adjusts the video and sound effects in real time based on that movement. This allows the user to enjoy a visual and auditory experience that makes them feel as if they are at the stadium.

[1494] User operations

[1495] First, users put on the device and log in to the system. They can freely change their viewpoint by moving their head, looking around in all 360 degrees to watch any scene they like in real time. For example, users can stand in the perspective of a specific player and experience important moments of the game up close. Users can also select different viewpoints to enjoy views from different locations within the stadium.

[1496] Specific examples

[1497] For example, when streaming a basketball game, the process goes like this: The server collects footage from multiple cameras installed within the venue, automatically recognizes overlapping areas of the footage, and seamlessly combines them. This combined video data is encoded in real time and sent to the device. The device receives the video data sent from the server and stores it in a temporary buffer. It then decodes it and renders the video in real time according to the user's head movements. The built-in gyro sensor adjusts the user's viewpoint, and sound effects are also adjusted in real time. For example, you can switch to a courtside view to experience the players' movements and plays in real time.

[1498] Prompt Sentence Examples

[1499] "I want to develop a VR application that allows users to freely watch soccer matches from a 360-degree perspective. Please tell me specifically how to collect video data from a server in real time and render the video on an HMD according to the user's viewpoint."

[1500] Thus, according to the present invention, the user can experience watching a game with a sense of real-time presence, and by freely changing the viewpoint, the user can watch the game from different viewpoints within the venue.

[1501] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1502] Step 1:

[1503] The server collects video data from multiple camera devices installed within the venue.

[1504] Input: Video data from fixed, mobile, and drone cameras.

[1505] Output: Collected multi-camera video data.

[1506] Specific operation: The server continuously receives video data from each imaging device and temporarily stores it in a buffer.

[1507] Step 2:

[1508] The server integrates the collected video data and generates 360-degree panoramic video data.

[1509] Input: Single camera video data (multiple).

[1510] Output: Integrated 360° video data.

[1511] Specific operation: The time series of video data is organized and overlapping parts are seamlessly combined using an automatic recognition system (computer vision algorithm).

[1512] Step 3:

[1513] The server encodes the generated 360-degree omnidirectional video data in real time, divides it into packets of a certain size, and transmits it to the terminal.

[1514] Input: Integrated 360° video data.

[1515] Output: The encoded data packet.

[1516] Specific operation: Video data is encoded and compressed, divided into packets, and sent to the terminal via a low-latency network.

[1517] Step 4:

[1518] The terminal receives real-time video data transmitted from the server and stores it in a temporary buffer.

[1519] Input: The encoded data packet.

[1520] Output: Data packets stored in a temporary buffer.

[1521] Specific operation: Starts the process of receiving data from the network and storing it in a temporary buffer.

[1522] Step 5:

[1523] The terminal decodes the data packets stored in the temporary buffer and assembles them into a video.

[1524] Input: Data packets stored in a temporary buffer.

[1525] Output: Decoded video data.

[1526] Specific operation: Using a decoding algorithm, the packets are converted into the original video data and a continuous video is reassembled.

[1527] Step 6:

[1528] The device uses a built-in gyro sensor to detect the user's head movements and adjusts visual and sound effects in real time based on those movements.

[1529] Input: Motion data from gyro sensor, decoded video data.

[1530] Output: Visual and sound effects adjusted to the user's point of view.

[1531] How it works: Using gyro sensor data, it tracks the user's head movements and dynamically adjusts the viewpoint and sound in real time.

[1532] Step 7:

[1533] The user wears the device and logs in to the system. While watching the captured video, the user can freely change the viewpoint by moving their head.

[1534] Input: Input data from the user interface.

[1535] Output: The visual and audio data that the user sees.

[1536] Specific operation: After the user wears the device and logs in to the system, they can watch the game while the video and audio are changed in real time based on their head movements.

[1537] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1538] The present invention relates to a system that uses VR technology and emotion recognition technology to enhance the sense of realism when watching a game. Hereinafter, an embodiment of the present invention will be described in detail.

[1539] Server Processing

[1540] The server collects video data from multiple camera systems within the venue, including fixed cameras, mobile cameras, and even drone cameras. The collected video data is buffered and the video data from each camera is sorted chronologically. A process then automatically recognizes overlapping areas of the video and seamlessly combines them to generate 360-degree video data.

[1541] The generated 360-degree video data is encoded in real time, divided into packets of a certain size, and transmitted to the device with low latency, allowing users to enjoy a smooth experience.

[1542] Terminal processing (VR goggle processing)

[1543] The device (VR goggles) receives real-time video data sent from the server. The device reassembles the received data packets and stores them in a temporary buffer as video data. Next, the device decodes the video data and renders the video in real time according to the user's viewpoint.

[1544] The device uses a built-in gyro sensor to capture the user's head movement data and adjusts the visual and audio effects in real time based on that movement, providing users with a visual and audio experience that makes them feel like they're at the stadium.

[1545] Furthermore, the device includes an emotion engine that uses a camera to detect the user's facial expressions and biometric signals (e.g., heart rate, brain waves, etc.) and recognizes the user's emotions in real time based on the detected emotions. The recognized emotion data is then sent to a server.

[1546] Emotion data processing on the server

[1547] The server receives the user's emotional data sent from the device. Based on this emotional data, the server can automatically adjust the video and sound effects. For example, if the user is excited, it can emphasize the display of important scenes and enhance the sound effects.

[1548] User operations

[1549] First, users put on VR goggles and log in to the system. Users can freely change their viewpoint by moving their head, looking around in all 360 degrees and watching their favorite scenes in real time. Furthermore, camera angles and content are automatically selected according to the user's emotions, giving users an even more immersive experience.

[1550] Specific examples

[1551] For example, if you were streaming a basketball game, it would look like this:

[1552] Server processing: Images are collected from multiple cameras installed within the venue, temporarily stored in a buffer, and then sorted chronologically. Next, overlapping areas of the images are automatically recognized and seamlessly combined to generate a 360-degree panoramic image. The image is then encoded, divided into packets, and sent to the terminal.

[1553] Device processing: The device receives and reassembles data packets and decodes the video data. It renders the video in real time according to the user's viewpoint and detects head movements using the built-in gyro sensor. It adjusts the video and sound effects in real time. The device's camera also detects the user's facial expressions and biometric signals, analyzes the data using the emotion engine, and sends it to the server as emotion data.

[1554] Server emotion data processing: The server receives the emotion data and adjusts the video and sound effects in real time according to the user's emotion. For example, if the user is excited, the video will be adjusted to emphasize particularly dynamic scenes or important plays.

[1555] User operation: The user puts on the VR goggles and logs in. They can freely change the viewpoint by moving their head, enjoying an interactive experience of watching specific scenes. Furthermore, camera angles and specific content are automatically selected according to the user's emotions, enhancing the sense of realism. For example, if the user is feeling nervous, they can receive feedback such as scenes and angles that will help them relax.

[1556] This allows users to experience the realistic sensation of being at a game venue from the comfort of their own home, and furthermore, to enjoy a visual and auditory experience that is customized to suit their individual emotions.

[1557] The processing flow will be explained below.

[1558] Server Processing

[1559] Step 1: Collect video data

[1560] The server receives video data from multiple camera devices installed within the venue, including fixed, mobile and drone cameras, each capturing the match from a different perspective or angle.

[1561] Step 2: Buffering video data

[1562] The server temporarily stores the received video data in a buffer, whereby the video data from each camera is temporarily stored.

[1563] Step 3: Synchronize and sort data

[1564] The server rearranges the video data in the temporary buffer in chronological order and synchronizes them based on the timestamps.

[1565] Step 4: Video data integration

[1566] The server then combines the sorted video data from multiple cameras, automatically recognizing overlapping areas and seamlessly joining them together.

[1567] Step 5: Encode the video data

[1568] The server encodes the integrated 360-degree panoramic video data in real time.

[1569] Step 6: Packetize the data

[1570] The server divides the encoded video data into packets of a fixed size and prepares them for transfer.

[1571] Step 7: Sending data

[1572] The server then transmits the divided data packets to the terminal via a high-speed network, aiming for low latency.

[1573] Step 8: Receiving emotion data

[1574] The server receives the user's emotional data transmitted from the device, including emotional information analyzed from the user's facial expressions and biometric signals.

[1575] Step 9: Adjust based on sentiment data

[1576] The server adjusts the visual and audio effects in real time based on the received emotion data, selecting specific camera angles and content to display according to the user's emotion.

[1577] Terminal (VR goggles) processing

[1578] Step 1: Receiving a data packet

[1579] The terminal receives the data packets sent from the server through a high-speed network.

[1580] Step 2: Reassembling the data packets

[1581] The terminal reassembles the received data packets into the original video data.

[1582] Step 3: Decoding the video data

[1583] The terminal decodes the reassembled video data and prepares it for display.

[1584] Step 4: Rendering the Perspective

[1585] Based on the decoded video data, the device renders a 360-degree panoramic image tailored to the user's viewpoint in real time.

[1586] Step 5: Acquire motion data

[1587] The device uses a built-in gyro sensor to detect the user's head movements.

[1588] Step 6: Adjusting the image and sound

[1589] The device adjusts the video and audio effects in real time based on the acquired motion data, providing a sense of realism in both visual and audio.

[1590] Step 7: Obtaining emotion data

[1591] The device uses cameras and sensors to detect the user's facial expressions and biometric signals, recognizing the user's emotions in real time.

[1592] Step 8: Sending Emotion Data

[1593] The terminal transmits the recognized emotion data of the user to the server.

[1594] User operations

[1595] Step 1: Put on the VR goggles and log in

[1596] The user puts on the VR goggles and logs into the system. To log in, the user must enter their account information.

[1597] Step 2: Change your perspective

[1598] Users can freely change their viewpoint by moving their head, enjoying a 360-degree view, and can also focus on a specific player or play.

[1599] Step 3: Interactive Experience

[1600] Users receive real-time visual and auditory feedback, making them feel as if they are at the stadium, with sounds such as the cheers of the crowd and the footsteps of the players reproduced in real time.

[1601] Step 4: Displaying emotions

[1602] Depending on the user's emotions, specific camera angles and content can be automatically selected and displayed. For example, if the user is excited, important scenes can be displayed more prominently and sound effects can be enhanced.

[1603] Example 2

[1604] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1605] Conventional game viewing systems have limited visual and auditory experiences, making it difficult to fully recreate the sense of presence of a game. Furthermore, it is not possible to optimize the content played based on the user's real-time emotions and movements, making it difficult to provide a personalized experience for each user.

[1606] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1607] In this invention, the server includes means for collecting video data from multiple camera devices within the venue, means for integrating the collected video data to generate 360-degree omnidirectional video data, and means for transmitting the generated 360-degree omnidirectional video data to the terminal in real time, thereby enabling the user to experience 360-degree omnidirectional real-time video, and for the terminal to receive the video data and provide a 360-degree omnidirectional field of view, and to adjust the video and sound effects according to the user's movements.

[1608] Furthermore, by including a means for the terminal to decode the video data sent from the server and render the video in real time according to the user's viewpoint, and a means for acquiring the user's emotional data and adjusting the video and sound effects based on that, it is possible to provide a more immersive and individually optimized visual and auditory experience.

[1609] A "game venue" is a specific location where a sport or event is played.

[1610] "Filming equipment" refers to equipment such as cameras and drone cameras used to collect video data.

[1611] "Video data" refers to video information collected from a camera and includes visual content.

[1612] "Collection" refers to the act of collecting video data from multiple imaging devices.

[1613] "Integration" is the process of bringing together collected video data.

[1614] "360-degree omnidirectional video data" is video data that provides visual information from all directions.

[1615] "Generation" is the act of creating new video data.

[1616] "Real-time" refers to almost instantaneous processing, with very little delay.

[1617] A "terminal" is a device that receives and displays video data, such as VR goggles.

[1618] "Transmitting" is the act of sending data to another device or system.

[1619] "Rendering" is the process of generating images or videos from digital data.

[1620] A "gyro sensor" is a sensor that detects the movement of the user's head.

[1621] "Emotion data" is data that indicates the emotional state of the user.

[1622] "Decoding" is the process of converting encoded data into a playable format.

[1623] "Adjustment" is the act of changing visual and sound effects based on the user's movements and emotions.

[1624] The present invention relates to a system that uses VR technology and emotion recognition technology to enhance the sense of realism when watching a game. Specific embodiments of the present invention will be described below.

[1625] The server collects video data from multiple camera devices within the venue. This includes fixed cameras, mobile cameras, and drone cameras. The video data from each camera is temporarily stored in a buffer and sorted chronologically. The server then automatically recognizes overlapping areas of the footage and seamlessly combines them to generate 360-degree panoramic video data.

[1626] The generated 360-degree video data is encoded in real time using codecs such as H.264 or H.265. The encoded data is divided into packets of a fixed size and transmitted to the device with low latency.

[1627] The device (VR goggles) receives video data packets sent from the server. The received data is stored in a buffer and decoded. The decoded video data is rendered in real time according to the user's viewpoint. The device uses a built-in gyro sensor to detect the user's head movement and adjusts the video and sound effects based on that movement.

[1628] Furthermore, the device is equipped with a camera and biometric sensors to detect the user's facial expressions and biometric signals (e.g., heart rate, brain waves). This data is analyzed by the emotion engine and sent to the server as user emotion data.

[1629] The server receives the emotion data sent from the terminal and adjusts the video and sound effects based on the user's emotional state. For example, if the user is excited, the server adjusts the video to emphasize important scenes or dynamic play.

[1630] Users can use the above services by wearing VR goggles and logging in to the system. Users can freely change their viewpoint by moving their head, enjoying a 360-degree panoramic view. In addition, camera angles and content are automatically selected based on the user's emotions, creating an even more immersive experience.

[1631] Specific examples

[1632] For example, when streaming a basketball game, the following steps are taken: Video data is collected from multiple cameras installed in the stadium and stored in a temporary buffer. The images from each camera are rearranged in chronological order and seamlessly combined to generate a 360-degree panoramic video. The video is then encoded, divided into packets, and sent to the device.

[1633] The device reassembles the received data packets and renders the decoded video data in real time. The built-in gyro sensor detects the user's head movements and adjusts the video and audio in real time. The device's camera and biometric sensors also detect the user's facial expressions and biometric signals, and the emotion engine analyzes the emotional data and transmits it to the server.

[1634] The server receives the emotion data and adjusts the video and sound effects according to the user's emotions. For example, if the user is excited, the video can be adjusted to emphasize dynamic scenes and important plays, allowing the user to enjoy the game more realistically.

[1635] Prompt Sentence Examples

[1636] "I would like to develop a system that uses VR technology to make watching basketball games more immersive, and also utilizes emotion recognition technology to adjust video and sound effects. This system will incorporate a mechanism to optimize video in real time according to the user's movements and emotions. Please tell me specifically how you will design the system and how you will use emotion data."

[1637] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1638] Step 1:

[1639] "Video data collection"

[1640] The server collects video data from fixed cameras, mobile cameras, drone cameras, etc. within the venue. The video data from each camera is temporarily stored in the server's buffer.

[1641] Input: Video streams from each imaging modality.

[1642] Output: Multiple video data stored in a buffer.

[1643] Specific operation: The server uses the stream URL of the camera registered in advance to pull in video data using a specified protocol (such as RTSP).

[1644] Step 2:

[1645] "Sorting time series data"

[1646] The server then sorts the video data stored in the buffer into chronological order based on the timestamps of each camera, thereby synchronizing the video data.

[1647] Input: Multiple video data stored in a buffer.

[1648] Output: Video data sorted in chronological order.

[1649] What it does: The server analyzes the timestamp information and runs an algorithm to arrange the video frames from each camera in the proper order.

[1650] Step 3:

[1651] "Seamless joining process"

[1652] The server automatically recognizes overlapping areas of the images and seamlessly combines them to generate 360-degree panoramic video data.

[1653] Input: Multiple video data sorted in chronological order.

[1654] Output: 360-degree panoramic video data.

[1655] How it works: The server detects overlapping areas and uses image processing algorithms to merge them together, creating a continuous omnidirectional image.

[1656] Step 4:

[1657] "Encoding 360-degree video"

[1658] The server encodes the generated 360-degree panoramic video data in real time using the H.264 or H.265 codec.

[1659] Input: 360-degree omnidirectional video data.

[1660] Output: Encoded video data.

[1661] Specific operation: The server inputs the video data into the codec, which performs compression and encoding processes, thereby reducing the data size appropriately and making transmission more efficient.

[1662] Step 5:

[1663] "Sending video packets"

[1664] The server divides the encoded video data into packets of a fixed size and transmits them to the terminal with low latency.

[1665] Input: Encoded video data.

[1666] Output: Split data packets.

[1667] Specific operation: The server packetizes the video data and transmits it to the terminal using the UDP or TCP protocol.

[1668] Step 6:

[1669] "Receiving Data Packets"

[1670] The terminal receives the video data packets sent from the server, checking for packet loss and duplication to ensure accurate reception.

[1671] Input: Data packet from the server.

[1672] Output: Data packets that have been received and processed.

[1673] Specific operation: The terminal performs packet reception processing at the network layer, and processes error checks and retransmission requests.

[1674] Step 7:

[1675] "Assembling Data Packets"

[1676] The received data packets are reassembled into the original video data and stored in a temporary buffer.

[1677] Input: Received data packets.

[1678] Output: Video data stored in a temporary buffer.

[1679] Specific operation: The terminal rearranges, reconstructs, and temporarily stores the data packets, referring to the timestamp information.

[1680] Step 8:

[1681] "Video data decoding"

[1682] The terminal decodes the video data in the buffer using an H.264 or H.265 decoder and acquires it as successive video frames.

[1683] Input: Video data stored in a temporary buffer.

[1684] Output: Decoded video frames.

[1685] Specific operation: The terminal decodes the video frame using dedicated decoding hardware or software.

[1686] Step 9:

[1687] "Video rendering"

[1688] The decoded video is rendered in real time according to the user's viewpoint and displayed in the VR goggles.

[1689] Input: Decoded video frames.

[1690] Output: Rendered video frames.

[1691] Specific operation: The terminal uses a rendering engine to generate video frames according to the user's viewpoint and displays them on the display.

[1692] Step 10:

[1693] "Acquiring head movement data"

[1694] The device's built-in gyro sensor detects the user's head movements and uses that data to adjust visual and sound effects in real time.

[1695] Input: User's head movement.

[1696] Output: Coordinated video and audio.

[1697] Specific operation: The device obtains data in real time from the gyro sensor and dynamically adjusts image rendering and audio filtering.

[1698] Step 11:

[1699] "Acquiring emotion data"

[1700] The device's camera and biometric sensors capture biometric data such as the user's facial expressions and heart rate, which is then analyzed by the emotion engine.

[1701] Input: User's facial expression data and biometric signals.

[1702] Output: Parsed emotion data.

[1703] Specific operation: The device's camera extracts feature points using a facial expression recognition algorithm, and biometric sensors measure heart rate and skin galvanic response, which are then comprehensively analyzed by the emotion engine.

[1704] Step 12:

[1705] "Sending emotional data"

[1706] The analyzed emotional data is sent to the server in real time.

[1707] Input: Parsed emotion data.

[1708] Output: Emotion data sent to the server.

[1709] Specific operation: The device encodes the emotion data into an appropriate format and makes a transmission request to the server.

[1710] Step 13:

[1711] "Receiving and adjusting emotional data"

[1712] The server receives the emotional data sent from the terminal and adjusts the visual and sound effects based on the user's emotional state.

[1713] Input: Emotion data sent from the device.

[1714] Output: Coordinated video and audio.

[1715] Specific operation: The server analyzes the emotional data and dynamically generates content optimized for the user's emotional state, emphasizing specific visual scenes and sound effects.

[1716] Through the above processing steps, the system can provide the user with a realistic sense of being at the game, and can also provide a customized visual and auditory experience according to the user's emotional state.

[1717] (Application example 2)

[1718] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1719] While conventional VR-based game viewing systems provide users with a high sense of realism, they have difficulty providing a personalized experience that reflects the emotional and physical state of each user. Therefore, a new system is needed that can provide appropriate visual and audio effects in real time according to the emotional state and preferences of various users.

[1720] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for collecting video data from multiple camera devices within the venue, means for integrating the collected video data to generate 360-degree omnidirectional video data, and means for transmitting the generated 360-degree omnidirectional video data to the terminal in real time. This allows the user to experience the immersive atmosphere of the venue in real time. The terminal also includes means for adjusting video and audio effects according to the user's movements, means for sensing the user's facial expressions and biometric signals and recognizing the user's emotions in real time based on the detected facial expressions and biometric signals, and means for automatically adjusting video and audio effects based on the recognized emotional data. This allows the user to enjoy a personalized visual and audio experience that is tailored to their emotions and physical state.

[1721] "Game venue" refers to a location where various sports and entertainment matches are played.

[1722] "Filming equipment" refers to equipment such as cameras and drones used to collect video data.

[1723] "Video data" refers to video data collected by an imaging device.

[1724] "360-degree omnidirectional video data" refers to video that can be viewed in all directions and is generated by integrating video data collected from multiple imaging devices.

[1725] "Terminal" refers to a device (e.g., VR goggles) that receives video data and provides a visual and auditory experience to the user.

[1726] "User movement" refers to the user's body movement, particularly head movement, detected by the terminal.

[1727] "Sound effects" refers to the audio data corresponding to the video and the method of reproducing it.

[1728] "Expression" refers to the facial expression of the user.

[1729] "Biological signals" refer to physiological data such as a user's heart rate and brain waves.

[1730] "Emotion data" refers to the emotional state recognized based on the user's facial expressions and biometric signals.

[1731] "Real-time" refers to near-instant processing and transmission.

[1732] A "gyro sensor" refers to an inertial sensor used to detect the movement of a device.

[1733] "Synthesis" refers to the process of combining collected data into one linked data set.

[1734] "Seamlessly combining" refers to combining different video data continuously without interruption.

[1735] "Automatic adjustment" refers to the system automatically adjusting the video and audio based on the user's emotional data.

[1736] This invention relates to a system that enhances the sense of realism when a user watches a game using VR goggles and automatically adjusts video and sound effects according to the emotions of each individual user. The specific configuration and operation of the system are described below.

[1737] Server Processing

[1738] The server collects video data from multiple camera devices within the venue. The camera devices include fixed cameras, mobile cameras, and drone cameras. The collected video data is first stored in a buffer, and the video data from each camera is sorted chronologically. The server then automatically recognizes overlapping areas of the video and seamlessly combines them to generate 360-degree panoramic video data. The generated video data is then encoded in real time, divided into packets of a fixed size, and transmitted to the device. This transmission is performed with low latency, ensuring a smooth user experience.

[1739] Terminal handling

[1740] The device (VR goggles) receives real-time video data sent from the server. It reassembles the received data packets and stores them in a temporary buffer. It then decodes the video data and renders the video in real time according to the user's viewpoint. The device uses a built-in gyro sensor to acquire data on the user's head movement and adjusts the video and sound effects in real time based on that movement. The device also uses a camera to detect the user's facial expressions and biometric signals (e.g., heart rate, brain waves, etc.), and includes an emotion engine that recognizes the user's emotions in real time based on this data. The recognized emotion data is sent to the server in real time.

[1741] Emotion data processing on the server

[1742] The server receives the user's emotional data sent from the device. Based on this emotional data, it can automatically adjust the visual and sound effects. For example, if the user is excited, it can highlight important scenes and thrilling plays and enhance the sound effects.

[1743] User operations

[1744] First, users put on VR goggles and log in to the system. Users can freely change their viewpoint by moving their head, and enjoy a 360-degree panoramic view of their favorite scenes in real time. Furthermore, camera angles and content are automatically selected based on the user's emotions, creating an even more immersive experience.

[1745] Hardware and software used

[1746] The system is implemented using the following hardware and software:

[1747] Hardware: VR goggles (e.g., Oculus Rift, HTC Vive), fixed cameras, movable cameras, drone cameras, heart rate monitors, electroencephalographs

[1748] Software: OpenCV (used for image acquisition and processing), Python socket library (used for sending and receiving image data), Unity or Unreal Engine (for VR rendering and user interface construction)

[1749] Specific examples

[1750] For example, when live streaming a soccer match, the following prompt sentence can be input into the generative AI model to generate video that highlights important scenes from the match.

[1751] "Analyze the user's emotion recognition data in real time during a soccer match. If the user is excited, highlight the goal scene and the play before the goal. If the user is nervous, switch to a more relaxing video."

[1752] This invention allows users to enjoy a highly realistic and personalized viewing experience.

[1753] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1754] Step 1:

[1755] The server collects video data from multiple camera devices (fixed cameras, mobile cameras, drone cameras) within the venue. The collected video data is temporarily stored in a buffer, and the video data from each camera is sorted chronologically. This ensures that the video data is arranged in the appropriate order.

[1756] Step 2:

[1757] The server automatically recognizes overlapping areas in the collected video data and seamlessly combines them. It uses OpenCV to detect overlapping areas in each frame and combines them continuously. The combined video data becomes 360-degree omnidirectional video data.

[1758] Step 3:

[1759] The server encodes the generated 360-degree omnidirectional video data in real time, divides it into packets of a fixed size, and transmits them to the device. This transmission is performed with low latency, and the data is sent to the device via the network. Video codecs such as H.264 and HEVC are used for encoding.

[1760] Step 4:

[1761] The device (VR goggles) receives real-time video data sent from the server, reassembles the received data packets, and stores them in a temporary buffer. Then, it decodes the video data. For decoding, it uses FFmpeg or other decoding libraries.

[1762] Step 5:

[1763] The device uses a built-in gyro sensor to capture data on the user's head movements, and uses this data to render video and sound effects in real time. Rendering is done using Unity or Unreal Engine, and the video is displayed according to the user's viewpoint.

[1764] Step 6:

[1765] The device uses a camera and biosensors to detect the user's facial expressions and biometric signals (heart rate, brain waves, etc.). It then uses an emotion engine to analyze this data and extract the user's emotional data. For example, it uses a machine learning model to infer emotions from facial expressions and biometric signals.

[1766] Step 7:

[1767] The device transmits the recognized emotion data to the server in real time with low latency, so the user's emotional state is immediately conveyed to the server.

[1768] Step 8:

[1769] The server automatically adjusts visual and audio effects based on the received user emotion data. For example, if the user is excited, it will highlight important scenes or thrilling action scenes and enhance audio effects. It uses a generative AI model to calculate the appropriate visual and audio settings in real time.

[1770] Step 9:

[1771] Users put on VR goggles, log in, and begin watching the game. They can freely change their viewpoint by moving their head, and enjoy a 360-degree panoramic view of the game to watch any scene they like. Furthermore, camera angles and content are automatically selected based on the user's emotions, creating an even more immersive experience.

[1772] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1773] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1774] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1775] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1776] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1777] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1778] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1779] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1780] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1781] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1782] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1783] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1784] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1785] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1786] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1787] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1788] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1789] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1790] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1791] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1792] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1793] The following is further disclosed regarding the above embodiment.

[1794] (Claim 1)

[1795] means for collecting video data from a plurality of image capture devices within the venue;

[1796] A means for integrating the collected video data and generating 360-degree video data;

[1797] A means for transmitting the generated 360-degree omnidirectional video data to a terminal in real time;

[1798] a means for receiving the video data from the terminal to provide a 360-degree field of view and adjust video and sound effects according to the user's movements;

[1799] A system including:

[1800] (Claim 2)

[1801] The system according to claim 1, further comprising means for automatically recognizing overlapping portions of video data from a plurality of camera devices within a game venue and performing processing to seamlessly combine the data.

[1802] (Claim 3)

[1803] 10. The system of claim 1, wherein the terminal includes means for detecting the user's head movement using an internal gyro sensor and adjusting the visual and sound effects in real time based on the movement.

[1804] "Example 1"

[1805] (Claim 1)

[1806] means for collecting video data from a plurality of image capture devices within the venue;

[1807] a means for storing the collected video data in a buffer and sorting the data in chronological order based on timestamps;

[1808] A means to automatically recognize overlapping parts of video data and seamlessly combine them;

[1809] A means for encoding the combined 360-degree omnidirectional video data in real time, dividing it into packets, and transmitting them to the terminal;

[1810] means for reconstructing data packets received by the terminal and storing the data in a temporary buffer as video data;

[1811] a means for the terminal to decode the reconstructed video data and render the video in real time according to the user's viewpoint;

[1812] a means for the device to acquire the user's head movement using a built-in gyro sensor and adjust the visual and sound effects in real time based on the movement;

[1813] A system including:

[1814] (Claim 2)

[1815] 2. The system according to claim 1, further comprising a means for automatically recognizing overlapping portions of the image data collected from the imaging devices and performing processing to seamlessly combine the data.

[1816] (Claim 3)

[1817] 10. The system of claim 1, wherein the terminal includes means for detecting the user's head movement using an internal gyro sensor and adjusting the visual and sound effects in real time based on the movement.

[1818] "Application Example 1"

[1819] (Claim 1)

[1820] means for collecting video data from a plurality of image capture devices within the venue;

[1821] A means for integrating the collected video data and generating 360-degree video data;

[1822] A means for transmitting the generated 360-degree omnidirectional video data to a terminal in real time;

[1823] a means for receiving the video data from the terminal to provide a 360-degree field of view and adjust video and sound effects according to the user's movements;

[1824] a means for detecting user movement using a built-in gyro sensor in the terminal and adjusting visual and sound effects in real time based on the movement;

[1825] A means for temporarily storing the collected video data in a buffer and enabling seamless playback;

[1826] A system including:

[1827] (Claim 2)

[1828] The system according to claim 1, further comprising means for automatically recognizing overlapping portions of video data from a plurality of camera devices within a game venue and performing processing to seamlessly combine the data.

[1829] (Claim 3)

[1830] 2. The system according to claim 1, further comprising means for enabling the user to freely change the viewpoint in response to the user's movements, thereby enabling the user to select different viewpoints within the stadium.

[1831] (Claim 4)

[1832] 10. The system of claim 1, wherein the terminal includes means for receiving video data transmitted from the server, storing it in a temporary buffer, decoding it, and rendering the video in real time.

[1833] "Example 2: Combining Emotion Engines"

[1834] (Claim 1)

[1835] means for collecting video data from a plurality of image capture devices within the venue;

[1836] A means for integrating the collected video data and generating 360-degree video data;

[1837] A means for transmitting the generated 360-degree omnidirectional video data to a terminal in real time;

[1838] a means for receiving the video data from the terminal to provide a 360-degree field of view and adjust video and sound effects according to the user's movements;

[1839] A device that decodes video data sent from the server and renders the video in real time according to the user's viewpoint;

[1840] means for acquiring user emotion data and adjusting visual and audio effects based thereon;

[1841] A system including:

[1842] (Claim 2)

[1843] The system according to claim 1, further comprising means for automatically recognizing overlapping portions of video data from a plurality of camera devices within a game venue and performing processing to seamlessly combine the data.

[1844] (Claim 3)

[1845] 10. The system of claim 1, wherein the terminal includes means for detecting the user's head movement using an internal gyro sensor and adjusting the visual and sound effects in real time based on the movement.

[1846] "Application example 2 when combining emotion engines"

[1847] (Claim 1)

[1848] means for collecting video data from a plurality of image capture devices within the venue;

[1849] A means for integrating the collected video data and generating 360-degree video data;

[1850] A means for transmitting the generated 360-degree omnidirectional video data to a terminal in real time;

[1851] a means for receiving the video data from the terminal to provide a 360-degree field of view and adjust video and sound effects according to the user's movements;

[1852] A means for detecting a user's facial expression and biometric signals and recognizing the user's emotions in real time based thereon;

[1853] means for automatically adjusting visual and audio effects based on the recognized emotion data;

[1854] A system including:

[1855] (Claim 2)

[1856] The system according to claim 1, further comprising means for automatically recognizing overlapping portions of video data from a plurality of camera devices within a game venue and performing processing to seamlessly combine the data.

[1857] (Claim 3)

[1858] 10. The system of claim 1, wherein the terminal includes means for detecting the user's head movement using an internal gyro sensor and adjusting the visual and sound effects in real time based on the movement. [Explanation of symbols]

[1859] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. means for collecting video data from a plurality of image capture devices within the venue; A means for integrating the collected video data and generating 360-degree video data; A means for transmitting the generated 360-degree omnidirectional video data to a terminal in real time; a means for receiving the video data from the terminal to provide a 360-degree field of view and adjust video and sound effects according to the user's movements; A system including:

2. 2. The system according to claim 1, further comprising means for automatically recognizing overlapping portions of video data from a plurality of camera devices within a game venue and performing processing to seamlessly combine the data.

3. 10. The system of claim 1, wherein the terminal includes means for detecting the user's head movements using an internal gyro sensor and for adjusting visual and sound effects in real time based on the movements.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A