system
The system addresses the challenge of providing a three-dimensional and interactive sports viewing experience by capturing and converting player movements into holographic videos, enabling viewers to freely change perspectives and enhance immersion.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-21
- Publication Date
- 2026-05-07
AI Technical Summary
Current video technologies struggle to provide a three-dimensional visual experience with a sense of presence, especially in real-time sports viewing, due to technical difficulties in capturing and reproducing freely changeable viewpoints, which limits the immersion and interaction for viewers.
A system that captures players' movements from multiple angles using high-precision cameras, converts the data into three-dimensional representations using artificial intelligence, integrates these into holographic videos, and projects them onto a display device allowing free viewpoint changes, providing a three-dimensional and interactive viewing experience.
Enables viewers to experience sports events in a visually rich and immersive manner, allowing them to intuitively understand the overall situation and individual player movements with the ability to freely change viewpoints, enhancing the viewing experience.
Smart Images

Figure 2026074876000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a method for controlling a persona chatbot performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In recent years, there has been an active movement to seek new forms of sports viewing. However, it is difficult for current video technologies to provide a three-dimensional visual experience with a sense of presence. In particular, real-time three-dimensional videos with freely changeable viewpoints have technical difficulties and require large-scale equipment. There is a need to provide an environment in which viewers can watch games in a more three-dimensional and interactive manner.
Means for Solving the Problems
[0005] This invention provides a system that captures players' movements from multiple angles and uses an artificial intelligence model to create a three-dimensional representation of that data. Furthermore, it integrates the three-dimensional data of all players to generate a holographic video that reproduces the entire field in real time. This allows viewers to watch the game in a three-dimensional and interactive way using a display device with an interface that allows for free changes in viewpoint.
[0006] A "camera" is a device used to record an athlete's movements in two dimensions from multiple angles.
[0007] An "artificial intelligence model" is an algorithm or learning system used to convert two-dimensional motion data into three-dimensional data.
[0008] A "holographic video" is a three-dimensional image created based on three-dimensional data.
[0009] "Integration means" refers to a process or device for combining three-dimensional data of multiple players to recreate a video of the entire field.
[0010] A "display device" is a device that projects generated holographic images three-dimensionally onto the viewer and has the function of allowing the viewer to freely adjust their viewpoint. [Brief explanation of the drawing]
[0011] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5]This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2, which incorporates an emotion engine. [Figure 14] This is a sequence diagram showing the processing flow of the data processing system in Application Example 2, which combines an emotion engine. [Modes for carrying out the invention]
[0012] Hereinafter, an example of an embodiment of the system relating to the technology of this disclosure will be described with reference to the attached drawings.
[0013] First, let's explain the terminology used in the following explanation.
[0014] In the following embodiments, the labeled processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0015] In the following embodiments, the labeled RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0016] In the following embodiments, the labeled storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, etc.
[0017] In the following embodiments, the labeled communication I / F (Interface) is an interface including a communication processor and an antenna, etc. The communication I / F manages communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark), etc.
[0018] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0019] [First Embodiment]
[0020] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0021] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0022] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0023] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0024] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0025] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0026] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0027] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0028] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0029] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0030] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0031] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0032] One embodiment of the present invention is a system for reproducing the movements of an athlete in three dimensions in real time. This system is comprised of a combination of a camera, an artificial intelligence model, a holographic video generation device, an integration means, and a display device.
[0033] The server first controls the camera equipment, capturing the players' movements on the sports field in two dimensions using multiple cameras. This allows for recording the players' movements from various angles, providing multifaceted data. For example, in a baseball game, the server captures the pitcher's throwing motion from multiple angles and synchronizes the footage from each camera.
[0034] Next, the server uses an artificial intelligence model to convert this two-dimensional data into three dimensions. The AI model utilizes deep learning techniques to analyze the position of the athlete's skeleton and joints during movement. This allows the server to create an accurate three-dimensional model and generate holographic images based on it.
[0035] The server constructs holographic videos of each player based on the generated three-dimensional models. These holographic videos are designed to visualize the individual movements of each player in a realistic way. Furthermore, an integration mechanism works to recreate the entire game situation on the field using these videos. The server uses this to create a video that recreates the entire field while maintaining the coordination of movements between players.
[0036] Ultimately, the terminal delivers integrated holographic video to the audience through the display device. This allows users to watch the game in 3D and freely change their viewpoint. For example, spectators can experience watching the game up close, even from the comfort of their own homes. This system will realize a new style of sports viewing.
[0037] The following describes the processing flow.
[0038] Step 1:
[0039] The server controls a multi-camera system to capture the players' movements in two dimensions. Each camera captures the players from a different angle and transmits the video data to the server in real time.
[0040] Step 2:
[0041] The server preprocesses the received two-dimensional video data, unifying the resolution and removing noise. This process makes the data suitable for three-dimensional conversion.
[0042] Step 3:
[0043] The server uses an artificial intelligence model to generate a three-dimensional model from pre-processed two-dimensional data. A deep learning algorithm analyzes the position of the athlete's skeleton and joints to construct an accurate three-dimensional model.
[0044] Step 4:
[0045] The server generates holographic videos based on three-dimensional models. This is intended to recreate the movements of each player in three dimensions, resulting in footage that can be observed from a 360-degree perspective.
[0046] Step 5:
[0047] The server integrates holographic videos of all players to create a unified video that recreates the entire match. It synchronizes the movements of the players and handles occlusion to accurately recreate the flow of the game.
[0048] Step 6:
[0049] The terminal projects integrated holographic video received from the server onto a display device. This device features an interface that allows the user to freely adjust their viewpoint, providing a visually three-dimensional video experience.
[0050] Step 7:
[0051] Users can watch the match footage displayed on the screen from various perspectives. By changing viewpoints and focusing on specific players, an immersive and interactive viewing experience is possible.
[0052] (Example 1)
[0053] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0054] In traditional sports viewing, the limited visual information resulted in a lack of realism, making it difficult for spectators to fully understand the overall situation on the field and the movements of individual players. This led to a decline in the quality of the viewing experience and a decrease in interest in sporting events.
[0055] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0056] In this invention, the server includes a recording device for recording the movements of players in two dimensions, a machine learning model for converting the recorded two-dimensional movement information into three dimensions, and means for generating a stereoscopic image using the three-dimensional information. This enables a new viewing experience for spectators that is visually rich and allows them to intuitively understand the movements of the players and the overall flow of the game.
[0057] A "recording device" is a device used to record an athlete's movements in two dimensions from multiple perspectives.
[0058] A "machine learning model" is a program that implements algorithms to analyze recorded two-dimensional motion information and convert it into three dimensions.
[0059] "3D images" are visually immersive and realistic images generated based on three-dimensional information.
[0060] "Three-dimensionalization" is the process of reconstructing three-dimensional shapes and movements based on data recorded in two dimensions.
[0061] A "presentation device" is a device used to visually present generated 3D images.
[0062] A "combination means" is a program or device that integrates 3D images of multiple players to recreate the overall situation.
[0063] An "operation interface" is an input method that allows the user to adjust the viewing angle and other settings via a display device.
[0064] This invention is a system that reproduces the movements of athletes in sports events in real time and in three dimensions. The central role of the system is played by a server. The server controls recording devices and records the movements of athletes in two dimensions from multiple viewpoints. The recording devices used are high-precision cameras, which enable the acquisition of multifaceted and highly accurate data.
[0065] Next, the server uses a machine learning model to analyze the recorded two-dimensional motion information and convert it into three dimensions. This process utilizes deep learning techniques to precisely model the player's skeleton and joint positions. This three-dimensional modeling makes it possible to realistically reproduce the player's movements.
[0066] Based on the generated three-dimensional data, the server creates a stereoscopic image. This stereoscopic image is visualized as a realistic and immersive visual experience using specialized holography generation technology. This provides viewers with a powerful and engaging sports viewing experience.
[0067] The integrated stereoscopic imagery is provided to the user via a display device. A high-definition holographic display is used as the display device. This allows the user to freely change their viewpoint, enabling them to observe the overall flow of the match and the individual plays of players in detail. For example, in a soccer match, the user can experience the game from the perspective of any player.
[0068] An example of a prompt message is, "Please provide instructions on how to recreate a baseball pitcher's throwing motion in three dimensions and display it as a holographic video." This invention can offer a new style of sports viewing, providing spectators with rich information and an entertaining experience.
[0069] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0070] Step 1:
[0071] The server controls the recording devices and records the players' movements in two dimensions. It receives real-time movements of the players on the sports field as input. The server synchronizes multiple high-precision cameras and captures the players' movements from each viewpoint. As output, it acquires multi-faceted two-dimensional video data, recording the players' positions and movements in detail.
[0072] Step 2:
[0073] The server inputs the acquired two-dimensional data into a machine learning model and converts it into three-dimensional data. The two-dimensional video data obtained in step 1 is used as input. The server utilizes deep learning technology to analyze the position of the athlete's skeleton and joints. As output, it generates an accurate three-dimensional model, allowing for a three-dimensional reproduction of the athlete's movements.
[0074] Step 3:
[0075] The server uses the generated 3D data to create a stereoscopic image. The 3D model generated in step 2 is used as input. The server employs holographic generation technology to visually represent the players' movements. The output is a powerful stereoscopic image, providing an immersive experience for the audience.
[0076] Step 4:
[0077] The server integrates multiple 3D images to recreate the entire game situation on the field. It uses 3D images of individual players as input. The server employs video editing technology to integrate the flow of the game into a single video while maintaining the coordination between players. The output provides a video that visually recreates the entire field.
[0078] Step 5:
[0079] The terminal provides the user with integrated 3D images through a display device. The integrated images from step 4 are used as input. The terminal utilizes a high-definition holographic display, allowing the user to freely change their viewing angle. As output, spectators can experience a free-viewpoint perspective, similar to watching the event live.
[0080] (Application Example 1)
[0081] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0082] In traditional sports viewing, spectators are limited to watching the game through two-dimensional footage from fixed camera angles, making it difficult to freely experience the sense of presence and the details of the players' movements. In particular, when spectators watch the game from a remote location such as their home, it is difficult to obtain a sense of presence that is close to the visual experience of being there in person.
[0083] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0084] In this invention, the server includes imaging means for acquiring the movements of players in two dimensions, an artificial intelligence system for converting the acquired two-dimensional movement information into three dimensions, and generation means for generating holographic audio and video information using the three-dimensional information. This makes it possible for spectators to experience a three-dimensional and immersive sports match from anywhere in the world, while freely changing their viewpoint.
[0085] "Imaging means" refers to a device for recording the movements of athletes in two dimensions from multiple viewpoints.
[0086] An "artificial intelligence system" is a technical means for analyzing two-dimensional motion information and converting it into three-dimensional data.
[0087] "Generation means" refers to an apparatus or method for creating holographic audio and video information using three-dimensional information.
[0088] "Integration means" refers to a technical method or machine for combining holographic video information of multiple athletes to recreate the entire stadium.
[0089] A "display means" is a device that visualizes the reproduced holographic image in three dimensions and provides it to the spectators.
[0090] To implement the present invention, the following configuration requirements are necessary. The server controls the imaging means and records the athlete's movements two-dimensionally from multiple directions. For example, multiple cameras are placed in a sports stadium to acquire the athlete's movements as video. This acquired two-dimensional video data is sent to an artificial intelligence system. This system uses deep learning technology to analyze the positions of the athlete's skeleton and joints in the video and converts it into three-dimensional data.
[0091] The generation means generates holographic audio and video information based on three-dimensional data, and the integration means combines holographic images of multiple athletes to recreate the entire stadium. A three-dimensional rendering engine could be used here. Finally, this recreated image is provided to spectators through the display means. Spectators can enjoy a more immersive sports viewing experience using devices such as head-mounted displays.
[0092] For example, in a soccer match, spectators can wear a head-mounted display at home and move freely around the field, experiencing the players' movements in real time. This entire process requires a low-latency network and high-speed data processing to achieve real-time performance.
[0093] An example of a prompt for a generative AI model is as follows: "Use a deep learning model to generate a three-dimensional holographic model from the following two-dimensional video data. The model should represent the movements of an athlete, and accurately analyze the skeletal structure and joint positions while maintaining a sense of realism."
[0094] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0095] Step 1:
[0096] The server controls the imaging equipment and acquires the players' movements as two-dimensional video. The input data consists of video from each camera, and the output is two-dimensional video data of the players' movements captured from multiple directions. The server receives this data in real time, ensuring that information from various angles is available.
[0097] Step 2:
[0098] The server supplies two-dimensional video data to the artificial intelligence system and performs the process of converting it into three-dimensional data. In this process, the input is two-dimensional video data from each camera, and the output is the three-dimensional position information and skeletal data of the players. Deep learning technology is used to analyze features within the images and generate three-dimensional data.
[0099] Step 3:
[0100] The generation method generates holographic audio and video information from three-dimensional data. The input here is the converted three-dimensional data, and the output is a three-dimensional holographic model. The server uses a rendering engine to convert the data into a visually understandable format.
[0101] Step 4:
[0102] The integration mechanism combines holographic images of multiple athletes to recreate the overall situation of the stadium. The input to this process is holographic models of individual athletes, and the output is a unified holographic image of the entire stadium. The server considers the coordination of movements between athletes to construct a natural-looking image.
[0103] Step 5:
[0104] The terminal uses a display device to provide spectators with three-dimensional holographic images. At this stage, the input is an integrated holographic image, and the output is a real-time image experienced by the spectator. Users receive the image through a head-mounted display or similar device and can freely change their viewpoint.
[0105] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0106] One embodiment of the present invention is a system that reproduces athlete movements in real time in three dimensions and further recognizes the user's emotions to optimize the video experience. This system consists of a shooting device, an artificial intelligence model, a holographic video generation device, an integration means, a display device, and an emotion engine.
[0107] First, the server controls the camera equipment to capture the movements of players on the sports field in two dimensions. By capturing players from different angles using multiple cameras and transmitting the data to the server in real time, detailed movement data is obtained. For example, in a soccer match, the server simultaneously captures a forward player's goal from multiple viewpoints.
[0108] Next, the server uses an artificial intelligence model to convert this two-dimensional data into three dimensions. Using deep learning techniques, it analyzes the position of the athlete's skeleton and joints during movement and generates a three-dimensional model. This three-dimensional model forms the basis of the holographic video.
[0109] The server creates holographic videos of each player based on the generated three-dimensional models. This provides a three-dimensional representation of the players' movements. The server then integrates the holographic videos of multiple players into a single integrated video that recreates the entire field. This integration process allows for a faithful reproduction of the flow of the game and the coordinated movements between players.
[0110] The device projects this integrated holographic video onto a display device, providing the audience with a visually immersive three-dimensional image. Furthermore, the device is equipped with an emotion engine that recognizes the user's emotions. This emotion engine analyzes the user's facial expressions and physical reactions to determine the user's emotional state.
[0111] Once the user's emotional state is determined, the device dynamically adjusts the video content based on that information. For example, if the user is excited, the video experience is optimized by increasing the frequency of replays or zooming in on the plays of specific players. Users can enjoy a viewing experience customized to their emotional state, allowing them to experience the game more deeply and interactively. This system embodies a new way of watching sports.
[0112] The following describes the processing flow.
[0113] Step 1:
[0114] The server controls the camera equipment, capturing the players' movements in two dimensions from multiple cameras. The collected video data is transmitted to the server in real time and stored in a database.
[0115] Step 2:
[0116] The server preprocesses the received video data. Specifically, it performs noise reduction and standardizes the resolution to prepare the data for analysis by artificial intelligence models.
[0117] Step 3:
[0118] The server uses an artificial intelligence model to generate a three-dimensional model from processed two-dimensional data. It analyzes the player's movements using deep learning techniques to create a highly accurate three-dimensional model.
[0119] Step 4:
[0120] The server generates holographic videos based on the generated three-dimensional models. This creates three-dimensional motion videos for each player, resulting in videos that can be observed from multiple angles.
[0121] Step 5:
[0122] The server integrates holographic videos of all players to generate a unified video that recreates the entire field. It synchronizes the players' position information and movements to recreate the coordinated flow of the game.
[0123] Step 6:
[0124] The terminal projects integrated holographic video onto a display device. Users can freely adjust their viewpoint through the display device, allowing them to watch the match in 3D from various perspectives.
[0125] Step 7:
[0126] The device uses an emotion engine to recognize the user's emotions. Through cameras and sensors, it analyzes the user's facial expressions, voice, heart rate, etc., to determine the user's emotional state.
[0127] Step 8:
[0128] The device adjusts the displayed content in real time based on recognized emotion data. For example, if the user is excited, it optimizes the video content by increasing replays or focusing on a specific player.
[0129] Step 9:
[0130] Users can enjoy an interactive viewing experience. The video content changes according to their emotions, creating a more personalized sports viewing experience.
[0131] (Example 2)
[0132] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0133] Conventional technologies have made it difficult to reproduce players' movements in real time and in three dimensions, and to dynamically optimize the video experience based on the user's emotions. In particular, there is a need for a system that can faithfully reproduce the movements of the entire match while adjusting the video according to the user's perspective and emotions. To solve this problem, technologies that capture players' movements from multiple angles and represent them in three dimensions, as well as video control technologies based on the user's emotions, are necessary.
[0134] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0135] In this invention, the server includes a device that provides a method for capturing a player's movements in two dimensions, a machine learning model for converting the acquired two-dimensional movement information into three dimensions, and means for generating a stereoscopic video using the three-dimensional information. This allows users to visualize the player's movements in three dimensions in real time. Furthermore, by using a device equipped with an emotion recognition function for analyzing the user's emotional state and dynamically adjusting the video, it becomes possible to provide a customized video experience for each individual user.
[0136] "Filming method" refers to the technique of capturing a player's movements precisely in two dimensions by filming from multiple viewpoints.
[0137] A "device" is a piece of equipment equipped with the hardware and software necessary to process the acquired two-dimensional video data.
[0138] A "machine learning model" is an artificial intelligence technology used to convert acquired two-dimensional motion information into three-dimensional data, and is an algorithm built on deep learning.
[0139] A "3D display video" is a video that uses three-dimensional information to create a three-dimensional representation of the players' movements.
[0140] "Integration means" refers to a technology used to recreate the movement of the entire stadium by combining 3D display videos of individual athletes.
[0141] A "display device" is a device that presents reproduced stereoscopic images to the user in three dimensions.
[0142] The "emotion recognition function" is an analysis system that analyzes the user's emotional state and adjusts the video content based on that information.
[0143] This invention is a system that reproduces athletes' movements in real time and in three dimensions, captures the user's emotions, and optimizes the video experience. This system consists of a complex including a shooting device, a machine learning model, a stereoscopic video generation means, an integration means, a display device, and an emotion recognition function.
[0144] The server first controls the camera equipment to capture images of the players' movements in two dimensions. This equipment includes multiple cameras, each capturing players on the sports field from a different viewpoint. For example, in a soccer match, capturing players' movements from multiple angles allows for the collection of more detailed motion data. This data is transmitted to the server in real time, preparing it for the next processing step.
[0145] Next, the server uses machine learning models to convert the acquired two-dimensional data into three-dimensional data. Typical software used includes deep learning frameworks such as TENSORFLOW® and PyTorch, which accurately analyze the position of the athlete's skeleton and joints and generate a three-dimensional model.
[0146] Using a 3D video generation system, the server creates 3D videos of each player based on the generated 3D models. This is done using a game engine such as Unity, enabling realistic movement simulations. The generated videos have the characteristic of being displayable flexibly, without being restricted to specific angles or viewing perspectives.
[0147] Following this, the server integrates the 3D video feeds of each player to recreate the overall movement of the stadium. By utilizing video processing technology, the coordination and positional relationships between players are accurately reproduced, and the movement of the entire match is provided as a single, unified video.
[0148] The device delivers this integrated video to the user through a display device. The display device enables three-dimensional visualization, providing a high level of visual immersion.
[0149] Furthermore, the device is equipped with emotion recognition capabilities to optimize the video experience by recognizing the user's emotional state. Using sensors and cameras, it analyzes the user's facial expressions and physical reactions in real time. Based on this analysis, it can dynamically adjust the video content. For example, if it determines that the user is excited, it can zoom in on a specific player's performance, providing an interactive video experience tailored to the user's state.
[0150] A concrete example of a prompt message might be, "Recreate the player's goal in detail in three dimensions, and replay it in slow motion to match the user's surprise." This system would allow users to experience sports viewing in a deeper, more interactive way.
[0151] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0152] Step 1:
[0153] The server controls the camera equipment to capture the players' movements in two dimensions. It uses video data from multiple cameras as input. Specifically, each camera captures the players from a different viewpoint and transmits the data to the server in real time. The data is processed by integrating the video from each viewpoint to capture the players' movements in detail. The output is the integrated two-dimensional video data.
[0154] Step 2:
[0155] The server generates a three-dimensional model using a machine learning model based on the acquired two-dimensional video data. The input is integrated two-dimensional video data. Specifically, it utilizes a deep learning framework to extract positional information of the athlete's skeleton and joints. As a data calculation, it reconstructs the three-dimensional structure from this information. The output is a model data that reproduces the athlete in three dimensions.
[0156] Step 3:
[0157] The server creates a 3D video based on the generated 3D model. The input is 3D model data. Specifically, it uses a game engine to add movement to the 3D model, resulting in a realistic video. Data processing involves adjusting smooth transitions between movements. The output is video data with a three-dimensional visual effect.
[0158] Step 4:
[0159] The server integrates 3D video data of multiple players to recreate their overall positional relationships. Individual 3D video data is used as input. The specific operation involves using video processing techniques to synchronize and align the movements of each player. Data processing combines different video clips to represent the movement of the entire stadium. The output is integrated video data of the entire field.
[0160] Step 5:
[0161] The terminal provides integrated video data to the user through a display device. Integrated holographic video data is used as input. Specifically, it provides the user with a visually interactive experience using a three-dimensional display device. The output is the three-dimensional image received by the user as a visual experience.
[0162] Step 6:
[0163] The device dynamically adjusts the video by analyzing the user's emotional state using emotion recognition technology. Inputs include user facial expressions and physical response data. Specifically, it collects data using cameras and sensors and determines the user's emotions using an emotion analysis algorithm. Data processing customizes the video content based on the emotional information. The output is a video experience optimized according to the user's state.
[0164] (Application Example 2)
[0165] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0166] In recent years, there has been a growing demand from sports spectators for an immersive, on-the-ground experience. However, conventional video technology has struggled to realistically reproduce athletes' movements in three dimensions and has been unable to customize the video experience to suit the individual emotions of each spectator. To address these challenges, there is a need for a system that can reproduce athletes' movements in real time and in three dimensions, while also dynamically adjusting the video experience based on the emotions of the spectators.
[0167] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0168] This invention includes a server that controls an imaging device for acquiring player movements in two dimensions, utilizes an intelligent model for converting the acquired two-dimensional movement information into three dimensions, generates an image using the three-dimensional information, an integration means for integrating the images of all players to reproduce the entire area, a display device for displaying the reproduced stereoscopic image, an emotion recognition engine for recognizing the user's emotional state, and an adjustment means for dynamically adjusting the image based on the user's emotional state. This enables a three-dimensional and immersive sports viewing experience, allowing spectators to enjoy a viewing experience optimized according to their individual emotions.
[0169] An "imaging device" is a device used to capture the movements of athletes in two dimensions.
[0170] An "intelligent model" is a machine learning-based model that converts acquired two-dimensional motion information into three-dimensional data.
[0171] "Image generation means" refers to technology and apparatus for generating images using three-dimensional information.
[0172] "Integration means" refers to methods and techniques for integrating video footage of all players to recreate the entire domain.
[0173] A "display device" is a device used to display reproduced 3D images.
[0174] An "emotion recognition engine" is software or hardware used to recognize the emotional state of a user.
[0175] "Adjustment means" refers to methods and techniques for dynamically adjusting the image based on the user's emotional state.
[0176] To implement this invention, multiple pieces of hardware and software must work in coordination with each other. The server first acquires the player's movements in two dimensions using an imaging device. This imaging device incorporates multiple cameras to capture the player from various angles. The acquired images are converted into three dimensions in real time using an intelligent model. This intelligent model is based on deep learning technology and analyzes the player's skeletal structure and movements in detail.
[0177] The three-dimensional model is converted into a holographic image by an image generation device. This image is then further integrated by an integration device to faithfully reproduce the entire field. The server transmits this integrated 3D image to display devices, providing spectators with an immersive viewing experience.
[0178] The device analyzes the audience's emotional state using an emotion recognition engine. The emotion recognition engine senses the audience's facial expressions and physical reactions to determine emotions such as excitement and joy. Based on this emotional information, the adjustment mechanism then customizes the video appropriately. For example, if the user's level of excitement is high, adjustments such as highlighting specific scenes may be made.
[0179] For example, in a soccer match, when a forward player makes an important play, the scene is zoomed in on, allowing spectators to experience the moment more vividly. By providing optimal visuals in response to the spectators' emotions in this way, a more personal and interactive viewing experience can be achieved.
[0180] Example prompt: "When you want to see an exciting moment in the game in more detail, recognize the emotion and zoom in on that scene."
[0181] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0182] Step 1:
[0183] The server controls the imaging device and uses multiple cameras to capture the players' movements in two dimensions. The input for this step is a video feed of the match, and the output is two-dimensional data capturing the players' movements. This data consists of frames taken from different angles by each camera.
[0184] Step 2:
[0185] The server inputs the acquired two-dimensional data into an intelligent model and converts it to three dimensions in real time. This conversion uses deep learning technology, and based on the two-dimensional input data, the skeleton and movements of each player are reconstructed in three-dimensional space. The output is a three-dimensional model of the player.
[0186] Step 3:
[0187] The server generates holographic images based on a three-dimensional model using image generation means. The input for this step is a three-dimensional model, and the output is a holographic image that realistically represents the player's movements. This image is optimized to provide a realistic visual experience.
[0188] Step 4:
[0189] The server integrates the generated holographic images using an integration mechanism to recreate the entire field of the match. The input is holographic images of each player, and the output is an integrated 3D image that faithfully reproduces the flow of the match and the coordination between players. This creates a realistic image that allows viewers to grasp the overall scene.
[0190] Step 5:
[0191] The terminal provides users with integrated stereoscopic images using a display device. The input is integrated holographic images, and the output is a visually immersive stereoscopic experience. The display device enables users to watch the match live.
[0192] Step 6:
[0193] An emotion recognition engine installed in the user's device operates, analyzing the user's facial expressions and movements to determine their emotional state. The input for this step is the user's real-time biometric data, and the output is the evaluation of their emotional state. Emotion recognition is performed using camera and sensor technology.
[0194] Step 7:
[0195] The device dynamically adjusts the video using adjustment mechanisms based on emotional state. The input is the emotional state obtained by the emotion recognition engine, and the output is video customized according to that emotion. An example of adjustment is zooming in on the main scene when the user is excited.
[0196] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0197] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0198] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0199] [Second Embodiment]
[0200] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0201] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0202] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0203] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0204] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0205] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0206] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0207] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0208] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0209] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0210] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0211] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0212] One embodiment of the present invention is a system for reproducing the movements of an athlete in three dimensions in real time. This system is comprised of a combination of a camera, an artificial intelligence model, a holographic video generation device, an integration means, and a display device.
[0213] The server first controls the camera equipment, capturing the players' movements on the sports field in two dimensions using multiple cameras. This allows for recording the players' movements from various angles, providing multifaceted data. For example, in a baseball game, the server captures the pitcher's throwing motion from multiple angles and synchronizes the footage from each camera.
[0214] Next, the server uses an artificial intelligence model to convert this two-dimensional data into three dimensions. The AI model utilizes deep learning techniques to analyze the position of the athlete's skeleton and joints during movement. This allows the server to create an accurate three-dimensional model and generate holographic images based on it.
[0215] The server constructs holographic videos of each player based on the generated three-dimensional models. These holographic videos are designed to visualize the individual movements of each player in a realistic way. Furthermore, an integration mechanism works to recreate the entire game situation on the field using these videos. The server uses this to create a video that recreates the entire field while maintaining the coordination of movements between players.
[0216] Ultimately, the terminal delivers integrated holographic video to the audience through the display device. This allows users to watch the game in 3D and freely change their viewpoint. For example, spectators can experience watching the game up close, even from the comfort of their own homes. This system will realize a new style of sports viewing.
[0217] The following describes the processing flow.
[0218] Step 1:
[0219] The server controls a multi-camera system to capture the players' movements in two dimensions. Each camera captures the players from a different angle and transmits the video data to the server in real time.
[0220] Step 2:
[0221] The server preprocesses the received two-dimensional video data, unifying the resolution and removing noise. This process makes the data suitable for three-dimensional conversion.
[0222] Step 3:
[0223] The server uses an artificial intelligence model to generate a three-dimensional model from pre-processed two-dimensional data. A deep learning algorithm analyzes the position of the athlete's skeleton and joints to construct an accurate three-dimensional model.
[0224] Step 4:
[0225] The server generates holographic videos based on three-dimensional models. This is intended to recreate the movements of each player in three dimensions, resulting in footage that can be observed from a 360-degree perspective.
[0226] Step 5:
[0227] The server integrates holographic videos of all players to create a unified video that recreates the entire match. It synchronizes the movements of the players and handles occlusion to accurately recreate the flow of the game.
[0228] Step 6:
[0229] The terminal projects integrated holographic video received from the server onto a display device. This device features an interface that allows the user to freely adjust their viewpoint, providing a visually three-dimensional video experience.
[0230] Step 7:
[0231] Users can watch the match footage displayed on the screen from various perspectives. By changing viewpoints and focusing on specific players, an immersive and interactive viewing experience is possible.
[0232] (Example 1)
[0233] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0234] In traditional sports viewing, the limited visual information resulted in a lack of realism, making it difficult for spectators to fully understand the overall situation on the field and the movements of individual players. This led to a decline in the quality of the viewing experience and a decrease in interest in sporting events.
[0235] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0236] In this invention, the server includes a recording device for recording the movements of players in two dimensions, a machine learning model for converting the recorded two-dimensional movement information into three dimensions, and means for generating a stereoscopic image using the three-dimensional information. This enables a new viewing experience for spectators that is visually rich and allows them to intuitively understand the movements of the players and the overall flow of the game.
[0237] A "recording device" is a device used to record an athlete's movements in two dimensions from multiple perspectives.
[0238] A "machine learning model" is a program that implements algorithms to analyze recorded two-dimensional motion information and convert it into three dimensions.
[0239] "3D images" are visually immersive and realistic images generated based on three-dimensional information.
[0240] "Three-dimensionalization" is the process of reconstructing three-dimensional shapes and movements based on data recorded in two dimensions.
[0241] A "presentation device" is a device used to visually present generated 3D images.
[0242] A "combination means" is a program or device that integrates 3D images of multiple players to recreate the overall situation.
[0243] An "operation interface" is an input method that allows the user to adjust the viewing angle and other settings via a display device.
[0244] This invention is a system that reproduces the movements of athletes in sports events in real time and in three dimensions. The central role of the system is played by a server. The server controls recording devices and records the movements of athletes in two dimensions from multiple viewpoints. The recording devices used are high-precision cameras, which enable the acquisition of multifaceted and highly accurate data.
[0245] Next, the server uses a machine learning model to analyze the recorded two-dimensional motion information and convert it into three dimensions. This process utilizes deep learning techniques to precisely model the player's skeleton and joint positions. This three-dimensional modeling makes it possible to realistically reproduce the player's movements.
[0246] Based on the generated three-dimensional data, the server creates a stereoscopic image. This stereoscopic image is visualized as a realistic and immersive visual experience using specialized holography generation technology. This provides viewers with a powerful and engaging sports viewing experience.
[0247] The integrated stereoscopic imagery is provided to the user via a display device. A high-definition holographic display is used as the display device. This allows the user to freely change their viewpoint, enabling them to observe the overall flow of the match and the individual plays of players in detail. For example, in a soccer match, the user can experience the game from the perspective of any player.
[0248] An example of a prompt message is, "Please provide instructions on how to recreate a baseball pitcher's throwing motion in three dimensions and display it as a holographic video." This invention can offer a new style of sports viewing, providing spectators with rich information and an entertaining experience.
[0249] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0250] Step 1:
[0251] The server controls the recording devices and records the players' movements in two dimensions. It receives real-time movements of the players on the sports field as input. The server synchronizes multiple high-precision cameras and captures the players' movements from each viewpoint. As output, it acquires multi-faceted two-dimensional video data, recording the players' positions and movements in detail.
[0252] Step 2:
[0253] The server inputs the acquired two-dimensional data into a machine learning model and converts it into three-dimensional data. The two-dimensional video data obtained in step 1 is used as input. The server utilizes deep learning technology to analyze the position of the athlete's skeleton and joints. As output, it generates an accurate three-dimensional model, allowing for a three-dimensional reproduction of the athlete's movements.
[0254] Step 3:
[0255] The server uses the generated 3D data to create a stereoscopic image. The 3D model generated in step 2 is used as input. The server employs holographic generation technology to visually represent the players' movements. The output is a powerful stereoscopic image, providing an immersive experience for the audience.
[0256] Step 4:
[0257] The server integrates multiple 3D images to recreate the entire game situation on the field. It uses 3D images of individual players as input. The server employs video editing technology to integrate the flow of the game into a single video while maintaining the coordination between players. The output provides a video that visually recreates the entire field.
[0258] Step 5:
[0259] The terminal provides the user with integrated 3D images through a display device. The integrated images from step 4 are used as input. The terminal utilizes a high-definition holographic display, allowing the user to freely change their viewing angle. As output, spectators can experience a free-viewpoint perspective, similar to watching the event live.
[0260] (Application Example 1)
[0261] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0262] In traditional sports viewing, spectators are limited to watching the game through two-dimensional footage from fixed camera angles, making it difficult to freely experience the sense of presence and the details of the players' movements. In particular, when spectators watch the game from a remote location such as their home, it is difficult to obtain a sense of presence that is close to the visual experience of being there in person.
[0263] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0264] In this invention, the server includes imaging means for acquiring the movements of players in two dimensions, an artificial intelligence system for converting the acquired two-dimensional movement information into three dimensions, and generation means for generating holographic audio and video information using the three-dimensional information. This makes it possible for spectators to experience a three-dimensional and immersive sports match from anywhere in the world, while freely changing their viewpoint.
[0265] "Imaging means" refers to a device for recording the movements of athletes in two dimensions from multiple viewpoints.
[0266] An "artificial intelligence system" is a technical means for analyzing two-dimensional motion information and converting it into three-dimensional data.
[0267] "Generation means" refers to an apparatus or method for creating holographic audio and video information using three-dimensional information.
[0268] "Integration means" refers to a technical method or machine for combining holographic video information of multiple athletes to recreate the entire stadium.
[0269] A "display means" is a device that visualizes the reproduced holographic image in three dimensions and provides it to the spectators.
[0270] To implement the present invention, the following configuration requirements are necessary. The server controls the imaging means and records the athlete's movements two-dimensionally from multiple directions. For example, multiple cameras are placed in a sports stadium to acquire the athlete's movements as video. This acquired two-dimensional video data is sent to an artificial intelligence system. This system uses deep learning technology to analyze the positions of the athlete's skeleton and joints in the video and converts it into three-dimensional data.
[0271] The generation means generates holographic audio and video information based on three-dimensional data, and the integration means combines holographic images of multiple athletes to recreate the entire stadium. A three-dimensional rendering engine could be used here. Finally, this recreated image is provided to spectators through the display means. Spectators can enjoy a more immersive sports viewing experience using devices such as head-mounted displays.
[0272] For example, in a soccer match, spectators can wear a head-mounted display at home and move freely around the field, experiencing the players' movements in real time. This entire process requires a low-latency network and high-speed data processing to achieve real-time performance.
[0273] An example of a prompt for a generative AI model is as follows: "Use a deep learning model to generate a three-dimensional holographic model from the following two-dimensional video data. The model should represent the movements of an athlete, and accurately analyze the skeletal structure and joint positions while maintaining a sense of realism."
[0274] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0275] Step 1:
[0276] The server controls the imaging equipment and acquires the players' movements as two-dimensional video. The input data consists of video from each camera, and the output is two-dimensional video data of the players' movements captured from multiple directions. The server receives this data in real time, ensuring that information from various angles is available.
[0277] Step 2:
[0278] The server supplies two-dimensional video data to the artificial intelligence system and performs the process of converting it into three-dimensional data. In this process, the input is two-dimensional video data from each camera, and the output is the three-dimensional position information and skeletal data of the players. Deep learning technology is used to analyze features within the images and generate three-dimensional data.
[0279] Step 3:
[0280] The generation method generates holographic audio and video information from three-dimensional data. The input here is the converted three-dimensional data, and the output is a three-dimensional holographic model. The server uses a rendering engine to convert the data into a visually understandable format.
[0281] Step 4:
[0282] The integration means combines the holographic images of multiple players to reproduce the situation of the entire stadium. The input for this process is the holographic models of individual players, and the output is the integrated holographic image of the entire stadium. The server constructs a natural image considering the coordination of movements among the players.
[0283] Step 5:
[0284] The terminal uses display means to provide the holographic image to the spectators in a three-dimensional manner. The input at this stage is the integrated holographic image, and the output is the real-time image experienced by the spectators. The user can receive the image through a head-mounted display or the like and can freely change the viewing point.
[0285] Furthermore, an emotion engine for estimating the user's emotion may be combined. That is, the specific processing unit 290 may estimate the user's emotion using the emotion recognition model 59 and perform specific processing using the user's emotion.
[0286] The embodiment for implementing the present invention is a system that three-dimensionally reproduces the actions of players in real time and further optimizes the video experience by recognizing the user's emotion. This system is composed of a photographing device, an artificial intelligence model, a holographic video generation device, an integration means, a display device, and an emotion engine.
[0287] First, the server controls the photographing device to two-dimensionally photograph the actions of players on the sports field. By capturing the players from different angles with multiple cameras and transmitting the data to the server in real time, detailed motion data is obtained. For example, in a soccer game, the server simultaneously captures the goal scene of a forward player from multiple viewpoints.
[0288] Next, the server uses an artificial intelligence model to convert this two-dimensional data into three dimensions. Using deep learning techniques, it analyzes the position of the athlete's skeleton and joints during movement and generates a three-dimensional model. This three-dimensional model forms the basis of the holographic video.
[0289] The server creates holographic videos of each player based on the generated three-dimensional models. This provides a three-dimensional representation of the players' movements. The server then integrates the holographic videos of multiple players into a single integrated video that recreates the entire field. This integration process allows for a faithful reproduction of the flow of the game and the coordinated movements between players.
[0290] The device projects this integrated holographic video onto a display device, providing the audience with a visually immersive three-dimensional image. Furthermore, the device is equipped with an emotion engine that recognizes the user's emotions. This emotion engine analyzes the user's facial expressions and physical reactions to determine the user's emotional state.
[0291] Once the user's emotional state is determined, the device dynamically adjusts the video content based on that information. For example, if the user is excited, the video experience is optimized by increasing the frequency of replays or zooming in on the plays of specific players. Users can enjoy a viewing experience customized to their emotional state, allowing them to experience the game more deeply and interactively. This system embodies a new way of watching sports.
[0292] The following describes the processing flow.
[0293] Step 1:
[0294] The server controls the camera equipment, capturing the players' movements in two dimensions from multiple cameras. The collected video data is transmitted to the server in real time and stored in a database.
[0295] Step 2:
[0296] The server preprocesses the received video data. Specifically, it performs noise removal and resolution unification to prepare the data in a state suitable for analysis by an artificial intelligence model.
[0297] Step 3:
[0298] The server uses an artificial intelligence model to generate a three-dimensional model from the processed two-dimensional data. It analyzes the movements of the players using deep learning technology to create a high-precision three-dimensional model.
[0299] Step 4:
[0300] The server generates a holographic video based on the generated three-dimensional model. Thereby, a stereoscopic motion video for each player is created, and a video that can be observed from multiple angles is obtained.
[0301] Step 5:
[0302] The server integrates the holographic videos of all the players to generate an integrated video that reproduces the entire field. It synchronizes the position information and movements of the players to reproduce the flow of the coordinated game.
[0303] Step 6:
[0304] The terminal projects the integrated holographic video onto a display device. The user can freely adjust the viewing point through the display device, enabling them to watch the game stereoscopically from various viewpoints.
[0305] Step 7:
[0306] The terminal uses an emotion engine to recognize the user's emotions. It analyzes the user's expressions, voice, heart rate, etc. through cameras and sensors to determine the user's emotional state.
[0307] Step 8:
[0308] The device adjusts the displayed content in real time based on recognized emotion data. For example, if the user is excited, it optimizes the video content by increasing replays or focusing on a specific player.
[0309] Step 9:
[0310] Users can enjoy an interactive viewing experience. The video content changes according to their emotions, creating a more personalized sports viewing experience.
[0311] (Example 2)
[0312] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0313] Conventional technologies have made it difficult to reproduce players' movements in real time and in three dimensions, and to dynamically optimize the video experience based on the user's emotions. In particular, there is a need for a system that can faithfully reproduce the movements of the entire match while adjusting the video according to the user's perspective and emotions. To solve this problem, technologies that capture players' movements from multiple angles and represent them in three dimensions, as well as video control technologies based on the user's emotions, are necessary.
[0314] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0315] In this invention, the server includes a device that provides a method for capturing a player's movements in two dimensions, a machine learning model for converting the acquired two-dimensional movement information into three dimensions, and means for generating a stereoscopic video using the three-dimensional information. This allows users to visualize the player's movements in three dimensions in real time. Furthermore, by using a device equipped with an emotion recognition function for analyzing the user's emotional state and dynamically adjusting the video, it becomes possible to provide a customized video experience for each individual user.
[0316] "Filming method" refers to the technique of capturing a player's movements precisely in two dimensions by filming from multiple viewpoints.
[0317] A "device" is a piece of equipment equipped with the hardware and software necessary to process the acquired two-dimensional video data.
[0318] A "machine learning model" is an artificial intelligence technology used to convert acquired two-dimensional motion information into three-dimensional data, and is an algorithm built on deep learning.
[0319] A "3D display video" is a video that uses three-dimensional information to create a three-dimensional representation of the players' movements.
[0320] "Integration means" refers to a technology used to recreate the movement of the entire stadium by combining 3D display videos of individual athletes.
[0321] A "display device" is a device that presents reproduced stereoscopic images to the user in three dimensions.
[0322] The "emotion recognition function" is an analysis system that analyzes the user's emotional state and adjusts the video content based on that information.
[0323] This invention is a system that reproduces athletes' movements in real time and in three dimensions, captures the user's emotions, and optimizes the video experience. This system consists of a complex including a shooting device, a machine learning model, a stereoscopic video generation means, an integration means, a display device, and an emotion recognition function.
[0324] The server first controls the camera equipment to capture images of the players' movements in two dimensions. This equipment includes multiple cameras, each capturing players on the sports field from a different viewpoint. For example, in a soccer match, capturing players' movements from multiple angles allows for the collection of more detailed motion data. This data is transmitted to the server in real time, preparing it for the next processing step.
[0325] Next, the server uses machine learning models to convert the acquired two-dimensional data into three-dimensional data. Typical software used includes deep learning frameworks such as TensorFlow and PyTorch, which accurately analyze the positions of the athlete's skeleton and joints and generate a three-dimensional model.
[0326] Using a 3D video generation system, the server creates 3D videos of each player based on the generated 3D models. This is done using a game engine such as Unity, enabling realistic movement simulations. The generated videos have the characteristic of being displayable flexibly, without being restricted to specific angles or viewing perspectives.
[0327] Following this, the server integrates the 3D video feeds of each player to recreate the overall movement of the stadium. By utilizing video processing technology, the coordination and positional relationships between players are accurately reproduced, and the movement of the entire match is provided as a single, unified video.
[0328] The device delivers this integrated video to the user through a display device. The display device enables three-dimensional visualization, providing a high level of visual immersion.
[0329] Furthermore, the device is equipped with emotion recognition capabilities to optimize the video experience by recognizing the user's emotional state. Using sensors and cameras, it analyzes the user's facial expressions and physical reactions in real time. Based on this analysis, it can dynamically adjust the video content. For example, if it determines that the user is excited, it can zoom in on a specific player's performance, providing an interactive video experience tailored to the user's state.
[0330] A concrete example of a prompt message might be, "Recreate the player's goal in detail in three dimensions, and replay it in slow motion to match the user's surprise." This system would allow users to experience sports viewing in a deeper, more interactive way.
[0331] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0332] Step 1:
[0333] The server controls the camera equipment to capture the players' movements in two dimensions. It uses video data from multiple cameras as input. Specifically, each camera captures the players from a different viewpoint and transmits the data to the server in real time. The data is processed by integrating the video from each viewpoint to capture the players' movements in detail. The output is the integrated two-dimensional video data.
[0334] Step 2:
[0335] The server generates a three-dimensional model using a machine learning model based on the acquired two-dimensional video data. The input is integrated two-dimensional video data. Specifically, it utilizes a deep learning framework to extract positional information of the athlete's skeleton and joints. As a data calculation, it reconstructs the three-dimensional structure from this information. The output is a model data that reproduces the athlete in three dimensions.
[0336] Step 3:
[0337] The server creates a 3D video based on the generated 3D model. The input is 3D model data. Specifically, it uses a game engine to add movement to the 3D model, resulting in a realistic video. Data processing involves adjusting smooth transitions between movements. The output is video data with a three-dimensional visual effect.
[0338] Step 4:
[0339] The server integrates 3D video data of multiple players to recreate their overall positional relationships. Individual 3D video data is used as input. The specific operation involves using video processing techniques to synchronize and align the movements of each player. Data processing combines different video clips to represent the movement of the entire stadium. The output is integrated video data of the entire field.
[0340] Step 5:
[0341] The terminal provides integrated video data to the user through a display device. Integrated holographic video data is used as input. Specifically, it provides the user with a visually interactive experience using a three-dimensional display device. The output is the three-dimensional image received by the user as a visual experience.
[0342] Step 6:
[0343] The device dynamically adjusts the video by analyzing the user's emotional state using emotion recognition technology. Inputs include user facial expressions and physical response data. Specifically, it collects data using cameras and sensors and determines the user's emotions using an emotion analysis algorithm. Data processing customizes the video content based on the emotional information. The output is a video experience optimized according to the user's state.
[0344] (Application Example 2)
[0345] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0346] In recent years, there has been a growing demand from sports spectators for an immersive, on-the-ground experience. However, conventional video technology has struggled to realistically reproduce athletes' movements in three dimensions and has been unable to customize the video experience to suit the individual emotions of each spectator. To address these challenges, there is a need for a system that can reproduce athletes' movements in real time and in three dimensions, while also dynamically adjusting the video experience based on the emotions of the spectators.
[0347] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0348] This invention includes a server that controls an imaging device for acquiring player movements in two dimensions, utilizes an intelligent model for converting the acquired two-dimensional movement information into three dimensions, generates an image using the three-dimensional information, an integration means for integrating the images of all players to reproduce the entire area, a display device for displaying the reproduced stereoscopic image, an emotion recognition engine for recognizing the user's emotional state, and an adjustment means for dynamically adjusting the image based on the user's emotional state. This enables a three-dimensional and immersive sports viewing experience, allowing spectators to enjoy a viewing experience optimized according to their individual emotions.
[0349] An "imaging device" is a device used to capture the movements of athletes in two dimensions.
[0350] An "intelligent model" is a machine learning-based model that converts acquired two-dimensional motion information into three-dimensional data.
[0351] "Image generation means" refers to technology and apparatus for generating images using three-dimensional information.
[0352] "Integration means" refers to methods and techniques for integrating video footage of all players to recreate the entire domain.
[0353] A "display device" is a device used to display reproduced 3D images.
[0354] An "emotion recognition engine" is software or hardware used to recognize the emotional state of a user.
[0355] "Adjustment means" refers to methods and techniques for dynamically adjusting the image based on the user's emotional state.
[0356] To implement this invention, multiple pieces of hardware and software must work in coordination with each other. The server first acquires the player's movements in two dimensions using an imaging device. This imaging device incorporates multiple cameras to capture the player from various angles. The acquired images are converted into three dimensions in real time using an intelligent model. This intelligent model is based on deep learning technology and analyzes the player's skeletal structure and movements in detail.
[0357] The three-dimensional model is converted into a holographic image by an image generation device. This image is then further integrated by an integration device to faithfully reproduce the entire field. The server transmits this integrated 3D image to display devices, providing spectators with an immersive viewing experience.
[0358] The device analyzes the audience's emotional state using an emotion recognition engine. The emotion recognition engine senses the audience's facial expressions and physical reactions to determine emotions such as excitement and joy. Based on this emotional information, the adjustment mechanism then customizes the video appropriately. For example, if the user's level of excitement is high, adjustments such as highlighting specific scenes may be made.
[0359] For example, in a soccer match, when a forward player makes an important play, the scene is zoomed in on, allowing spectators to experience the moment more vividly. By providing optimal visuals in response to the spectators' emotions in this way, a more personal and interactive viewing experience can be achieved.
[0360] Example prompt: "When you want to see an exciting moment in the game in more detail, recognize the emotion and zoom in on that scene."
[0361] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0362] Step 1:
[0363] The server controls the imaging device and uses multiple cameras to capture the players' movements in two dimensions. The input for this step is a video feed of the match, and the output is two-dimensional data capturing the players' movements. This data consists of frames taken from different angles by each camera.
[0364] Step 2:
[0365] The server inputs the acquired two-dimensional data into an intelligent model and converts it to three dimensions in real time. This conversion uses deep learning technology, and based on the two-dimensional input data, the skeleton and movements of each player are reconstructed in three-dimensional space. The output is a three-dimensional model of the player.
[0366] Step 3:
[0367] The server generates holographic images based on a three-dimensional model using image generation means. The input for this step is a three-dimensional model, and the output is a holographic image that realistically represents the player's movements. This image is optimized to provide a realistic visual experience.
[0368] Step 4:
[0369] The server integrates the generated holographic images using an integration mechanism to recreate the entire field of the match. The input is holographic images of each player, and the output is an integrated 3D image that faithfully reproduces the flow of the match and the coordination between players. This creates a realistic image that allows viewers to grasp the overall scene.
[0370] Step 5:
[0371] The terminal provides users with integrated stereoscopic images using a display device. The input is integrated holographic images, and the output is a visually immersive stereoscopic experience. The display device enables users to watch the match live.
[0372] Step 6:
[0373] An emotion recognition engine installed in the user's device operates, analyzing the user's facial expressions and movements to determine their emotional state. The input for this step is the user's real-time biometric data, and the output is the evaluation of their emotional state. Emotion recognition is performed using camera and sensor technology.
[0374] Step 7:
[0375] The device dynamically adjusts the video using adjustment mechanisms based on emotional state. The input is the emotional state obtained by the emotion recognition engine, and the output is video customized according to that emotion. An example of adjustment is zooming in on the main scene when the user is excited.
[0376] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0377] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0378] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0379] [Third Embodiment]
[0380] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0381] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0382] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0383] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0384] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0385] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0386] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0387] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0388] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0389] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0390] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0391] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0392] One embodiment of the present invention is a system for reproducing the movements of an athlete in three dimensions in real time. This system is comprised of a combination of a camera, an artificial intelligence model, a holographic video generation device, an integration means, and a display device.
[0393] The server first controls the camera equipment, capturing the players' movements on the sports field in two dimensions using multiple cameras. This allows for recording the players' movements from various angles, providing multifaceted data. For example, in a baseball game, the server captures the pitcher's throwing motion from multiple angles and synchronizes the footage from each camera.
[0394] Next, the server uses an artificial intelligence model to convert this two-dimensional data into three dimensions. The AI model utilizes deep learning techniques to analyze the position of the athlete's skeleton and joints during movement. This allows the server to create an accurate three-dimensional model and generate holographic images based on it.
[0395] The server constructs holographic videos of each player based on the generated three-dimensional models. These holographic videos are designed to visualize the individual movements of each player in a realistic way. Furthermore, an integration mechanism works to recreate the entire game situation on the field using these videos. The server uses this to create a video that recreates the entire field while maintaining the coordination of movements between players.
[0396] Ultimately, the terminal delivers integrated holographic video to the audience through the display device. This allows users to watch the game in 3D and freely change their viewpoint. For example, spectators can experience watching the game up close, even from the comfort of their own homes. This system will realize a new style of sports viewing.
[0397] The following describes the processing flow.
[0398] Step 1:
[0399] The server controls a multi-camera system to capture the players' movements in two dimensions. Each camera captures the players from a different angle and transmits the video data to the server in real time.
[0400] Step 2:
[0401] The server preprocesses the received two-dimensional video data, unifying the resolution and removing noise. This process makes the data suitable for three-dimensional conversion.
[0402] Step 3:
[0403] The server uses an artificial intelligence model to generate a three-dimensional model from pre-processed two-dimensional data. A deep learning algorithm analyzes the position of the athlete's skeleton and joints to construct an accurate three-dimensional model.
[0404] Step 4:
[0405] The server generates holographic videos based on three-dimensional models. This is intended to recreate the movements of each player in three dimensions, resulting in footage that can be observed from a 360-degree perspective.
[0406] Step 5:
[0407] The server integrates holographic videos of all players to create a unified video that recreates the entire match. It synchronizes the movements of the players and handles occlusion to accurately recreate the flow of the game.
[0408] Step 6:
[0409] The terminal projects integrated holographic video received from the server onto a display device. This device features an interface that allows the user to freely adjust their viewpoint, providing a visually three-dimensional video experience.
[0410] Step 7:
[0411] Users can watch the match footage displayed on the screen from various perspectives. By changing viewpoints and focusing on specific players, an immersive and interactive viewing experience is possible.
[0412] (Example 1)
[0413] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0414] In traditional sports viewing, the limited visual information resulted in a lack of realism, making it difficult for spectators to fully understand the overall situation on the field and the movements of individual players. This led to a decline in the quality of the viewing experience and a decrease in interest in sporting events.
[0415] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0416] In this invention, the server includes a recording device for recording the movements of players in two dimensions, a machine learning model for converting the recorded two-dimensional movement information into three dimensions, and means for generating a stereoscopic image using the three-dimensional information. This enables a new viewing experience for spectators that is visually rich and allows them to intuitively understand the movements of the players and the overall flow of the game.
[0417] A "recording device" is a device used to record an athlete's movements in two dimensions from multiple perspectives.
[0418] A "machine learning model" is a program that implements algorithms to analyze recorded two-dimensional motion information and convert it into three dimensions.
[0419] "3D images" are visually immersive and realistic images generated based on three-dimensional information.
[0420] "Three-dimensionalization" is the process of reconstructing three-dimensional shapes and movements based on data recorded in two dimensions.
[0421] A "presentation device" is a device used to visually present generated 3D images.
[0422] A "combination means" is a program or device that integrates 3D images of multiple players to recreate the overall situation.
[0423] An "operation interface" is an input method that allows the user to adjust the viewing angle and other settings via a display device.
[0424] This invention is a system that reproduces the movements of athletes in sports events in real time and in three dimensions. The central role of the system is played by a server. The server controls recording devices and records the movements of athletes in two dimensions from multiple viewpoints. The recording devices used are high-precision cameras, which enable the acquisition of multifaceted and highly accurate data.
[0425] Next, the server uses a machine learning model to analyze the recorded two-dimensional motion information and convert it into three dimensions. This process utilizes deep learning techniques to precisely model the player's skeleton and joint positions. This three-dimensional modeling makes it possible to realistically reproduce the player's movements.
[0426] Based on the generated three-dimensional data, the server creates a stereoscopic image. This stereoscopic image is visualized as a realistic and immersive visual experience using specialized holography generation technology. This provides viewers with a powerful and engaging sports viewing experience.
[0427] The integrated stereoscopic imagery is provided to the user via a display device. A high-definition holographic display is used as the display device. This allows the user to freely change their viewpoint, enabling them to observe the overall flow of the match and the individual plays of players in detail. For example, in a soccer match, the user can experience the game from the perspective of any player.
[0428] An example of a prompt message is, "Please provide instructions on how to recreate a baseball pitcher's throwing motion in three dimensions and display it as a holographic video." This invention can offer a new style of sports viewing, providing spectators with rich information and an entertaining experience.
[0429] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0430] Step 1:
[0431] The server controls the recording devices and records the players' movements in two dimensions. It receives real-time movements of the players on the sports field as input. The server synchronizes multiple high-precision cameras and captures the players' movements from each viewpoint. As output, it acquires multi-faceted two-dimensional video data, recording the players' positions and movements in detail.
[0432] Step 2:
[0433] The server inputs the acquired two-dimensional data into a machine learning model and converts it into three-dimensional data. The two-dimensional video data obtained in step 1 is used as input. The server utilizes deep learning technology to analyze the position of the athlete's skeleton and joints. As output, it generates an accurate three-dimensional model, allowing for a three-dimensional reproduction of the athlete's movements.
[0434] Step 3:
[0435] The server uses the generated 3D data to create a stereoscopic image. The 3D model generated in step 2 is used as input. The server employs holographic generation technology to visually represent the players' movements. The output is a powerful stereoscopic image, providing an immersive experience for the audience.
[0436] Step 4:
[0437] The server integrates multiple 3D images to recreate the entire game situation on the field. It uses 3D images of individual players as input. The server employs video editing technology to integrate the flow of the game into a single video while maintaining the coordination between players. The output provides a video that visually recreates the entire field.
[0438] Step 5:
[0439] The terminal provides the user with integrated 3D images through a display device. The integrated images from step 4 are used as input. The terminal utilizes a high-definition holographic display, allowing the user to freely change their viewing angle. As output, spectators can experience a free-viewpoint perspective, similar to watching the event live.
[0440] (Application Example 1)
[0441] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0442] In traditional sports viewing, spectators are limited to watching the game through two-dimensional footage from fixed camera angles, making it difficult to freely experience the sense of presence and the details of the players' movements. In particular, when spectators watch the game from a remote location such as their home, it is difficult to obtain a sense of presence that is close to the visual experience of being there in person.
[0443] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0444] In this invention, the server includes imaging means for acquiring the movements of players in two dimensions, an artificial intelligence system for converting the acquired two-dimensional movement information into three dimensions, and generation means for generating holographic audio and video information using the three-dimensional information. This makes it possible for spectators to experience a three-dimensional and immersive sports match from anywhere in the world, while freely changing their viewpoint.
[0445] "Imaging means" refers to a device for recording the movements of athletes in two dimensions from multiple viewpoints.
[0446] An "artificial intelligence system" is a technical means for analyzing two-dimensional motion information and converting it into three-dimensional data.
[0447] "Generation means" refers to an apparatus or method for creating holographic audio and video information using three-dimensional information.
[0448] "Integration means" refers to a technical method or machine for combining holographic video information of multiple athletes to recreate the entire stadium.
[0449] A "display means" is a device that visualizes the reproduced holographic image in three dimensions and provides it to the spectators.
[0450] To implement the present invention, the following configuration requirements are necessary. The server controls the imaging means and records the athlete's movements two-dimensionally from multiple directions. For example, multiple cameras are placed in a sports stadium to acquire the athlete's movements as video. This acquired two-dimensional video data is sent to an artificial intelligence system. This system uses deep learning technology to analyze the positions of the athlete's skeleton and joints in the video and converts it into three-dimensional data.
[0451] The generation means generates holographic audio and video information based on three-dimensional data, and the integration means combines holographic images of multiple athletes to recreate the entire stadium. A three-dimensional rendering engine could be used here. Finally, this recreated image is provided to spectators through the display means. Spectators can enjoy a more immersive sports viewing experience using devices such as head-mounted displays.
[0452] For example, in a soccer match, spectators can wear a head-mounted display at home and move freely around the field, experiencing the players' movements in real time. This entire process requires a low-latency network and high-speed data processing to achieve real-time performance.
[0453] An example of a prompt for a generative AI model is as follows: "Use a deep learning model to generate a three-dimensional holographic model from the following two-dimensional video data. The model should represent the movements of an athlete, and accurately analyze the skeletal structure and joint positions while maintaining a sense of realism."
[0454] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0455] Step 1:
[0456] The server controls the imaging equipment and acquires the players' movements as two-dimensional video. The input data consists of video from each camera, and the output is two-dimensional video data of the players' movements captured from multiple directions. The server receives this data in real time, ensuring that information from various angles is available.
[0457] Step 2:
[0458] The server supplies two-dimensional video data to the artificial intelligence system and performs the process of converting it into three-dimensional data. In this process, the input is two-dimensional video data from each camera, and the output is the three-dimensional position information and skeletal data of the players. Deep learning technology is used to analyze features within the images and generate three-dimensional data.
[0459] Step 3:
[0460] The generation method generates holographic audio and video information from three-dimensional data. The input here is the converted three-dimensional data, and the output is a three-dimensional holographic model. The server uses a rendering engine to convert the data into a visually understandable format.
[0461] Step 4:
[0462] The integration mechanism combines holographic images of multiple athletes to recreate the overall situation of the stadium. The input to this process is holographic models of individual athletes, and the output is a unified holographic image of the entire stadium. The server considers the coordination of movements between athletes to construct a natural-looking image.
[0463] Step 5:
[0464] The terminal uses a display device to provide spectators with three-dimensional holographic images. At this stage, the input is an integrated holographic image, and the output is a real-time image experienced by the spectator. Users receive the image through a head-mounted display or similar device and can freely change their viewpoint.
[0465] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0466] One embodiment of the present invention is a system that reproduces athlete movements in real time in three dimensions and further recognizes the user's emotions to optimize the video experience. This system consists of a shooting device, an artificial intelligence model, a holographic video generation device, an integration means, a display device, and an emotion engine.
[0467] First, the server controls the camera equipment to capture the movements of players on the sports field in two dimensions. By capturing players from different angles using multiple cameras and transmitting the data to the server in real time, detailed movement data is obtained. For example, in a soccer match, the server simultaneously captures a forward player's goal from multiple viewpoints.
[0468] Next, the server uses an artificial intelligence model to convert this two-dimensional data into three dimensions. Using deep learning techniques, it analyzes the position of the athlete's skeleton and joints during movement and generates a three-dimensional model. This three-dimensional model forms the basis of the holographic video.
[0469] The server creates holographic videos of each player based on the generated three-dimensional models. This provides a three-dimensional representation of the players' movements. The server then integrates the holographic videos of multiple players into a single integrated video that recreates the entire field. This integration process allows for a faithful reproduction of the flow of the game and the coordinated movements between players.
[0470] The device projects this integrated holographic video onto a display device, providing the audience with a visually immersive three-dimensional image. Furthermore, the device is equipped with an emotion engine that recognizes the user's emotions. This emotion engine analyzes the user's facial expressions and physical reactions to determine the user's emotional state.
[0471] Once the user's emotional state is determined, the device dynamically adjusts the video content based on that information. For example, if the user is excited, the video experience is optimized by increasing the frequency of replays or zooming in on the plays of specific players. Users can enjoy a viewing experience customized to their emotional state, allowing them to experience the game more deeply and interactively. This system embodies a new way of watching sports.
[0472] The following describes the processing flow.
[0473] Step 1:
[0474] The server controls the camera equipment, capturing the players' movements in two dimensions from multiple cameras. The collected video data is transmitted to the server in real time and stored in a database.
[0475] Step 2:
[0476] The server preprocesses the received video data. Specifically, it performs noise reduction and standardizes the resolution to prepare the data for analysis by artificial intelligence models.
[0477] Step 3:
[0478] The server uses an artificial intelligence model to generate a three-dimensional model from processed two-dimensional data. It analyzes the player's movements using deep learning techniques to create a highly accurate three-dimensional model.
[0479] Step 4:
[0480] The server generates holographic videos based on the generated three-dimensional models. This creates three-dimensional motion videos for each player, resulting in videos that can be observed from multiple angles.
[0481] Step 5:
[0482] The server integrates holographic videos of all players to generate a unified video that recreates the entire field. It synchronizes the players' position information and movements to recreate the coordinated flow of the game.
[0483] Step 6:
[0484] The terminal projects integrated holographic video onto a display device. Users can freely adjust their viewpoint through the display device, allowing them to watch the match in 3D from various perspectives.
[0485] Step 7:
[0486] The device uses an emotion engine to recognize the user's emotions. Through cameras and sensors, it analyzes the user's facial expressions, voice, heart rate, etc., to determine the user's emotional state.
[0487] Step 8:
[0488] The device adjusts the displayed content in real time based on recognized emotion data. For example, if the user is excited, it optimizes the video content by increasing replays or focusing on a specific player.
[0489] Step 9:
[0490] Users can enjoy an interactive viewing experience. The video content changes according to their emotions, creating a more personalized sports viewing experience.
[0491] (Example 2)
[0492] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0493] Conventional technologies have made it difficult to reproduce players' movements in real time and in three dimensions, and to dynamically optimize the video experience based on the user's emotions. In particular, there is a need for a system that can faithfully reproduce the movements of the entire match while adjusting the video according to the user's perspective and emotions. To solve this problem, technologies that capture players' movements from multiple angles and represent them in three dimensions, as well as video control technologies based on the user's emotions, are necessary.
[0494] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0495] In this invention, the server includes a device that provides a method for capturing a player's movements in two dimensions, a machine learning model for converting the acquired two-dimensional movement information into three dimensions, and means for generating a stereoscopic video using the three-dimensional information. This allows users to visualize the player's movements in three dimensions in real time. Furthermore, by using a device equipped with an emotion recognition function for analyzing the user's emotional state and dynamically adjusting the video, it becomes possible to provide a customized video experience for each individual user.
[0496] "Filming method" refers to the technique of capturing a player's movements precisely in two dimensions by filming from multiple viewpoints.
[0497] A "device" is a piece of equipment equipped with the hardware and software necessary to process the acquired two-dimensional video data.
[0498] A "machine learning model" is an artificial intelligence technology used to convert acquired two-dimensional motion information into three-dimensional data, and is an algorithm built on deep learning.
[0499] A "3D display video" is a video that uses three-dimensional information to create a three-dimensional representation of the players' movements.
[0500] "Integration means" refers to a technology used to recreate the movement of the entire stadium by combining 3D display videos of individual athletes.
[0501] A "display device" is a device that presents reproduced stereoscopic images to the user in three dimensions.
[0502] The "emotion recognition function" is an analysis system that analyzes the user's emotional state and adjusts the video content based on that information.
[0503] This invention is a system that reproduces athletes' movements in real time and in three dimensions, captures the user's emotions, and optimizes the video experience. This system consists of a complex including a shooting device, a machine learning model, a stereoscopic video generation means, an integration means, a display device, and an emotion recognition function.
[0504] The server first controls the camera equipment to capture images of the players' movements in two dimensions. This equipment includes multiple cameras, each capturing players on the sports field from a different viewpoint. For example, in a soccer match, capturing players' movements from multiple angles allows for the collection of more detailed motion data. This data is transmitted to the server in real time, preparing it for the next processing step.
[0505] Next, the server uses machine learning models to convert the acquired two-dimensional data into three-dimensional data. Typical software used includes deep learning frameworks such as TensorFlow and PyTorch, which accurately analyze the positions of the athlete's skeleton and joints and generate a three-dimensional model.
[0506] Using a 3D video generation system, the server creates 3D videos of each player based on the generated 3D models. This is done using a game engine such as Unity, enabling realistic movement simulations. The generated videos have the characteristic of being displayable flexibly, without being restricted to specific angles or viewing perspectives.
[0507] Following this, the server integrates the 3D video feeds of each player to recreate the overall movement of the stadium. By utilizing video processing technology, the coordination and positional relationships between players are accurately reproduced, and the movement of the entire match is provided as a single, unified video.
[0508] The device delivers this integrated video to the user through a display device. The display device enables three-dimensional visualization, providing a high level of visual immersion.
[0509] Furthermore, the device is equipped with emotion recognition capabilities to optimize the video experience by recognizing the user's emotional state. Using sensors and cameras, it analyzes the user's facial expressions and physical reactions in real time. Based on this analysis, it can dynamically adjust the video content. For example, if it determines that the user is excited, it can zoom in on a specific player's performance, providing an interactive video experience tailored to the user's state.
[0510] A concrete example of a prompt message might be, "Recreate the player's goal in detail in three dimensions, and replay it in slow motion to match the user's surprise." This system would allow users to experience sports viewing in a deeper, more interactive way.
[0511] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0512] Step 1:
[0513] The server controls the camera equipment to capture the players' movements in two dimensions. It uses video data from multiple cameras as input. Specifically, each camera captures the players from a different viewpoint and transmits the data to the server in real time. The data is processed by integrating the video from each viewpoint to capture the players' movements in detail. The output is the integrated two-dimensional video data.
[0514] Step 2:
[0515] The server generates a three-dimensional model using a machine learning model based on the acquired two-dimensional video data. The input is integrated two-dimensional video data. Specifically, it utilizes a deep learning framework to extract positional information of the athlete's skeleton and joints. As a data calculation, it reconstructs the three-dimensional structure from this information. The output is a model data that reproduces the athlete in three dimensions.
[0516] Step 3:
[0517] The server creates a 3D video based on the generated 3D model. The input is 3D model data. Specifically, it uses a game engine to add movement to the 3D model, resulting in a realistic video. Data processing involves adjusting smooth transitions between movements. The output is video data with a three-dimensional visual effect.
[0518] Step 4:
[0519] The server integrates 3D video data of multiple players to recreate their overall positional relationships. Individual 3D video data is used as input. The specific operation involves using video processing techniques to synchronize and align the movements of each player. Data processing combines different video clips to represent the movement of the entire stadium. The output is integrated video data of the entire field.
[0520] Step 5:
[0521] The terminal provides integrated video data to the user through a display device. Integrated holographic video data is used as input. Specifically, it provides the user with a visually interactive experience using a three-dimensional display device. The output is the three-dimensional image received by the user as a visual experience.
[0522] Step 6:
[0523] The device dynamically adjusts the video by analyzing the user's emotional state using emotion recognition technology. Inputs include user facial expressions and physical response data. Specifically, it collects data using cameras and sensors and determines the user's emotions using an emotion analysis algorithm. Data processing customizes the video content based on the emotional information. The output is a video experience optimized according to the user's state.
[0524] (Application Example 2)
[0525] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0526] In recent years, there has been a growing demand from sports spectators for an immersive, on-the-ground experience. However, conventional video technology has struggled to realistically reproduce athletes' movements in three dimensions and has been unable to customize the video experience to suit the individual emotions of each spectator. To address these challenges, there is a need for a system that can reproduce athletes' movements in real time and in three dimensions, while also dynamically adjusting the video experience based on the emotions of the spectators.
[0527] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0528] This invention includes a server that controls an imaging device for acquiring player movements in two dimensions, utilizes an intelligent model for converting the acquired two-dimensional movement information into three dimensions, generates an image using the three-dimensional information, an integration means for integrating the images of all players to reproduce the entire area, a display device for displaying the reproduced stereoscopic image, an emotion recognition engine for recognizing the user's emotional state, and an adjustment means for dynamically adjusting the image based on the user's emotional state. This enables a three-dimensional and immersive sports viewing experience, allowing spectators to enjoy a viewing experience optimized according to their individual emotions.
[0529] An "imaging device" is a device used to capture the movements of athletes in two dimensions.
[0530] An "intelligent model" is a machine learning-based model that converts acquired two-dimensional motion information into three-dimensional data.
[0531] "Image generation means" refers to technology and apparatus for generating images using three-dimensional information.
[0532] "Integration means" refers to methods and techniques for integrating video footage of all players to recreate the entire domain.
[0533] A "display device" is a device used to display reproduced 3D images.
[0534] An "emotion recognition engine" is software or hardware used to recognize the emotional state of a user.
[0535] "Adjustment means" refers to methods and techniques for dynamically adjusting the image based on the user's emotional state.
[0536] To implement this invention, multiple pieces of hardware and software must work in coordination with each other. The server first acquires the player's movements in two dimensions using an imaging device. This imaging device incorporates multiple cameras to capture the player from various angles. The acquired images are converted into three dimensions in real time using an intelligent model. This intelligent model is based on deep learning technology and analyzes the player's skeletal structure and movements in detail.
[0537] The three-dimensional model is converted into a holographic image by an image generation device. This image is then further integrated by an integration device to faithfully reproduce the entire field. The server transmits this integrated 3D image to display devices, providing spectators with an immersive viewing experience.
[0538] The device analyzes the audience's emotional state using an emotion recognition engine. The emotion recognition engine senses the audience's facial expressions and physical reactions to determine emotions such as excitement and joy. Based on this emotional information, the adjustment mechanism then customizes the video appropriately. For example, if the user's level of excitement is high, adjustments such as highlighting specific scenes may be made.
[0539] For example, in a soccer match, when a forward player makes an important play, the scene is zoomed in on, allowing spectators to experience the moment more vividly. By providing optimal visuals in response to the spectators' emotions in this way, a more personal and interactive viewing experience can be achieved.
[0540] Example prompt: "When you want to see an exciting moment in the game in more detail, recognize the emotion and zoom in on that scene."
[0541] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0542] Step 1:
[0543] The server controls the imaging device and uses multiple cameras to capture the players' movements in two dimensions. The input for this step is a video feed of the match, and the output is two-dimensional data capturing the players' movements. This data consists of frames taken from different angles by each camera.
[0544] Step 2:
[0545] The server inputs the acquired two-dimensional data into an intelligent model and converts it to three dimensions in real time. This conversion uses deep learning technology, and based on the two-dimensional input data, the skeleton and movements of each player are reconstructed in three-dimensional space. The output is a three-dimensional model of the player.
[0546] Step 3:
[0547] The server generates holographic images based on a three-dimensional model using image generation means. The input for this step is a three-dimensional model, and the output is a holographic image that realistically represents the player's movements. This image is optimized to provide a realistic visual experience.
[0548] Step 4:
[0549] The server integrates the generated holographic images using an integration mechanism to recreate the entire field of the match. The input is holographic images of each player, and the output is an integrated 3D image that faithfully reproduces the flow of the match and the coordination between players. This creates a realistic image that allows viewers to grasp the overall scene.
[0550] Step 5:
[0551] The terminal provides users with integrated stereoscopic images using a display device. The input is integrated holographic images, and the output is a visually immersive stereoscopic experience. The display device enables users to watch the match live.
[0552] Step 6:
[0553] An emotion recognition engine installed in the user's device operates, analyzing the user's facial expressions and movements to determine their emotional state. The input for this step is the user's real-time biometric data, and the output is the evaluation of their emotional state. Emotion recognition is performed using camera and sensor technology.
[0554] Step 7:
[0555] The device dynamically adjusts the video using adjustment mechanisms based on emotional state. The input is the emotional state obtained by the emotion recognition engine, and the output is video customized according to that emotion. An example of adjustment is zooming in on the main scene when the user is excited.
[0556] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0557] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0558] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[0559] [Fourth Embodiment]
[0560] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0561] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0562] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0563] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0564] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0565] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0566] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0567] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0568] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0569] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0570] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0571] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0572] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0573] One embodiment of the present invention is a system for reproducing the movements of an athlete in three dimensions in real time. This system is comprised of a combination of a camera, an artificial intelligence model, a holographic video generation device, an integration means, and a display device.
[0574] The server first controls the camera equipment, capturing the players' movements on the sports field in two dimensions using multiple cameras. This allows for recording the players' movements from various angles, providing multifaceted data. For example, in a baseball game, the server captures the pitcher's throwing motion from multiple angles and synchronizes the footage from each camera.
[0575] Next, the server uses an artificial intelligence model to convert this two-dimensional data into three dimensions. The AI model utilizes deep learning techniques to analyze the position of the athlete's skeleton and joints during movement. This allows the server to create an accurate three-dimensional model and generate holographic images based on it.
[0576] The server constructs holographic videos of each player based on the generated three-dimensional models. These holographic videos are designed to visualize the individual movements of each player in a realistic way. Furthermore, an integration mechanism works to recreate the entire game situation on the field using these videos. The server uses this to create a video that recreates the entire field while maintaining the coordination of movements between players.
[0577] Ultimately, the terminal delivers integrated holographic video to the audience through the display device. This allows users to watch the game in 3D and freely change their viewpoint. For example, spectators can experience watching the game up close, even from the comfort of their own homes. This system will realize a new style of sports viewing.
[0578] The following describes the processing flow.
[0579] Step 1:
[0580] The server controls a multi-camera system to capture the players' movements in two dimensions. Each camera captures the players from a different angle and transmits the video data to the server in real time.
[0581] Step 2:
[0582] The server preprocesses the received two-dimensional video data, unifying the resolution and removing noise. This process makes the data suitable for three-dimensional conversion.
[0583] Step 3:
[0584] The server uses an artificial intelligence model to generate a three-dimensional model from pre-processed two-dimensional data. A deep learning algorithm analyzes the position of the athlete's skeleton and joints to construct an accurate three-dimensional model.
[0585] Step 4:
[0586] The server generates holographic videos based on three-dimensional models. This is intended to recreate the movements of each player in three dimensions, resulting in footage that can be observed from a 360-degree perspective.
[0587] Step 5:
[0588] The server integrates holographic videos of all players to create a unified video that recreates the entire match. It synchronizes the movements of the players and handles occlusion to accurately recreate the flow of the game.
[0589] Step 6:
[0590] The terminal projects integrated holographic video received from the server onto a display device. This device features an interface that allows the user to freely adjust their viewpoint, providing a visually three-dimensional video experience.
[0591] Step 7:
[0592] Users can watch the match footage displayed on the screen from various perspectives. By changing viewpoints and focusing on specific players, an immersive and interactive viewing experience is possible.
[0593] (Example 1)
[0594] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0595] In traditional sports viewing, the limited visual information resulted in a lack of realism, making it difficult for spectators to fully understand the overall situation on the field and the movements of individual players. This led to a decline in the quality of the viewing experience and a decrease in interest in sporting events.
[0596] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0597] In this invention, the server includes a recording device for recording the movements of players in two dimensions, a machine learning model for converting the recorded two-dimensional movement information into three dimensions, and means for generating a stereoscopic image using the three-dimensional information. This enables a new viewing experience for spectators that is visually rich and allows them to intuitively understand the movements of the players and the overall flow of the game.
[0598] A "recording device" is a device used to record an athlete's movements in two dimensions from multiple perspectives.
[0599] A "machine learning model" is a program that implements algorithms to analyze recorded two-dimensional motion information and convert it into three dimensions.
[0600] "3D images" are visually immersive and realistic images generated based on three-dimensional information.
[0601] "Three-dimensionalization" is the process of reconstructing three-dimensional shapes and movements based on data recorded in two dimensions.
[0602] A "presentation device" is a device used to visually present generated 3D images.
[0603] A "combination means" is a program or device that integrates 3D images of multiple players to recreate the overall situation.
[0604] An "operation interface" is an input method that allows the user to adjust the viewing angle and other settings via a display device.
[0605] This invention is a system that reproduces the movements of athletes in sports events in real time and in three dimensions. The central role of the system is played by a server. The server controls recording devices and records the movements of athletes in two dimensions from multiple viewpoints. The recording devices used are high-precision cameras, which enable the acquisition of multifaceted and highly accurate data.
[0606] Next, the server uses a machine learning model to analyze the recorded two-dimensional motion information and convert it into three dimensions. This process utilizes deep learning techniques to precisely model the player's skeleton and joint positions. This three-dimensional modeling makes it possible to realistically reproduce the player's movements.
[0607] Based on the generated three-dimensional data, the server creates a stereoscopic image. This stereoscopic image is visualized as a realistic and immersive visual experience using specialized holography generation technology. This provides viewers with a powerful and engaging sports viewing experience.
[0608] The integrated stereoscopic imagery is provided to the user via a display device. A high-definition holographic display is used as the display device. This allows the user to freely change their viewpoint, enabling them to observe the overall flow of the match and the individual plays of players in detail. For example, in a soccer match, the user can experience the game from the perspective of any player.
[0609] An example of a prompt message is, "Please provide instructions on how to recreate a baseball pitcher's throwing motion in three dimensions and display it as a holographic video." This invention can offer a new style of sports viewing, providing spectators with rich information and an entertaining experience.
[0610] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0611] Step 1:
[0612] The server controls the recording devices and records the players' movements in two dimensions. It receives real-time movements of the players on the sports field as input. The server synchronizes multiple high-precision cameras and captures the players' movements from each viewpoint. As output, it acquires multi-faceted two-dimensional video data, recording the players' positions and movements in detail.
[0613] Step 2:
[0614] The server inputs the acquired two-dimensional data into a machine learning model and converts it into three-dimensional data. The two-dimensional video data obtained in step 1 is used as input. The server utilizes deep learning technology to analyze the position of the athlete's skeleton and joints. As output, it generates an accurate three-dimensional model, allowing for a three-dimensional reproduction of the athlete's movements.
[0615] Step 3:
[0616] The server uses the generated 3D data to create a stereoscopic image. The 3D model generated in step 2 is used as input. The server employs holographic generation technology to visually represent the players' movements. The output is a powerful stereoscopic image, providing an immersive experience for the audience.
[0617] Step 4:
[0618] The server integrates multiple 3D images to recreate the entire game situation on the field. It uses 3D images of individual players as input. The server employs video editing technology to integrate the flow of the game into a single video while maintaining the coordination between players. The output provides a video that visually recreates the entire field.
[0619] Step 5:
[0620] The terminal provides the user with integrated 3D images through a display device. The integrated images from step 4 are used as input. The terminal utilizes a high-definition holographic display, allowing the user to freely change their viewing angle. As output, spectators can experience a free-viewpoint perspective, similar to watching the event live.
[0621] (Application Example 1)
[0622] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0623] In traditional sports viewing, spectators are limited to watching the game through two-dimensional footage from fixed camera angles, making it difficult to freely experience the sense of presence and the details of the players' movements. In particular, when spectators watch the game from a remote location such as their home, it is difficult to obtain a sense of presence that is close to the visual experience of being there in person.
[0624] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0625] In this invention, the server includes imaging means for acquiring the movements of players in two dimensions, an artificial intelligence system for converting the acquired two-dimensional movement information into three dimensions, and generation means for generating holographic audio and video information using the three-dimensional information. This makes it possible for spectators to experience a three-dimensional and immersive sports match from anywhere in the world, while freely changing their viewpoint.
[0626] "Imaging means" refers to a device for recording the movements of athletes in two dimensions from multiple viewpoints.
[0627] An "artificial intelligence system" is a technical means for analyzing two-dimensional motion information and converting it into three-dimensional data.
[0628] "Generation means" refers to an apparatus or method for creating holographic audio and video information using three-dimensional information.
[0629] "Integration means" refers to a technical method or machine for combining holographic video information of multiple athletes to recreate the entire stadium.
[0630] A "display means" is a device that visualizes the reproduced holographic image in three dimensions and provides it to the spectators.
[0631] To implement the present invention, the following configuration requirements are necessary. The server controls the imaging means and records the athlete's movements two-dimensionally from multiple directions. For example, multiple cameras are placed in a sports stadium to acquire the athlete's movements as video. This acquired two-dimensional video data is sent to an artificial intelligence system. This system uses deep learning technology to analyze the positions of the athlete's skeleton and joints in the video and converts it into three-dimensional data.
[0632] The generation means generates holographic audio and video information based on three-dimensional data, and the integration means combines holographic images of multiple athletes to recreate the entire stadium. A three-dimensional rendering engine could be used here. Finally, this recreated image is provided to spectators through the display means. Spectators can enjoy a more immersive sports viewing experience using devices such as head-mounted displays.
[0633] For example, in a soccer match, spectators can wear a head-mounted display at home and move freely around the field, experiencing the players' movements in real time. This entire process requires a low-latency network and high-speed data processing to achieve real-time performance.
[0634] An example of a prompt for a generative AI model is as follows: "Use a deep learning model to generate a three-dimensional holographic model from the following two-dimensional video data. The model should represent the movements of an athlete, and accurately analyze the skeletal structure and joint positions while maintaining a sense of realism."
[0635] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0636] Step 1:
[0637] The server controls the imaging equipment and acquires the players' movements as two-dimensional video. The input data consists of video from each camera, and the output is two-dimensional video data of the players' movements captured from multiple directions. The server receives this data in real time, ensuring that information from various angles is available.
[0638] Step 2:
[0639] The server supplies two-dimensional video data to the artificial intelligence system and performs the process of converting it into three-dimensional data. In this process, the input is two-dimensional video data from each camera, and the output is the three-dimensional position information and skeletal data of the players. Deep learning technology is used to analyze features within the images and generate three-dimensional data.
[0640] Step 3:
[0641] The generation method generates holographic audio and video information from three-dimensional data. The input here is the converted three-dimensional data, and the output is a three-dimensional holographic model. The server uses a rendering engine to convert the data into a visually understandable format.
[0642] Step 4:
[0643] The integration mechanism combines holographic images of multiple athletes to recreate the overall situation of the stadium. The input to this process is holographic models of individual athletes, and the output is a unified holographic image of the entire stadium. The server considers the coordination of movements between athletes to construct a natural-looking image.
[0644] Step 5:
[0645] The terminal uses a display device to provide spectators with three-dimensional holographic images. At this stage, the input is an integrated holographic image, and the output is a real-time image experienced by the spectator. Users receive the image through a head-mounted display or similar device and can freely change their viewpoint.
[0646] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0647] One embodiment of the present invention is a system that reproduces athlete movements in real time in three dimensions and further recognizes the user's emotions to optimize the video experience. This system consists of a shooting device, an artificial intelligence model, a holographic video generation device, an integration means, a display device, and an emotion engine.
[0648] First, the server controls the camera equipment to capture the movements of players on the sports field in two dimensions. By capturing players from different angles using multiple cameras and transmitting the data to the server in real time, detailed movement data is obtained. For example, in a soccer match, the server simultaneously captures a forward player's goal from multiple viewpoints.
[0649] Next, the server uses an artificial intelligence model to convert this two-dimensional data into three dimensions. Using deep learning techniques, it analyzes the position of the athlete's skeleton and joints during movement and generates a three-dimensional model. This three-dimensional model forms the basis of the holographic video.
[0650] The server creates holographic videos of each player based on the generated three-dimensional models. This provides a three-dimensional representation of the players' movements. The server then integrates the holographic videos of multiple players into a single integrated video that recreates the entire field. This integration process allows for a faithful reproduction of the flow of the game and the coordinated movements between players.
[0651] The device projects this integrated holographic video onto a display device, providing the audience with a visually immersive three-dimensional image. Furthermore, the device is equipped with an emotion engine that recognizes the user's emotions. This emotion engine analyzes the user's facial expressions and physical reactions to determine the user's emotional state.
[0652] Once the user's emotional state is determined, the device dynamically adjusts the video content based on that information. For example, if the user is excited, the video experience is optimized by increasing the frequency of replays or zooming in on the plays of specific players. Users can enjoy a viewing experience customized to their emotional state, allowing them to experience the game more deeply and interactively. This system embodies a new way of watching sports.
[0653] The following describes the processing flow.
[0654] Step 1:
[0655] The server controls the camera equipment, capturing the players' movements in two dimensions from multiple cameras. The collected video data is transmitted to the server in real time and stored in a database.
[0656] Step 2:
[0657] The server preprocesses the received video data. Specifically, it performs noise reduction and standardizes the resolution to prepare the data for analysis by artificial intelligence models.
[0658] Step 3:
[0659] The server uses an artificial intelligence model to generate a three-dimensional model from processed two-dimensional data. It analyzes the player's movements using deep learning techniques to create a highly accurate three-dimensional model.
[0660] Step 4:
[0661] The server generates holographic videos based on the generated three-dimensional models. This creates three-dimensional motion videos for each player, resulting in videos that can be observed from multiple angles.
[0662] Step 5:
[0663] The server integrates holographic videos of all players to generate a unified video that recreates the entire field. It synchronizes the players' position information and movements to recreate the coordinated flow of the game.
[0664] Step 6:
[0665] The terminal projects integrated holographic video onto a display device. Users can freely adjust their viewpoint through the display device, allowing them to watch the match in 3D from various perspectives.
[0666] Step 7:
[0667] The device uses an emotion engine to recognize the user's emotions. Through cameras and sensors, it analyzes the user's facial expressions, voice, heart rate, etc., to determine the user's emotional state.
[0668] Step 8:
[0669] The device adjusts the displayed content in real time based on recognized emotion data. For example, if the user is excited, it optimizes the video content by increasing replays or focusing on a specific player.
[0670] Step 9:
[0671] Users can enjoy an interactive viewing experience. The video content changes according to their emotions, creating a more personalized sports viewing experience.
[0672] (Example 2)
[0673] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0674] Conventional technologies have made it difficult to reproduce players' movements in real time and in three dimensions, and to dynamically optimize the video experience based on the user's emotions. In particular, there is a need for a system that can faithfully reproduce the movements of the entire match while adjusting the video according to the user's perspective and emotions. To solve this problem, technologies that capture players' movements from multiple angles and represent them in three dimensions, as well as video control technologies based on the user's emotions, are necessary.
[0675] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0676] In this invention, the server includes a device that provides a method for capturing a player's movements in two dimensions, a machine learning model for converting the acquired two-dimensional movement information into three dimensions, and means for generating a stereoscopic video using the three-dimensional information. This allows users to visualize the player's movements in three dimensions in real time. Furthermore, by using a device equipped with an emotion recognition function for analyzing the user's emotional state and dynamically adjusting the video, it becomes possible to provide a customized video experience for each individual user.
[0677] "Filming method" refers to the technique of capturing a player's movements precisely in two dimensions by filming from multiple viewpoints.
[0678] A "device" is a piece of equipment equipped with the hardware and software necessary to process the acquired two-dimensional video data.
[0679] A "machine learning model" is an artificial intelligence technology used to convert acquired two-dimensional motion information into three-dimensional data, and is an algorithm built on deep learning.
[0680] A "3D display video" is a video that uses three-dimensional information to create a three-dimensional representation of the players' movements.
[0681] "Integration means" refers to a technology used to recreate the movement of the entire stadium by combining 3D display videos of individual athletes.
[0682] A "display device" is a device that presents reproduced stereoscopic images to the user in three dimensions.
[0683] The "emotion recognition function" is an analysis system that analyzes the user's emotional state and adjusts the video content based on that information.
[0684] This invention is a system that reproduces athletes' movements in real time and in three dimensions, captures the user's emotions, and optimizes the video experience. This system consists of a complex including a shooting device, a machine learning model, a stereoscopic video generation means, an integration means, a display device, and an emotion recognition function.
[0685] The server first controls the camera equipment to capture images of the players' movements in two dimensions. This equipment includes multiple cameras, each capturing players on the sports field from a different viewpoint. For example, in a soccer match, capturing players' movements from multiple angles allows for the collection of more detailed motion data. This data is transmitted to the server in real time, preparing it for the next processing step.
[0686] Next, the server uses machine learning models to convert the acquired two-dimensional data into three-dimensional data. Typical software used includes deep learning frameworks such as TensorFlow and PyTorch, which accurately analyze the positions of the athlete's skeleton and joints and generate a three-dimensional model.
[0687] Using a 3D video generation system, the server creates 3D videos of each player based on the generated 3D models. This is done using a game engine such as Unity, enabling realistic movement simulations. The generated videos have the characteristic of being displayable flexibly, without being restricted to specific angles or viewing perspectives.
[0688] Following this, the server integrates the 3D video feeds of each player to recreate the overall movement of the stadium. By utilizing video processing technology, the coordination and positional relationships between players are accurately reproduced, and the movement of the entire match is provided as a single, unified video.
[0689] The device delivers this integrated video to the user through a display device. The display device enables three-dimensional visualization, providing a high level of visual immersion.
[0690] Furthermore, the device is equipped with emotion recognition capabilities to optimize the video experience by recognizing the user's emotional state. Using sensors and cameras, it analyzes the user's facial expressions and physical reactions in real time. Based on this analysis, it can dynamically adjust the video content. For example, if it determines that the user is excited, it can zoom in on a specific player's performance, providing an interactive video experience tailored to the user's state.
[0691] A concrete example of a prompt message might be, "Recreate the player's goal in detail in three dimensions, and replay it in slow motion to match the user's surprise." This system would allow users to experience sports viewing in a deeper, more interactive way.
[0692] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0693] Step 1:
[0694] The server controls the camera equipment to capture the players' movements in two dimensions. It uses video data from multiple cameras as input. Specifically, each camera captures the players from a different viewpoint and transmits the data to the server in real time. The data is processed by integrating the video from each viewpoint to capture the players' movements in detail. The output is the integrated two-dimensional video data.
[0695] Step 2:
[0696] The server generates a three-dimensional model using a machine learning model based on the acquired two-dimensional video data. The input is integrated two-dimensional video data. Specifically, it utilizes a deep learning framework to extract positional information of the athlete's skeleton and joints. As a data calculation, it reconstructs the three-dimensional structure from this information. The output is a model data that reproduces the athlete in three dimensions.
[0697] Step 3:
[0698] The server creates a 3D video based on the generated 3D model. The input is 3D model data. Specifically, it uses a game engine to add movement to the 3D model, resulting in a realistic video. Data processing involves adjusting smooth transitions between movements. The output is video data with a three-dimensional visual effect.
[0699] Step 4:
[0700] The server integrates 3D video data of multiple players to recreate their overall positional relationships. Individual 3D video data is used as input. The specific operation involves using video processing techniques to synchronize and align the movements of each player. Data processing combines different video clips to represent the movement of the entire stadium. The output is integrated video data of the entire field.
[0701] Step 5:
[0702] The terminal provides integrated video data to the user through a display device. Integrated holographic video data is used as input. Specifically, it provides the user with a visually interactive experience using a three-dimensional display device. The output is the three-dimensional image received by the user as a visual experience.
[0703] Step 6:
[0704] The device dynamically adjusts the video by analyzing the user's emotional state using emotion recognition technology. Inputs include user facial expressions and physical response data. Specifically, it collects data using cameras and sensors and determines the user's emotions using an emotion analysis algorithm. Data processing customizes the video content based on the emotional information. The output is a video experience optimized according to the user's state.
[0705] (Application Example 2)
[0706] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0707] In recent years, there has been a growing demand from sports spectators for an immersive, on-the-ground experience. However, conventional video technology has struggled to realistically reproduce athletes' movements in three dimensions and has been unable to customize the video experience to suit the individual emotions of each spectator. To address these challenges, there is a need for a system that can reproduce athletes' movements in real time and in three dimensions, while also dynamically adjusting the video experience based on the emotions of the spectators.
[0708] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0709] This invention includes a server that controls an imaging device for acquiring player movements in two dimensions, utilizes an intelligent model for converting the acquired two-dimensional movement information into three dimensions, generates an image using the three-dimensional information, an integration means for integrating the images of all players to reproduce the entire area, a display device for displaying the reproduced stereoscopic image, an emotion recognition engine for recognizing the user's emotional state, and an adjustment means for dynamically adjusting the image based on the user's emotional state. This enables a three-dimensional and immersive sports viewing experience, allowing spectators to enjoy a viewing experience optimized according to their individual emotions.
[0710] An "imaging device" is a device used to capture the movements of athletes in two dimensions.
[0711] An "intelligent model" is a machine learning-based model that converts acquired two-dimensional motion information into three-dimensional data.
[0712] "Image generation means" refers to technology and apparatus for generating images using three-dimensional information.
[0713] "Integration means" refers to methods and techniques for integrating video footage of all players to recreate the entire domain.
[0714] A "display device" is a device used to display reproduced 3D images.
[0715] An "emotion recognition engine" is software or hardware used to recognize the emotional state of a user.
[0716] "Adjustment means" refers to methods and techniques for dynamically adjusting the image based on the user's emotional state.
[0717] To implement this invention, multiple pieces of hardware and software must work in coordination with each other. The server first acquires the player's movements in two dimensions using an imaging device. This imaging device incorporates multiple cameras to capture the player from various angles. The acquired images are converted into three dimensions in real time using an intelligent model. This intelligent model is based on deep learning technology and analyzes the player's skeletal structure and movements in detail.
[0718] The three-dimensional model is converted into a holographic image by an image generation device. This image is then further integrated by an integration device to faithfully reproduce the entire field. The server transmits this integrated 3D image to display devices, providing spectators with an immersive viewing experience.
[0719] The device analyzes the audience's emotional state using an emotion recognition engine. The emotion recognition engine senses the audience's facial expressions and physical reactions to determine emotions such as excitement and joy. Based on this emotional information, the adjustment mechanism then customizes the video appropriately. For example, if the user's level of excitement is high, adjustments such as highlighting specific scenes may be made.
[0720] For example, in a soccer match, when a forward player makes an important play, the scene is zoomed in on, allowing spectators to experience the moment more vividly. By providing optimal visuals in response to the spectators' emotions in this way, a more personal and interactive viewing experience can be achieved.
[0721] Example prompt: "When you want to see an exciting moment in the game in more detail, recognize the emotion and zoom in on that scene."
[0722] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0723] Step 1:
[0724] The server controls the imaging device and uses multiple cameras to capture the players' movements in two dimensions. The input for this step is a video feed of the match, and the output is two-dimensional data capturing the players' movements. This data consists of frames taken from different angles by each camera.
[0725] Step 2:
[0726] The server inputs the acquired two-dimensional data into an intelligent model and converts it to three dimensions in real time. This conversion uses deep learning technology, and based on the two-dimensional input data, the skeleton and movements of each player are reconstructed in three-dimensional space. The output is a three-dimensional model of the player.
[0727] Step 3:
[0728] The server generates holographic images based on a three-dimensional model using image generation means. The input for this step is a three-dimensional model, and the output is a holographic image that realistically represents the player's movements. This image is optimized to provide a realistic visual experience.
[0729] Step 4:
[0730] The server integrates the generated holographic images using an integration mechanism to recreate the entire field of the match. The input is holographic images of each player, and the output is an integrated 3D image that faithfully reproduces the flow of the match and the coordination between players. This creates a realistic image that allows viewers to grasp the overall scene.
[0731] Step 5:
[0732] The terminal provides users with integrated stereoscopic images using a display device. The input is integrated holographic images, and the output is a visually immersive stereoscopic experience. The display device enables users to watch the match live.
[0733] Step 6:
[0734] An emotion recognition engine installed in the user's device operates, analyzing the user's facial expressions and movements to determine their emotional state. The input for this step is the user's real-time biometric data, and the output is the evaluation of their emotional state. Emotion recognition is performed using camera and sensor technology.
[0735] Step 7:
[0736] The device dynamically adjusts the video using adjustment mechanisms based on emotional state. The input is the emotional state obtained by the emotion recognition engine, and the output is video customized according to that emotion. An example of adjustment is zooming in on the main scene when the user is excited.
[0737] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0738] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0739] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[0740] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0741] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0742] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0743] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0744] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0745] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0746] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0747] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[0748] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[0749] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0750] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0751] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0752] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0753] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0754] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0755] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0756] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0757] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[0758] The following is further disclosed regarding the embodiments described above.
[0759] (Claim 1)
[0760] A camera for capturing the movements of athletes in two dimensions,
[0761] An artificial intelligence model for converting acquired two-dimensional motion data into three dimensions,
[0762] A means for generating holographic videos using three-dimensional data,
[0763] An integration method for combining holographic videos of all players to recreate the entire field,
[0764] A system including a display device for displaying reproduced holographic images in three dimensions.
[0765] (Claim 2)
[0766] The system according to claim 1, wherein the camera is configured to acquire the player's movements from multiple angles.
[0767] (Claim 3)
[0768] The system according to claim 1, wherein the display device is configured to provide an interface that allows the viewpoint to be freely changed.
[0769] "Example 1"
[0770] (Claim 1)
[0771] A recording device for recording the movements of athletes in two dimensions,
[0772] A machine learning model for converting recorded two-dimensional motion information into three dimensions,
[0773] A means for generating a stereoscopic image using three-dimensional information,
[0774] A means of integrating 3D images of all players to recreate the overall situation,
[0775] A system including a display device for presenting visualized stereoscopic images in three dimensions.
[0776] (Claim 2)
[0777] The system according to claim 1, wherein the recording device is configured to record the player's movements in a multifaceted manner.
[0778] (Claim 3)
[0779] The system according to claim 1, wherein the display device is configured to provide an operating interface that allows the viewing angle to be arbitrarily adjusted.
[0780] "Application Example 1"
[0781] (Claim 1)
[0782] An imaging means for acquiring the movements of athletes in two dimensions,
[0783] An artificial intelligence system for converting acquired two-dimensional motion information into three dimensions,
[0784] A generation means for generating holographic audio and video information using three-dimensional information,
[0785] An integration method for combining holographic video information from multiple players to recreate the entire stadium,
[0786] A system including a display means for displaying reproduced holographic images in three dimensions.
[0787] (Claim 2)
[0788] The system according to claim 1, wherein the imaging means is configured to acquire the movements of the player from multiple angles.
[0789] (Claim 3)
[0790] The system according to claim 1, wherein the display means is configured to provide a user interface that allows the viewer to freely change their viewpoint.
[0791] "Example 2 of combining an emotion engine"
[0792] (Claim 1)
[0793] A device that provides a method for capturing the movements of athletes in two dimensions,
[0794] A machine learning model for converting acquired two-dimensional motion information into three dimensions,
[0795] A means for generating a stereoscopic video using three-dimensional information,
[0796] An integration method to recreate the entire stadium by combining 3D display videos of all players,
[0797] A display device for presenting reproduced stereoscopic images in three dimensions,
[0798] A device equipped with an emotion recognition function to analyze the user's emotional state and dynamically adjust the video,
[0799] ...
[0800] A system that includes this.
[0801] (Claim 2)
[0802] The system according to claim 1, wherein the shooting method is configured to acquire the player's movements from multiple viewpoints.
[0803] (Claim 3)
[0804] The system according to claim 1, configured to provide a display method in which the display device can be selectively changed to change the viewpoint.
[0805] "Application example 2 when combining with an emotional engine"
[0806] (Claim 1)
[0807] An imaging device for acquiring the movements of athletes in two dimensions,
[0808] An intelligent model for converting acquired two-dimensional motion information into three dimensions,
[0809] A means of generating images using three-dimensional information,
[0810] An integration method to recreate the entire area by combining video footage of all players,
[0811] Display equipment for displaying reproduced 3D images,
[0812] An emotion recognition engine for recognizing the user's emotional state,
[0813] An adjustment mechanism for dynamically adjusting the video based on the user's emotional state,
[0814] A system that includes this.
[0815] (Claim 2)
[0816] The system according to claim 1, wherein the imaging device is configured to acquire the player's movements from multiple angles.
[0817] (Claim 3)
[0818] The system according to claim 1, wherein the display device is configured to provide an interface that allows the viewpoint to be freely changed. [Explanation of symbols]
[0819] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. A camera for capturing the movements of athletes in two dimensions, An artificial intelligence model for converting acquired two-dimensional motion data into three dimensions, A means for generating holographic videos using three-dimensional data, An integration method for combining holographic videos of all players to recreate the entire field, A system including a display device for displaying reproduced holographic images in three dimensions.
2. The system according to claim 1, wherein the camera is configured to acquire the player's movements from multiple angles.
3. The system according to claim 1, wherein the display device is configured to provide an interface that allows the viewpoint to be freely changed.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A