System
The system addresses the challenge of uniform commentary by analyzing game footage in real-time, using generative AI to generate customized commentary based on user knowledge and team support, improving the sports viewing experience.
Patent Information
- Application Number
- JP2024123969
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-30
- Publication Date
- 2026-02-12
AI Technical Summary
Conventional sports commentary systems fail to provide customized commentary tailored to individual viewers' knowledge levels and their supported teams, lacking real-time game progress recognition and commentary generation capabilities, leading to a uniform viewing experience that may not satisfy diverse audience needs.
A system that inputs users' sports knowledge level and supported team, analyzes game footage in real-time to recognize player positions and actions, generates commentary using generative AI, converts it to audio with a text-to-speech engine, and delivers it to users via terminals.
Enables real-time, customized sports commentary that matches viewers' knowledge levels and team support, enhancing the sports viewing experience by providing tailored commentary.
Smart Images

Figure 2026022452000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] Because live sports commentary depends on the announcer's knowledge and expressiveness, insufficient knowledge or biased support attitudes can affect the commentary and ruin the viewer's experience. Furthermore, because viewers' levels of knowledge about sports and the teams they support vary, it is difficult to satisfy all viewers with the current standard commentary. Therefore, it is necessary to provide customized commentary tailored to each viewer's individual needs in order to enhance the sports viewing experience. [Means for solving the problem]
[0005] We propose a system that includes a means for inputting a user's sports knowledge level and the person they are rooting for, a means for analyzing game footage in real time to recognize the positions, movements, and score of players, a means for generating commentary that matches the user's profile based on the analysis results, a means for converting the commentary into audio data using a text-to-speech engine, and a means for providing the audio data to the user. This system makes it possible to provide appropriate commentary in real time that matches the user's sports knowledge level and the person they are rooting for, thereby improving viewer satisfaction.
[0006] A "user" is a viewer who watches sports and uses the system to set his or her own level of sports knowledge and the sport he or she supports.
[0007] The "sports knowledge level" is an index that indicates the user's level of knowledge and understanding of sports, and has stages such as beginner, intermediate, and advanced.
[0008] The term "target of support" refers to the team or player that the user supports in a particular game.
[0009] "Game Footage" means video footage of a sports game captured on camera and distributed via live streaming.
[0010] "Analyzing in real time" means receiving game footage and processing it to instantly recognize the players' positions, movements, scoring status, etc.
[0011] "Player position" refers to the current location of each player within the competition area.
[0012] "Player actions" refers to a series of actions or movements performed by a player, including specific actions such as dribbling and shooting.
[0013] "Score status" refers to information showing the current score or points in a match.
[0014] "Analysis results" refers to the specific information and data obtained after analyzing game footage in real time.
[0015] "Live commentary" refers to commentary in text format that is generated as the game progresses, and is the source of the commentary audio provided to the user.
[0016] A "text-to-speech engine" refers to software or hardware for analyzing text data and converting it into speech data.
[0017] "Audio Data" refers to the digital audio files generated as a result of voice synthesis of commentary.
[0018] "Terminal" refers to an electronic device used by a user, such as a smartphone, PC, or tablet, and refers to equipment that receives game footage and plays audio data.
[0019] "Server" refers to the central computer system that handles tasks such as analyzing game footage, managing user profiles, and generating commentary. [Brief explanation of the drawings]
[0020] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0021] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0022] First, the terms used in the following description will be explained.
[0023] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0024] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0025] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0026] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0027] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0028] [First embodiment]
[0029] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0030] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0031] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0032] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0033] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0034] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0035] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0036] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0037] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0038] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0039] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0040] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0041] This invention relates to a system that uses generative AI to generate live sports commentary audio in real time. This allows for customizable live commentary based on the user's level of sports knowledge and the sport they are rooting for. The detailed program processing of this system is explained in natural language.
[0042] User Profile Settings
[0043] User: Enter information
[0044] Users use their devices to set their sports knowledge level (e.g., beginner, intermediate, advanced) and the team or player they support. This information forms the basis for customization settings used throughout the system.
[0045] Terminal: Data transmission
[0046] The device generates and sends an API request to the server to transmit the information entered by the user, including the sports knowledge level and support target selected by the user.
[0047] Server: Data storage
[0048] The server stores the received user information in a database and creates a profile for each user, which specifies the user's level of sports knowledge and the sport they support.
[0049] Real-time analysis of game footage
[0050] Terminal: Video reception
[0051] The device receives game footage from a live streaming service, which is then streamed to a server.
[0052] Terminal:Video Stream
[0053] The device then streams the received game footage to a server, where the streaming video becomes the input data for real-time analysis.
[0054] Server: Video analysis
[0055] The server analyzes the received video using a video analysis algorithm, specifically an object recognition algorithm that tracks the player's position in real time, recognizes each player's actions (e.g., dribbling, passing, shooting), and determines the goal situation and ball position.
[0056] Generating live commentary audio
[0057] Server: Customization information integration
[0058] The server combines user profile data with video analysis results, and this information is used to generate commentary.
[0059] Server: Explanation generation
[0060] The server then inputs the integrated data into a generative AI model to generate commentary tailored to the user. This commentary varies in detail depending on the user's level of sports knowledge. For example, commentary for beginners starts with basic explanations, while commentary for advanced players focuses on technical details.
[0061] Server: Speech synthesis
[0062] The server converts the generated commentary into audio data using a text-to-speech (TTS) engine, which is ready to be served to the user.
[0063] Server: Send data
[0064] The server transmits the generated audio data to the terminal, which uses the data to provide audio commentary to the user.
[0065] Customized commentary
[0066] Device: Audio playback
[0067] The device then provides the received audio data to the user through a playback device (e.g., speaker or headphones).The user watches the game while listening to customized commentary based on the user's sports knowledge level and the sport they are rooting for.
[0068] Specific examples
[0069] For example, consider a user who has a "beginner level of soccer knowledge" and "supports a specific team." In this case, the user uses the device to set their own knowledge level and the team they support. When the game starts, the device receives live video and streams it to the server. The server analyzes the video and generates appropriate commentary based on the user's profile. The commentary is converted into audio data using a text-to-speech engine and sent to the device. Finally, the device plays the generated audio data and provides it to the user.
[0070] This system allows users to receive real-time commentary tailored to their level of sports knowledge and the sport they are rooting for, improving the sports viewing experience.
[0071] The processing flow will be explained below.
[0072] Step 1: User inputs their sports knowledge level and the sport they support
[0073] Using the terminal, the user sets his / her sports knowledge level (e.g., beginner, intermediate, advanced) and the team or player he / she supports.
[0074] This information is entered through the terminal's user interface.
[0075] Step 2: The device sends the data to the server
[0076] The terminal receives the sports knowledge level and the information of the target of support input by the user.
[0077] The device generates an API request containing this information and sends it to the server.
[0078] Step 3: The server stores the user data
[0079] The server receives the user information sent from the terminal.
[0080] The server stores the received information in a database and creates a profile for each user.
[0081] Step 4: Receiving game footage
[0082] The device receives game footage from a live streaming service.
[0083] The received match footage is ready to be processed in real time.
[0084] Step 5: Stream the game
[0085] The device streams the received game footage to the server in real time.
[0086] The streamed video is analyzed by the server.
[0087] Step 6: The server analyzes the video in real time
[0088] The server uses a video analysis algorithm to analyze the streamed video in real time.
[0089] The server uses object recognition algorithms to recognize players' positions and actions (e.g., dribbling, passing, shooting).
[0090] The server also keeps track of the score and the position of the ball.
[0091] Step 7: Integrating user information and analysis results
[0092] The server integrates the user's profile data with the video analysis results.
[0093] Compile materials for commentary based on the user's level of sports knowledge and the sport they support.
[0094] Step 8: Generate explanatory text
[0095] The server inputs the integrated data into the generative AI model and generates a commentary appropriate for the user.
[0096] The explanations are tailored to the user's level of knowledge, with the level of detail varying between beginners and advanced users.
[0097] Step 9: Generate audio data
[0098] The server converts the generated explanatory text into audio data using a text-to-speech (TTS) engine.
[0099] This audio data is then formatted so that it can be heard by the user.
[0100] Step 10: Send audio data to the device
[0101] The server transmits the generated voice data to the terminal.
[0102] The terminal prepares to receive the audio data.
[0103] Step 11: Playing back audio data
[0104] The terminal plays the received audio data through a playback device (e.g., speaker, headphones).
[0105] Users can watch the game while listening to customized commentary in real time.
[0106] Example 1
[0107] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0108] Conventional sports commentary systems have difficulty meeting the diverse needs of users, and have been unable to provide customized commentary based on individual sports knowledge levels or the sports fans they support. Furthermore, they lack technology that can accurately recognize the progress of a game in real time and instantly generate appropriate commentary. The present invention aims to solve these problems and provide users with real-time, customized sports commentary.
[0109] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0110] In this invention, the server includes means for inputting a user's sports knowledge level and a target player, means for transmitting the input information to the server, means for storing the input information in a database, means for receiving game footage and streaming the footage to the server, means for analyzing the game footage in real time and recognizing player positions, movements, and scores, means for generating commentary text tailored to the user's profile based on the analysis results and the user's sports knowledge level and target player, means for converting the generated commentary text into audio data using a text-to-speech engine, and means for transmitting the audio data from the server to a terminal and providing it to the user via a playback device. This enables real-time, customized sports commentary tailored to the diverse needs of users.
[0111] A "user" is an individual who watches sports and enters information into the system to receive customized commentary.
[0112] The "sports knowledge level" is an index that indicates the depth of a user's knowledge about sports, and is classified into categories such as beginner, intermediate, and advanced.
[0113] The term "target of support" refers to a team or player in which the user has a particular interest and supports.
[0114] A "terminal" is an electronic device that a user uses to provide input information, receive game footage, and play audio data.
[0115] The "server" is a centralized management system that stores information sent by users, analyzes game footage, generates commentary, and distributes audio data.
[0116] A "database" is a system located within a server that stores and manages user information and analysis results.
[0117] "Game footage" refers to video data of sports games distributed via live streaming services.
[0118] "Streaming" refers to the act of transmitting or receiving game footage in real time.
[0119] "Video analysis" refers to the process of processing game footage in real time and recognizing players' positions, movements, scoring situations, etc.
[0120] A "user profile" is a data set that includes individually set information such as the user's level of sports knowledge and the sport they support.
[0121] "Live commentary" is a narration text that is generated based on the progress of the match.
[0122] A "text-to-speech engine" is a software technology for converting text data into speech data.
[0123] "Audio data" refers to commentary data in audio format that is generated by a text-to-speech synthesis engine and provided to the user.
[0124] A "playback device" refers to a device such as a speaker or headphones that is connected to a terminal and allows the user to listen to audio data.
[0125] This invention relates to a system that uses generative AI to generate live sports commentary audio in real time. This system can provide customized commentary according to the user's sports knowledge level and the sport they are rooting for. Each processing step of the system is described in detail below.
[0126] User Profile Settings
[0127] User: Enter information
[0128] Users use the terminal to input their sports knowledge level (beginner, intermediate, advanced) and information about the team or player they support. The input information is used as the basis for customization settings within the system.
[0129] Terminal: Data transmission
[0130] The device sends the information entered by the user to the server as an API request. Specifically, the device packages the input information in JSON format and sends an HTTP request to the server.
[0131] Server: Data storage
[0132] The server stores the data of the received API request in a database. For example, the database stores information such as the user ID, knowledge level, and support team.
[0133] Real-time analysis of game footage
[0134] Terminal: Video reception
[0135] The device receives game footage from a live streaming service, for example, by obtaining live footage in HLS format from a streaming URL.
[0136] Terminal:Video Stream
[0137] The device then streams the received game footage to the server in real time, using the RTMP or WebRTC protocol to send the video data to the server.
[0138] Server: Video analysis
[0139] The server analyzes the streamed game footage using a video analysis algorithm (e.g., YOLO or OpenCV) to recognize the positions and movements of players. As a result of the analysis, the score and ball position can be determined.
[0140] Generating live commentary audio
[0141] Server: Customization information integration
[0142] The server integrates the user profile information with the video analysis results. For example, the analysis results might link information such as "Player A is dribbling" with information such as "the user is a beginner."
[0143] Server: Explanation generation
[0144] The server inputs the integrated data using a generative AI model and generates commentary appropriate for the user. As a specific example, for beginners, a commentary such as "Player A broke through the opponent with a great dribble" is generated. An example of a prompt is as follows:
[0145] "Generate a play-by-play commentary of a soccer game, given that the user has a beginner level of knowledge and identifies the team they support."
[0146] Server: Speech synthesis
[0147] The server converts the generated commentary into voice data using a text-to-speech engine (TTS engine). Specifically, it uses the API of the TTS engine to convert the text "Player A made a great dribble..." into voice data.
[0148] Server: Sends audio data
[0149] The server sends the generated audio data to the user's device by returning the audio data as an HTTP response.
[0150] Customized commentary
[0151] Device: Audio playback
[0152] The terminal provides the received audio data to the user through a playback device (speaker or headphones), specifically by decoding the audio data and sending it to the audio system.
[0153] In this way, the system can provide real-time, customized sports commentary that meets the diverse needs of users.
[0154] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0155] Step 1:
[0156] User: Enter sports knowledge level and support target
[0157] Using the terminal, users input their sports knowledge level (beginner, intermediate, advanced) and information about the team and player they support.
[0158] Input: User knowledge level, target of support
[0159] Output: User input data (knowledge level and support target)
[0160] Specific actions: The user operates text fields and selection menus on the device screen to enter the required information.
[0161] Step 2:
[0162] Terminal: Data transmission
[0163] The terminal sends the information entered by the user to the server as an API request.
[0164] Input: User input data (knowledge level and support target)
[0165] Output: API request
[0166] Specific operation: The terminal converts the user's input data into JSON format and sends it to the server as an HTTP POST request.
[0167] Step 3:
[0168] Server: Data storage
[0169] The server analyzes the received API requests and stores them in a database.
[0170] Input: API request
[0171] Output: User profile stored in the database
[0172] Specific operation: The server parses the received JSON data and saves the corresponding user profile in the database.
[0173] Step 4:
[0174] Terminal: Receiving game footage
[0175] The device receives game footage from a live streaming service.
[0176] Input: Live Streaming URL
[0177] Output: Received game video
[0178] Specific operation: The device retrieves live game footage in HLS or DASH format from the specified URL.
[0179] Step 5:
[0180] Terminal:Video Stream
[0181] The device streams the received game footage to the server in real time.
[0182] Input: Received game footage
[0183] Output: Streamed game footage
[0184] Specific operation: The device sends video data to the server using the RTMP or WebRTC protocol.
[0185] Step 6:
[0186] Server: Video analysis
[0187] The server analyzes the received video in real time and recognizes the players' positions and movements.
[0188] Input: Streamed game footage
[0189] Output: Analysis results (player positions, actions, scoring status)
[0190] Specific operation: The server performs analysis processing using a video analysis algorithm (e.g., YOLO or OpenCV).
[0191] Step 7:
[0192] Server: Customization information integration
[0193] The server integrates the user profile information with the video analysis results.
[0194] Input: User profile, analysis results
[0195] Output: Integrated data
[0196] Specific operation: The server links the analysis results with the user's knowledge level and support target information, and integrates them into an appropriate data structure.
[0197] Step 8:
[0198] Server: Live commentary generation
[0199] The server generates commentary using a generative AI model.
[0200] Input: Integrated data
[0201] Output: Commentary
[0202] Specific operation: The server inputs the integrated data into the generative AI model and generates a commentary appropriate for the user.
[0203] Example prompt: "Generate a play-by-play commentary of a soccer game given the user's beginner level of knowledge and which team they support."
[0204] Step 9:
[0205] Server: Speech synthesis
[0206] The server converts the generated commentary into voice data using a text-to-speech synthesis engine.
[0207] Input: Commentary
[0208] Output: Audio data
[0209] Specific operation: The server uses a TTS engine (e.g., Google TTS API) to convert the text into audio data.
[0210] Step 10:
[0211] Server: Sends audio data
[0212] The server transmits the generated voice data to the terminal.
[0213] Input: Audio data
[0214] Output: HTTP response
[0215] Specific operation: The server sends the audio data to the terminal as an HTTP response.
[0216] Step 11:
[0217] Device: Audio playback
[0218] The terminal provides the received audio data to the user on a playback device.
[0219] Input: Audio data
[0220] Output: Played audio
[0221] Specific operation: The device decodes the audio data and plays it through speakers or headphones.
[0222] Through the above steps, users can enjoy real-time commentary that is customized to their level of sports knowledge and the sport they are rooting for.
[0223] (Application example 1)
[0224] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0225] Conventional sports commentary systems have difficulty customizing commentary to suit individual users' knowledge levels or the teams they support, making it impossible to provide appropriate commentary in real time. Furthermore, the process of converting the generated commentary text into audio data and providing it to users' devices in real time was lacking. As a result, users only received the same commentary, which led to a problem of degrading the quality of each individual viewing experience.
[0226] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0227] In this invention, the server includes means for inputting a user's sports knowledge level and a team the user supports, means for analyzing game footage in real time to recognize player positions, movements, and scores, means for generating commentary according to the user's profile based on the analysis results, means for converting the commentary into voice data using a text-to-speech engine, means for providing the voice data to the user, means for receiving game footage via live streaming and transmitting it to the server, means for generating appropriate commentary from prompts using a generative AI model, and means for transmitting the generated voice data to the user's device in real time and playing it back. This makes it possible to provide appropriate and customized commentary in real time according to the user's sports knowledge level and the team the user supports.
[0228] The "user's sports knowledge level" indicates the depth of knowledge and understanding of the user about sports, and is usually classified into stages such as beginner, intermediate, and advanced.
[0229] The term "target of support" refers to a team or player that the user particularly supports in a particular sporting event or game.
[0230] "Game footage" refers to video data that records the progress of a sporting event or game in real time.
[0231] "Real-time analysis" means receiving game video data, processing it instantly, and extracting the necessary information and features.
[0232] "Player position" is information indicating where each player is located on the playing field at a specific time during the game.
[0233] "Action" refers to the specific actions or movements that a player makes during a game or competition, such as dribbling, passing, and shooting.
[0234] "Scoring Status" refers to information about the scores of both teams during the match, including the current score, the players who scored, and the method of scoring.
[0235] "Live commentary" is commentary based on the progress of the match and the actions of the players, and is intended to explain the content of the match to viewers in an easy-to-understand manner.
[0236] A "text-to-speech engine" is software or a process for converting input text data into speech data.
[0237] "Audio data" means a data file in audio format generated by a text-to-speech engine.
[0238] "Live streaming reception" is the process of instantly acquiring video and audio that is being streamed in real time over the Internet.
[0239] "Send to server" means sending data from a terminal or the like to a central server.
[0240] A "generative AI model" is an algorithm or framework for generating text or content using artificial intelligence.
[0241] A "prompt" is an instruction that is input into a generative AI model, and is text that causes the model to generate appropriate text or content based on that instruction.
[0242] "Terminal" means a device used by a user, including a smartphone, tablet, computer, etc.
[0243] "Playback" means aurally reproducing the generated audio data using an audio output device.
[0244] The system based on this invention utilizes generative AI models to provide users with customized, real-time commentary while watching sports.
[0245] 1. Setting up your user profile
[0246] User:
[0247] Users input their sports knowledge level (e.g., beginner, intermediate, advanced) and their favorite teams and athletes via their terminals. This information forms the basis for customization settings used throughout the system.
[0248] Device:
[0249] The information entered by the user is sent from the device to the server. Here, the user interface is built using React Native, and the device generates an API request to send the data.
[0250] server:
[0251] The server stores the received user information in a database (e.g., PostgreSQL) and creates a profile for each user, which includes the user's level of sports knowledge and the sport they support.
[0252] 2. Real-time analysis of game footage
[0253] Device:
[0254] The device uses WebRTC to receive game footage in real time from a live streaming service and streams the video data to a server.
[0255] server:
[0256] The server uses OpenCV to analyze the received video in real time, and uses object recognition algorithms to identify the player's position and recognize their actions (e.g., dribbling, passing, shooting), as well as to track the goal situation and ball position.
[0257] 3. Generation of commentary audio
[0258] server:
[0259] The server combines the user profile data with the video analysis results and inputs prompts into the generative AI model to generate appropriate commentary. For example, the following prompts are input into the generative AI model:
[0260] "Generate beginner-friendly commentary: Game status {Score}, player {Name} performed {Action}."
[0261] The generated commentary is converted into audio data using the Google Cloud Text-to-Speech API.
[0262] 4. Customized commentary
[0263] server:
[0264] The generated voice data is transmitted from the server to the terminal.
[0265] Device:
[0266] The device then provides the received audio data to the user via a playback device (e.g., smartphone, head-mounted display), allowing users to watch the game while listening to commentary tailored to their level of sports knowledge and the sport they are rooting for.
[0267] Specific examples
[0268] For example, consider a case where a user is at an "intermediate level" and is cheering for "Team A." If Player X is recognized as dribbling during a match and reaching the goal, the generative AI model will generate an appropriate commentary as follows:
[0269] Example prompt sentence:
[0270] "During a game, Player X dribbles the ball and reaches the goal. Please provide appropriate commentary for intermediate level users."
[0271] The generated commentary is converted into audio data using the Google Cloud Text-to-Speech API and sent to the user's device in real time, where they can enjoy this customized commentary via a head-mounted display or smartphone.
[0272] In this way, by implementing the invention, users can receive appropriate and customized commentary in real time according to their level of sports knowledge and the team they support.
[0273] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0274] Program processing steps and their specific explanations
[0275] Step 1: Configure your user profile
[0276] Users use a smartphone app to input their sports knowledge level (beginner, intermediate, advanced) and the team or player they support. The input data is then saved on the device.
[0277] Input: The user enters their sports knowledge level and the team or player they support into the input form.
[0278] Data processing: Convert input data into API request format
[0279] Output: Sent to the server as an API request
[0280] Step 2: Send and store user information
[0281] The terminal sends the information entered by the user to the server, where it is saved in a database that stores user profiles.
[0282] Input: User's sports knowledge level and favorite team and player information
[0283] Data processing: Converting transmitted data into a database format
[0284] Output: User profile saved in database
[0285] Step 3: Receiving and streaming game footage
[0286] The device receives game footage from the live streaming service using the WebRTC protocol and transmits it to the server.
[0287] Input: Real-time game footage from a live streaming service
[0288] Data processing: Converts video data into a format suitable for the server in real time
[0289] Output: Streaming data is sent to the server
[0290] Step 4: Video analysis
[0291] The server analyzes the received game footage in real time using OpenCV to recognize the players' positions, movements, and scoring status.
[0292] Input: Streamed game footage
[0293] Data processing: Apply object recognition algorithms to recognize player positions, movements, and scoring situations
[0294] Output: Analysis results: player positions, movements, and score information
[0295] Step 5: Generate commentary
[0296] Based on the analysis results and the user profile, the server inputs prompt text into the generative AI model to generate appropriate commentary.
[0297] Input: Analysis results and user profile
[0298] Data processing: Input prompt text into the generative AI model to generate commentary
[0299] Output: Generated commentary
[0300] Example of specific behavior: Send the following prompt to the generative AI model: "Generate commentary appropriate for beginners: Game situation {Score}, Player {Name} performed {Action}."
[0301] Step 6: Convert the commentary into audio
[0302] The server converts the generated commentary into audio data using the Google Cloud Text-to-Speech API.
[0303] Input: Generated commentary
[0304] Data processing: Convert text data into audio data
[0305] Output: Audio data
[0306] Step 7: Sending audio data
[0307] The server transmits the converted voice data to the user's terminal.
[0308] Input: Audio data
[0309] Data processing: Real-time transmission of voice data
[0310] Output: Audio data sent to the user's device
[0311] Step 8: Play commentary audio
[0312] The terminal plays the received audio data on a playback device (e.g., smartphone, head-mounted display) and provides it to the user.
[0313] Input: Audio data sent from the server
[0314] Data processing: Conversion to playback format
[0315] Output: Audio description played from the playback device
[0316] By having each processing step work in conjunction with each other in this way, users can receive live commentary in real time that is individually customized according to their level of sports knowledge and the sport they are rooting for.
[0317] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0318] This invention relates to a system that combines generative AI and an emotion engine to generate live sports commentary audio in real time. This allows for customizable live commentary based on the user's sports knowledge level, the sport they are rooting for, and even their emotional state. The detailed program processing of this system is explained in natural language.
[0319] User Profile Settings
[0320] User: Enter information
[0321] Users use their devices to set their sports knowledge level (e.g., beginner, intermediate, advanced) and the team or player they support. This information forms the basis for customization settings used throughout the system.
[0322] Terminal: Data transmission
[0323] The device receives the sports knowledge level and the information of the target sport entered by the user, and then generates an API request including this information and sends it to the server.
[0324] Server: Data storage
[0325] The server stores the received user information in a database and creates a profile for each user, which specifies the user's level of sports knowledge and the sport they support.
[0326] Real-time analysis of game footage
[0327] Terminal: Video reception
[0328] The device receives game footage from a live streaming service, which prepares it for processing in real time.
[0329] Device: Video streaming
[0330] The device streams the received game video to the server in real time, where it is analyzed.
[0331] Server: Video analysis
[0332] The server uses video analysis algorithms to analyze the streamed video in real time, and object recognition algorithms to recognize player positions and actions (e.g., dribbling, passing, shooting), as well as to determine goal situations and ball position.
[0333] Recognizing user emotions
[0334] Device: Emotion data acquisition
[0335] The device uses an emotion engine to analyze the user's facial expressions and voice tone to recognize the user's emotional state.
[0336] Server: Sends and stores emotion data
[0337] The device sends the recognized emotion data to the server, which then integrates it into the user profile.
[0338] Generating live commentary audio
[0339] Server: Customization information integration
[0340] The server integrates the user's sports knowledge level, favorites, and emotional data, which are used to generate commentary.
[0341] Server: Explanation generation
[0342] The server inputs the integrated data into a generative AI model to generate commentary tailored to the user, using different expressions and tones depending on the user's knowledge level and emotional state.
[0343] Server: Generates voice data
[0344] The server converts the generated commentary into audio data using a text-to-speech (TTS) engine, which is then formatted for listening by the user.
[0345] Server: Sending audio data
[0346] The server transmits the generated audio data to the terminal, which uses this data to provide audio commentary to the user.
[0347] Customized commentary
[0348] Device: Audio playback
[0349] The device then plays the received audio data through a playback device (e.g., speaker, headphones). Users can watch the game while listening to a customized commentary in real time. The commentary is tailored to the user's sports knowledge level, favorite sport, and emotional state, providing a more personalized experience.
[0350] Specific examples
[0351] For example, consider a user with a beginner level of soccer knowledge who is rooting for a specific team and is excited during a match. The user uses their device to set their own knowledge level and the team they are rooting for. When the game starts, the device receives the footage and streams it to the server. The server analyzes the footage and receives the user's emotional state from the device. The emotion engine detects the user's level of excitement, and the server generates a commentary based on this information. The commentary is written in a tone that matches the user's excitement, such as "Great goal, your team is leading!" The generated commentary is converted into audio data, sent to the device, and then provided to the user. This system allows users to receive commentary appropriate to their sports knowledge level, the team they are rooting for, and their real-time emotional state.
[0352] The processing flow will be explained below.
[0353] Step 1: User inputs their sports knowledge level and the sport they support
[0354] Using the terminal, the user sets his / her sports knowledge level (e.g., beginner, intermediate, advanced) and the team or player he / she supports.
[0355] This provides basic information that the system needs to provide the user with the most appropriate explanation.
[0356] Step 2: The device sends the data to the server
[0357] The terminal receives the sports knowledge level and the information of the target of support input by the user.
[0358] The device then generates an API request containing this information and sends it to the server.
[0359] Step 3: The server stores the user data
[0360] The server receives the user information sent from the terminal.
[0361] The received user information is stored in a database and a profile is created for each user.
[0362] Step 4: Receiving game footage
[0363] The device receives game footage from a live streaming service.
[0364] The received match footage is ready to be processed in real time.
[0365] Step 5: Stream the game
[0366] The device streams the received game footage to the server in real time.
[0367] The streamed video is analyzed by the server.
[0368] Step 6: The server analyzes the video in real time
[0369] The server uses a video analysis algorithm to analyze the streamed video in real time.
[0370] The server uses object recognition algorithms to recognize players' positions and actions (e.g., dribbling, passing, shooting), as well as to track the scoring situation and ball position.
[0371] Step 7: Recognize user emotions
[0372] The device uses an emotion engine to analyze the user's facial expressions and voice tone to recognize the user's emotional state.
[0373] The device transmits the emotional state as data to the server.
[0374] Step 8: The server stores the emotion data
[0375] The server receives the emotion data sent from the terminal.
[0376] The emotion data is integrated into the user profile and the profile information is updated.
[0377] Step 9: Generate explanatory text
[0378] The server integrates the user's sports knowledge level, favorite target, and emotion data.
[0379] Based on this, a generative AI model is used to generate live commentary that is appropriate for the user.
[0380] The content of the commentary is adjusted according to the user's knowledge level and emotional state.
[0381] Step 10: Creating audio data
[0382] The server converts the generated explanatory text into audio data using a text-to-speech (TTS) engine.
[0383] The converted audio data is ready to be provided to the user.
[0384] Step 11: Send audio data to the device
[0385] The server transmits the generated voice data to the terminal.
[0386] The terminal prepares to play the received audio data.
[0387] Step 12: Playing back audio data
[0388] The terminal plays the received audio data through a playback device (e.g., speaker, headphones).
[0389] Users can watch the game while listening to customized commentary in real time.
[0390] The commentary is customized to the user's level of sports knowledge, their roots, and their real-time emotional state.
[0391] Example 2
[0392] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0393] Current live sports commentary systems often provide uniform commentary without taking into account the individual user's knowledge level or emotional state. As a result, they are unable to address the different needs and expectations of each user, making it difficult to provide a personalized experience that matches each user's level of excitement and understanding. Furthermore, current systems often lack the accuracy of real-time video and emotion analysis, making it difficult to generate appropriate commentary based on user profiles. A system that solves these issues is needed.
[0394] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0395] In this invention, the server includes means for inputting a user's sports knowledge level and a target sport, means for analyzing game footage in real time to recognize the positions, movements, and score of players, means for analyzing the user's facial expressions and tone of voice to recognize the user's emotional state, means for generating commentary according to the user's profile based on the analysis results and the user's emotional state, means for converting the commentary into voice data using a text-to-speech engine, and means for providing the voice data to the user, thereby enabling the provision of customizable commentary in real time according to the user's individual knowledge level and emotional state.
[0396] The "user's sports knowledge level" indicates the user's knowledge and understanding of sports, and is expressed as a level such as beginner, intermediate, or advanced.
[0397] The term "target of support" refers to a specific team or player that the user supports during a game.
[0398] "Input means" includes devices or interfaces that allow a user to provide information to a system, such as a keyboard, touch screen, or voice input.
[0399] "Game footage" refers to video data showing a sports game being played, including live streaming and recorded footage.
[0400] "Real-time analysis" refers to the process of analyzing game video data as it is received and making the results immediately available.
[0401] "Player position" is information indicating the current location of each player within the competition area.
[0402] "Action" refers to specific actions that players perform during a game, such as dribbling, passing, shooting, etc.
[0403] "Scoring status" is information showing the distribution of scores in the current match and the latest scoring results.
[0404] "Means for analyzing" includes algorithms and software for processing the obtained data and extracting useful information, such as object recognition algorithms.
[0405] "Facial expressions" and "vocal tones" are physical and vocal characteristics that indicate the user's emotional state, and analyzing them allows the user's emotions to be inferred.
[0406] An "emotional state" refers to the psychological state that a user is feeling at a given moment, such as excitement, joy, sadness, etc.
[0407] "Means for recognizing emotional states" includes technologies and devices that analyze user characteristics such as facial expressions and vocal tone to estimate the emotions at that time.
[0408] A "user profile" is a data structure for centrally managing information related to a user, and includes information such as the user's level of sports knowledge, the sport they support, and their emotional state.
[0409] "Live commentary" is text that explains the progress and events of the match to users.
[0410] A "text-to-speech engine (TTS)" is a technology or software for converting text data into audio data.
[0411] "Audio data" refers to data in an audio format that is generated by a text-to-speech synthesis engine and can be heard by a user.
[0412] The "means for providing" includes devices and software for actually playing back the audio data generated by the system to the user, such as speakers and headphones.
[0413] This invention is a system that provides customized real-time commentary based on a user's sports knowledge level, emotional state, and favorite sport. The system consists of a terminal, a server, and related software and hardware.
[0414] User Profile Settings
[0415] First, a user uses their device to input their sports knowledge level (e.g., beginner, intermediate, advanced) and the team or player they support. The device receives this information, generates an API request, and sends it to the server. The server stores the received user information in a database and creates a profile for each user. The profile specifies the user's sports knowledge level and the team or player they support.
[0416] Real-time analysis of game footage
[0417] Game footage is received by the device from the live streaming service. The device then streams the received footage to the server in real time. The server then uses a video analysis algorithm to analyze the streamed footage. The server uses an object recognition algorithm to recognize player positions, actions (e.g., dribbling, passing, shooting), scoring situations, and ball position.
[0418] Recognizing user emotions
[0419] The device captures the user's facial expressions and voice tone using a camera and microphone, and analyzes them using an emotion engine to recognize the user's emotional state (e.g., excitement, joy, sadness, etc.). The recognized emotion data is sent from the device to a server and integrated into the user profile.
[0420] Generating live commentary audio
[0421] The server then combines the user's sports knowledge level, the sport they support, and their emotional data and inputs it into a generative AI model. Based on this combined data, the generative AI model generates a commentary tailored to the user. The commentary content is adjusted according to the user's knowledge level and emotional state. For example, an excited user will receive detailed commentary in an emotional tone.
[0422] Generating and providing voice data
[0423] The server converts the generated commentary into audio data using a text-to-speech (TTS) engine. This audio data is then sent to the device, which then plays the received audio data through a playback device (e.g., speaker, headphones). Users can watch the game while listening to the customized commentary in real time.
[0424] Specific examples
[0425] For example, consider a case where a user has a "beginner level of soccer knowledge" and is "supporting a specific team," becoming excited during a game. The user uses their device to set their own knowledge level and the team they are supporting. When the game video begins, the device receives the video and streams it to the server. The server analyzes the video and also receives the user's emotional state from the device. The emotion engine detects the user's level of excitement, and the server generates a commentary based on this. The commentary is in a tone that matches the user's excitement, such as "Great goal, your team is leading!" The generated commentary is converted into audio data, sent to the device, and then provided to the user. This system allows users to receive commentary appropriate to their sports knowledge level, the team they are supporting, and their real-time emotional state.
[0426] Prompt Sentence Examples
[0427] "Generate a commentary of a soccer game in which a beginner-level user who is rooting for a specific team gets excited when a goal is scored."
[0428] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0429] Step 1:
[0430] The user uses the device's input interface to set their sports knowledge level (e.g., beginner, intermediate, advanced) and the team or player they support. This becomes the input data. The device receives this information and generates an API request. The information included in this request (the user's sports knowledge level and the team they support) is sent to the server.
[0431] Step 2:
[0432] The server processes the received API request and stores the user information in a database. This creates a profile for each user, including the user's sports knowledge level and favorite sports. The output data is added to the database as a user profile.
[0433] Step 3:
[0434] When the match begins, the device receives the match video from the live streaming service. This video data becomes the input data, and is ready to be processed in real time.
[0435] Step 4:
[0436] The device receives the game video in real time and streams it to the server. The streamed video becomes the input data. The server prepares this video data for analysis.
[0437] Step 5:
[0438] The server analyzes the streamed video in real time using an object recognition algorithm. Using the video data as input, it recognizes the player's position, actions (e.g., dribbling, passing, shooting), scoring status, and ball position. The analysis results are output data.
[0439] Step 6:
[0440] The device captures the user's facial expressions and voice tone using a camera and microphone, which serve as input data. The emotion engine analyzes this data to recognize the user's emotional state, and the recognized emotional state is generated as output data.
[0441] Step 7:
[0442] The device sends the recognized emotion data to the server, which then integrates it into the user profile and keeps it up to date. This updates the user profile.
[0443] Step 8:
[0444] The server inputs the integrated user profile into the generative AI model. The input data includes the user's sports knowledge level, the people they support, and emotional data. Based on this data, the generative AI model generates a commentary appropriate for the user. This becomes the output data.
[0445] Step 9:
[0446] The server inputs the generated commentary into a text-to-speech (TTS) engine. The commentary text is used as input data. The TTS engine converts it into voice data. The voice data is generated as output data.
[0447] Step 10:
[0448] The server sends the generated audio data to the device, which becomes the input data. The device receives this data and plays it through a playback device (e.g., speaker, headphones). This allows the user to listen to a customized commentary in real time.
[0449] (Application example 2)
[0450] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0451] Conventional sports commentary systems have difficulty providing personalized commentary in real time based on the user's sports knowledge level or the team they support. They also lack the ability to provide customized commentary based on the user's emotional state. Therefore, multifaceted customization is needed to provide users with a more engaging and immersive viewing experience.
[0452] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0453] In this invention, the server includes: means for inputting a user's sports knowledge level and a target player; means for analyzing game footage in real time to recognize player positions, movements, and scores; means for analyzing the user's facial expressions and vocal tone to recognize the user's emotional state; means for generating prompt sentences based on the analysis results and the user's emotional state; means for inputting the prompt sentences into a generative AI model to generate commentary according to the user's profile and emotional state; means for converting the commentary into voice data using a text-to-speech engine; and means for providing the voice data to the user. This makes it possible to provide highly personalized commentary in real time based on the user's sports knowledge level, target player, and real-time emotional state.
[0454] "User's sports knowledge level" is information indicating the depth of knowledge and understanding of the user regarding sports.
[0455] "Support target" is information about a team or player that the user particularly supports.
[0456] "Game video" refers to real-time video data of a sports game.
[0457] "Player position" refers to information about the current position of each player within the competition venue.
[0458] "Action" refers to a series of movements that a player makes during a game (e.g., dribbling, passing, shooting, etc.).
[0459] "Scoring status" refers to the scoring status and score information during a match.
[0460] "User's facial expression" is information for reading emotions from the user's facial features and movements.
[0461] "Voice tone" is information that indicates the tone or pitch of the voice uttered by the user.
[0462] "User's emotional state" refers to the type of emotion the user is currently feeling (e.g., excitement, joy, sadness, etc.).
[0463] A "prompt sentence" is a sentence that contains prior information for generating an explanatory sentence and is input into the generative AI model.
[0464] A "generative AI model" is an artificial intelligence system that generates text for a specific task (in this case, commentary) based on input data.
[0465] A "text-to-speech synthesis engine" is a system that converts text into speech data.
[0466] "Audio data" is audio information that has been converted into a format that can be heard by a user.
[0467] This invention relates to a system for generating and providing play-by-play commentary audio that is customized based on the user's sports knowledge level and the sport they support, and also in response to the user's real-time emotional state. To achieve this, the following specific means and processes are required.
[0468] User Profile Settings
[0469] Users can enter their sports knowledge level and favorite teams and athletes via their device. This information is sent from the device to the server, which then receives it and stores it in a database, creating a customized profile for each user.
[0470] Real-time analysis of game footage
[0471] The device receives game footage from a live streaming service and streams it in real time to a server, which then analyzes the footage using object recognition algorithms to identify player positions, actions (e.g., dribbling, passing, shooting), and scoring situations.
[0472] Recognizing user emotions
[0473] The device uses the user's facial expressions and tone of voice to recognize the user's emotional state through an emotion engine, which then sends the emotion data to the server and stores it as a user profile.
[0474] Generating live commentary audio
[0475] The server combines the user's sports knowledge level, the team they are rooting for, and their emotional state to generate a prompt. The generated prompt is then input into a generative AI model to generate a play-by-play commentary appropriate for the user. For example, a prompt might be in the format: "The user's knowledge level is beginner, and the team they are rooting for is Team A. Their current emotional state is excited. The analysis results of the game are as follows: Team A has scored a goal!"
[0476] The generated commentary is converted into audio data using a text-to-speech (TTS) engine, which is then formatted for the user to hear.
[0477] Customized commentary
[0478] The terminal plays the audio data received from the server through a playback device (e.g., speaker, headphones), allowing the user to watch the game while listening to customized commentary in real time.
[0479] Specific examples
[0480] For example, if a user has a beginner level of soccer knowledge and supports a specific team (Team A), and the server detects that the user is excited, the server will generate the following prompt:
[0481] The user's knowledge level is beginner, and the team they support is Team A. Their current emotional state is excited. The analysis result of the match is as follows: Team A has scored a goal!
[0482] Based on this prompt, the generative AI model generates a commentary such as "Great goal, Team A is in the lead!", which is then converted into audio using a text-to-speech engine and sent to the device. The device then plays this audio through a playback device and provides it to the user.
[0483] In this way, the present invention is able to provide a customized commentary that is tailored to the user's profile and real-time emotional state.
[0484] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0485] Step 1:
[0486] A user inputs their level of sports knowledge and the teams and players they support via a terminal. The input information is data that indicates the user's sports viewing preferences. The terminal receives this data and sends it to a server. The server stores the received data in a database and generates a user profile. This profile includes information about the user's level of sports knowledge and the teams and players they support.
[0487] Step 2:
[0488] The device receives game footage from a live streaming service. The received video data is a real-time video stream showing the specific situation of the game. The device streams this video data to a server, which analyzes it in real time using video analysis algorithms. The analysis results include player positions, actions (e.g., dribbling, passing, shooting), and scoring status.
[0489] Step 3:
[0490] The device analyzes the user's facial expressions and vocal tone using an emotion engine. The analyzed emotion data indicates the user's real-time emotional state (e.g., excitement, joy, sadness, etc.). The device transmits this emotion data to the server, which then integrates the received emotion data into the user profile.
[0491] Step 4:
[0492] The server generates a prompt sentence based on the user's sports knowledge level, the team they are rooting for, and their emotional state. This prompt sentence contains prior information to be input into the generative AI model. For example, a prompt sentence of the form "The user's knowledge level is beginner level, and the team they are rooting for is Team A. Their current emotional state is excited. The analysis results of the game are as follows: Team A has scored a goal!" is generated.
[0493] Step 5:
[0494] The server inputs the generated prompt into a generative AI model to generate a commentary appropriate for the user. The generative AI model is an artificial intelligence system that generates commentary based on the prompt. The content of the generated commentary is adjusted according to the user's knowledge level and emotional state. As a result, a commentary appropriate for the user is obtained.
[0495] Step 6:
[0496] The server converts the generated commentary into audio data using a text-to-speech (TTS) engine, which then formats the audio data in a user-friendly format. Users can listen to this audio data in real time while watching the game.
[0497] Step 7:
[0498] The terminal plays the audio data received from the server through a playback device (e.g., speaker, headphones). This allows the user to listen to commentary customized according to the situation of the game. For example, if the user is excited, commentary such as "Great goal, Team A is in the lead!" is provided in an excited tone.
[0499] In this way, users can experience a personalized play-by-play commentary based on their level of sports knowledge, their favorite team, and their real-time emotional state.
[0500] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0501] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0502] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0503] [Second embodiment]
[0504] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0505] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0506] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0507] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0508] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0509] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0510] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0511] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0512] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0513] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0514] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0515] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0516] This invention relates to a system that uses generative AI to generate live sports commentary audio in real time. This allows for customizable live commentary based on the user's level of sports knowledge and the sport they are rooting for. The detailed program processing of this system is explained in natural language.
[0517] User Profile Settings
[0518] User: Enter information
[0519] Users use their devices to set their sports knowledge level (e.g., beginner, intermediate, advanced) and the team or player they support. This information forms the basis for customization settings used throughout the system.
[0520] Terminal: Data transmission
[0521] The device generates and sends an API request to the server to transmit the information entered by the user, including the sports knowledge level and support target selected by the user.
[0522] Server: Data storage
[0523] The server stores the received user information in a database and creates a profile for each user, which specifies the user's level of sports knowledge and the sport they support.
[0524] Real-time analysis of game footage
[0525] Terminal: Video reception
[0526] The device receives game footage from a live streaming service, which is then streamed to a server.
[0527] Terminal:Video Stream
[0528] The device then streams the received game footage to a server, where the streaming video becomes the input data for real-time analysis.
[0529] Server: Video analysis
[0530] The server analyzes the received video using a video analysis algorithm, specifically an object recognition algorithm that tracks the player's position in real time, recognizes each player's actions (e.g., dribbling, passing, shooting), and determines the goal situation and ball position.
[0531] Generating live commentary audio
[0532] Server: Customization information integration
[0533] The server combines user profile data with video analysis results, and this information is used to generate commentary.
[0534] Server: Explanation generation
[0535] The server then inputs the integrated data into a generative AI model to generate commentary tailored to the user. This commentary varies in detail depending on the user's level of sports knowledge. For example, commentary for beginners starts with basic explanations, while commentary for advanced players focuses on technical details.
[0536] Server: Speech synthesis
[0537] The server converts the generated commentary into audio data using a text-to-speech (TTS) engine, which is ready to be served to the user.
[0538] Server: Send data
[0539] The server transmits the generated audio data to the terminal, which uses the data to provide audio commentary to the user.
[0540] Customized commentary
[0541] Device: Audio playback
[0542] The device then provides the received audio data to the user through a playback device (e.g., speaker or headphones).The user watches the game while listening to customized commentary based on the user's sports knowledge level and the sport they are rooting for.
[0543] Specific examples
[0544] For example, consider a user who has a "beginner level of soccer knowledge" and "supports a specific team." In this case, the user uses the device to set their own knowledge level and the team they support. When the game starts, the device receives live video and streams it to the server. The server analyzes the video and generates appropriate commentary based on the user's profile. The commentary is converted into audio data using a text-to-speech engine and sent to the device. Finally, the device plays the generated audio data and provides it to the user.
[0545] This system allows users to receive real-time commentary tailored to their level of sports knowledge and the sport they are rooting for, improving the sports viewing experience.
[0546] The processing flow will be explained below.
[0547] Step 1: User inputs their sports knowledge level and the sport they support
[0548] Using the terminal, the user sets his / her sports knowledge level (e.g., beginner, intermediate, advanced) and the team or player he / she supports.
[0549] This information is entered through the terminal's user interface.
[0550] Step 2: The device sends the data to the server
[0551] The terminal receives the sports knowledge level and the information of the target of support input by the user.
[0552] The device generates an API request containing this information and sends it to the server.
[0553] Step 3: The server stores the user data
[0554] The server receives the user information sent from the terminal.
[0555] The server stores the received information in a database and creates a profile for each user.
[0556] Step 4: Receiving game footage
[0557] The device receives game footage from a live streaming service.
[0558] The received match footage is ready to be processed in real time.
[0559] Step 5: Stream the game
[0560] The device streams the received game footage to the server in real time.
[0561] The streamed video is analyzed by the server.
[0562] Step 6: The server analyzes the video in real time
[0563] The server uses a video analysis algorithm to analyze the streamed video in real time.
[0564] The server uses object recognition algorithms to recognize players' positions and actions (e.g., dribbling, passing, shooting).
[0565] The server also keeps track of the score and the position of the ball.
[0566] Step 7: Integrating user information and analysis results
[0567] The server integrates the user's profile data with the video analysis results.
[0568] Compile materials for commentary based on the user's level of sports knowledge and the sport they support.
[0569] Step 8: Generate explanatory text
[0570] The server inputs the integrated data into the generative AI model and generates a commentary appropriate for the user.
[0571] The explanations are tailored to the user's level of knowledge, with the level of detail varying between beginners and advanced users.
[0572] Step 9: Generate audio data
[0573] The server converts the generated explanatory text into audio data using a text-to-speech (TTS) engine.
[0574] This audio data is then formatted so that it can be heard by the user.
[0575] Step 10: Send audio data to the device
[0576] The server transmits the generated voice data to the terminal.
[0577] The terminal prepares to receive the audio data.
[0578] Step 11: Playing back audio data
[0579] The terminal plays the received audio data through a playback device (e.g., speaker, headphones).
[0580] Users can watch the game while listening to customized commentary in real time.
[0581] Example 1
[0582] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0583] Conventional sports commentary systems have difficulty meeting the diverse needs of users, and have been unable to provide customized commentary based on individual sports knowledge levels or the sports fans they support. Furthermore, they lack technology that can accurately recognize the progress of a game in real time and instantly generate appropriate commentary. The present invention aims to solve these problems and provide users with real-time, customized sports commentary.
[0584] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0585] In this invention, the server includes means for inputting a user's sports knowledge level and a target player, means for transmitting the input information to the server, means for storing the input information in a database, means for receiving game footage and streaming the footage to the server, means for analyzing the game footage in real time and recognizing player positions, movements, and scores, means for generating commentary text tailored to the user's profile based on the analysis results and the user's sports knowledge level and target player, means for converting the generated commentary text into audio data using a text-to-speech engine, and means for transmitting the audio data from the server to a terminal and providing it to the user via a playback device. This enables real-time, customized sports commentary tailored to the diverse needs of users.
[0586] A "user" is an individual who watches sports and enters information into the system to receive customized commentary.
[0587] The "sports knowledge level" is an index that indicates the depth of a user's knowledge about sports, and is classified into categories such as beginner, intermediate, and advanced.
[0588] The term "target of support" refers to a team or player in which the user has a particular interest and supports.
[0589] A "terminal" is an electronic device that a user uses to provide input information, receive game footage, and play audio data.
[0590] The "server" is a centralized management system that stores information sent by users, analyzes game footage, generates commentary, and distributes audio data.
[0591] A "database" is a system located within a server that stores and manages user information and analysis results.
[0592] "Game footage" refers to video data of sports games distributed via live streaming services.
[0593] "Streaming" refers to the act of transmitting or receiving game footage in real time.
[0594] "Video analysis" refers to the process of processing game footage in real time and recognizing players' positions, movements, scoring situations, etc.
[0595] A "user profile" is a data set that includes individually set information such as the user's level of sports knowledge and the sport they support.
[0596] "Live commentary" is a narration text that is generated based on the progress of the match.
[0597] A "text-to-speech engine" is a software technology for converting text data into speech data.
[0598] "Audio data" refers to commentary data in audio format that is generated by a text-to-speech synthesis engine and provided to the user.
[0599] A "playback device" refers to a device such as a speaker or headphones that is connected to a terminal and allows the user to listen to audio data.
[0600] This invention relates to a system that uses generative AI to generate live sports commentary audio in real time. This system can provide customized commentary according to the user's sports knowledge level and the sport they are rooting for. Each processing step of the system is described in detail below.
[0601] User Profile Settings
[0602] User: Enter information
[0603] Users use the terminal to input their sports knowledge level (beginner, intermediate, advanced) and information about the team or player they support. The input information is used as the basis for customization settings within the system.
[0604] Terminal: Data transmission
[0605] The device sends the information entered by the user to the server as an API request. Specifically, the device packages the input information in JSON format and sends an HTTP request to the server.
[0606] Server: Data storage
[0607] The server stores the data of the received API request in a database. For example, the database stores information such as the user ID, knowledge level, and support team.
[0608] Real-time analysis of game footage
[0609] Terminal: Video reception
[0610] The device receives game footage from a live streaming service, for example, by obtaining live footage in HLS format from a streaming URL.
[0611] Terminal:Video Stream
[0612] The device then streams the received game footage to the server in real time, using the RTMP or WebRTC protocol to send the video data to the server.
[0613] Server: Video analysis
[0614] The server analyzes the streamed game footage using a video analysis algorithm (e.g., YOLO or OpenCV) to recognize the positions and movements of players. As a result of the analysis, the score and ball position can be determined.
[0615] Generating live commentary audio
[0616] Server: Customization information integration
[0617] The server integrates the user profile information with the video analysis results. For example, the analysis results might link information such as "Player A is dribbling" with information such as "the user is a beginner."
[0618] Server: Explanation generation
[0619] The server inputs the integrated data using a generative AI model and generates commentary appropriate for the user. As a specific example, for beginners, a commentary such as "Player A broke through the opponent with a great dribble" is generated. An example of a prompt is as follows:
[0620] "Generate a play-by-play commentary of a soccer game, given that the user has a beginner level of knowledge and identifies the team they support."
[0621] Server: Speech synthesis
[0622] The server converts the generated commentary into voice data using a text-to-speech engine (TTS engine). Specifically, it uses the API of the TTS engine to convert the text "Player A made a great dribble..." into voice data.
[0623] Server: Sends audio data
[0624] The server sends the generated audio data to the user's device by returning the audio data as an HTTP response.
[0625] Customized commentary
[0626] Device: Audio playback
[0627] The terminal provides the received audio data to the user through a playback device (speaker or headphones), specifically by decoding the audio data and sending it to the audio system.
[0628] In this way, the system can provide real-time, customized sports commentary that meets the diverse needs of users.
[0629] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0630] Step 1:
[0631] User: Enter sports knowledge level and support target
[0632] Using the terminal, users input their sports knowledge level (beginner, intermediate, advanced) and information about the team and player they support.
[0633] Input: User knowledge level, target of support
[0634] Output: User input data (knowledge level and support target)
[0635] Specific actions: The user operates text fields and selection menus on the device screen to enter the required information.
[0636] Step 2:
[0637] Terminal: Data transmission
[0638] The terminal sends the information entered by the user to the server as an API request.
[0639] Input: User input data (knowledge level and support target)
[0640] Output: API request
[0641] Specific operation: The terminal converts the user's input data into JSON format and sends it to the server as an HTTP POST request.
[0642] Step 3:
[0643] Server: Data storage
[0644] The server analyzes the received API requests and stores them in a database.
[0645] Input: API request
[0646] Output: User profile stored in the database
[0647] Specific operation: The server parses the received JSON data and saves the corresponding user profile in the database.
[0648] Step 4:
[0649] Terminal: Receiving game footage
[0650] The device receives game footage from a live streaming service.
[0651] Input: Live Streaming URL
[0652] Output: Received game video
[0653] Specific operation: The device retrieves live game footage in HLS or DASH format from the specified URL.
[0654] Step 5:
[0655] Terminal:Video Stream
[0656] The device streams the received game footage to the server in real time.
[0657] Input: Received game footage
[0658] Output: Streamed game footage
[0659] Specific operation: The device sends video data to the server using the RTMP or WebRTC protocol.
[0660] Step 6:
[0661] Server: Video analysis
[0662] The server analyzes the received video in real time and recognizes the players' positions and movements.
[0663] Input: Streamed game footage
[0664] Output: Analysis results (player positions, actions, scoring status)
[0665] Specific operation: The server performs analysis processing using a video analysis algorithm (e.g., YOLO or OpenCV).
[0666] Step 7:
[0667] Server: Customization information integration
[0668] The server integrates the user profile information with the video analysis results.
[0669] Input: User profile, analysis results
[0670] Output: Integrated data
[0671] Specific operation: The server links the analysis results with the user's knowledge level and support target information, and integrates them into an appropriate data structure.
[0672] Step 8:
[0673] Server: Live commentary generation
[0674] The server generates commentary using a generative AI model.
[0675] Input: Integrated data
[0676] Output: Commentary
[0677] Specific operation: The server inputs the integrated data into the generative AI model and generates a commentary appropriate for the user.
[0678] Example prompt: "Generate a play-by-play commentary of a soccer game given the user's beginner level of knowledge and which team they support."
[0679] Step 9:
[0680] Server: Speech synthesis
[0681] The server converts the generated commentary into voice data using a text-to-speech synthesis engine.
[0682] Input: Commentary
[0683] Output: Audio data
[0684] Specific operation: The server uses a TTS engine (e.g., Google TTS API) to convert the text into audio data.
[0685] Step 10:
[0686] Server: Sends audio data
[0687] The server transmits the generated voice data to the terminal.
[0688] Input: Audio data
[0689] Output: HTTP response
[0690] Specific operation: The server sends the audio data to the terminal as an HTTP response.
[0691] Step 11:
[0692] Device: Audio playback
[0693] The terminal provides the received audio data to the user on a playback device.
[0694] Input: Audio data
[0695] Output: Played audio
[0696] Specific operation: The device decodes the audio data and plays it through speakers or headphones.
[0697] Through the above steps, users can enjoy real-time commentary that is customized to their level of sports knowledge and the sport they are rooting for.
[0698] (Application example 1)
[0699] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0700] Conventional sports commentary systems have difficulty customizing commentary to suit individual users' knowledge levels or the teams they support, making it impossible to provide appropriate commentary in real time. Furthermore, the process of converting the generated commentary text into audio data and providing it to users' devices in real time was lacking. As a result, users only received the same commentary, which led to a problem of degrading the quality of each individual viewing experience.
[0701] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0702] In this invention, the server includes means for inputting a user's sports knowledge level and a team the user supports, means for analyzing game footage in real time to recognize player positions, movements, and scores, means for generating commentary according to the user's profile based on the analysis results, means for converting the commentary into voice data using a text-to-speech engine, means for providing the voice data to the user, means for receiving game footage via live streaming and transmitting it to the server, means for generating appropriate commentary from prompts using a generative AI model, and means for transmitting the generated voice data to the user's device in real time and playing it back. This makes it possible to provide appropriate and customized commentary in real time according to the user's sports knowledge level and the team the user supports.
[0703] The "user's sports knowledge level" indicates the depth of knowledge and understanding of the user about sports, and is usually classified into stages such as beginner, intermediate, and advanced.
[0704] The term "target of support" refers to a team or player that the user particularly supports in a particular sporting event or game.
[0705] "Game footage" refers to video data that records the progress of a sporting event or game in real time.
[0706] "Real-time analysis" means receiving game video data, processing it instantly, and extracting the necessary information and features.
[0707] "Player position" is information indicating where each player is located on the playing field at a specific time during the game.
[0708] "Action" refers to the specific actions or movements that a player makes during a game or competition, such as dribbling, passing, and shooting.
[0709] "Scoring Status" refers to information about the scores of both teams during the match, including the current score, the players who scored, and the method of scoring.
[0710] "Live commentary" is commentary based on the progress of the match and the actions of the players, and is intended to explain the content of the match to viewers in an easy-to-understand manner.
[0711] A "text-to-speech engine" is software or a process for converting input text data into speech data.
[0712] "Audio data" means a data file in audio format generated by a text-to-speech engine.
[0713] "Live streaming reception" is the process of instantly acquiring video and audio that is being streamed in real time over the Internet.
[0714] "Send to server" means sending data from a terminal or the like to a central server.
[0715] A "generative AI model" is an algorithm or framework for generating text or content using artificial intelligence.
[0716] A "prompt" is an instruction that is input into a generative AI model, and is text that causes the model to generate appropriate text or content based on that instruction.
[0717] "Terminal" means a device used by a user, including a smartphone, tablet, computer, etc.
[0718] "Playback" means aurally reproducing the generated audio data using an audio output device.
[0719] The system based on this invention utilizes generative AI models to provide users with customized, real-time commentary while watching sports.
[0720] 1. Setting up your user profile
[0721] User:
[0722] Users input their sports knowledge level (e.g., beginner, intermediate, advanced) and their favorite teams and athletes via their terminals. This information forms the basis for customization settings used throughout the system.
[0723] Device:
[0724] The information entered by the user is sent from the device to the server. Here, the user interface is built using React Native, and the device generates an API request to send the data.
[0725] server:
[0726] The server stores the received user information in a database (e.g., PostgreSQL) and creates a profile for each user, which includes the user's level of sports knowledge and the sport they support.
[0727] 2. Real-time analysis of game footage
[0728] Device:
[0729] The device uses WebRTC to receive game footage in real time from a live streaming service and streams the video data to a server.
[0730] server:
[0731] The server uses OpenCV to analyze the received video in real time, and uses object recognition algorithms to identify the player's position and recognize their actions (e.g., dribbling, passing, shooting), as well as to track the goal situation and ball position.
[0732] 3. Generation of commentary audio
[0733] server:
[0734] The server combines the user profile data with the video analysis results and inputs prompts into the generative AI model to generate appropriate commentary. For example, the following prompts are input into the generative AI model:
[0735] "Generate beginner-friendly commentary: Game status {Score}, player {Name} performed {Action}."
[0736] The generated commentary is converted into audio data using the Google Cloud Text-to-Speech API.
[0737] 4. Customized commentary
[0738] server:
[0739] The generated voice data is transmitted from the server to the terminal.
[0740] Device:
[0741] The device then provides the received audio data to the user via a playback device (e.g., smartphone, head-mounted display), allowing users to watch the game while listening to commentary tailored to their level of sports knowledge and the sport they are rooting for.
[0742] Specific examples
[0743] For example, consider a case where a user is at an "intermediate level" and is cheering for "Team A." If Player X is recognized as dribbling during a match and reaching the goal, the generative AI model will generate an appropriate commentary as follows:
[0744] Example prompt sentence:
[0745] "During a game, Player X dribbles the ball and reaches the goal. Please provide appropriate commentary for intermediate level users."
[0746] The generated commentary is converted into audio data using the Google Cloud Text-to-Speech API and sent to the user's device in real time, where they can enjoy this customized commentary via a head-mounted display or smartphone.
[0747] In this way, by implementing the invention, users can receive appropriate and customized commentary in real time according to their level of sports knowledge and the team they support.
[0748] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0749] Program processing steps and their specific explanations
[0750] Step 1: Configure your user profile
[0751] Users use a smartphone app to input their sports knowledge level (beginner, intermediate, advanced) and the team or player they support. The input data is then saved on the device.
[0752] Input: The user enters their sports knowledge level and the team or player they support into the input form.
[0753] Data processing: Convert input data into API request format
[0754] Output: Sent to the server as an API request
[0755] Step 2: Send and store user information
[0756] The terminal sends the information entered by the user to the server, where it is saved in a database that stores user profiles.
[0757] Input: User's sports knowledge level and favorite team and player information
[0758] Data processing: Converting transmitted data into a database format
[0759] Output: User profile saved in database
[0760] Step 3: Receiving and streaming game footage
[0761] The device receives game footage from the live streaming service using the WebRTC protocol and transmits it to the server.
[0762] Input: Real-time game footage from a live streaming service
[0763] Data processing: Converts video data into a format suitable for the server in real time
[0764] Output: Streaming data is sent to the server
[0765] Step 4: Video analysis
[0766] The server analyzes the received game footage in real time using OpenCV to recognize the players' positions, movements, and scoring status.
[0767] Input: Streamed game footage
[0768] Data processing: Apply object recognition algorithms to recognize player positions, movements, and scoring situations
[0769] Output: Analysis results: player positions, movements, and score information
[0770] Step 5: Generate commentary
[0771] Based on the analysis results and the user profile, the server inputs prompt text into the generative AI model to generate appropriate commentary.
[0772] Input: Analysis results and user profile
[0773] Data processing: Input prompt text into the generative AI model to generate commentary
[0774] Output: Generated commentary
[0775] Example of specific behavior: Send the following prompt to the generative AI model: "Generate commentary appropriate for beginners: Game situation {Score}, Player {Name} performed {Action}."
[0776] Step 6: Convert the commentary into audio
[0777] The server converts the generated commentary into audio data using the Google Cloud Text-to-Speech API.
[0778] Input: Generated commentary
[0779] Data processing: Convert text data into audio data
[0780] Output: Audio data
[0781] Step 7: Sending audio data
[0782] The server transmits the converted voice data to the user's terminal.
[0783] Input: Audio data
[0784] Data processing: Real-time transmission of voice data
[0785] Output: Audio data sent to the user's device
[0786] Step 8: Play commentary audio
[0787] The terminal plays the received audio data on a playback device (e.g., smartphone, head-mounted display) and provides it to the user.
[0788] Input: Audio data sent from the server
[0789] Data processing: Conversion to playback format
[0790] Output: Audio description played from the playback device
[0791] By having each processing step work in conjunction with each other in this way, users can receive live commentary in real time that is individually customized according to their level of sports knowledge and the sport they are rooting for.
[0792] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0793] This invention relates to a system that combines generative AI and an emotion engine to generate live sports commentary audio in real time. This allows for customizable live commentary based on the user's sports knowledge level, the sport they are rooting for, and even their emotional state. The detailed program processing of this system is explained in natural language.
[0794] User Profile Settings
[0795] User: Enter information
[0796] Users use their devices to set their sports knowledge level (e.g., beginner, intermediate, advanced) and the team or player they support. This information forms the basis for customization settings used throughout the system.
[0797] Terminal: Data transmission
[0798] The device receives the sports knowledge level and the information of the target sport entered by the user, and then generates an API request including this information and sends it to the server.
[0799] Server: Data storage
[0800] The server stores the received user information in a database and creates a profile for each user, which specifies the user's level of sports knowledge and the sport they support.
[0801] Real-time analysis of game footage
[0802] Terminal: Video reception
[0803] The device receives game footage from a live streaming service, which prepares it for processing in real time.
[0804] Device: Video streaming
[0805] The device streams the received game video to the server in real time, where it is analyzed.
[0806] Server: Video analysis
[0807] The server uses video analysis algorithms to analyze the streamed video in real time, and object recognition algorithms to recognize player positions and actions (e.g., dribbling, passing, shooting), as well as to determine goal situations and ball position.
[0808] Recognizing user emotions
[0809] Device: Emotion data acquisition
[0810] The device uses an emotion engine to analyze the user's facial expressions and voice tone to recognize the user's emotional state.
[0811] Server: Sends and stores emotion data
[0812] The device sends the recognized emotion data to the server, which then integrates it into the user profile.
[0813] Generating live commentary audio
[0814] Server: Customization information integration
[0815] The server integrates the user's sports knowledge level, favorites, and emotional data, which are used to generate commentary.
[0816] Server: Explanation generation
[0817] The server inputs the integrated data into a generative AI model to generate commentary tailored to the user, using different expressions and tones depending on the user's knowledge level and emotional state.
[0818] Server: Generates voice data
[0819] The server converts the generated commentary into audio data using a text-to-speech (TTS) engine, which is then formatted for listening by the user.
[0820] Server: Sending audio data
[0821] The server transmits the generated audio data to the terminal, which uses this data to provide audio commentary to the user.
[0822] Customized commentary
[0823] Device: Audio playback
[0824] The device then plays the received audio data through a playback device (e.g., speaker, headphones). Users can watch the game while listening to a customized commentary in real time. The commentary is tailored to the user's sports knowledge level, favorite sport, and emotional state, providing a more personalized experience.
[0825] Specific examples
[0826] For example, consider a user with a beginner level of soccer knowledge who is rooting for a specific team and is excited during a match. The user uses their device to set their own knowledge level and the team they are rooting for. When the game starts, the device receives the footage and streams it to the server. The server analyzes the footage and receives the user's emotional state from the device. The emotion engine detects the user's level of excitement, and the server generates a commentary based on this information. The commentary is written in a tone that matches the user's excitement, such as "Great goal, your team is leading!" The generated commentary is converted into audio data, sent to the device, and then provided to the user. This system allows users to receive commentary appropriate to their sports knowledge level, the team they are rooting for, and their real-time emotional state.
[0827] The processing flow will be explained below.
[0828] Step 1: User inputs their sports knowledge level and the sport they support
[0829] Using the terminal, the user sets his / her sports knowledge level (e.g., beginner, intermediate, advanced) and the team or player he / she supports.
[0830] This provides basic information that the system needs to provide the user with the most appropriate explanation.
[0831] Step 2: The device sends the data to the server
[0832] The terminal receives the sports knowledge level and the information of the target of support input by the user.
[0833] The device then generates an API request containing this information and sends it to the server.
[0834] Step 3: The server stores the user data
[0835] The server receives the user information sent from the terminal.
[0836] The received user information is stored in a database and a profile is created for each user.
[0837] Step 4: Receiving game footage
[0838] The device receives game footage from a live streaming service.
[0839] The received match footage is ready to be processed in real time.
[0840] Step 5: Stream the game
[0841] The device streams the received game footage to the server in real time.
[0842] The streamed video is analyzed by the server.
[0843] Step 6: The server analyzes the video in real time
[0844] The server uses a video analysis algorithm to analyze the streamed video in real time.
[0845] The server uses object recognition algorithms to recognize players' positions and actions (e.g., dribbling, passing, shooting), as well as to track the scoring situation and ball position.
[0846] Step 7: Recognize user emotions
[0847] The device uses an emotion engine to analyze the user's facial expressions and voice tone to recognize the user's emotional state.
[0848] The device transmits the emotional state as data to the server.
[0849] Step 8: The server stores the emotion data
[0850] The server receives the emotion data sent from the terminal.
[0851] The emotion data is integrated into the user profile and the profile information is updated.
[0852] Step 9: Generate explanatory text
[0853] The server integrates the user's sports knowledge level, favorite target, and emotion data.
[0854] Based on this, a generative AI model is used to generate live commentary that is appropriate for the user.
[0855] The content of the commentary is adjusted according to the user's knowledge level and emotional state.
[0856] Step 10: Creating audio data
[0857] The server converts the generated explanatory text into audio data using a text-to-speech (TTS) engine.
[0858] The converted audio data is ready to be provided to the user.
[0859] Step 11: Send audio data to the device
[0860] The server transmits the generated voice data to the terminal.
[0861] The terminal prepares to play the received audio data.
[0862] Step 12: Playing back audio data
[0863] The terminal plays the received audio data through a playback device (e.g., speaker, headphones).
[0864] Users can watch the game while listening to customized commentary in real time.
[0865] The commentary is customized to the user's level of sports knowledge, their roots, and their real-time emotional state.
[0866] Example 2
[0867] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0868] Current live sports commentary systems often provide uniform commentary without taking into account the individual user's knowledge level or emotional state. As a result, they are unable to address the different needs and expectations of each user, making it difficult to provide a personalized experience that matches each user's level of excitement and understanding. Furthermore, current systems often lack the accuracy of real-time video and emotion analysis, making it difficult to generate appropriate commentary based on user profiles. A system that solves these issues is needed.
[0869] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0870] In this invention, the server includes means for inputting a user's sports knowledge level and a target sport, means for analyzing game footage in real time to recognize the positions, movements, and score of players, means for analyzing the user's facial expressions and tone of voice to recognize the user's emotional state, means for generating commentary according to the user's profile based on the analysis results and the user's emotional state, means for converting the commentary into voice data using a text-to-speech engine, and means for providing the voice data to the user, thereby enabling the provision of customizable commentary in real time according to the user's individual knowledge level and emotional state.
[0871] The "user's sports knowledge level" indicates the user's knowledge and understanding of sports, and is expressed as a level such as beginner, intermediate, or advanced.
[0872] The term "target of support" refers to a specific team or player that the user supports during a game.
[0873] "Input means" includes devices or interfaces that allow a user to provide information to a system, such as a keyboard, touch screen, or voice input.
[0874] "Game footage" refers to video data showing a sports game being played, including live streaming and recorded footage.
[0875] "Real-time analysis" refers to the process of analyzing game video data as it is received and making the results immediately available.
[0876] "Player position" is information indicating the current location of each player within the competition area.
[0877] "Action" refers to specific actions that players perform during a game, such as dribbling, passing, shooting, etc.
[0878] "Scoring status" is information showing the distribution of scores in the current match and the latest scoring results.
[0879] "Means for analyzing" includes algorithms and software for processing the obtained data and extracting useful information, such as object recognition algorithms.
[0880] "Facial expressions" and "vocal tones" are physical and vocal characteristics that indicate the user's emotional state, and analyzing them allows the user's emotions to be inferred.
[0881] An "emotional state" refers to the psychological state that a user is feeling at a given moment, such as excitement, joy, sadness, etc.
[0882] "Means for recognizing emotional states" includes technologies and devices that analyze user characteristics such as facial expressions and vocal tone to estimate the emotions at that time.
[0883] A "user profile" is a data structure for centrally managing information related to a user, and includes information such as the user's level of sports knowledge, the sport they support, and their emotional state.
[0884] "Live commentary" is text that explains the progress and events of the match to users.
[0885] A "text-to-speech engine (TTS)" is a technology or software for converting text data into audio data.
[0886] "Audio data" refers to data in an audio format that is generated by a text-to-speech synthesis engine and can be heard by a user.
[0887] The "means for providing" includes devices and software for actually playing back the audio data generated by the system to the user, such as speakers and headphones.
[0888] This invention is a system that provides customized real-time commentary based on a user's sports knowledge level, emotional state, and favorite sport. The system consists of a terminal, a server, and related software and hardware.
[0889] User Profile Settings
[0890] First, a user uses their device to input their sports knowledge level (e.g., beginner, intermediate, advanced) and the team or player they support. The device receives this information, generates an API request, and sends it to the server. The server stores the received user information in a database and creates a profile for each user. The profile specifies the user's sports knowledge level and the team or player they support.
[0891] Real-time analysis of game footage
[0892] Game footage is received by the device from the live streaming service. The device then streams the received footage to the server in real time. The server then uses a video analysis algorithm to analyze the streamed footage. The server uses an object recognition algorithm to recognize player positions, actions (e.g., dribbling, passing, shooting), scoring situations, and ball position.
[0893] Recognizing user emotions
[0894] The device captures the user's facial expressions and voice tone using a camera and microphone, and analyzes them using an emotion engine to recognize the user's emotional state (e.g., excitement, joy, sadness, etc.). The recognized emotion data is sent from the device to a server and integrated into the user profile.
[0895] Generating live commentary audio
[0896] The server then combines the user's sports knowledge level, the sport they support, and their emotional data and inputs it into a generative AI model. Based on this combined data, the generative AI model generates a commentary tailored to the user. The commentary content is adjusted according to the user's knowledge level and emotional state. For example, an excited user will receive detailed commentary in an emotional tone.
[0897] Generating and providing voice data
[0898] The server converts the generated commentary into audio data using a text-to-speech (TTS) engine. This audio data is then sent to the device, which then plays the received audio data through a playback device (e.g., speaker, headphones). Users can watch the game while listening to the customized commentary in real time.
[0899] Specific examples
[0900] For example, consider a case where a user has a "beginner level of soccer knowledge" and is "supporting a specific team," becoming excited during a game. The user uses their device to set their own knowledge level and the team they are supporting. When the game video begins, the device receives the video and streams it to the server. The server analyzes the video and also receives the user's emotional state from the device. The emotion engine detects the user's level of excitement, and the server generates a commentary based on this. The commentary is in a tone that matches the user's excitement, such as "Great goal, your team is leading!" The generated commentary is converted into audio data, sent to the device, and then provided to the user. This system allows users to receive commentary appropriate to their sports knowledge level, the team they are supporting, and their real-time emotional state.
[0901] Prompt Sentence Examples
[0902] "Generate a commentary of a soccer game in which a beginner-level user who is rooting for a specific team gets excited when a goal is scored."
[0903] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0904] Step 1:
[0905] The user uses the device's input interface to set their sports knowledge level (e.g., beginner, intermediate, advanced) and the team or player they support. This becomes the input data. The device receives this information and generates an API request. The information included in this request (the user's sports knowledge level and the team they support) is sent to the server.
[0906] Step 2:
[0907] The server processes the received API request and stores the user information in a database. This creates a profile for each user, including the user's sports knowledge level and favorite sports. The output data is added to the database as a user profile.
[0908] Step 3:
[0909] When the match begins, the device receives the match video from the live streaming service. This video data becomes the input data, and is ready to be processed in real time.
[0910] Step 4:
[0911] The device receives the game video in real time and streams it to the server. The streamed video becomes the input data. The server prepares this video data for analysis.
[0912] Step 5:
[0913] The server analyzes the streamed video in real time using an object recognition algorithm. Using the video data as input, it recognizes the player's position, actions (e.g., dribbling, passing, shooting), scoring status, and ball position. The analysis results are output data.
[0914] Step 6:
[0915] The device captures the user's facial expressions and voice tone using a camera and microphone, which serve as input data. The emotion engine analyzes this data to recognize the user's emotional state, and the recognized emotional state is generated as output data.
[0916] Step 7:
[0917] The device sends the recognized emotion data to the server, which then integrates it into the user profile and keeps it up to date. This updates the user profile.
[0918] Step 8:
[0919] The server inputs the integrated user profile into the generative AI model. The input data includes the user's sports knowledge level, the people they support, and emotional data. Based on this data, the generative AI model generates a commentary appropriate for the user. This becomes the output data.
[0920] Step 9:
[0921] The server inputs the generated commentary into a text-to-speech (TTS) engine. The commentary text is used as input data. The TTS engine converts it into voice data. The voice data is generated as output data.
[0922] Step 10:
[0923] The server sends the generated audio data to the device, which becomes the input data. The device receives this data and plays it through a playback device (e.g., speaker, headphones). This allows the user to listen to a customized commentary in real time.
[0924] (Application example 2)
[0925] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0926] Conventional sports commentary systems have difficulty providing personalized commentary in real time based on the user's sports knowledge level or the team they support. They also lack the ability to provide customized commentary based on the user's emotional state. Therefore, multifaceted customization is needed to provide users with a more engaging and immersive viewing experience.
[0927] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0928] In this invention, the server includes: means for inputting a user's sports knowledge level and a target player; means for analyzing game footage in real time to recognize player positions, movements, and scores; means for analyzing the user's facial expressions and vocal tone to recognize the user's emotional state; means for generating prompt sentences based on the analysis results and the user's emotional state; means for inputting the prompt sentences into a generative AI model to generate commentary according to the user's profile and emotional state; means for converting the commentary into voice data using a text-to-speech engine; and means for providing the voice data to the user. This makes it possible to provide highly personalized commentary in real time based on the user's sports knowledge level, target player, and real-time emotional state.
[0929] "User's sports knowledge level" is information indicating the depth of knowledge and understanding of the user regarding sports.
[0930] "Support target" is information about a team or player that the user particularly supports.
[0931] "Game video" refers to real-time video data of a sports game.
[0932] "Player position" refers to information about the current position of each player within the competition venue.
[0933] "Action" refers to a series of movements that a player makes during a game (e.g., dribbling, passing, shooting, etc.).
[0934] "Scoring status" refers to the scoring status and score information during a match.
[0935] "User's facial expression" is information for reading emotions from the user's facial features and movements.
[0936] "Voice tone" is information that indicates the tone or pitch of the voice uttered by the user.
[0937] "User's emotional state" refers to the type of emotion the user is currently feeling (e.g., excitement, joy, sadness, etc.).
[0938] A "prompt sentence" is a sentence that contains prior information for generating an explanatory sentence and is input into the generative AI model.
[0939] A "generative AI model" is an artificial intelligence system that generates text for a specific task (in this case, commentary) based on input data.
[0940] A "text-to-speech synthesis engine" is a system that converts text into speech data.
[0941] "Audio data" is audio information that has been converted into a format that can be heard by a user.
[0942] This invention relates to a system for generating and providing play-by-play commentary audio that is customized based on the user's sports knowledge level and the sport they support, and also in response to the user's real-time emotional state. To achieve this, the following specific means and processes are required.
[0943] User Profile Settings
[0944] Users can enter their sports knowledge level and favorite teams and athletes via their device. This information is sent from the device to the server, which then receives it and stores it in a database, creating a customized profile for each user.
[0945] Real-time analysis of game footage
[0946] The device receives game footage from a live streaming service and streams it in real time to a server, which then analyzes the footage using object recognition algorithms to identify player positions, actions (e.g., dribbling, passing, shooting), and scoring situations.
[0947] Recognizing user emotions
[0948] The device uses the user's facial expressions and tone of voice to recognize the user's emotional state through an emotion engine, which then sends the emotion data to the server and stores it as a user profile.
[0949] Generating live commentary audio
[0950] The server combines the user's sports knowledge level, the team they are rooting for, and their emotional state to generate a prompt. The generated prompt is then input into a generative AI model to generate a play-by-play commentary appropriate for the user. For example, a prompt might be in the format: "The user's knowledge level is beginner, and the team they are rooting for is Team A. Their current emotional state is excited. The analysis results of the game are as follows: Team A has scored a goal!"
[0951] The generated commentary is converted into audio data using a text-to-speech (TTS) engine, which is then formatted for the user to hear.
[0952] Customized commentary
[0953] The terminal plays the audio data received from the server through a playback device (e.g., speaker, headphones), allowing the user to watch the game while listening to customized commentary in real time.
[0954] Specific examples
[0955] For example, if a user has a beginner level of soccer knowledge and supports a specific team (Team A), and the server detects that the user is excited, the server will generate the following prompt:
[0956] The user's knowledge level is beginner, and the team they support is Team A. Their current emotional state is excited. The analysis result of the match is as follows: Team A has scored a goal!
[0957] Based on this prompt, the generative AI model generates a commentary such as "Great goal, Team A is in the lead!", which is then converted into audio using a text-to-speech engine and sent to the device. The device then plays this audio through a playback device and provides it to the user.
[0958] In this way, the present invention is able to provide a customized commentary that is tailored to the user's profile and real-time emotional state.
[0959] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0960] Step 1:
[0961] A user inputs their level of sports knowledge and the teams and players they support via a terminal. The input information is data that indicates the user's sports viewing preferences. The terminal receives this data and sends it to a server. The server stores the received data in a database and generates a user profile. This profile includes information about the user's level of sports knowledge and the teams and players they support.
[0962] Step 2:
[0963] The device receives game footage from a live streaming service. The received video data is a real-time video stream showing the specific situation of the game. The device streams this video data to a server, which analyzes it in real time using video analysis algorithms. The analysis results include player positions, actions (e.g., dribbling, passing, shooting), and scoring status.
[0964] Step 3:
[0965] The device analyzes the user's facial expressions and vocal tone using an emotion engine. The analyzed emotion data indicates the user's real-time emotional state (e.g., excitement, joy, sadness, etc.). The device transmits this emotion data to the server, which then integrates the received emotion data into the user profile.
[0966] Step 4:
[0967] The server generates a prompt sentence based on the user's sports knowledge level, the team they are rooting for, and their emotional state. This prompt sentence contains prior information to be input into the generative AI model. For example, a prompt sentence of the form "The user's knowledge level is beginner level, and the team they are rooting for is Team A. Their current emotional state is excited. The analysis results of the game are as follows: Team A has scored a goal!" is generated.
[0968] Step 5:
[0969] The server inputs the generated prompt into a generative AI model to generate a commentary appropriate for the user. The generative AI model is an artificial intelligence system that generates commentary based on the prompt. The content of the generated commentary is adjusted according to the user's knowledge level and emotional state. As a result, a commentary appropriate for the user is obtained.
[0970] Step 6:
[0971] The server converts the generated commentary into audio data using a text-to-speech (TTS) engine, which then formats the audio data in a user-friendly format. Users can listen to this audio data in real time while watching the game.
[0972] Step 7:
[0973] The terminal plays the audio data received from the server through a playback device (e.g., speaker, headphones). This allows the user to listen to commentary customized according to the situation of the game. For example, if the user is excited, commentary such as "Great goal, Team A is in the lead!" is provided in an excited tone.
[0974] In this way, users can experience a personalized play-by-play commentary based on their level of sports knowledge, their favorite team, and their real-time emotional state.
[0975] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0976] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0977] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0978] [Third embodiment]
[0979] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0980] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[0981] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0982] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0983] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0984] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0985] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0986] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0987] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0988] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0989] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0990] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0991] This invention relates to a system that uses generative AI to generate live sports commentary audio in real time. This allows for customizable live commentary based on the user's level of sports knowledge and the sport they are rooting for. The detailed program processing of this system is explained in natural language.
[0992] User Profile Settings
[0993] User: Enter information
[0994] Users use their devices to set their sports knowledge level (e.g., beginner, intermediate, advanced) and the team or player they support. This information forms the basis for customization settings used throughout the system.
[0995] Terminal: Data transmission
[0996] The device generates and sends an API request to the server to transmit the information entered by the user, including the sports knowledge level and support target selected by the user.
[0997] Server: Data storage
[0998] The server stores the received user information in a database and creates a profile for each user, which specifies the user's level of sports knowledge and the sport they support.
[0999] Real-time analysis of game footage
[1000] Terminal: Video reception
[1001] The device receives game footage from a live streaming service, which is then streamed to a server.
[1002] Terminal:Video Stream
[1003] The device then streams the received game footage to a server, where the streaming video becomes the input data for real-time analysis.
[1004] Server: Video analysis
[1005] The server analyzes the received video using a video analysis algorithm, specifically an object recognition algorithm that tracks the player's position in real time, recognizes each player's actions (e.g., dribbling, passing, shooting), and determines the goal situation and ball position.
[1006] Generating live commentary audio
[1007] Server: Customization information integration
[1008] The server combines user profile data with video analysis results, and this information is used to generate commentary.
[1009] Server: Explanation generation
[1010] The server then inputs the integrated data into a generative AI model to generate commentary tailored to the user. This commentary varies in detail depending on the user's level of sports knowledge. For example, commentary for beginners starts with basic explanations, while commentary for advanced players focuses on technical details.
[1011] Server: Speech synthesis
[1012] The server converts the generated commentary into audio data using a text-to-speech (TTS) engine, which is ready to be served to the user.
[1013] Server: Send data
[1014] The server transmits the generated audio data to the terminal, which uses the data to provide audio commentary to the user.
[1015] Customized commentary
[1016] Device: Audio playback
[1017] The device then provides the received audio data to the user through a playback device (e.g., speaker or headphones).The user watches the game while listening to customized commentary based on the user's sports knowledge level and the sport they are rooting for.
[1018] Specific examples
[1019] For example, consider a user who has a "beginner level of soccer knowledge" and "supports a specific team." In this case, the user uses the device to set their own knowledge level and the team they support. When the game starts, the device receives live video and streams it to the server. The server analyzes the video and generates appropriate commentary based on the user's profile. The commentary is converted into audio data using a text-to-speech engine and sent to the device. Finally, the device plays the generated audio data and provides it to the user.
[1020] This system allows users to receive real-time commentary tailored to their level of sports knowledge and the sport they are rooting for, improving the sports viewing experience.
[1021] The processing flow will be explained below.
[1022] Step 1: User inputs their sports knowledge level and the sport they support
[1023] Using the terminal, the user sets his / her sports knowledge level (e.g., beginner, intermediate, advanced) and the team or player he / she supports.
[1024] This information is entered through the terminal's user interface.
[1025] Step 2: The device sends the data to the server
[1026] The terminal receives the sports knowledge level and the information of the target of support input by the user.
[1027] The device generates an API request containing this information and sends it to the server.
[1028] Step 3: The server stores the user data
[1029] The server receives the user information sent from the terminal.
[1030] The server stores the received information in a database and creates a profile for each user.
[1031] Step 4: Receiving game footage
[1032] The device receives game footage from a live streaming service.
[1033] The received match footage is ready to be processed in real time.
[1034] Step 5: Stream the game
[1035] The device streams the received game footage to the server in real time.
[1036] The streamed video is analyzed by the server.
[1037] Step 6: The server analyzes the video in real time
[1038] The server uses a video analysis algorithm to analyze the streamed video in real time.
[1039] The server uses object recognition algorithms to recognize players' positions and actions (e.g., dribbling, passing, shooting).
[1040] The server also keeps track of the score and the position of the ball.
[1041] Step 7: Integrating user information and analysis results
[1042] The server integrates the user's profile data with the video analysis results.
[1043] Compile materials for commentary based on the user's level of sports knowledge and the sport they support.
[1044] Step 8: Generate explanatory text
[1045] The server inputs the integrated data into the generative AI model and generates a commentary appropriate for the user.
[1046] The explanations are tailored to the user's level of knowledge, with the level of detail varying between beginners and advanced users.
[1047] Step 9: Generate audio data
[1048] The server converts the generated explanatory text into audio data using a text-to-speech (TTS) engine.
[1049] This audio data is then formatted so that it can be heard by the user.
[1050] Step 10: Send audio data to the device
[1051] The server transmits the generated voice data to the terminal.
[1052] The terminal prepares to receive the audio data.
[1053] Step 11: Playing back audio data
[1054] The terminal plays the received audio data through a playback device (e.g., speaker, headphones).
[1055] Users can watch the game while listening to customized commentary in real time.
[1056] Example 1
[1057] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1058] Conventional sports commentary systems have difficulty meeting the diverse needs of users, and have been unable to provide customized commentary based on individual sports knowledge levels or the sports fans they support. Furthermore, they lack technology that can accurately recognize the progress of a game in real time and instantly generate appropriate commentary. The present invention aims to solve these problems and provide users with real-time, customized sports commentary.
[1059] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1060] In this invention, the server includes means for inputting a user's sports knowledge level and a target player, means for transmitting the input information to the server, means for storing the input information in a database, means for receiving game footage and streaming the footage to the server, means for analyzing the game footage in real time and recognizing player positions, movements, and scores, means for generating commentary text tailored to the user's profile based on the analysis results and the user's sports knowledge level and target player, means for converting the generated commentary text into audio data using a text-to-speech engine, and means for transmitting the audio data from the server to a terminal and providing it to the user via a playback device. This enables real-time, customized sports commentary tailored to the diverse needs of users.
[1061] A "user" is an individual who watches sports and enters information into the system to receive customized commentary.
[1062] The "sports knowledge level" is an index that indicates the depth of a user's knowledge about sports, and is classified into categories such as beginner, intermediate, and advanced.
[1063] The term "target of support" refers to a team or player in which the user has a particular interest and supports.
[1064] A "terminal" is an electronic device that a user uses to provide input information, receive game footage, and play audio data.
[1065] The "server" is a centralized management system that stores information sent by users, analyzes game footage, generates commentary, and distributes audio data.
[1066] A "database" is a system located within a server that stores and manages user information and analysis results.
[1067] "Game footage" refers to video data of sports games distributed via live streaming services.
[1068] "Streaming" refers to the act of transmitting or receiving game footage in real time.
[1069] "Video analysis" refers to the process of processing game footage in real time and recognizing players' positions, movements, scoring situations, etc.
[1070] A "user profile" is a data set that includes individually set information such as the user's level of sports knowledge and the sport they support.
[1071] "Live commentary" is a narration text that is generated based on the progress of the match.
[1072] A "text-to-speech engine" is a software technology for converting text data into speech data.
[1073] "Audio data" refers to commentary data in audio format that is generated by a text-to-speech synthesis engine and provided to the user.
[1074] A "playback device" refers to a device such as a speaker or headphones that is connected to a terminal and allows the user to listen to audio data.
[1075] This invention relates to a system that uses generative AI to generate live sports commentary audio in real time. This system can provide customized commentary according to the user's sports knowledge level and the sport they are rooting for. Each processing step of the system is described in detail below.
[1076] User Profile Settings
[1077] User: Enter information
[1078] Users use the terminal to input their sports knowledge level (beginner, intermediate, advanced) and information about the team or player they support. The input information is used as the basis for customization settings within the system.
[1079] Terminal: Data transmission
[1080] The device sends the information entered by the user to the server as an API request. Specifically, the device packages the input information in JSON format and sends an HTTP request to the server.
[1081] Server: Data storage
[1082] The server stores the data of the received API request in a database. For example, the database stores information such as the user ID, knowledge level, and support team.
[1083] Real-time analysis of game footage
[1084] Terminal: Video reception
[1085] The device receives game footage from a live streaming service, for example, by obtaining live footage in HLS format from a streaming URL.
[1086] Terminal:Video Stream
[1087] The device then streams the received game footage to the server in real time, using the RTMP or WebRTC protocol to send the video data to the server.
[1088] Server: Video analysis
[1089] The server analyzes the streamed game footage using a video analysis algorithm (e.g., YOLO or OpenCV) to recognize the positions and movements of players. As a result of the analysis, the score and ball position can be determined.
[1090] Generating live commentary audio
[1091] Server: Customization information integration
[1092] The server integrates the user profile information with the video analysis results. For example, the analysis results might link information such as "Player A is dribbling" with information such as "the user is a beginner."
[1093] Server: Explanation generation
[1094] The server inputs the integrated data using a generative AI model and generates commentary appropriate for the user. As a specific example, for beginners, a commentary such as "Player A broke through the opponent with a great dribble" is generated. An example of a prompt is as follows:
[1095] "Generate a play-by-play commentary of a soccer game, given that the user has a beginner level of knowledge and identifies the team they support."
[1096] Server: Speech synthesis
[1097] The server converts the generated commentary into voice data using a text-to-speech engine (TTS engine). Specifically, it uses the API of the TTS engine to convert the text "Player A made a great dribble..." into voice data.
[1098] Server: Sends audio data
[1099] The server sends the generated audio data to the user's device by returning the audio data as an HTTP response.
[1100] Customized commentary
[1101] Device: Audio playback
[1102] The terminal provides the received audio data to the user through a playback device (speaker or headphones), specifically by decoding the audio data and sending it to the audio system.
[1103] In this way, the system can provide real-time, customized sports commentary that meets the diverse needs of users.
[1104] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1105] Step 1:
[1106] User: Enter sports knowledge level and support target
[1107] Using the terminal, users input their sports knowledge level (beginner, intermediate, advanced) and information about the team and player they support.
[1108] Input: User knowledge level, target of support
[1109] Output: User input data (knowledge level and support target)
[1110] Specific actions: The user operates text fields and selection menus on the device screen to enter the required information.
[1111] Step 2:
[1112] Terminal: Data transmission
[1113] The terminal sends the information entered by the user to the server as an API request.
[1114] Input: User input data (knowledge level and support target)
[1115] Output: API request
[1116] Specific operation: The terminal converts the user's input data into JSON format and sends it to the server as an HTTP POST request.
[1117] Step 3:
[1118] Server: Data storage
[1119] The server analyzes the received API requests and stores them in a database.
[1120] Input: API request
[1121] Output: User profile stored in the database
[1122] Specific operation: The server parses the received JSON data and saves the corresponding user profile in the database.
[1123] Step 4:
[1124] Terminal: Receiving game footage
[1125] The device receives game footage from a live streaming service.
[1126] Input: Live Streaming URL
[1127] Output: Received game video
[1128] Specific operation: The device retrieves live game footage in HLS or DASH format from the specified URL.
[1129] Step 5:
[1130] Terminal:Video Stream
[1131] The device streams the received game footage to the server in real time.
[1132] Input: Received game footage
[1133] Output: Streamed game footage
[1134] Specific operation: The device sends video data to the server using the RTMP or WebRTC protocol.
[1135] Step 6:
[1136] Server: Video analysis
[1137] The server analyzes the received video in real time and recognizes the players' positions and movements.
[1138] Input: Streamed game footage
[1139] Output: Analysis results (player positions, actions, scoring status)
[1140] Specific operation: The server performs analysis processing using a video analysis algorithm (e.g., YOLO or OpenCV).
[1141] Step 7:
[1142] Server: Customization information integration
[1143] The server integrates the user profile information with the video analysis results.
[1144] Input: User profile, analysis results
[1145] Output: Integrated data
[1146] Specific operation: The server links the analysis results with the user's knowledge level and support target information, and integrates them into an appropriate data structure.
[1147] Step 8:
[1148] Server: Live commentary generation
[1149] The server generates commentary using a generative AI model.
[1150] Input: Integrated data
[1151] Output: Commentary
[1152] Specific operation: The server inputs the integrated data into the generative AI model and generates a commentary appropriate for the user.
[1153] Example prompt: "Generate a play-by-play commentary of a soccer game given the user's beginner level of knowledge and which team they support."
[1154] Step 9:
[1155] Server: Speech synthesis
[1156] The server converts the generated commentary into voice data using a text-to-speech synthesis engine.
[1157] Input: Commentary
[1158] Output: Audio data
[1159] Specific operation: The server uses a TTS engine (e.g., Google TTS API) to convert the text into audio data.
[1160] Step 10:
[1161] Server: Sends audio data
[1162] The server transmits the generated voice data to the terminal.
[1163] Input: Audio data
[1164] Output: HTTP response
[1165] Specific operation: The server sends the audio data to the terminal as an HTTP response.
[1166] Step 11:
[1167] Device: Audio playback
[1168] The terminal provides the received audio data to the user on a playback device.
[1169] Input: Audio data
[1170] Output: Played audio
[1171] Specific operation: The device decodes the audio data and plays it through speakers or headphones.
[1172] Through the above steps, users can enjoy real-time commentary that is customized to their level of sports knowledge and the sport they are rooting for.
[1173] (Application example 1)
[1174] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1175] Conventional sports commentary systems have difficulty customizing commentary to suit individual users' knowledge levels or the teams they support, making it impossible to provide appropriate commentary in real time. Furthermore, the process of converting the generated commentary text into audio data and providing it to users' devices in real time was lacking. As a result, users only received the same commentary, which led to a problem of degrading the quality of each individual viewing experience.
[1176] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1177] In this invention, the server includes means for inputting a user's sports knowledge level and a team the user supports, means for analyzing game footage in real time to recognize player positions, movements, and scores, means for generating commentary according to the user's profile based on the analysis results, means for converting the commentary into voice data using a text-to-speech engine, means for providing the voice data to the user, means for receiving game footage via live streaming and transmitting it to the server, means for generating appropriate commentary from prompts using a generative AI model, and means for transmitting the generated voice data to the user's device in real time and playing it back. This makes it possible to provide appropriate and customized commentary in real time according to the user's sports knowledge level and the team the user supports.
[1178] The "user's sports knowledge level" indicates the depth of knowledge and understanding of the user about sports, and is usually classified into stages such as beginner, intermediate, and advanced.
[1179] The term "target of support" refers to a team or player that the user particularly supports in a particular sporting event or game.
[1180] "Game footage" refers to video data that records the progress of a sporting event or game in real time.
[1181] "Real-time analysis" means receiving game video data, processing it instantly, and extracting the necessary information and features.
[1182] "Player position" is information indicating where each player is located on the playing field at a specific time during the game.
[1183] "Action" refers to the specific actions or movements that a player makes during a game or competition, such as dribbling, passing, and shooting.
[1184] "Scoring Status" refers to information about the scores of both teams during the match, including the current score, the players who scored, and the method of scoring.
[1185] "Live commentary" is commentary based on the progress of the match and the actions of the players, and is intended to explain the content of the match to viewers in an easy-to-understand manner.
[1186] A "text-to-speech engine" is software or a process for converting input text data into speech data.
[1187] "Audio data" means a data file in audio format generated by a text-to-speech engine.
[1188] "Live streaming reception" is the process of instantly acquiring video and audio that is being streamed in real time over the Internet.
[1189] "Send to server" means sending data from a terminal or the like to a central server.
[1190] A "generative AI model" is an algorithm or framework for generating text or content using artificial intelligence.
[1191] A "prompt" is an instruction that is input into a generative AI model, and is text that causes the model to generate appropriate text or content based on that instruction.
[1192] "Terminal" means a device used by a user, including a smartphone, tablet, computer, etc.
[1193] "Playback" means aurally reproducing the generated audio data using an audio output device.
[1194] The system based on this invention utilizes generative AI models to provide users with customized, real-time commentary while watching sports.
[1195] 1. Setting up your user profile
[1196] User:
[1197] Users input their sports knowledge level (e.g., beginner, intermediate, advanced) and their favorite teams and athletes via their terminals. This information forms the basis for customization settings used throughout the system.
[1198] Device:
[1199] The information entered by the user is sent from the device to the server. Here, the user interface is built using React Native, and the device generates an API request to send the data.
[1200] server:
[1201] The server stores the received user information in a database (e.g., PostgreSQL) and creates a profile for each user, which includes the user's level of sports knowledge and the sport they support.
[1202] 2. Real-time analysis of game footage
[1203] Device:
[1204] The device uses WebRTC to receive game footage in real time from a live streaming service and streams the video data to a server.
[1205] server:
[1206] The server uses OpenCV to analyze the received video in real time, and uses object recognition algorithms to identify the player's position and recognize their actions (e.g., dribbling, passing, shooting), as well as to track the goal situation and ball position.
[1207] 3. Generation of commentary audio
[1208] server:
[1209] The server combines the user profile data with the video analysis results and inputs prompts into the generative AI model to generate appropriate commentary. For example, the following prompts are input into the generative AI model:
[1210] "Generate beginner-friendly commentary: Game status {Score}, player {Name} performed {Action}."
[1211] The generated commentary is converted into audio data using the Google Cloud Text-to-Speech API.
[1212] 4. Customized commentary
[1213] server:
[1214] The generated voice data is transmitted from the server to the terminal.
[1215] Device:
[1216] The device then provides the received audio data to the user via a playback device (e.g., smartphone, head-mounted display), allowing users to watch the game while listening to commentary tailored to their level of sports knowledge and the sport they are rooting for.
[1217] Specific examples
[1218] For example, consider a case where a user is at an "intermediate level" and is cheering for "Team A." If Player X is recognized as dribbling during a match and reaching the goal, the generative AI model will generate an appropriate commentary as follows:
[1219] Example prompt sentence:
[1220] "During a game, Player X dribbles the ball and reaches the goal. Please provide appropriate commentary for intermediate level users."
[1221] The generated commentary is converted into audio data using the Google Cloud Text-to-Speech API and sent to the user's device in real time, where they can enjoy this customized commentary via a head-mounted display or smartphone.
[1222] In this way, by implementing the invention, users can receive appropriate and customized commentary in real time according to their level of sports knowledge and the team they support.
[1223] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1224] Program processing steps and their specific explanations
[1225] Step 1: Configure your user profile
[1226] Users use a smartphone app to input their sports knowledge level (beginner, intermediate, advanced) and the team or player they support. The input data is then saved on the device.
[1227] Input: The user enters their sports knowledge level and the team or player they support into the input form.
[1228] Data processing: Convert input data into API request format
[1229] Output: Sent to the server as an API request
[1230] Step 2: Send and store user information
[1231] The terminal sends the information entered by the user to the server, where it is saved in a database that stores user profiles.
[1232] Input: User's sports knowledge level and favorite team and player information
[1233] Data processing: Converting transmitted data into a database format
[1234] Output: User profile saved in database
[1235] Step 3: Receiving and streaming game footage
[1236] The device receives game footage from the live streaming service using the WebRTC protocol and transmits it to the server.
[1237] Input: Real-time game footage from a live streaming service
[1238] Data processing: Converts video data into a format suitable for the server in real time
[1239] Output: Streaming data is sent to the server
[1240] Step 4: Video analysis
[1241] The server analyzes the received game footage in real time using OpenCV to recognize the players' positions, movements, and scoring status.
[1242] Input: Streamed game footage
[1243] Data processing: Apply object recognition algorithms to recognize player positions, movements, and scoring situations
[1244] Output: Analysis results: player positions, movements, and score information
[1245] Step 5: Generate commentary
[1246] Based on the analysis results and the user profile, the server inputs prompt text into the generative AI model to generate appropriate commentary.
[1247] Input: Analysis results and user profile
[1248] Data processing: Input prompt text into the generative AI model to generate commentary
[1249] Output: Generated commentary
[1250] Example of specific behavior: Send the following prompt to the generative AI model: "Generate commentary appropriate for beginners: Game situation {Score}, Player {Name} performed {Action}."
[1251] Step 6: Convert the commentary into audio
[1252] The server converts the generated commentary into audio data using the Google Cloud Text-to-Speech API.
[1253] Input: Generated commentary
[1254] Data processing: Convert text data into audio data
[1255] Output: Audio data
[1256] Step 7: Sending audio data
[1257] The server transmits the converted voice data to the user's terminal.
[1258] Input: Audio data
[1259] Data processing: Real-time transmission of voice data
[1260] Output: Audio data sent to the user's device
[1261] Step 8: Play commentary audio
[1262] The terminal plays the received audio data on a playback device (e.g., smartphone, head-mounted display) and provides it to the user.
[1263] Input: Audio data sent from the server
[1264] Data processing: Conversion to playback format
[1265] Output: Audio description played from the playback device
[1266] By having each processing step work in conjunction with each other in this way, users can receive live commentary in real time that is individually customized according to their level of sports knowledge and the sport they are rooting for.
[1267] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1268] This invention relates to a system that combines generative AI and an emotion engine to generate live sports commentary audio in real time. This allows for customizable live commentary based on the user's sports knowledge level, the sport they are rooting for, and even their emotional state. The detailed program processing of this system is explained in natural language.
[1269] User Profile Settings
[1270] User: Enter information
[1271] Users use their devices to set their sports knowledge level (e.g., beginner, intermediate, advanced) and the team or player they support. This information forms the basis for customization settings used throughout the system.
[1272] Terminal: Data transmission
[1273] The device receives the sports knowledge level and the information of the target sport entered by the user, and then generates an API request including this information and sends it to the server.
[1274] Server: Data storage
[1275] The server stores the received user information in a database and creates a profile for each user, which specifies the user's level of sports knowledge and the sport they support.
[1276] Real-time analysis of game footage
[1277] Terminal: Video reception
[1278] The device receives game footage from a live streaming service, which prepares it for processing in real time.
[1279] Device: Video streaming
[1280] The device streams the received game video to the server in real time, where it is analyzed.
[1281] Server: Video analysis
[1282] The server uses video analysis algorithms to analyze the streamed video in real time, and object recognition algorithms to recognize player positions and actions (e.g., dribbling, passing, shooting), as well as to determine goal situations and ball position.
[1283] Recognizing user emotions
[1284] Device: Emotion data acquisition
[1285] The device uses an emotion engine to analyze the user's facial expressions and voice tone to recognize the user's emotional state.
[1286] Server: Sends and stores emotion data
[1287] The device sends the recognized emotion data to the server, which then integrates it into the user profile.
[1288] Generating live commentary audio
[1289] Server: Customization information integration
[1290] The server integrates the user's sports knowledge level, favorites, and emotional data, which are used to generate commentary.
[1291] Server: Explanation generation
[1292] The server inputs the integrated data into a generative AI model to generate commentary tailored to the user, using different expressions and tones depending on the user's knowledge level and emotional state.
[1293] Server: Generates voice data
[1294] The server converts the generated commentary into audio data using a text-to-speech (TTS) engine, which is then formatted for listening by the user.
[1295] Server: Sending audio data
[1296] The server transmits the generated audio data to the terminal, which uses this data to provide audio commentary to the user.
[1297] Customized commentary
[1298] Device: Audio playback
[1299] The device then plays the received audio data through a playback device (e.g., speaker, headphones). Users can watch the game while listening to a customized commentary in real time. The commentary is tailored to the user's sports knowledge level, favorite sport, and emotional state, providing a more personalized experience.
[1300] Specific examples
[1301] For example, consider a user with a beginner level of soccer knowledge who is rooting for a specific team and is excited during a match. The user uses their device to set their own knowledge level and the team they are rooting for. When the game starts, the device receives the footage and streams it to the server. The server analyzes the footage and receives the user's emotional state from the device. The emotion engine detects the user's level of excitement, and the server generates a commentary based on this information. The commentary is written in a tone that matches the user's excitement, such as "Great goal, your team is leading!" The generated commentary is converted into audio data, sent to the device, and then provided to the user. This system allows users to receive commentary appropriate to their sports knowledge level, the team they are rooting for, and their real-time emotional state.
[1302] The processing flow will be explained below.
[1303] Step 1: User inputs their sports knowledge level and the sport they support
[1304] Using the terminal, the user sets his / her sports knowledge level (e.g., beginner, intermediate, advanced) and the team or player he / she supports.
[1305] This provides basic information that the system needs to provide the user with the most appropriate explanation.
[1306] Step 2: The device sends the data to the server
[1307] The terminal receives the sports knowledge level and the information of the target of support input by the user.
[1308] The device then generates an API request containing this information and sends it to the server.
[1309] Step 3: The server stores the user data
[1310] The server receives the user information sent from the terminal.
[1311] The received user information is stored in a database and a profile is created for each user.
[1312] Step 4: Receiving game footage
[1313] The device receives game footage from a live streaming service.
[1314] The received match footage is ready to be processed in real time.
[1315] Step 5: Stream the game
[1316] The device streams the received game footage to the server in real time.
[1317] The streamed video is analyzed by the server.
[1318] Step 6: The server analyzes the video in real time
[1319] The server uses a video analysis algorithm to analyze the streamed video in real time.
[1320] The server uses object recognition algorithms to recognize players' positions and actions (e.g., dribbling, passing, shooting), as well as to track the scoring situation and ball position.
[1321] Step 7: Recognize user emotions
[1322] The device uses an emotion engine to analyze the user's facial expressions and voice tone to recognize the user's emotional state.
[1323] The device transmits the emotional state as data to the server.
[1324] Step 8: The server stores the emotion data
[1325] The server receives the emotion data sent from the terminal.
[1326] The emotion data is integrated into the user profile and the profile information is updated.
[1327] Step 9: Generate explanatory text
[1328] The server integrates the user's sports knowledge level, favorite target, and emotion data.
[1329] Based on this, a generative AI model is used to generate live commentary that is appropriate for the user.
[1330] The content of the commentary is adjusted according to the user's knowledge level and emotional state.
[1331] Step 10: Creating audio data
[1332] The server converts the generated explanatory text into audio data using a text-to-speech (TTS) engine.
[1333] The converted audio data is ready to be provided to the user.
[1334] Step 11: Send audio data to the device
[1335] The server transmits the generated voice data to the terminal.
[1336] The terminal prepares to play the received audio data.
[1337] Step 12: Playing back audio data
[1338] The terminal plays the received audio data through a playback device (e.g., speaker, headphones).
[1339] Users can watch the game while listening to customized commentary in real time.
[1340] The commentary is customized to the user's level of sports knowledge, their roots, and their real-time emotional state.
[1341] Example 2
[1342] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1343] Current live sports commentary systems often provide uniform commentary without taking into account the individual user's knowledge level or emotional state. As a result, they are unable to address the different needs and expectations of each user, making it difficult to provide a personalized experience that matches each user's level of excitement and understanding. Furthermore, current systems often lack the accuracy of real-time video and emotion analysis, making it difficult to generate appropriate commentary based on user profiles. A system that solves these issues is needed.
[1344] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1345] In this invention, the server includes means for inputting a user's sports knowledge level and a target sport, means for analyzing game footage in real time to recognize the positions, movements, and score of players, means for analyzing the user's facial expressions and tone of voice to recognize the user's emotional state, means for generating commentary according to the user's profile based on the analysis results and the user's emotional state, means for converting the commentary into voice data using a text-to-speech engine, and means for providing the voice data to the user, thereby enabling the provision of customizable commentary in real time according to the user's individual knowledge level and emotional state.
[1346] The "user's sports knowledge level" indicates the user's knowledge and understanding of sports, and is expressed as a level such as beginner, intermediate, or advanced.
[1347] The term "target of support" refers to a specific team or player that the user supports during a game.
[1348] "Input means" includes devices or interfaces that allow a user to provide information to a system, such as a keyboard, touch screen, or voice input.
[1349] "Game footage" refers to video data showing a sports game being played, including live streaming and recorded footage.
[1350] "Real-time analysis" refers to the process of analyzing game video data as it is received and making the results immediately available.
[1351] "Player position" is information indicating the current location of each player within the competition area.
[1352] "Action" refers to specific actions that players perform during a game, such as dribbling, passing, shooting, etc.
[1353] "Scoring status" is information showing the distribution of scores in the current match and the latest scoring results.
[1354] "Means for analyzing" includes algorithms and software for processing the obtained data and extracting useful information, such as object recognition algorithms.
[1355] "Facial expressions" and "vocal tones" are physical and vocal characteristics that indicate the user's emotional state, and analyzing them allows the user's emotions to be inferred.
[1356] An "emotional state" refers to the psychological state that a user is feeling at a given moment, such as excitement, joy, sadness, etc.
[1357] "Means for recognizing emotional states" includes technologies and devices that analyze user characteristics such as facial expressions and vocal tone to estimate the emotions at that time.
[1358] A "user profile" is a data structure for centrally managing information related to a user, and includes information such as the user's level of sports knowledge, the sport they support, and their emotional state.
[1359] "Live commentary" is text that explains the progress and events of the match to users.
[1360] A "text-to-speech engine (TTS)" is a technology or software for converting text data into audio data.
[1361] "Audio data" refers to data in an audio format that is generated by a text-to-speech synthesis engine and can be heard by a user.
[1362] The "means for providing" includes devices and software for actually playing back the audio data generated by the system to the user, such as speakers and headphones.
[1363] This invention is a system that provides customized real-time commentary based on a user's sports knowledge level, emotional state, and favorite sport. The system consists of a terminal, a server, and related software and hardware.
[1364] User Profile Settings
[1365] First, a user uses their device to input their sports knowledge level (e.g., beginner, intermediate, advanced) and the team or player they support. The device receives this information, generates an API request, and sends it to the server. The server stores the received user information in a database and creates a profile for each user. The profile specifies the user's sports knowledge level and the team or player they support.
[1366] Real-time analysis of game footage
[1367] Game footage is received by the device from the live streaming service. The device then streams the received footage to the server in real time. The server then uses a video analysis algorithm to analyze the streamed footage. The server uses an object recognition algorithm to recognize player positions, actions (e.g., dribbling, passing, shooting), scoring situations, and ball position.
[1368] Recognizing user emotions
[1369] The device captures the user's facial expressions and voice tone using a camera and microphone, and analyzes them using an emotion engine to recognize the user's emotional state (e.g., excitement, joy, sadness, etc.). The recognized emotion data is sent from the device to a server and integrated into the user profile.
[1370] Generating live commentary audio
[1371] The server then combines the user's sports knowledge level, the sport they support, and their emotional data and inputs it into a generative AI model. Based on this combined data, the generative AI model generates a commentary tailored to the user. The commentary content is adjusted according to the user's knowledge level and emotional state. For example, an excited user will receive detailed commentary in an emotional tone.
[1372] Generating and providing voice data
[1373] The server converts the generated commentary into audio data using a text-to-speech (TTS) engine. This audio data is then sent to the device, which then plays the received audio data through a playback device (e.g., speaker, headphones). Users can watch the game while listening to the customized commentary in real time.
[1374] Specific examples
[1375] For example, consider a case where a user has a "beginner level of soccer knowledge" and is "supporting a specific team," becoming excited during a game. The user uses their device to set their own knowledge level and the team they are supporting. When the game video begins, the device receives the video and streams it to the server. The server analyzes the video and also receives the user's emotional state from the device. The emotion engine detects the user's level of excitement, and the server generates a commentary based on this. The commentary is in a tone that matches the user's excitement, such as "Great goal, your team is leading!" The generated commentary is converted into audio data, sent to the device, and then provided to the user. This system allows users to receive commentary appropriate to their sports knowledge level, the team they are supporting, and their real-time emotional state.
[1376] Prompt Sentence Examples
[1377] "Generate a commentary of a soccer game in which a beginner-level user who is rooting for a specific team gets excited when a goal is scored."
[1378] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1379] Step 1:
[1380] The user uses the device's input interface to set their sports knowledge level (e.g., beginner, intermediate, advanced) and the team or player they support. This becomes the input data. The device receives this information and generates an API request. The information included in this request (the user's sports knowledge level and the team they support) is sent to the server.
[1381] Step 2:
[1382] The server processes the received API request and stores the user information in a database. This creates a profile for each user, including the user's sports knowledge level and favorite sports. The output data is added to the database as a user profile.
[1383] Step 3:
[1384] When the match begins, the device receives the match video from the live streaming service. This video data becomes the input data, and is ready to be processed in real time.
[1385] Step 4:
[1386] The device receives the game video in real time and streams it to the server. The streamed video becomes the input data. The server prepares this video data for analysis.
[1387] Step 5:
[1388] The server analyzes the streamed video in real time using an object recognition algorithm. Using the video data as input, it recognizes the player's position, actions (e.g., dribbling, passing, shooting), scoring status, and ball position. The analysis results are output data.
[1389] Step 6:
[1390] The device captures the user's facial expressions and voice tone using a camera and microphone, which serve as input data. The emotion engine analyzes this data to recognize the user's emotional state, and the recognized emotional state is generated as output data.
[1391] Step 7:
[1392] The device sends the recognized emotion data to the server, which then integrates it into the user profile and keeps it up to date. This updates the user profile.
[1393] Step 8:
[1394] The server inputs the integrated user profile into the generative AI model. The input data includes the user's sports knowledge level, the people they support, and emotional data. Based on this data, the generative AI model generates a commentary appropriate for the user. This becomes the output data.
[1395] Step 9:
[1396] The server inputs the generated commentary into a text-to-speech (TTS) engine. The commentary text is used as input data. The TTS engine converts it into voice data. The voice data is generated as output data.
[1397] Step 10:
[1398] The server sends the generated audio data to the device, which becomes the input data. The device receives this data and plays it through a playback device (e.g., speaker, headphones). This allows the user to listen to a customized commentary in real time.
[1399] (Application example 2)
[1400] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1401] Conventional sports commentary systems have difficulty providing personalized commentary in real time based on the user's sports knowledge level or the team they support. They also lack the ability to provide customized commentary based on the user's emotional state. Therefore, multifaceted customization is needed to provide users with a more engaging and immersive viewing experience.
[1402] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1403] In this invention, the server includes: means for inputting a user's sports knowledge level and a target player; means for analyzing game footage in real time to recognize player positions, movements, and scores; means for analyzing the user's facial expressions and vocal tone to recognize the user's emotional state; means for generating prompt sentences based on the analysis results and the user's emotional state; means for inputting the prompt sentences into a generative AI model to generate commentary according to the user's profile and emotional state; means for converting the commentary into voice data using a text-to-speech engine; and means for providing the voice data to the user. This makes it possible to provide highly personalized commentary in real time based on the user's sports knowledge level, target player, and real-time emotional state.
[1404] "User's sports knowledge level" is information indicating the depth of knowledge and understanding of the user regarding sports.
[1405] "Support target" is information about a team or player that the user particularly supports.
[1406] "Game video" refers to real-time video data of a sports game.
[1407] "Player position" refers to information about the current position of each player within the competition venue.
[1408] "Action" refers to a series of movements that a player makes during a game (e.g., dribbling, passing, shooting, etc.).
[1409] "Scoring status" refers to the scoring status and score information during a match.
[1410] "User's facial expression" is information for reading emotions from the user's facial features and movements.
[1411] "Voice tone" is information that indicates the tone or pitch of the voice uttered by the user.
[1412] "User's emotional state" refers to the type of emotion the user is currently feeling (e.g., excitement, joy, sadness, etc.).
[1413] A "prompt sentence" is a sentence that contains prior information for generating an explanatory sentence and is input into the generative AI model.
[1414] A "generative AI model" is an artificial intelligence system that generates text for a specific task (in this case, commentary) based on input data.
[1415] A "text-to-speech synthesis engine" is a system that converts text into speech data.
[1416] "Audio data" is audio information that has been converted into a format that can be heard by a user.
[1417] This invention relates to a system for generating and providing play-by-play commentary audio that is customized based on the user's sports knowledge level and the sport they support, and also in response to the user's real-time emotional state. To achieve this, the following specific means and processes are required.
[1418] User Profile Settings
[1419] Users can enter their sports knowledge level and favorite teams and athletes via their device. This information is sent from the device to the server, which then receives it and stores it in a database, creating a customized profile for each user.
[1420] Real-time analysis of game footage
[1421] The device receives game footage from a live streaming service and streams it in real time to a server, which then analyzes the footage using object recognition algorithms to identify player positions, actions (e.g., dribbling, passing, shooting), and scoring situations.
[1422] Recognizing user emotions
[1423] The device uses the user's facial expressions and tone of voice to recognize the user's emotional state through an emotion engine, which then sends the emotion data to the server and stores it as a user profile.
[1424] Generating live commentary audio
[1425] The server combines the user's sports knowledge level, the team they are rooting for, and their emotional state to generate a prompt. The generated prompt is then input into a generative AI model to generate a play-by-play commentary appropriate for the user. For example, a prompt might be in the format: "The user's knowledge level is beginner, and the team they are rooting for is Team A. Their current emotional state is excited. The analysis results of the game are as follows: Team A has scored a goal!"
[1426] The generated commentary is converted into audio data using a text-to-speech (TTS) engine, which is then formatted for the user to hear.
[1427] Customized commentary
[1428] The terminal plays the audio data received from the server through a playback device (e.g., speaker, headphones), allowing the user to watch the game while listening to customized commentary in real time.
[1429] Specific examples
[1430] For example, if a user has a beginner level of soccer knowledge and supports a specific team (Team A), and the server detects that the user is excited, the server will generate the following prompt:
[1431] The user's knowledge level is beginner, and the team they support is Team A. Their current emotional state is excited. The analysis result of the match is as follows: Team A has scored a goal!
[1432] Based on this prompt, the generative AI model generates a commentary such as "Great goal, Team A is in the lead!", which is then converted into audio using a text-to-speech engine and sent to the device. The device then plays this audio through a playback device and provides it to the user.
[1433] In this way, the present invention is able to provide a customized commentary that is tailored to the user's profile and real-time emotional state.
[1434] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1435] Step 1:
[1436] A user inputs their level of sports knowledge and the teams and players they support via a terminal. The input information is data that indicates the user's sports viewing preferences. The terminal receives this data and sends it to a server. The server stores the received data in a database and generates a user profile. This profile includes information about the user's level of sports knowledge and the teams and players they support.
[1437] Step 2:
[1438] The device receives game footage from a live streaming service. The received video data is a real-time video stream showing the specific situation of the game. The device streams this video data to a server, which analyzes it in real time using video analysis algorithms. The analysis results include player positions, actions (e.g., dribbling, passing, shooting), and scoring status.
[1439] Step 3:
[1440] The device analyzes the user's facial expressions and vocal tone using an emotion engine. The analyzed emotion data indicates the user's real-time emotional state (e.g., excitement, joy, sadness, etc.). The device transmits this emotion data to the server, which then integrates the received emotion data into the user profile.
[1441] Step 4:
[1442] The server generates a prompt sentence based on the user's sports knowledge level, the team they are rooting for, and their emotional state. This prompt sentence contains prior information to be input into the generative AI model. For example, a prompt sentence of the form "The user's knowledge level is beginner level, and the team they are rooting for is Team A. Their current emotional state is excited. The analysis results of the game are as follows: Team A has scored a goal!" is generated.
[1443] Step 5:
[1444] The server inputs the generated prompt into a generative AI model to generate a commentary appropriate for the user. The generative AI model is an artificial intelligence system that generates commentary based on the prompt. The content of the generated commentary is adjusted according to the user's knowledge level and emotional state. As a result, a commentary appropriate for the user is obtained.
[1445] Step 6:
[1446] The server converts the generated commentary into audio data using a text-to-speech (TTS) engine, which then formats the audio data in a user-friendly format. Users can listen to this audio data in real time while watching the game.
[1447] Step 7:
[1448] The terminal plays the audio data received from the server through a playback device (e.g., speaker, headphones). This allows the user to listen to commentary customized according to the situation of the game. For example, if the user is excited, commentary such as "Great goal, Team A is in the lead!" is provided in an excited tone.
[1449] In this way, users can experience a personalized play-by-play commentary based on their level of sports knowledge, their favorite team, and their real-time emotional state.
[1450] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1451] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1452] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1453] [Fourth embodiment]
[1454] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1455] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1456] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1457] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1458] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1459] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1460] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1461] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1462] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1463] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1464] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1465] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1466] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1467] This invention relates to a system that uses generative AI to generate live sports commentary audio in real time. This allows for customizable live commentary based on the user's level of sports knowledge and the sport they are rooting for. The detailed program processing of this system is explained in natural language.
[1468] User Profile Settings
[1469] User: Enter information
[1470] Users use their devices to set their sports knowledge level (e.g., beginner, intermediate, advanced) and the team or player they support. This information forms the basis for customization settings used throughout the system.
[1471] Terminal: Data transmission
[1472] The device generates and sends an API request to the server to transmit the information entered by the user, including the sports knowledge level and support target selected by the user.
[1473] Server: Data storage
[1474] The server stores the received user information in a database and creates a profile for each user, which specifies the user's level of sports knowledge and the sport they support.
[1475] Real-time analysis of game footage
[1476] Terminal: Video reception
[1477] The device receives game footage from a live streaming service, which is then streamed to a server.
[1478] Terminal:Video Stream
[1479] The device then streams the received game footage to a server, where the streaming video becomes the input data for real-time analysis.
[1480] Server: Video analysis
[1481] The server analyzes the received video using a video analysis algorithm, specifically an object recognition algorithm that tracks the player's position in real time, recognizes each player's actions (e.g., dribbling, passing, shooting), and determines the goal situation and ball position.
[1482] Generating live commentary audio
[1483] Server: Customization information integration
[1484] The server combines user profile data with video analysis results, and this information is used to generate commentary.
[1485] Server: Explanation generation
[1486] The server then inputs the integrated data into a generative AI model to generate commentary tailored to the user. This commentary varies in detail depending on the user's level of sports knowledge. For example, commentary for beginners starts with basic explanations, while commentary for advanced players focuses on technical details.
[1487] Server: Speech synthesis
[1488] The server converts the generated commentary into audio data using a text-to-speech (TTS) engine, which is ready to be served to the user.
[1489] Server: Send data
[1490] The server transmits the generated audio data to the terminal, which uses the data to provide audio commentary to the user.
[1491] Customized commentary
[1492] Device: Audio playback
[1493] The device then provides the received audio data to the user through a playback device (e.g., speaker or headphones).The user watches the game while listening to customized commentary based on the user's sports knowledge level and the sport they are rooting for.
[1494] Specific examples
[1495] For example, consider a user who has a "beginner level of soccer knowledge" and "supports a specific team." In this case, the user uses the device to set their own knowledge level and the team they support. When the game starts, the device receives live video and streams it to the server. The server analyzes the video and generates appropriate commentary based on the user's profile. The commentary is converted into audio data using a text-to-speech engine and sent to the device. Finally, the device plays the generated audio data and provides it to the user.
[1496] This system allows users to receive real-time commentary tailored to their level of sports knowledge and the sport they are rooting for, improving the sports viewing experience.
[1497] The processing flow will be explained below.
[1498] Step 1: User inputs their sports knowledge level and the sport they support
[1499] Using the terminal, the user sets his / her sports knowledge level (e.g., beginner, intermediate, advanced) and the team or player he / she supports.
[1500] This information is entered through the terminal's user interface.
[1501] Step 2: The device sends the data to the server
[1502] The terminal receives the sports knowledge level and the information of the target of support input by the user.
[1503] The device generates an API request containing this information and sends it to the server.
[1504] Step 3: The server stores the user data
[1505] The server receives the user information sent from the terminal.
[1506] The server stores the received information in a database and creates a profile for each user.
[1507] Step 4: Receiving game footage
[1508] The device receives game footage from a live streaming service.
[1509] The received match footage is ready to be processed in real time.
[1510] Step 5: Stream the game
[1511] The device streams the received game footage to the server in real time.
[1512] The streamed video is analyzed by the server.
[1513] Step 6: The server analyzes the video in real time
[1514] The server uses a video analysis algorithm to analyze the streamed video in real time.
[1515] The server uses object recognition algorithms to recognize players' positions and actions (e.g., dribbling, passing, shooting).
[1516] The server also keeps track of the score and the position of the ball.
[1517] Step 7: Integrating user information and analysis results
[1518] The server integrates the user's profile data with the video analysis results.
[1519] Compile materials for commentary based on the user's level of sports knowledge and the sport they support.
[1520] Step 8: Generate explanatory text
[1521] The server inputs the integrated data into the generative AI model and generates a commentary appropriate for the user.
[1522] The explanations are tailored to the user's level of knowledge, with the level of detail varying between beginners and advanced users.
[1523] Step 9: Generate audio data
[1524] The server converts the generated explanatory text into audio data using a text-to-speech (TTS) engine.
[1525] This audio data is then formatted so that it can be heard by the user.
[1526] Step 10: Send audio data to the device
[1527] The server transmits the generated voice data to the terminal.
[1528] The terminal prepares to receive the audio data.
[1529] Step 11: Playing back audio data
[1530] The terminal plays the received audio data through a playback device (e.g., speaker, headphones).
[1531] Users can watch the game while listening to customized commentary in real time.
[1532] Example 1
[1533] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1534] Conventional sports commentary systems have difficulty meeting the diverse needs of users, and have been unable to provide customized commentary based on individual sports knowledge levels or the sports fans they support. Furthermore, they lack technology that can accurately recognize the progress of a game in real time and instantly generate appropriate commentary. The present invention aims to solve these problems and provide users with real-time, customized sports commentary.
[1535] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1536] In this invention, the server includes means for inputting a user's sports knowledge level and a target player, means for transmitting the input information to the server, means for storing the input information in a database, means for receiving game footage and streaming the footage to the server, means for analyzing the game footage in real time and recognizing player positions, movements, and scores, means for generating commentary text tailored to the user's profile based on the analysis results and the user's sports knowledge level and target player, means for converting the generated commentary text into audio data using a text-to-speech engine, and means for transmitting the audio data from the server to a terminal and providing it to the user via a playback device. This enables real-time, customized sports commentary tailored to the diverse needs of users.
[1537] A "user" is an individual who watches sports and enters information into the system to receive customized commentary.
[1538] The "sports knowledge level" is an index that indicates the depth of a user's knowledge about sports, and is classified into categories such as beginner, intermediate, and advanced.
[1539] The term "target of support" refers to a team or player in which the user has a particular interest and supports.
[1540] A "terminal" is an electronic device that a user uses to provide input information, receive game footage, and play audio data.
[1541] The "server" is a centralized management system that stores information sent by users, analyzes game footage, generates commentary, and distributes audio data.
[1542] A "database" is a system located within a server that stores and manages user information and analysis results.
[1543] "Game footage" refers to video data of sports games distributed via live streaming services.
[1544] "Streaming" refers to the act of transmitting or receiving game footage in real time.
[1545] "Video analysis" refers to the process of processing game footage in real time and recognizing players' positions, movements, scoring situations, etc.
[1546] A "user profile" is a data set that includes individually set information such as the user's level of sports knowledge and the sport they support.
[1547] "Live commentary" is a narration text that is generated based on the progress of the match.
[1548] A "text-to-speech engine" is a software technology for converting text data into speech data.
[1549] "Audio data" refers to commentary data in audio format that is generated by a text-to-speech synthesis engine and provided to the user.
[1550] A "playback device" refers to a device such as a speaker or headphones that is connected to a terminal and allows the user to listen to audio data.
[1551] This invention relates to a system that uses generative AI to generate live sports commentary audio in real time. This system can provide customized commentary according to the user's sports knowledge level and the sport they are rooting for. Each processing step of the system is described in detail below.
[1552] User Profile Settings
[1553] User: Enter information
[1554] Users use the terminal to input their sports knowledge level (beginner, intermediate, advanced) and information about the team or player they support. The input information is used as the basis for customization settings within the system.
[1555] Terminal: Data transmission
[1556] The device sends the information entered by the user to the server as an API request. Specifically, the device packages the input information in JSON format and sends an HTTP request to the server.
[1557] Server: Data storage
[1558] The server stores the data of the received API request in a database. For example, the database stores information such as the user ID, knowledge level, and support team.
[1559] Real-time analysis of game footage
[1560] Terminal: Video reception
[1561] The device receives game footage from a live streaming service, for example, by obtaining live footage in HLS format from a streaming URL.
[1562] Terminal:Video Stream
[1563] The device then streams the received game footage to the server in real time, using the RTMP or WebRTC protocol to send the video data to the server.
[1564] Server: Video analysis
[1565] The server analyzes the streamed game footage using a video analysis algorithm (e.g., YOLO or OpenCV) to recognize the positions and movements of players. As a result of the analysis, the score and ball position can be determined.
[1566] Generating live commentary audio
[1567] Server: Customization information integration
[1568] The server integrates the user profile information with the video analysis results. For example, the analysis results might link information such as "Player A is dribbling" with information such as "the user is a beginner."
[1569] Server: Explanation generation
[1570] The server inputs the integrated data using a generative AI model and generates commentary appropriate for the user. As a specific example, for beginners, a commentary such as "Player A broke through the opponent with a great dribble" is generated. An example of a prompt is as follows:
[1571] "Generate a play-by-play commentary of a soccer game, given that the user has a beginner level of knowledge and identifies the team they support."
[1572] Server: Speech synthesis
[1573] The server converts the generated commentary into voice data using a text-to-speech engine (TTS engine). Specifically, it uses the API of the TTS engine to convert the text "Player A made a great dribble..." into voice data.
[1574] Server: Sends audio data
[1575] The server sends the generated audio data to the user's device by returning the audio data as an HTTP response.
[1576] Customized commentary
[1577] Device: Audio playback
[1578] The terminal provides the received audio data to the user through a playback device (speaker or headphones), specifically by decoding the audio data and sending it to the audio system.
[1579] In this way, the system can provide real-time, customized sports commentary that meets the diverse needs of users.
[1580] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1581] Step 1:
[1582] User: Enter sports knowledge level and support target
[1583] Using the terminal, users input their sports knowledge level (beginner, intermediate, advanced) and information about the team and player they support.
[1584] Input: User knowledge level, target of support
[1585] Output: User input data (knowledge level and support target)
[1586] Specific actions: The user operates text fields and selection menus on the device screen to enter the required information.
[1587] Step 2:
[1588] Terminal: Data transmission
[1589] The terminal sends the information entered by the user to the server as an API request.
[1590] Input: User input data (knowledge level and support target)
[1591] Output: API request
[1592] Specific operation: The terminal converts the user's input data into JSON format and sends it to the server as an HTTP POST request.
[1593] Step 3:
[1594] Server: Data storage
[1595] The server analyzes the received API requests and stores them in a database.
[1596] Input: API request
[1597] Output: User profile stored in the database
[1598] Specific operation: The server parses the received JSON data and saves the corresponding user profile in the database.
[1599] Step 4:
[1600] Terminal: Receiving game footage
[1601] The device receives game footage from a live streaming service.
[1602] Input: Live Streaming URL
[1603] Output: Received game video
[1604] Specific operation: The device retrieves live game footage in HLS or DASH format from the specified URL.
[1605] Step 5:
[1606] Terminal:Video Stream
[1607] The device streams the received game footage to the server in real time.
[1608] Input: Received game footage
[1609] Output: Streamed game footage
[1610] Specific operation: The device sends video data to the server using the RTMP or WebRTC protocol.
[1611] Step 6:
[1612] Server: Video analysis
[1613] The server analyzes the received video in real time and recognizes the players' positions and movements.
[1614] Input: Streamed game footage
[1615] Output: Analysis results (player positions, actions, scoring status)
[1616] Specific operation: The server performs analysis processing using a video analysis algorithm (e.g., YOLO or OpenCV).
[1617] Step 7:
[1618] Server: Customization information integration
[1619] The server integrates the user profile information with the video analysis results.
[1620] Input: User profile, analysis results
[1621] Output: Integrated data
[1622] Specific operation: The server links the analysis results with the user's knowledge level and support target information, and integrates them into an appropriate data structure.
[1623] Step 8:
[1624] Server: Live commentary generation
[1625] The server generates commentary using a generative AI model.
[1626] Input: Integrated data
[1627] Output: Commentary
[1628] Specific operation: The server inputs the integrated data into the generative AI model and generates a commentary appropriate for the user.
[1629] Example prompt: "Generate a play-by-play commentary of a soccer game given the user's beginner level of knowledge and which team they support."
[1630] Step 9:
[1631] Server: Speech synthesis
[1632] The server converts the generated commentary into voice data using a text-to-speech synthesis engine.
[1633] Input: Commentary
[1634] Output: Audio data
[1635] Specific operation: The server uses a TTS engine (e.g., Google TTS API) to convert the text into audio data.
[1636] Step 10:
[1637] Server: Sends audio data
[1638] The server transmits the generated voice data to the terminal.
[1639] Input: Audio data
[1640] Output: HTTP response
[1641] Specific operation: The server sends the audio data to the terminal as an HTTP response.
[1642] Step 11:
[1643] Device: Audio playback
[1644] The terminal provides the received audio data to the user on a playback device.
[1645] Input: Audio data
[1646] Output: Played audio
[1647] Specific operation: The device decodes the audio data and plays it through speakers or headphones.
[1648] Through the above steps, users can enjoy real-time commentary that is customized to their level of sports knowledge and the sport they are rooting for.
[1649] (Application example 1)
[1650] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1651] Conventional sports commentary systems have difficulty customizing commentary to suit individual users' knowledge levels or the teams they support, making it impossible to provide appropriate commentary in real time. Furthermore, the process of converting the generated commentary text into audio data and providing it to users' devices in real time was lacking. As a result, users only received the same commentary, which led to a problem of degrading the quality of each individual viewing experience.
[1652] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1653] In this invention, the server includes means for inputting a user's sports knowledge level and a team the user supports, means for analyzing game footage in real time to recognize player positions, movements, and scores, means for generating commentary according to the user's profile based on the analysis results, means for converting the commentary into voice data using a text-to-speech engine, means for providing the voice data to the user, means for receiving game footage via live streaming and transmitting it to the server, means for generating appropriate commentary from prompts using a generative AI model, and means for transmitting the generated voice data to the user's device in real time and playing it back. This makes it possible to provide appropriate and customized commentary in real time according to the user's sports knowledge level and the team the user supports.
[1654] The "user's sports knowledge level" indicates the depth of knowledge and understanding of the user about sports, and is usually classified into stages such as beginner, intermediate, and advanced.
[1655] The term "target of support" refers to a team or player that the user particularly supports in a particular sporting event or game.
[1656] "Game footage" refers to video data that records the progress of a sporting event or game in real time.
[1657] "Real-time analysis" means receiving game video data, processing it instantly, and extracting the necessary information and features.
[1658] "Player position" is information indicating where each player is located on the playing field at a specific time during the game.
[1659] "Action" refers to the specific actions or movements that a player makes during a game or competition, such as dribbling, passing, and shooting.
[1660] "Scoring Status" refers to information about the scores of both teams during the match, including the current score, the players who scored, and the method of scoring.
[1661] "Live commentary" is commentary based on the progress of the match and the actions of the players, and is intended to explain the content of the match to viewers in an easy-to-understand manner.
[1662] A "text-to-speech engine" is software or a process for converting input text data into speech data.
[1663] "Audio data" means a data file in audio format generated by a text-to-speech engine.
[1664] "Live streaming reception" is the process of instantly acquiring video and audio that is being streamed in real time over the Internet.
[1665] "Send to server" means sending data from a terminal or the like to a central server.
[1666] A "generative AI model" is an algorithm or framework for generating text or content using artificial intelligence.
[1667] A "prompt" is an instruction that is input into a generative AI model, and is text that causes the model to generate appropriate text or content based on that instruction.
[1668] "Terminal" means a device used by a user, including a smartphone, tablet, computer, etc.
[1669] "Playback" means aurally reproducing the generated audio data using an audio output device.
[1670] The system based on this invention utilizes generative AI models to provide users with customized, real-time commentary while watching sports.
[1671] 1. Setting up your user profile
[1672] User:
[1673] Users input their sports knowledge level (e.g., beginner, intermediate, advanced) and their favorite teams and athletes via their terminals. This information forms the basis for customization settings used throughout the system.
[1674] Device:
[1675] The information entered by the user is sent from the device to the server. Here, the user interface is built using React Native, and the device generates an API request to send the data.
[1676] server:
[1677] The server stores the received user information in a database (e.g., PostgreSQL) and creates a profile for each user, which includes the user's level of sports knowledge and the sport they support.
[1678] 2. Real-time analysis of game footage
[1679] Device:
[1680] The device uses WebRTC to receive game footage in real time from a live streaming service and streams the video data to a server.
[1681] server:
[1682] The server uses OpenCV to analyze the received video in real time, and uses object recognition algorithms to identify the player's position and recognize their actions (e.g., dribbling, passing, shooting), as well as to track the goal situation and ball position.
[1683] 3. Generation of commentary audio
[1684] server:
[1685] The server combines the user profile data with the video analysis results and inputs prompts into the generative AI model to generate appropriate commentary. For example, the following prompts are input into the generative AI model:
[1686] "Generate beginner-friendly commentary: Game status {Score}, player {Name} performed {Action}."
[1687] The generated commentary is converted into audio data using the Google Cloud Text-to-Speech API.
[1688] 4. Customized commentary
[1689] server:
[1690] The generated voice data is transmitted from the server to the terminal.
[1691] Device:
[1692] The device then provides the received audio data to the user via a playback device (e.g., smartphone, head-mounted display), allowing users to watch the game while listening to commentary tailored to their level of sports knowledge and the sport they are rooting for.
[1693] Specific examples
[1694] For example, consider a case where a user is at an "intermediate level" and is cheering for "Team A." If Player X is recognized as dribbling during a match and reaching the goal, the generative AI model will generate an appropriate commentary as follows:
[1695] Example prompt sentence:
[1696] "During a game, Player X dribbles the ball and reaches the goal. Please provide appropriate commentary for intermediate level users."
[1697] The generated commentary is converted into audio data using the Google Cloud Text-to-Speech API and sent to the user's device in real time, where they can enjoy this customized commentary via a head-mounted display or smartphone.
[1698] In this way, by implementing the invention, users can receive appropriate and customized commentary in real time according to their level of sports knowledge and the team they support.
[1699] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1700] Program processing steps and their specific explanations
[1701] Step 1: Configure your user profile
[1702] Users use a smartphone app to input their sports knowledge level (beginner, intermediate, advanced) and the team or player they support. The input data is then saved on the device.
[1703] Input: The user enters their sports knowledge level and the team or player they support into the input form.
[1704] Data processing: Convert input data into API request format
[1705] Output: Sent to the server as an API request
[1706] Step 2: Send and store user information
[1707] The terminal sends the information entered by the user to the server, where it is saved in a database that stores user profiles.
[1708] Input: User's sports knowledge level and favorite team and player information
[1709] Data processing: Converting transmitted data into a database format
[1710] Output: User profile saved in database
[1711] Step 3: Receiving and streaming game footage
[1712] The device receives game footage from the live streaming service using the WebRTC protocol and transmits it to the server.
[1713] Input: Real-time game footage from a live streaming service
[1714] Data processing: Converts video data into a format suitable for the server in real time
[1715] Output: Streaming data is sent to the server
[1716] Step 4: Video analysis
[1717] The server analyzes the received game footage in real time using OpenCV to recognize the players' positions, movements, and scoring status.
[1718] Input: Streamed game footage
[1719] Data processing: Apply object recognition algorithms to recognize player positions, movements, and scoring situations
[1720] Output: Analysis results: player positions, movements, and score information
[1721] Step 5: Generate commentary
[1722] Based on the analysis results and the user profile, the server inputs prompt text into the generative AI model to generate appropriate commentary.
[1723] Input: Analysis results and user profile
[1724] Data processing: Input prompt text into the generative AI model to generate commentary
[1725] Output: Generated commentary
[1726] Example of specific behavior: Send the following prompt to the generative AI model: "Generate commentary appropriate for beginners: Game situation {Score}, Player {Name} performed {Action}."
[1727] Step 6: Convert the commentary into audio
[1728] The server converts the generated commentary into audio data using the Google Cloud Text-to-Speech API.
[1729] Input: Generated commentary
[1730] Data processing: Convert text data into audio data
[1731] Output: Audio data
[1732] Step 7: Sending audio data
[1733] The server transmits the converted voice data to the user's terminal.
[1734] Input: Audio data
[1735] Data processing: Real-time transmission of voice data
[1736] Output: Audio data sent to the user's device
[1737] Step 8: Play commentary audio
[1738] The terminal plays the received audio data on a playback device (e.g., smartphone, head-mounted display) and provides it to the user.
[1739] Input: Audio data sent from the server
[1740] Data processing: Conversion to playback format
[1741] Output: Audio description played from the playback device
[1742] By having each processing step work in conjunction with each other in this way, users can receive live commentary in real time that is individually customized according to their level of sports knowledge and the sport they are rooting for.
[1743] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1744] This invention relates to a system that combines generative AI and an emotion engine to generate live sports commentary audio in real time. This allows for customizable live commentary based on the user's sports knowledge level, the sport they are rooting for, and even their emotional state. The detailed program processing of this system is explained in natural language.
[1745] User Profile Settings
[1746] User: Enter information
[1747] Users use their devices to set their sports knowledge level (e.g., beginner, intermediate, advanced) and the team or player they support. This information forms the basis for customization settings used throughout the system.
[1748] Terminal: Data transmission
[1749] The device receives the sports knowledge level and the information of the target sport entered by the user, and then generates an API request including this information and sends it to the server.
[1750] Server: Data storage
[1751] The server stores the received user information in a database and creates a profile for each user, which specifies the user's level of sports knowledge and the sport they support.
[1752] Real-time analysis of game footage
[1753] Terminal: Video reception
[1754] The device receives game footage from a live streaming service, which prepares it for processing in real time.
[1755] Device: Video streaming
[1756] The device streams the received game video to the server in real time, where it is analyzed.
[1757] Server: Video analysis
[1758] The server uses video analysis algorithms to analyze the streamed video in real time, and object recognition algorithms to recognize player positions and actions (e.g., dribbling, passing, shooting), as well as to determine goal situations and ball position.
[1759] Recognizing user emotions
[1760] Device: Emotion data acquisition
[1761] The device uses an emotion engine to analyze the user's facial expressions and voice tone to recognize the user's emotional state.
[1762] Server: Sends and stores emotion data
[1763] The device sends the recognized emotion data to the server, which then integrates it into the user profile.
[1764] Generating live commentary audio
[1765] Server: Customization information integration
[1766] The server integrates the user's sports knowledge level, favorites, and emotional data, which are used to generate commentary.
[1767] Server: Explanation generation
[1768] The server inputs the integrated data into a generative AI model to generate commentary tailored to the user, using different expressions and tones depending on the user's knowledge level and emotional state.
[1769] Server: Generates voice data
[1770] The server converts the generated commentary into audio data using a text-to-speech (TTS) engine, which is then formatted for listening by the user.
[1771] Server: Sending audio data
[1772] The server transmits the generated audio data to the terminal, which uses this data to provide audio commentary to the user.
[1773] Customized commentary
[1774] Device: Audio playback
[1775] The device then plays the received audio data through a playback device (e.g., speaker, headphones). Users can watch the game while listening to a customized commentary in real time. The commentary is tailored to the user's sports knowledge level, favorite sport, and emotional state, providing a more personalized experience.
[1776] Specific examples
[1777] For example, consider a user with a beginner level of soccer knowledge who is rooting for a specific team and is excited during a match. The user uses their device to set their own knowledge level and the team they are rooting for. When the game starts, the device receives the footage and streams it to the server. The server analyzes the footage and receives the user's emotional state from the device. The emotion engine detects the user's level of excitement, and the server generates a commentary based on this information. The commentary is written in a tone that matches the user's excitement, such as "Great goal, your team is leading!" The generated commentary is converted into audio data, sent to the device, and then provided to the user. This system allows users to receive commentary appropriate to their sports knowledge level, the team they are rooting for, and their real-time emotional state.
[1778] The processing flow will be explained below.
[1779] Step 1: User inputs their sports knowledge level and the sport they support
[1780] Using the terminal, the user sets his / her sports knowledge level (e.g., beginner, intermediate, advanced) and the team or player he / she supports.
[1781] This provides basic information that the system needs to provide the user with the most appropriate explanation.
[1782] Step 2: The device sends the data to the server
[1783] The terminal receives the sports knowledge level and the information of the target of support input by the user.
[1784] The device then generates an API request containing this information and sends it to the server.
[1785] Step 3: The server stores the user data
[1786] The server receives the user information sent from the terminal.
[1787] The received user information is stored in a database and a profile is created for each user.
[1788] Step 4: Receiving game footage
[1789] The device receives game footage from a live streaming service.
[1790] The received match footage is ready to be processed in real time.
[1791] Step 5: Stream the game
[1792] The device streams the received game footage to the server in real time.
[1793] The streamed video is analyzed by the server.
[1794] Step 6: The server analyzes the video in real time
[1795] The server uses a video analysis algorithm to analyze the streamed video in real time.
[1796] The server uses object recognition algorithms to recognize players' positions and actions (e.g., dribbling, passing, shooting), as well as to track the scoring situation and ball position.
[1797] Step 7: Recognize user emotions
[1798] The device uses an emotion engine to analyze the user's facial expressions and voice tone to recognize the user's emotional state.
[1799] The device transmits the emotional state as data to the server.
[1800] Step 8: The server stores the emotion data
[1801] The server receives the emotion data sent from the terminal.
[1802] The emotion data is integrated into the user profile and the profile information is updated.
[1803] Step 9: Generate explanatory text
[1804] The server integrates the user's sports knowledge level, favorite target, and emotion data.
[1805] Based on this, a generative AI model is used to generate live commentary that is appropriate for the user.
[1806] The content of the commentary is adjusted according to the user's knowledge level and emotional state.
[1807] Step 10: Creating audio data
[1808] The server converts the generated explanatory text into audio data using a text-to-speech (TTS) engine.
[1809] The converted audio data is ready to be provided to the user.
[1810] Step 11: Send audio data to the device
[1811] The server transmits the generated voice data to the terminal.
[1812] The terminal prepares to play the received audio data.
[1813] Step 12: Playing back audio data
[1814] The terminal plays the received audio data through a playback device (e.g., speaker, headphones).
[1815] Users can watch the game while listening to customized commentary in real time.
[1816] The commentary is customized to the user's level of sports knowledge, their roots, and their real-time emotional state.
[1817] Example 2
[1818] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1819] Current live sports commentary systems often provide uniform commentary without taking into account the individual user's knowledge level or emotional state. As a result, they are unable to address the different needs and expectations of each user, making it difficult to provide a personalized experience that matches each user's level of excitement and understanding. Furthermore, current systems often lack the accuracy of real-time video and emotion analysis, making it difficult to generate appropriate commentary based on user profiles. A system that solves these issues is needed.
[1820] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1821] In this invention, the server includes means for inputting a user's sports knowledge level and a target sport, means for analyzing game footage in real time to recognize the positions, movements, and score of players, means for analyzing the user's facial expressions and tone of voice to recognize the user's emotional state, means for generating commentary according to the user's profile based on the analysis results and the user's emotional state, means for converting the commentary into voice data using a text-to-speech engine, and means for providing the voice data to the user, thereby enabling the provision of customizable commentary in real time according to the user's individual knowledge level and emotional state.
[1822] The "user's sports knowledge level" indicates the user's knowledge and understanding of sports, and is expressed as a level such as beginner, intermediate, or advanced.
[1823] The term "target of support" refers to a specific team or player that the user supports during a game.
[1824] "Input means" includes devices or interfaces that allow a user to provide information to a system, such as a keyboard, touch screen, or voice input.
[1825] "Game footage" refers to video data showing a sports game being played, including live streaming and recorded footage.
[1826] "Real-time analysis" refers to the process of analyzing game video data as it is received and making the results immediately available.
[1827] "Player position" is information indicating the current location of each player within the competition area.
[1828] "Action" refers to specific actions that players perform during a game, such as dribbling, passing, shooting, etc.
[1829] "Scoring status" is information showing the distribution of scores in the current match and the latest scoring results.
[1830] "Means for analyzing" includes algorithms and software for processing the obtained data and extracting useful information, such as object recognition algorithms.
[1831] "Facial expressions" and "vocal tones" are physical and vocal characteristics that indicate the user's emotional state, and analyzing them allows the user's emotions to be inferred.
[1832] An "emotional state" refers to the psychological state that a user is feeling at a given moment, such as excitement, joy, sadness, etc.
[1833] "Means for recognizing emotional states" includes technologies and devices that analyze user characteristics such as facial expressions and vocal tone to estimate the emotions at that time.
[1834] A "user profile" is a data structure for centrally managing information related to a user, and includes information such as the user's level of sports knowledge, the sport they support, and their emotional state.
[1835] "Live commentary" is text that explains the progress and events of the match to users.
[1836] A "text-to-speech engine (TTS)" is a technology or software for converting text data into audio data.
[1837] "Audio data" refers to data in an audio format that is generated by a text-to-speech synthesis engine and can be heard by a user.
[1838] The "means for providing" includes devices and software for actually playing back the audio data generated by the system to the user, such as speakers and headphones.
[1839] This invention is a system that provides customized real-time commentary based on a user's sports knowledge level, emotional state, and favorite sport. The system consists of a terminal, a server, and related software and hardware.
[1840] User Profile Settings
[1841] First, a user uses their device to input their sports knowledge level (e.g., beginner, intermediate, advanced) and the team or player they support. The device receives this information, generates an API request, and sends it to the server. The server stores the received user information in a database and creates a profile for each user. The profile specifies the user's sports knowledge level and the team or player they support.
[1842] Real-time analysis of game footage
[1843] Game footage is received by the device from the live streaming service. The device then streams the received footage to the server in real time. The server then uses a video analysis algorithm to analyze the streamed footage. The server uses an object recognition algorithm to recognize player positions, actions (e.g., dribbling, passing, shooting), scoring situations, and ball position.
[1844] Recognizing user emotions
[1845] The device captures the user's facial expressions and voice tone using a camera and microphone, and analyzes them using an emotion engine to recognize the user's emotional state (e.g., excitement, joy, sadness, etc.). The recognized emotion data is sent from the device to a server and integrated into the user profile.
[1846] Generating live commentary audio
[1847] The server then combines the user's sports knowledge level, the sport they support, and their emotional data and inputs it into a generative AI model. Based on this combined data, the generative AI model generates a commentary tailored to the user. The commentary content is adjusted according to the user's knowledge level and emotional state. For example, an excited user will receive detailed commentary in an emotional tone.
[1848] Generating and providing voice data
[1849] The server converts the generated commentary into audio data using a text-to-speech (TTS) engine. This audio data is then sent to the device, which then plays the received audio data through a playback device (e.g., speaker, headphones). Users can watch the game while listening to the customized commentary in real time.
[1850] Specific examples
[1851] For example, consider a case where a user has a "beginner level of soccer knowledge" and is "supporting a specific team," becoming excited during a game. The user uses their device to set their own knowledge level and the team they are supporting. When the game video begins, the device receives the video and streams it to the server. The server analyzes the video and also receives the user's emotional state from the device. The emotion engine detects the user's level of excitement, and the server generates a commentary based on this. The commentary is in a tone that matches the user's excitement, such as "Great goal, your team is leading!" The generated commentary is converted into audio data, sent to the device, and then provided to the user. This system allows users to receive commentary appropriate to their sports knowledge level, the team they are supporting, and their real-time emotional state.
[1852] Prompt Sentence Examples
[1853] "Generate a commentary of a soccer game in which a beginner-level user who is rooting for a specific team gets excited when a goal is scored."
[1854] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1855] Step 1:
[1856] The user uses the device's input interface to set their sports knowledge level (e.g., beginner, intermediate, advanced) and the team or player they support. This becomes the input data. The device receives this information and generates an API request. The information included in this request (the user's sports knowledge level and the team they support) is sent to the server.
[1857] Step 2:
[1858] The server processes the received API request and stores the user information in a database. This creates a profile for each user, including the user's sports knowledge level and favorite sports. The output data is added to the database as a user profile.
[1859] Step 3:
[1860] When the match begins, the device receives the match video from the live streaming service. This video data becomes the input data, and is ready to be processed in real time.
[1861] Step 4:
[1862] The device receives the game video in real time and streams it to the server. The streamed video becomes the input data. The server prepares this video data for analysis.
[1863] Step 5:
[1864] The server analyzes the streamed video in real time using an object recognition algorithm. Using the video data as input, it recognizes the player's position, actions (e.g., dribbling, passing, shooting), scoring status, and ball position. The analysis results are output data.
[1865] Step 6:
[1866] The device captures the user's facial expressions and voice tone using a camera and microphone, which serve as input data. The emotion engine analyzes this data to recognize the user's emotional state, and the recognized emotional state is generated as output data.
[1867] Step 7:
[1868] The device sends the recognized emotion data to the server, which then integrates it into the user profile and keeps it up to date. This updates the user profile.
[1869] Step 8:
[1870] The server inputs the integrated user profile into the generative AI model. The input data includes the user's sports knowledge level, the people they support, and emotional data. Based on this data, the generative AI model generates a commentary appropriate for the user. This becomes the output data.
[1871] Step 9:
[1872] The server inputs the generated commentary into a text-to-speech (TTS) engine. The commentary text is used as input data. The TTS engine converts it into voice data. The voice data is generated as output data.
[1873] Step 10:
[1874] The server sends the generated audio data to the device, which becomes the input data. The device receives this data and plays it through a playback device (e.g., speaker, headphones). This allows the user to listen to a customized commentary in real time.
[1875] (Application example 2)
[1876] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1877] Conventional sports commentary systems have difficulty providing personalized commentary in real time based on the user's sports knowledge level or the team they support. They also lack the ability to provide customized commentary based on the user's emotional state. Therefore, multifaceted customization is needed to provide users with a more engaging and immersive viewing experience.
[1878] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1879] In this invention, the server includes: means for inputting a user's sports knowledge level and a target player; means for analyzing game footage in real time to recognize player positions, movements, and scores; means for analyzing the user's facial expressions and vocal tone to recognize the user's emotional state; means for generating prompt sentences based on the analysis results and the user's emotional state; means for inputting the prompt sentences into a generative AI model to generate commentary according to the user's profile and emotional state; means for converting the commentary into voice data using a text-to-speech engine; and means for providing the voice data to the user. This makes it possible to provide highly personalized commentary in real time based on the user's sports knowledge level, target player, and real-time emotional state.
[1880] "User's sports knowledge level" is information indicating the depth of knowledge and understanding of the user regarding sports.
[1881] "Support target" is information about a team or player that the user particularly supports.
[1882] "Game video" refers to real-time video data of a sports game.
[1883] "Player position" refers to information about the current position of each player within the competition venue.
[1884] "Action" refers to a series of movements that a player makes during a game (e.g., dribbling, passing, shooting, etc.).
[1885] "Scoring status" refers to the scoring status and score information during a match.
[1886] "User's facial expression" is information for reading emotions from the user's facial features and movements.
[1887] "Voice tone" is information that indicates the tone or pitch of the voice uttered by the user.
[1888] "User's emotional state" refers to the type of emotion the user is currently feeling (e.g., excitement, joy, sadness, etc.).
[1889] A "prompt sentence" is a sentence that contains prior information for generating an explanatory sentence and is input into the generative AI model.
[1890] A "generative AI model" is an artificial intelligence system that generates text for a specific task (in this case, commentary) based on input data.
[1891] A "text-to-speech synthesis engine" is a system that converts text into speech data.
[1892] "Audio data" is audio information that has been converted into a format that can be heard by a user.
[1893] This invention relates to a system for generating and providing play-by-play commentary audio that is customized based on the user's sports knowledge level and the sport they support, and also in response to the user's real-time emotional state. To achieve this, the following specific means and processes are required.
[1894] User Profile Settings
[1895] Users can enter their sports knowledge level and favorite teams and athletes via their device. This information is sent from the device to the server, which then receives it and stores it in a database, creating a customized profile for each user.
[1896] Real-time analysis of game footage
[1897] The device receives game footage from a live streaming service and streams it in real time to a server, which then analyzes the footage using object recognition algorithms to identify player positions, actions (e.g., dribbling, passing, shooting), and scoring situations.
[1898] Recognizing user emotions
[1899] The device uses the user's facial expressions and tone of voice to recognize the user's emotional state through an emotion engine, which then sends the emotion data to the server and stores it as a user profile.
[1900] Generating live commentary audio
[1901] The server combines the user's sports knowledge level, the team they are rooting for, and their emotional state to generate a prompt. The generated prompt is then input into a generative AI model to generate a play-by-play commentary appropriate for the user. For example, a prompt might be in the format: "The user's knowledge level is beginner, and the team they are rooting for is Team A. Their current emotional state is excited. The analysis results of the game are as follows: Team A has scored a goal!"
[1902] The generated commentary is converted into audio data using a text-to-speech (TTS) engine, which is then formatted for the user to hear.
[1903] Customized commentary
[1904] The terminal plays the audio data received from the server through a playback device (e.g., speaker, headphones), allowing the user to watch the game while listening to customized commentary in real time.
[1905] Specific examples
[1906] For example, if a user has a beginner level of soccer knowledge and supports a specific team (Team A), and the server detects that the user is excited, the server will generate the following prompt:
[1907] The user's knowledge level is beginner, and the team they support is Team A. Their current emotional state is excited. The analysis result of the match is as follows: Team A has scored a goal!
[1908] Based on this prompt, the generative AI model generates a commentary such as "Great goal, Team A is in the lead!", which is then converted into audio using a text-to-speech engine and sent to the device. The device then plays this audio through a playback device and provides it to the user.
[1909] In this way, the present invention is able to provide a customized commentary that is tailored to the user's profile and real-time emotional state.
[1910] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1911] Step 1:
[1912] A user inputs their level of sports knowledge and the teams and players they support via a terminal. The input information is data that indicates the user's sports viewing preferences. The terminal receives this data and sends it to a server. The server stores the received data in a database and generates a user profile. This profile includes information about the user's level of sports knowledge and the teams and players they support.
[1913] Step 2:
[1914] The device receives game footage from a live streaming service. The received video data is a real-time video stream showing the specific situation of the game. The device streams this video data to a server, which analyzes it in real time using video analysis algorithms. The analysis results include player positions, actions (e.g., dribbling, passing, shooting), and scoring status.
[1915] Step 3:
[1916] The device analyzes the user's facial expressions and vocal tone using an emotion engine. The analyzed emotion data indicates the user's real-time emotional state (e.g., excitement, joy, sadness, etc.). The device transmits this emotion data to the server, which then integrates the received emotion data into the user profile.
[1917] Step 4:
[1918] The server generates a prompt sentence based on the user's sports knowledge level, the team they are rooting for, and their emotional state. This prompt sentence contains prior information to be input into the generative AI model. For example, a prompt sentence of the form "The user's knowledge level is beginner level, and the team they are rooting for is Team A. Their current emotional state is excited. The analysis results of the game are as follows: Team A has scored a goal!" is generated.
[1919] Step 5:
[1920] The server inputs the generated prompt into a generative AI model to generate a commentary appropriate for the user. The generative AI model is an artificial intelligence system that generates commentary based on the prompt. The content of the generated commentary is adjusted according to the user's knowledge level and emotional state. As a result, a commentary appropriate for the user is obtained.
[1921] Step 6:
[1922] The server converts the generated commentary into audio data using a text-to-speech (TTS) engine, which then formats the audio data in a user-friendly format. Users can listen to this audio data in real time while watching the game.
[1923] Step 7:
[1924] The terminal plays the audio data received from the server through a playback device (e.g., speaker, headphones). This allows the user to listen to commentary customized according to the situation of the game. For example, if the user is excited, commentary such as "Great goal, Team A is in the lead!" is provided in an excited tone.
[1925] In this way, users can experience a personalized play-by-play commentary based on their level of sports knowledge, their favorite team, and their real-time emotional state.
[1926] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1927] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1928] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1929] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1930] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1931] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1932] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1933] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1934] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1935] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1936] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1937] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1938] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1939] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1940] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1941] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1942] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1943] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1944] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1945] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1946] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1947] The following is further disclosed regarding the above embodiment.
[1948] (Claim 1)
[1949] A means for inputting the user's sports knowledge level and the sport they support;
[1950] A means of analyzing game footage in real time to recognize players' positions, movements, and scoring situations,
[1951] means for generating commentary according to a user profile based on the analysis results;
[1952] means for converting the commentary into voice data using a text-to-speech synthesis engine;
[1953] a means for providing said audio data to a user.
[1954] (Claim 2)
[1955] 10. The system of claim 1, further comprising a terminal that receives the game video and streams it to a server.
[1956] (Claim 3)
[1957] The system according to claim 1, wherein commentary text of different levels is generated according to a user's profile based on the analysis result.
[1958] "Example 1"
[1959] (Claim 1)
[1960] A means for inputting the user's sports knowledge level and the sport they support;
[1961] means for transmitting the input information to a server;
[1962] means for storing the input information in a database;
[1963] means for receiving game footage and streaming the footage to a server;
[1964] A means of analyzing game footage in real time to recognize players' positions, movements, and scoring situations,
[1965] means for generating commentary according to a user's profile based on the analysis results, the user's sports knowledge level, and the sport the user supports;
[1966] A means for converting the generated commentary into voice data using a text-to-speech synthesis engine;
[1967] The system includes a means for transmitting the audio data from the server to the terminal, and for the terminal to provide the audio data to the user through a playback device.
[1968] (Claim 2)
[1969] 10. The system of claim 1, including a terminal that receives game footage and streams it to a server.
[1970] (Claim 3)
[1971] The system according to claim 1, wherein different levels of commentary are generated based on the analysis results, the user's sports knowledge level, and the target of support.
[1972] "Application Example 1"
[1973] (Claim 1)
[1974] A means for inputting the user's sports knowledge level and the sport they support;
[1975] A means of analyzing game footage in real time to recognize players' positions, movements, and scoring situations,
[1976] means for generating commentary according to a user profile based on the analysis results;
[1977] means for converting the commentary into voice data using a text-to-speech synthesis engine;
[1978] means for providing said audio data to a user;
[1979] means for receiving live streaming video of the match and transmitting it to a server;
[1980] a means for generating appropriate commentary from the prompt using a generative AI model; and
[1981] The system includes a means for transmitting the generated voice data to a user's terminal in real time and playing it back.
[1982] (Claim 2)
[1983] 10. The system of claim 1, further comprising a terminal that receives the game video and streams it to a server.
[1984] (Claim 3)
[1985] The system according to claim 1, wherein commentary text of different levels is generated according to a user's profile based on the analysis result.
[1986] "Example 2: Combining Emotion Engines"
[1987] (Claim 1)
[1988] A means for inputting the user's sports knowledge level and the sport they support;
[1989] A means of analyzing game footage in real time to recognize players' positions, movements, and scoring situations,
[1990] means for analyzing a user's facial expression and vocal tone to recognize the user's emotional state;
[1991] means for generating commentary according to a user profile based on the analysis result and the user's emotional state;
[1992] means for converting the commentary into voice data using a text-to-speech synthesis engine;
[1993] a means for providing said audio data to a user.
[1994] (Claim 2)
[1995] 10. The system of claim 1, further comprising a terminal that receives the game video and streams it to a server.
[1996] (Claim 3)
[1997] The system of claim 1, wherein different levels of commentary are generated according to a user's profile based on the analysis results and emotional state.
[1998] "Application example 2 when combining emotion engines"
[1999] (Claim 1)
[2000] A means for inputting the user's sports knowledge level and the sport they support;
[2001] A means of analyzing game footage in real time to recognize players' positions, movements, and scoring situations,
[2002] means for recognizing the emotional state of the user by analyzing the user's facial expressions and tone of voice;
[2003] means for generating a prompt sentence based on the analysis result and the user's emotional state;
[2004] a means for inputting the prompt sentence into a generative AI model to generate a commentary sentence according to a user's profile and emotional state;
[2005] means for converting the commentary into voice data using a text-to-speech synthesis engine;
[2006] a means for providing said audio data to a user.
[2007] (Claim 2)
[2008] 10. The system of claim 1, further comprising a terminal that receives the game video and streams it to a server.
[2009] (Claim 3)
[2010] The system according to claim 1, wherein different levels of commentary are generated according to a user profile based on the analysis results and the user's emotional state. [Explanation of symbols]
[2011] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. A means for inputting the user's sports knowledge level and the sport they support; A means of analyzing game footage in real time to recognize players' positions, movements, and scoring situations, means for generating commentary according to a user profile based on the analysis results; means for converting the commentary into voice data using a text-to-speech synthesis engine; a means for providing said audio data to a user.
2. The system of claim 1 further comprising a terminal that receives the game video and streams it to a server.
3. The system according to claim 1 , wherein different levels of commentary are generated according to a user profile based on the analysis results.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A