System
The system automates soccer match footage editing and commentary generation, addressing the challenge of enjoying and analyzing games by reducing filming and editing effort, and providing professional-quality videos through cloud-based processing.
Patent Information
- Application Number
- JP2024130351
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-08-06
- Publication Date
- 2026-02-19
AI Technical Summary
Cameras recording soccer matches face challenges in providing an enjoyable visual experience during the game, and editing the footage is time-consuming and laborious, with limited accessibility for families and individuals to create review and analysis tools.
A system that uploads video from fixed cameras to the cloud, detects the ball and players, automatically extracts important scenes, generates commentary, and provides edited videos to users.
Reduces the effort required for filming and editing, enabling users to easily create professional-quality videos with commentary, accessible through cloud storage and viewing.
Smart Images

Figure 2026028053000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] When recording soccer match footage, the cameraman faces the problem of being unable to enjoy the game visually. Editing the footage is also time-consuming and laborious, and there is a lack of editing services that are easily accessible to many families and individuals. Therefore, there is a need to meet the needs of people who want to review and analyze the game while watching it later. [Means for solving the problem]
[0005] The present invention solves the above problem by providing a system that includes a means for uploading video captured by a fixed camera to the cloud, a means for receiving the uploaded video and detecting the ball and players, a means for automatically extracting important scenes based on the detected information, a means for adding generated commentary to the video, and a means for providing the edited video to the user.
[0006] A "fixed camera" is a device for capturing images from a fixed position.
[0007] The "cloud" is a collection of remote servers for storing, managing, and processing data over the Internet.
[0008] "Upload" is an operation for sending data from a user's terminal to a server.
[0009] "Means for detecting the ball and players" refers to algorithms that use video analysis technology to identify the location of the ball and players within the video.
[0010] "Means for automatically extracting important scenes" refers to technology that automatically selects noteworthy parts of a game, such as scoring scenes or decisive plays, from video footage.
[0011] The "means for generating commentary" is a technology that uses artificial intelligence technology to automatically create commentary content in the style of a specific commentator.
[0012] The "means for providing edited video to users" is a service that allows users to access, view, or download edited video stored on the cloud. [Brief explanation of the drawings]
[0013] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2]1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0014] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0015] First, the terms used in the following description will be explained.
[0016] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0017] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0018] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0019] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0020] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0021] [First embodiment]
[0022] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0023] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0024] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0025] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0026] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0027] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0028] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0029] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0030] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0031] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0032] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0033] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0034] This invention is a system that automatically edits soccer match footage on the cloud and provides videos with commentary, eliminating the need for users to record and edit the footage while watching the match.
[0035] System configuration
[0036] The system includes the following main components:
[0037] 1. Photography Method:
[0038] Users set up fixed cameras to film the entire soccer match. The fixed cameras are fixed on the sidelines or behind the goals, and use wide-angle lenses to capture the entire match.
[0039] 2. Upload method:
[0040] Users upload the footage they have taken to the cloud platform, and the video files are sent to the cloud server via the internet.
[0041] 3. Receiving and preprocessing means:
[0042] The server receives the video uploaded to the cloud, checks the format and quality of the video, and performs any necessary pre-processing (e.g., adjusting the resolution or converting the frame rate).
[0043] 4. Analysis method:
[0044] The AI analysis system on the server analyzes the video data frame by frame. The AI uses image processing algorithms to detect the position and movement of the ball and players in each frame. Based on this information, the system tracks the ball and players in real time.
[0045] 5. Extraction means:
[0046] The server analyzes the tracking data to detect when specific events (e.g., goals, saves, fouls) occur, and automatically extracts key moments based on these events to generate highlight clips.
[0047] 6. Explanation Generation Method:
[0048] The server uses a generative AI model to generate commentary in the style of a specific commentator, synthesizes the generated commentary to create a commentary audio, and integrates the commentary audio and subtitles into the video to create an edited highlight video.
[0049] 7. Means of provision:
[0050] The server stores the edited video in cloud storage, and users can access the cloud platform to view or download the edited video.
[0051] Program processing and specific examples
[0052] Natural language explanation of the process
[0053] Users film the match with fixed cameras and upload the footage to the cloud after the match ends. The server preprocesses the received footage and uses an AI analysis system to detect the ball and players. It then extracts key scenes, generates commentary using a generative AI model, and adds commentary audio and subtitles to the video. The edited footage is then stored in the cloud and made available for users to watch or download.
[0054] Specific examples
[0055] For example, suppose a user films their child's soccer game and uploads the footage to the cloud. The server receives the footage and performs preprocessing, and the AI analysis system detects and tracks the movement of the ball and players. Goals and important plays are detected, and highlights are automatically generated based on them. The generative AI model generates commentary such as, "Player A made a decisive pass here, and Player B scored a shot!" and adds it to the video as audio and subtitles. Finally, the edited footage is stored in the cloud, where the user can enjoy it with their family.
[0056] In this way, the present invention is a system that can significantly reduce the effort required for filming and automatically provide professional-quality video.
[0057] The processing flow will be explained below.
[0058] Step 1:
[0059] A user films a soccer match with a fixed camera. The camera is set up and recording begins. The video is set to capture the entire match, using a wide-angle lens, etc.
[0060] Step 2:
[0061] The video footage taken by the user is uploaded to the cloud server. After the match, the user connects the camera's memory card to a computer and uploads the video files to the cloud platform.
[0062] Step 3:
[0063] The server receives the video uploaded to the cloud, checks the format and quality of the received video file, and performs pre-processing such as adjusting the resolution and converting the frame rate as necessary.
[0064] Step 4:
[0065] The AI analysis system on the server analyzes the received video data frame by frame, detecting the ball and players in the video and saving their positional information as tracking data.
[0066] Step 5:
[0067] The server automatically extracts important scenes (goals, fouls, saves, etc.) based on the tracking data, and compiles the extracted scenes to generate a highlight clip.
[0068] Step 6:
[0069] The server uses the generative AI model to generate commentary in the style of a specific commentator, and then performs voice synthesis based on the generated commentary to create the commentary audio.
[0070] Step 7:
[0071] The server integrates audio commentary and subtitles into the video, and edits and completes the highlight video with commentary.
[0072] Step 8:
[0073] The server saves the edited video to cloud storage and notifies the user that the video is ready.
[0074] Step 9:
[0075] Users log in to the cloud platform to view or download the edited footage, and share it with family and team members.
[0076] Example 1
[0077] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0078] Previously, shooting, editing, and creating a video with commentary of a soccer game required a great deal of time and effort. Especially for amateur and family games, where professional video editing skills and commentators are not required, there was a demand for a method to easily create high-quality video with commentary, but no system existed that could achieve this. Furthermore, manual editing and the use of specialized software required specialized knowledge, making it a significant burden for average users.
[0079] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0080] In this invention, the server includes a means for uploading video to the cloud, a means for preprocessing the received video, a means for detecting the ball and players using an image processing algorithm, a means for automatically extracting important scenes and generating highlight clips, a means for creating commentary audio using a generative AI model for generating commentary text and speech synthesis technology, and a means for integrating the commentary audio and subtitles into the video, and a means for saving the edited video in cloud storage and providing it to users, thereby enabling users to easily create highlight videos with professional quality commentary and easily watch and share them.
[0081] A "fixed camera" is a camera that is fixed in a certain position and continuously captures a specific range or scene.
[0082] The "cloud" is a service platform for storing and processing data on remote servers on the Internet.
[0083] "Uploading" is the act of transferring data stored on a local device to a remote server over the Internet.
[0084] "Preprocessing" refers to the initial processing performed on received data, primarily to improve data analysis and algorithm performance.
[0085] An "image processing algorithm" is a computational method for extracting and analyzing specific features or information from image data.
[0086] "Ball and player detection" refers to identifying the position of the ball and players in each frame of the video and tracking their movements.
[0087] A "key moment" is a particularly noteworthy event in a soccer match (e.g., a goal, a save, a foul, etc.).
[0088] A "highlight clip" is a shortened, edited video clip that selects particularly important scenes from the entire game.
[0089] A "generative AI model" is an artificial intelligence model that generates natural-looking explanatory text and content based on given data and prompt text.
[0090] "Speech synthesis technology" is a technology that converts text data into actual speech.
[0091] "Commentary audio" is audio data that has been converted using speech synthesis technology from explanatory text created by a generative AI model.
[0092] "Subtitles" are textual information that is visually displayed within a video and complements the audio content.
[0093] "Cloud storage" is a data storage service on the Internet that stores data on a remote server and allows it to be accessed as needed.
[0094] "Providing to the user" means supplying the edited video in a form that the user can access.
[0095] This invention is a system that automatically edits soccer match footage on the cloud and provides videos with commentary, eliminating the need for users to record and edit the footage while watching the match.
[0096] System configuration
[0097] The system includes the following main components:
[0098] 1. Photography Method:
[0099] Users set up fixed cameras to film the entire soccer match. The fixed cameras are fixed on the sidelines or behind the goals, and use wide-angle lenses to capture the entire match.
[0100] 2. Upload method:
[0101] Users upload the footage they have taken to a cloud platform. The video files are sent to a cloud server via the Internet. The cloud platform uses commonly used storage services (e.g., Google Drive or Dropbox).
[0102] 3. Receiving and preprocessing means:
[0103] The server receives the video uploaded to the cloud. After receiving it, it checks the format and quality of the video and performs any necessary pre-processing (e.g., adjusting the resolution or converting the frame rate) using a video processing library (e.g., FFmpeg).
[0104] 4. Analysis method:
[0105] The AI analysis system on the server analyzes the video data frame by frame. The AI uses an image processing algorithm (e.g., YOLO (You Only Look Once)) to detect the position and movement of the ball and players in each frame. Based on this detected information, the ball and players are tracked in real time.
[0106] 5. Extraction means:
[0107] The server analyzes the tracking data to detect when specific events (e.g., goals, saves, fouls) occur, and automatically extracts key moments based on these events to generate highlight clips.
[0108] 6. Explanation Generation Method:
[0109] The server uses a generative AI model to generate commentary in the style of a specific commentator. Using a prompt as input, the generative AI model generates appropriate commentary. For example, a prompt such as "Player A made a pass and Player B scored a goal" is used. Based on the generated commentary, a voice synthesis technology (e.g., Google Text-to-Speech API) is used to create a commentary voice.
[0110] 7. Integration Methods:
[0111] The server integrates the commentary audio and subtitles into the video to create an edited highlight video. A video editing library (e.g., OpenCV or MoviePy) is used to integrate the audio file with the video and overlay the subtitle text.
[0112] 8. Means of provision:
[0113] The server stores the edited video in cloud storage (e.g., AWS S3), and users can access the cloud platform to view or download the edited video.
[0114] Specific examples
[0115] For example, suppose a user films their child's soccer game and uploads the footage to the cloud. The server receives the footage and performs preprocessing, and the AI analysis system detects and tracks the movement of the ball and players. Goals and important plays are detected, and highlights are automatically generated based on them. The generative AI model generates commentary such as, "Player A made a decisive pass here, and Player B scored a shot!" and adds it to the video as audio and subtitles. Finally, the edited footage is stored in the cloud, where the user can enjoy it with their family.
[0116] In this way, the present invention is a system that can significantly reduce the effort required for filming and automatically provide professional-quality video.
[0117] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0118] Step 1:
[0119] The user sets up a fixed camera and captures footage of the match.
[0120] Specifically, the user fixes a fixed camera on the sideline or behind the goal and adjusts the camera settings, for example, using a wide-angle lens to capture the entire game in the frame.
[0121] Input: The physical camera equipment and the soccer match to be filmed.
[0122] Output: Footage files of the captured match.
[0123] Step 2:
[0124] The user uploads the footage they have taken to a cloud platform.
[0125] Users use their home computer or smartphone to access the designated upload form on the cloud platform, select and submit the video file.
[0126] Input: The captured video file.
[0127] Output: Video files uploaded to the cloud platform.
[0128] Step 3:
[0129] The server receives the video uploaded to the cloud platform and performs pre-processing.
[0130] The server uses a video processing library (e.g., FFmpeg) to check the format and quality of the received video file and make any necessary resolution adjustments and frame rate conversions, for example, converting 4K video to 1080p and changing the frame rate from 60 fps to 30 fps.
[0131] Input: Video files on the cloud.
[0132] Output: Pre-processed video files.
[0133] Step 4:
[0134] The server uses an AI analysis system to analyze the pre-processed video data frame by frame.
[0135] The server uses an image processing algorithm (e.g., YOLO) to detect the positions of the ball and players in each frame. Specifically, it obtains the coordinate data of the bounding boxes of the ball and players included in each frame.
[0136] Input: Preprocessed video files.
[0137] Output: Ball and player position data.
[0138] Step 5:
[0139] The server analyzes the tracking data to detect when a particular event (e.g., goal, save, foul) occurs.
[0140] The server records the timestamps of the frames corresponding to these events and automatically extracts important scenes, such as five seconds before and after a goal.
[0141] Input: Ball and player position data.
[0142] Output: timestamps and footage clips of key scenes.
[0143] Step 6:
[0144] The server uses a generative AI model to generate commentary in the style of a specific commentator.
[0145] The server inputs a prompt, and the generative AI model generates an appropriate commentary. For example, a prompt such as "Player A passes the ball, and Player B scores a goal" is used.
[0146] Input: prompt text and video clips of key scenes.
[0147] Output: The generated description.
[0148] Step 7:
[0149] The server creates a commentary audio based on the generated commentary text using speech synthesis technology (e.g., Google Text-to-Speech API).
[0150] The commentary is converted into an audio file and is also prepared as subtitle text. In detail, the server sends the commentary to the speech synthesis API and obtains the audio file.
[0151] Input: The generated description.
[0152] Output: Descriptive audio file and subtitle text.
[0153] Step 8:
[0154] The server integrates audio commentary and subtitles into the video clips to create an edited highlight video.
[0155] Use a video editing library (e.g., OpenCV or MoviePy) to integrate the audio file and subtitle text into the video clip. Specifically, the audio description is overlaid on the video and the subtitle text is displayed in the appropriate position.
[0156] Input: Descriptive audio files, subtitle text, and video clips of key scenes.
[0157] Output: Edited highlight video.
[0158] Step 9:
[0159] The server saves the completed highlight video to cloud storage (e.g. AWS S3).
[0160] Generate a link that allows users to access the saved video and provide that link to the user. For example, generate a URL like "https: / / cloudstorage.example.com / highlight_video.mp4".
[0161] Input: edited highlight video.
[0162] Output: Highlight video URL on cloud storage.
[0163] Step 10:
[0164] The user accesses the cloud platform and clicks on the provided link to view the edited footage.
[0165] Users can use a browser or a dedicated app to play or download videos from cloud storage.
[0166] Input: Highlight video URL on cloud storage.
[0167] Output: Highlights viewed or downloaded.
[0168] Through the above process, the system automatically edits soccer match footage shot by the user and provides a professional highlight video with commentary.
[0169] (Application example 1)
[0170] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0171] Shooting soccer game footage, editing the footage, and creating a highlight video with commentary has traditionally required a great deal of effort. Editing the footage and creating the commentary, in particular, requires time and effort, making it difficult to easily achieve professional-quality results. The objective of this invention is to solve this problem and provide a system that enables users to easily create high-quality highlight videos with commentary.
[0172] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0173] In this invention, the server includes means for uploading video captured by a fixed camera to the cloud, means for receiving the uploaded video and detecting the ball and players, means for automatically extracting key scenes based on the detected information, means for adding automatically generated commentary to the video using a generative AI model, means for integrating commentary audio and subtitles to complete the edited video, and means for saving the edited video in cloud storage and enabling users to view or share it. This allows users to easily generate, view, and share professional-quality highlight videos with commentary simply by uploading their game video to the cloud.
[0174] A "fixed camera" is a camera that captures video from a fixed position.
[0175] The "cloud" is a collection of servers for storing, managing, and processing data over the Internet.
[0176] "Uploading" is the act of sending data from a local device to a cloud server.
[0177] "Reception" refers to the act of the cloud server receiving data sent from the user.
[0178] "Detection" is the process of identifying a specific object (the ball or a player) through video analysis.
[0179] A "Key Moment" is a specific key moment during a match (e.g., a goal, save, foul).
[0180] A "generative AI model" is an artificial intelligence model that uses machine learning algorithms to generate content such as text and images.
[0181] "Explanation" is written or audio that explains a particular event.
[0182] The "explanatory voice" is the generated explanatory text output as voice through voice synthesis.
[0183] "Subtitles" are textual information displayed on video.
[0184] "Editing" is the process of adding audio commentary and subtitles to the footage to complete the final video.
[0185] "Cloud storage" is a storage service for storing data on the cloud.
[0186] "Viewing" refers to the act of a user playing and watching a video.
[0187] "Sharing" refers to the act of a user sharing video or data with other people.
[0188] This invention is a system that automatically edits soccer match footage taken with fixed cameras and provides videos with commentary, eliminating the need for users to record and edit the footage while watching the match.
[0189] System configuration
[0190] The system includes the following main components:
[0191] 1. Fixed camera
[0192] Users can set up fixed cameras on the sidelines or behind the goals and use wide-angle lenses to capture the entire game.
[0193] 2. Cloud Upload
[0194] Users upload the footage they have taken to the cloud platform, and the video files are sent to the cloud server via the internet.
[0195] 3. Reception and Preprocessing
[0196] The server receives the video uploaded to the cloud. After receiving the video, it checks the format and quality of the video and performs any necessary preprocessing (e.g., adjusting the resolution and converting the frame rate). For preprocessing, it uses an open-source video processing library such as OpenCV.
[0197] 4. AI analysis
[0198] The AI analysis system on the server analyzes the video data frame by frame, and uses OpenCV and deep learning models to detect the position of the ball and the positions and movements of players in each frame.
[0199] 5. Extracting important scenes
[0200] The server analyzes the tracking data to detect when specific events (e.g., goals, saves, fouls) occur, and automatically extracts key moments based on these events to generate highlight clips.
[0201] 6. Explanation Generation
[0202] The server uses a generative AI model (e.g., GPT-4) to generate commentary in the style of a specific commentator. It then uses voice synthesis based on the generated commentary to create the commentary audio. For voice synthesis, it uses a service like Amazon Polly.
[0203] 7. Video Editing
[0204] The commentary audio and subtitles are integrated into the video, and video editing libraries such as FFmpeg are used for video editing. This completes the edited highlight video.
[0205] 8. Cloud Storage
[0206] The edited footage is saved to cloud storage, allowing users to view or share it.
[0207] Specific examples
[0208] For example, a user can film their child's soccer game with their smartphone and upload the footage to the cloud. The server receives the footage and performs pre-processing. The AI analysis system detects the movement of the ball and players and extracts goals and other important plays. Based on this, the generative AI model generates commentary such as, "Player A made a decisive pass here, and Player B scored a shot!" and adds it to the video as audio and subtitles. Finally, the edited video is stored in the cloud, and users can share it with family and friends.
[0209] Prompt Sentence Examples
[0210] For example, when generating explanatory text using a generative AI model, the following prompt is used:
[0211] "I uploaded a video of my child's soccer game. Analyze the ball position and player movements to detect goals and important plays. Based on that, generate commentary such as, 'Player A made the decisive pass, and Player B scored the shot!' and add it to the video as audio and subtitles."
[0212] In this way, the present invention is a system that can significantly reduce the effort required for filming and automatically provide professional-quality video.
[0213] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0214] Step 1:
[0215] A user films a soccer match with a fixed camera. The fixed camera uses a wide-angle lens to capture the entire game, and is fixed on the sidelines or behind the goal. The input is the video data captured by the camera, and the output is a video file stored in local storage.
[0216] Step 2:
[0217] Users upload the videos they have taken from their smartphones to the cloud platform. The uploaded videos are then sent to the cloud server via the Internet. The input is the video file stored on the smartphone, and the output is the video data on the cloud storage.
[0218] Step 3:
[0219] The server receives the video uploaded to the cloud and performs preprocessing. This includes adjusting the resolution and converting the frame rate, and uses a video processing library such as OpenCV. The input is the video data stored in the cloud storage, and the output is the preprocessed video data converted into an analyzable format.
[0220] Step 4:
[0221] The AI analysis system in the server analyzes the preprocessed video frame by frame to detect the position and movement of the ball and players. This analysis uses OpenCV and deep learning models. The input is the preprocessed video data, and the output is tracking data including the position information of the ball and players.
[0222] Step 5:
[0223] The server automatically extracts important scenes from the tracking data. Specifically, it detects events such as goals, saves, and fouls and extracts those moments as highlight clips. The input is the tracking data, and the output is a list of important scenes and their clip data.
[0224] Step 6:
[0225] The server uses a generative AI model (e.g., GPT-4) to create automatically generated commentary. The generated commentary is converted into speech using speech synthesis technology (e.g., Amazon Polly). The input is a list of important scenes and prompts, and the output is commentary audio and subtitle data.
[0226] Step 7:
[0227] The server integrates the commentary audio and subtitles into the video to create an edited highlight video. This editing is done using a video editing library such as FFmpeg. The input is the clip data, commentary audio, and subtitle data, and the output is the edited video.
[0228] Step 8:
[0229] The server saves the completed edited video in cloud storage. Users can access the cloud platform to view and share the video. The input is the edited video data, and the output is the video data saved in cloud storage.
[0230] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0231] This invention is a system that automatically edits soccer match footage on the cloud and provides videos with commentary using an emotion engine. This system saves users the trouble of recording and editing the game while watching it. It also adjusts the commentary content based on the user's emotions, providing a more personalized video experience.
[0232] System configuration
[0233] The system includes the following main components:
[0234] 1. Photography Method:
[0235] Users set up fixed cameras to film the entire soccer match. The fixed cameras are fixed on the sidelines or behind the goals, and use wide-angle lenses to capture the entire match.
[0236] 2. Upload method:
[0237] Users upload the footage they have taken to the cloud platform, and the video files are sent to the cloud server via the internet.
[0238] 3. Receiving and preprocessing means:
[0239] The server receives the video uploaded to the cloud, checks the format and quality of the video, and performs any necessary pre-processing (e.g., adjusting the resolution or converting the frame rate).
[0240] 4. Analysis method:
[0241] The AI analysis system on the server analyzes the received video data frame by frame, detecting the ball and players in the video and saving their positional information as tracking data.
[0242] 5. Extraction means:
[0243] The server automatically extracts important scenes (goals, fouls, saves, etc.) based on the tracking data, and compiles the extracted scenes to generate highlight clips.
[0244] 6. Emotion Engine:
[0245] The server is equipped with an emotion engine that recognizes the user's emotions. It analyzes the user's emotion data (e.g., facial expression recognition and biometric data) and grasps the user's emotional state in real time.
[0246] 7. Explanation Generation Method:
[0247] The server uses a generative AI model to generate commentary in the style of a specific commentator, adjusts the commentary content based on the user's emotions recognized by the emotion engine, and performs voice synthesis based on the generated commentary to create a commentary voice.
[0248] 8. Means of provision:
[0249] The server stores the edited video in cloud storage, and users can access the cloud platform to view or download the edited video.
[0250] Program processing and specific examples
[0251] Natural language explanation of the process
[0252] Users film the match with fixed cameras and upload the footage to the cloud after the match ends. The server preprocesses the received footage and uses an AI analysis system to detect the ball and players. An emotion engine then collects and analyzes the user's emotions in real time. Key scenes are then extracted, and a generative AI model generates commentary based on the emotion data and integrates it into the video. Finally, the edited video is stored in the cloud and made available for users to watch or download.
[0253] Specific examples
[0254] For example, suppose a user films their child's soccer game and uploads the footage to the cloud. The server receives the footage and performs preprocessing, while the AI analysis system detects and tracks the movement of the ball and players. At the same time, while the user is watching the game, the emotion engine analyzes the user's facial expressions, heart rate, and other data to collect emotional data. Goals and important plays are detected, and highlights are automatically generated based on them. The generative AI model generates commentary such as, "Player A made a decisive pass here, and Player B scored a shot!" If it determines that the user is happy, the commentary will be written in a positive tone. Finally, the edited video is saved in the cloud, where the user can enjoy it with their family.
[0255] In this way, the present invention not only significantly reduces the effort required for filming and automatically provides professional-quality video, but also provides a more personalized video experience based on the user's emotions.
[0256] The processing flow will be explained below.
[0257] Step 1:
[0258] A user films a soccer match with a fixed camera. The camera is set up and recording begins. The video is set to capture the entire match, using a wide-angle lens, etc.
[0259] Step 2:
[0260] The video footage taken by the user is uploaded to the cloud server. After the match, the user connects the camera's memory card to a computer and uploads the video files to the cloud platform.
[0261] Step 3:
[0262] The server receives the video uploaded to the cloud, checks the format and quality of the received video file, and performs pre-processing such as adjusting the resolution and converting the frame rate as necessary.
[0263] Step 4:
[0264] The AI analysis system on the server analyzes the received video data frame by frame, detecting the ball and players in the video and saving their positional information as tracking data.
[0265] Step 5:
[0266] The server automatically extracts important scenes (goals, fouls, saves, etc.) based on the tracking data, and compiles the extracted scenes to generate a highlight clip.
[0267] Step 6:
[0268] As users watch a game, the emotion engine collects their facial expressions and heart rate in real time to generate emotion data, which is then sent to a cloud server.
[0269] Step 7:
[0270] The server receives and analyzes the emotion data sent from the emotion engine to understand the user's emotional state (happiness, excitement, sadness, etc.).
[0271] Step 8:
[0272] The server uses a generative AI model to generate commentary in the style of a specific commentator, adjusting the commentary based on emotion data, for example, using positive expressions if the user is happy.
[0273] Step 9:
[0274] The server integrates audio commentary and subtitles into the video, and edits and completes the highlight video with commentary.
[0275] Step 10:
[0276] The server saves the edited video to cloud storage and notifies the user that the video is ready.
[0277] Step 11:
[0278] Users log in to the cloud platform to view or download the edited footage, and share it with family and team members.
[0279] Example 2
[0280] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0281] Conventional soccer match filming and editing systems require lengthy editing work after filming, and adding commentary is labor-intensive. Furthermore, they are unable to reflect the user's emotional state while watching the video, making it difficult to provide a personalized viewing experience. There is a need for a system that can solve these problems and easily provide high-quality, personalized video.
[0282] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0283] In this invention, the server includes means for uploading video captured by a fixed camera to the cloud, means for receiving the uploaded video and pre-processing the video, means for detecting the ball and players from the received video, means for automatically extracting important scenes based on the detected information, means for collecting user emotion data, means for analyzing the collected emotion data and generating commentary using a generative AI model, means for adding the generated commentary to the video, and means for providing the edited video to the user, thereby enabling users to watch high-quality video containing match highlights and personalized commentary without any hassle.
[0284] A "fixed camera" is a camera that is fixed at a specific location and captures a wide range of images using a wide-angle lens.
[0285] "Cloud" is a general term for an online platform that provides data storage and computing resources by accessing remote servers via the Internet.
[0286] "Uploading means" refers to the method or technology used to transfer digital data from a local device to a remote server, such as cloud storage.
[0287] "Means of receiving" refers to the method or technology by which the remote server obtains data uploaded to the cloud.
[0288] "Preprocessing means" refers to processing carried out in advance, such as converting data format or adjusting quality, to ensure smooth subsequent analysis and editing processes.
[0289] "Detection means" refers to methods or technologies that automatically recognize specific objects (e.g., balls or players) from video data and track their location information.
[0290] "Means for automatically extracting important scenes" refers to methods and technologies that use detected information to identify specific events such as goals and fouls and extract them as highlights from the video.
[0291] "Means for collecting emotional data" refers to methods and technologies for acquiring a user's facial expressions and biometric data in real time and analyzing their emotional state.
[0292] A "generative AI model" is a technology that uses a pre-trained artificial intelligence model to generate explanatory text in natural language from given data and prompts.
[0293] "Means for adding commentary" refers to methods or techniques for integrating the generated commentary text and audio into the video data and providing it as the final edited video.
[0294] "Means for providing edited video" refers to the methods and technologies for storing the final edited video in cloud storage and allowing users to access, view, or download it.
[0295] This system automatically edits soccer match footage on the cloud and provides videos with commentary using an emotion engine. This allows users to concentrate on watching the game, while automating the editing and commentary tasks. It also provides a personalized video experience based on the user's emotions.
[0296] First, the user sets up a fixed camera to film a soccer match. The fixed camera uses a wide-angle lens and is installed on the sidelines or behind the goal. After the match ends, the user uploads the filmed video files to a cloud platform. This can be done using cloud storage services such as Google Drive or Dropbox.
[0297] The server then receives the uploaded video on the cloud, checks the format and quality of the video file, and performs any necessary pre-processing, such as converting AVI video to MP4, reducing the resolution from 1080p to 720p, and converting the frame rate from 30fps to 24fps.
[0298] Once preprocessing is complete, the server uses an AI analysis system to analyze the received video data frame by frame. The AI analysis system uses deep learning models and libraries such as OpenCV to detect the ball and players in the video and save their location information as tracking data. This allows the progress of the game and the movements of the players to be understood.
[0299] The server then automatically extracts important moments (e.g., goals, fouls, saves, etc.) from the tracking data. In this step, AI identifies specific events and cuts them out into highlight videos.
[0300] Furthermore, the server uses an emotion engine to collect user emotional data, including biometric data such as facial expressions and heart rate, and analyzes it in real time, allowing the server to understand the user's emotional state during a match.
[0301] The server then uses a generative AI model to generate commentary in the style of a particular commentator. The generative AI model is given the following prompt:
[0302] "Generate commentary on goal scenes."
[0303] "Please provide an inspiring commentary based on the emotional data."
[0304] "Describe a highlight scene of a key play."
[0305] The generated commentary text is used to synthesize speech, creating a commentary voice, which adds a commentary tailored to the user's emotions to the video.
[0306] Finally, the server integrates the generated commentary audio into the highlight video. Specifically, it uses video editing software such as FFmpeg to combine the highlight footage and commentary audio into a single file. The edited video file is then stored in cloud storage, and users can access the cloud platform to watch or download it.
[0307] For example, consider the case where a user films a child's soccer game and uploads the video to the cloud. The server receives the video and performs preprocessing, and the AI analysis system detects and tracks the movement of the ball and players. Then, while the user is watching the game, the emotion engine analyzes the user's facial expressions, heart rate, etc., and collects emotional data. Based on this data, the generative AI model generates commentary, providing commentary in positive terms, allowing the user to enjoy watching the highlights of the game.
[0308] In this way, the present invention not only significantly reduces the effort required for filming and editing, and automatically provides professional-quality video, but also enables a personally tailored video experience that reflects the user's emotions.
[0309] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0310] Step 1:
[0311] The user sets up a fixed camera to film a soccer match. The fixed camera used for recording uses a wide-angle lens and is fixed to the sideline or behind the goal. After the match ends, the user transfers the filmed footage to a computer. The input data is the filmed video file, and the output data is the video file saved on the computer.
[0312] Step 2:
[0313] Users upload video files stored on their computers to a cloud platform. Specifically, they use cloud storage services such as Google Drive and Dropbox to upload files. The input data is the video file stored on their computer, and the output data is the video file stored on the cloud.
[0314] Step 3:
[0315] The server receives the video uploaded to the cloud. After receiving it, it checks the format and quality of the video file and performs any necessary pre-processing. Pre-processing includes converting AVI format video to MP4 format, reducing the resolution from 1080p to 720p, and converting the frame rate from 30fps to 24fps. The input data is the video file on the cloud, and the output data is the video file after pre-processing. Specific operations involve the use of video conversion software such as FFmpeg.
[0316] Step 4:
[0317] The server uses an AI analysis system to analyze the pre-processed video data frame by frame. The AI analysis system uses a deep learning model and libraries such as OpenCV to detect the ball and players in the video. It then saves this position information as tracking data. The input data is the pre-processed video file, and the output data is the analysis results, including the tracking data.
[0318] Step 5:
[0319] The server automatically extracts key moments (e.g., goals, fouls, saves, etc.) from the tracking data. This involves AI identifying specific events and extracting them into highlight videos. The input data is the tracking data, and the output data is video clips of key moments.
[0320] Step 6:
[0321] The server uses an emotion engine to collect user emotion data in real time. It also uses devices such as webcams and smartwatches to acquire biometric data such as the user's facial expressions and heart rate. The input data is the user's biometric data, and the output data is the analyzed emotion data.
[0322] Step 7:
[0323] The server uses a generative AI model to generate commentary in the style of a specific commentator. The generative AI model is given the following prompt:
[0324] "Generate commentary on goal scenes."
[0325] "Please provide an inspiring commentary based on the emotional data."
[0326] "Describe a highlight scene of a key play."
[0327] The input data is emotion data and prompt sentences, and the output data is the generated commentary. Specifically, a generative AI model such as GPT-3 is used.
[0328] Step 8:
[0329] The server synthesizes the commentary based on the generated commentary text and creates the commentary audio. The input data is the generated commentary text, and the output data is the commentary audio file. Specifically, speech synthesis software is used.
[0330] Step 9:
[0331] The server then integrates the generated commentary audio into the highlight video. Specifically, it uses video editing software such as FFmpeg to combine the highlight video and commentary audio into a single file. The input data is the commentary audio file and the video clips of key scenes, and the output data is the edited video file.
[0332] Step 10:
[0333] The server stores the edited video files in cloud storage, allowing users to access and watch or download them. The input data is the edited video files, and the output data is the video files stored on the cloud. Specifically, a cloud storage platform is used.
[0334] (Application example 2)
[0335] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0336] Conventional video editing systems require manual editing after shooting footage, which requires a great deal of time and effort. Furthermore, the generated commentary is standardized, making it difficult to provide a personalized experience that reflects each user's individual emotions and circumstances. This requires a lot of effort for users to record and edit events, and the viewing experience is not personalized enough.
[0337] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for uploading video captured by a fixed camera to the cloud, means for receiving the uploaded video and detecting objects and people, means for automatically extracting important scenes based on the detected information, generation AI model means for analyzing user emotion data and generating customized commentary based on the analysis results, means for adding the generated commentary to the video, and means for providing the edited video to the user. This eliminates the need for manual editing of the video and makes it possible to provide a personalized viewing experience based on the user's emotions.
[0338] A "fixed camera" is a camera that continuously captures a specific area from a fixed position.
[0339] The "cloud" refers to computing resources and data storage provided via the Internet.
[0340] "Uploading" is the act of sending local data to a remote server over the Internet or a network.
[0341] "Object" is a general term that refers to any concrete thing that can be recognized in an image.
[0342] "Person" is a general term that refers to a person recognized in the video.
[0343] "Detection" is the act of using specific algorithms to identify and track objects or people in a video.
[0344] "Key scenes" refer to moments or events that are particularly noteworthy in footage of sporting events or everyday life.
[0345] "Extraction" is the act of extracting a specific part from the whole data.
[0346] "Emotional data" refers to the emotional state obtained from the user's facial expressions, biometric data, etc.
[0347] A "generative AI model" is an algorithm that uses machine learning and natural language processing to automatically generate explanatory text and other content.
[0348] "Commentary" refers to text or audio that adds explanations or comments to the content of the video.
[0349] "Customization" is the act of tailoring content to a particular user or situation.
[0350] "Edited footage" refers to the final video file in which processing and modification have been applied to the original video data.
[0351] "Providing" refers to the act of making a video available for viewing or acquisition by a user through the cloud or other media.
[0352] The following system is used as an embodiment of the present invention: The main components of the system and their respective roles will be described below.
[0353] System Components
[0354] 1. Fixed camera:
[0355] Used by users to capture events, fixed cameras are fixed in specific locations and use wide-angle lenses to capture a wide range of footage.
[0356] 2. Cloud Platform:
[0357] Users upload footage taken with fixed cameras, and the cloud platform receives the video files via the internet and performs the analysis and editing processes described below.
[0358] 3. Receiving and Preprocessing Server:
[0359] The server receives the uploaded video data. It checks the format and quality of the video and performs any necessary preprocessing (adjusting the resolution and converting the frame rate).
[0360] 4. AI analysis system:
[0361] Video data is analyzed frame by frame, objects and people are detected in the video, and their location information is saved as tracking data.
[0362] 5. Sentiment Analysis System:
[0363] Analyze the user's emotional state. Collect user emotional data (facial recognition and biometric data) using sensors on smartphones and smart glasses, and analyze it in real time.
[0364] 6. Generative AI Models:
[0365] Based on the collected emotional data, a commentary in the style of a specific commentator is generated. Speech synthesis is performed based on the generated commentary to create a commentary voice. The software used here includes a natural language processing library and a speech synthesis module.
[0366] 7. Editing System:
[0367] Based on the analysis results and audio commentary, the video is edited to generate highlight clips. This editing system uses video editing libraries (e.g., OpenCV and MoviePy).
[0368] 8. Delivery system:
[0369] The edited video is stored in cloud storage, and users can view or download the edited video from the cloud platform.
[0370] A description of what the program does
[0371] In this system, a user first captures an event with a fixed camera and uploads the footage to a cloud platform. The cloud server receives the video data, checks its format and quality, and performs any necessary preprocessing. The AI analysis system then analyzes the received video data frame by frame to obtain information for detecting and tracking objects and people.
[0372] At the same time, devices such as smartphones and smart glasses are used to collect the user's emotional data (e.g., facial expressions, heart rate, etc.), which is then analyzed in real time by an emotion analysis system to determine the user's emotional state.
[0373] Based on the results of the sentiment analysis and AI analysis, the generative AI model generates commentary in the style of a specific commentator. For example, if it determines that the user is happy, a positive commentary will be selected. The generated commentary is converted into an audio file using text-to-speech software. These commentary voices are then integrated into existing footage, and the editing system generates a highlight clip.
[0374] Finally, the edited video is stored in cloud storage. Users can access the cloud platform and watch or download the generated personalized video. This process allows users to easily obtain high-quality edited video and provides a personalized experience based on their emotions.
[0375] Specific examples
[0376] For example, if a user films a family birthday party, the app uploads the video to the cloud, and the AI performs emotion analysis and video editing as follows: When the emotion analysis system determines that the user is happy at the climax of the party, the generative AI model generates a caption such as, "Here, the whole family has fun and cut the cake!"
[0377] Prompt Sentence Examples
[0378] Here is a video of a family birthday party. Highlight the most exciting scenes and generate commentary based on the emotion data.
[0379] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0380] Step 1:
[0381] A user captures an event with a fixed camera or a mobile information terminal. As input, the user obtains video data. As output, a video file is generated.
[0382] Step 2:
[0383] The user uploads the video they have taken to the cloud platform. The input is the video file they have taken, which the cloud server receives. The output is the video data stored on the cloud.
[0384] Step 3:
[0385] The server receives the uploaded video data and checks the format and quality. The input is the video data stored on the cloud server, and the server performs preprocessing such as adjusting the resolution and converting the frame rate. The output is the preprocessed video data.
[0386] Step 4:
[0387] The server's AI analysis system analyzes the preprocessed video data frame by frame to detect objects and people. The input is the preprocessed video data, and the server applies detection algorithms to identify objects and people, saving their location information as tracking data. The output is tracking data.
[0388] Step 5:
[0389] The device (smartphone or smart glasses) collects emotion data. The input is sensor data such as the user's facial expression and heart rate, and the device sends this data to the emotion analysis system. The output is the user's emotion data as an analysis result.
[0390] Step 6:
[0391] The server uses a generative AI model to generate an explanatory text based on the user's emotional data and the results of AI analysis. The inputs are emotional data and tracking data, and the generative AI model generates a prompt based on this data and then generates an explanatory text (e.g., "That was a great goal!" if the emotion is positive). The output is an explanatory text.
[0392] Step 7:
[0393] The server converts the generated commentary into an audio file using a speech synthesis module. The commentary is input, the server synthesizes the speech, and obtains the audio file. The audio file is obtained as output.
[0394] Step 8:
[0395] The server's editing system integrates the video and audio commentary to generate a highlight clip. The inputs are the pre-processed video data, tracking data, and audio commentary files, and the server uses a video editing library to integrate them. The output is an edited highlight video.
[0396] Step 9:
[0397] The server stores the edited footage in cloud storage for users to access. As input, there is the edited highlight footage, which the server stores in cloud storage. As output, users can watch or download the footage from the cloud platform.
[0398] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0399] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0400] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0401] [Second embodiment]
[0402] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0403] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0404] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0405] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0406] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0407] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0408] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0409] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0410] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0411] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0412] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0413] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0414] This invention is a system that automatically edits soccer match footage on the cloud and provides videos with commentary, eliminating the need for users to record and edit the footage while watching the match.
[0415] System configuration
[0416] The system includes the following main components:
[0417] 1. Photography Method:
[0418] Users set up fixed cameras to film the entire soccer match. The fixed cameras are fixed on the sidelines or behind the goals, and use wide-angle lenses to capture the entire match.
[0419] 2. Upload method:
[0420] Users upload the footage they have taken to the cloud platform, and the video files are sent to the cloud server via the internet.
[0421] 3. Receiving and preprocessing means:
[0422] The server receives the video uploaded to the cloud, checks the format and quality of the video, and performs any necessary pre-processing (e.g., adjusting the resolution or converting the frame rate).
[0423] 4. Analysis method:
[0424] The AI analysis system on the server analyzes the video data frame by frame. The AI uses image processing algorithms to detect the position and movement of the ball and players in each frame. Based on this information, the system tracks the ball and players in real time.
[0425] 5. Extraction means:
[0426] The server analyzes the tracking data to detect when specific events (e.g., goals, saves, fouls) occur, and automatically extracts key moments based on these events to generate highlight clips.
[0427] 6. Explanation Generation Method:
[0428] The server uses a generative AI model to generate commentary in the style of a specific commentator, synthesizes the generated commentary to create a commentary audio, and integrates the commentary audio and subtitles into the video to create an edited highlight video.
[0429] 7. Means of provision:
[0430] The server stores the edited video in cloud storage, and users can access the cloud platform to view or download the edited video.
[0431] Program processing and specific examples
[0432] Natural language explanation of the process
[0433] Users film the match with fixed cameras and upload the footage to the cloud after the match ends. The server preprocesses the received footage and uses an AI analysis system to detect the ball and players. It then extracts key scenes, generates commentary using a generative AI model, and adds commentary audio and subtitles to the video. The edited footage is then stored in the cloud and made available for users to watch or download.
[0434] Specific examples
[0435] For example, suppose a user films their child's soccer game and uploads the footage to the cloud. The server receives the footage and performs preprocessing, and the AI analysis system detects and tracks the movement of the ball and players. Goals and important plays are detected, and highlights are automatically generated based on them. The generative AI model generates commentary such as, "Player A made a decisive pass here, and Player B scored a shot!" and adds it to the video as audio and subtitles. Finally, the edited footage is stored in the cloud, where the user can enjoy it with their family.
[0436] In this way, the present invention is a system that can significantly reduce the effort required for filming and automatically provide professional-quality video.
[0437] The processing flow will be explained below.
[0438] Step 1:
[0439] A user films a soccer match with a fixed camera. The camera is set up and recording begins. The video is set to capture the entire match, using a wide-angle lens, etc.
[0440] Step 2:
[0441] The video footage taken by the user is uploaded to the cloud server. After the match, the user connects the camera's memory card to a computer and uploads the video files to the cloud platform.
[0442] Step 3:
[0443] The server receives the video uploaded to the cloud, checks the format and quality of the received video file, and performs pre-processing such as adjusting the resolution and converting the frame rate as necessary.
[0444] Step 4:
[0445] The AI analysis system on the server analyzes the received video data frame by frame, detecting the ball and players in the video and saving their positional information as tracking data.
[0446] Step 5:
[0447] The server automatically extracts important scenes (goals, fouls, saves, etc.) based on the tracking data, and compiles the extracted scenes to generate a highlight clip.
[0448] Step 6:
[0449] The server uses the generative AI model to generate commentary in the style of a specific commentator, and then performs voice synthesis based on the generated commentary to create the commentary audio.
[0450] Step 7:
[0451] The server integrates audio commentary and subtitles into the video, and edits and completes the highlight video with commentary.
[0452] Step 8:
[0453] The server saves the edited video to cloud storage and notifies the user that the video is ready.
[0454] Step 9:
[0455] Users log in to the cloud platform to view or download the edited footage, and share it with family and team members.
[0456] Example 1
[0457] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0458] Previously, shooting, editing, and creating a video with commentary of a soccer game required a great deal of time and effort. Especially for amateur and family games, where professional video editing skills and commentators are not required, there was a demand for a method to easily create high-quality video with commentary, but no system existed that could achieve this. Furthermore, manual editing and the use of specialized software required specialized knowledge, making it a significant burden for average users.
[0459] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0460] In this invention, the server includes a means for uploading video to the cloud, a means for preprocessing the received video, a means for detecting the ball and players using an image processing algorithm, a means for automatically extracting important scenes and generating highlight clips, a means for creating commentary audio using a generative AI model for generating commentary text and speech synthesis technology, and a means for integrating the commentary audio and subtitles into the video, and a means for saving the edited video in cloud storage and providing it to users, thereby enabling users to easily create highlight videos with professional quality commentary and easily watch and share them.
[0461] A "fixed camera" is a camera that is fixed in a certain position and continuously captures a specific range or scene.
[0462] The "cloud" is a service platform for storing and processing data on remote servers on the Internet.
[0463] "Uploading" is the act of transferring data stored on a local device to a remote server over the Internet.
[0464] "Preprocessing" refers to the initial processing performed on received data, primarily to improve data analysis and algorithm performance.
[0465] An "image processing algorithm" is a computational method for extracting and analyzing specific features or information from image data.
[0466] "Ball and player detection" refers to identifying the position of the ball and players in each frame of the video and tracking their movements.
[0467] A "key moment" is a particularly noteworthy event in a soccer match (e.g., a goal, a save, a foul, etc.).
[0468] A "highlight clip" is a shortened, edited video clip that selects particularly important scenes from the entire game.
[0469] A "generative AI model" is an artificial intelligence model that generates natural-looking explanatory text and content based on given data and prompt text.
[0470] "Speech synthesis technology" is a technology that converts text data into actual speech.
[0471] "Commentary audio" is audio data that has been converted using speech synthesis technology from explanatory text created by a generative AI model.
[0472] "Subtitles" are textual information that is visually displayed within a video and complements the audio content.
[0473] "Cloud storage" is a data storage service on the Internet that stores data on a remote server and allows it to be accessed as needed.
[0474] "Providing to the user" means supplying the edited video in a form that the user can access.
[0475] This invention is a system that automatically edits soccer match footage on the cloud and provides videos with commentary, eliminating the need for users to record and edit the footage while watching the match.
[0476] System configuration
[0477] The system includes the following main components:
[0478] 1. Photography Method:
[0479] Users set up fixed cameras to film the entire soccer match. The fixed cameras are fixed on the sidelines or behind the goals, and use wide-angle lenses to capture the entire match.
[0480] 2. Upload method:
[0481] Users upload the footage they have taken to a cloud platform. The video files are sent to a cloud server via the Internet. The cloud platform uses commonly used storage services (e.g., Google Drive or Dropbox).
[0482] 3. Receiving and preprocessing means:
[0483] The server receives the video uploaded to the cloud. After receiving it, it checks the format and quality of the video and performs any necessary pre-processing (e.g., adjusting the resolution or converting the frame rate) using a video processing library (e.g., FFmpeg).
[0484] 4. Analysis method:
[0485] The AI analysis system on the server analyzes the video data frame by frame. The AI uses an image processing algorithm (e.g., YOLO (You Only Look Once)) to detect the position and movement of the ball and players in each frame. Based on this detected information, the ball and players are tracked in real time.
[0486] 5. Extraction means:
[0487] The server analyzes the tracking data to detect when specific events (e.g., goals, saves, fouls) occur, and automatically extracts key moments based on these events to generate highlight clips.
[0488] 6. Explanation Generation Method:
[0489] The server uses a generative AI model to generate commentary in the style of a specific commentator. Using a prompt as input, the generative AI model generates appropriate commentary. For example, a prompt such as "Player A made a pass and Player B scored a goal" is used. Based on the generated commentary, a voice synthesis technology (e.g., Google Text-to-Speech API) is used to create a commentary voice.
[0490] 7. Integration Methods:
[0491] The server integrates the commentary audio and subtitles into the video to create an edited highlight video. A video editing library (e.g., OpenCV or MoviePy) is used to integrate the audio file with the video and overlay the subtitle text.
[0492] 8. Means of provision:
[0493] The server stores the edited video in cloud storage (e.g., AWS S3), and users can access the cloud platform to view or download the edited video.
[0494] Specific examples
[0495] For example, suppose a user films their child's soccer game and uploads the footage to the cloud. The server receives the footage and performs preprocessing, and the AI analysis system detects and tracks the movement of the ball and players. Goals and important plays are detected, and highlights are automatically generated based on them. The generative AI model generates commentary such as, "Player A made a decisive pass here, and Player B scored a shot!" and adds it to the video as audio and subtitles. Finally, the edited footage is stored in the cloud, where the user can enjoy it with their family.
[0496] In this way, the present invention is a system that can significantly reduce the effort required for filming and automatically provide professional-quality video.
[0497] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0498] Step 1:
[0499] The user sets up a fixed camera and captures footage of the match.
[0500] Specifically, the user fixes a fixed camera on the sideline or behind the goal and adjusts the camera settings, for example, using a wide-angle lens to capture the entire game in the frame.
[0501] Input: The physical camera equipment and the soccer match to be filmed.
[0502] Output: Footage files of the captured match.
[0503] Step 2:
[0504] The user uploads the footage they have taken to a cloud platform.
[0505] Users use their home computer or smartphone to access the designated upload form on the cloud platform, select and submit the video file.
[0506] Input: The captured video file.
[0507] Output: Video files uploaded to the cloud platform.
[0508] Step 3:
[0509] The server receives the video uploaded to the cloud platform and performs pre-processing.
[0510] The server uses a video processing library (e.g., FFmpeg) to check the format and quality of the received video file and make any necessary resolution adjustments and frame rate conversions, for example, converting 4K video to 1080p and changing the frame rate from 60 fps to 30 fps.
[0511] Input: Video files on the cloud.
[0512] Output: Pre-processed video files.
[0513] Step 4:
[0514] The server uses an AI analysis system to analyze the pre-processed video data frame by frame.
[0515] The server uses an image processing algorithm (e.g., YOLO) to detect the positions of the ball and players in each frame. Specifically, it obtains the coordinate data of the bounding boxes of the ball and players included in each frame.
[0516] Input: Preprocessed video files.
[0517] Output: Ball and player position data.
[0518] Step 5:
[0519] The server analyzes the tracking data to detect when a particular event (e.g., goal, save, foul) occurs.
[0520] The server records the timestamps of the frames corresponding to these events and automatically extracts important scenes, such as five seconds before and after a goal.
[0521] Input: Ball and player position data.
[0522] Output: timestamps and footage clips of key scenes.
[0523] Step 6:
[0524] The server uses a generative AI model to generate commentary in the style of a specific commentator.
[0525] The server inputs a prompt, and the generative AI model generates an appropriate commentary. For example, a prompt such as "Player A passes the ball, and Player B scores a goal" is used.
[0526] Input: prompt text and video clips of key scenes.
[0527] Output: The generated description.
[0528] Step 7:
[0529] The server creates a commentary audio based on the generated commentary text using speech synthesis technology (e.g., Google Text-to-Speech API).
[0530] The commentary is converted into an audio file and is also prepared as subtitle text. In detail, the server sends the commentary to the speech synthesis API and obtains the audio file.
[0531] Input: The generated description.
[0532] Output: Descriptive audio file and subtitle text.
[0533] Step 8:
[0534] The server integrates audio commentary and subtitles into the video clips to create an edited highlight video.
[0535] Use a video editing library (e.g., OpenCV or MoviePy) to integrate the audio file and subtitle text into the video clip. Specifically, the audio description is overlaid on the video and the subtitle text is displayed in the appropriate position.
[0536] Input: Descriptive audio files, subtitle text, and video clips of key scenes.
[0537] Output: Edited highlight video.
[0538] Step 9:
[0539] The server saves the completed highlight video to cloud storage (e.g. AWS S3).
[0540] Generate a link that allows users to access the saved video and provide that link to the user. For example, generate a URL like "https: / / cloudstorage.example.com / highlight_video.mp4".
[0541] Input: edited highlight video.
[0542] Output: Highlight video URL on cloud storage.
[0543] Step 10:
[0544] The user accesses the cloud platform and clicks on the provided link to view the edited footage.
[0545] Users can use a browser or a dedicated app to play or download videos from cloud storage.
[0546] Input: Highlight video URL on cloud storage.
[0547] Output: Highlights viewed or downloaded.
[0548] Through the above process, the system automatically edits soccer match footage shot by the user and provides a professional highlight video with commentary.
[0549] (Application example 1)
[0550] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0551] Shooting soccer game footage, editing the footage, and creating a highlight video with commentary has traditionally required a great deal of effort. Editing the footage and creating the commentary, in particular, requires time and effort, making it difficult to easily achieve professional-quality results. The objective of this invention is to solve this problem and provide a system that enables users to easily create high-quality highlight videos with commentary.
[0552] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0553] In this invention, the server includes means for uploading video captured by a fixed camera to the cloud, means for receiving the uploaded video and detecting the ball and players, means for automatically extracting key scenes based on the detected information, means for adding automatically generated commentary to the video using a generative AI model, means for integrating commentary audio and subtitles to complete the edited video, and means for saving the edited video in cloud storage and enabling users to view or share it. This allows users to easily generate, view, and share professional-quality highlight videos with commentary simply by uploading their game video to the cloud.
[0554] A "fixed camera" is a camera that captures video from a fixed position.
[0555] The "cloud" is a collection of servers for storing, managing, and processing data over the Internet.
[0556] "Uploading" is the act of sending data from a local device to a cloud server.
[0557] "Reception" refers to the act of the cloud server receiving data sent from the user.
[0558] "Detection" is the process of identifying a specific object (the ball or a player) through video analysis.
[0559] A "Key Moment" is a specific key moment during a match (e.g., a goal, save, foul).
[0560] A "generative AI model" is an artificial intelligence model that uses machine learning algorithms to generate content such as text and images.
[0561] "Explanation" is written or audio that explains a particular event.
[0562] The "explanatory voice" is the generated explanatory text output as voice through voice synthesis.
[0563] "Subtitles" are textual information displayed on video.
[0564] "Editing" is the process of adding audio commentary and subtitles to the footage to complete the final video.
[0565] "Cloud storage" is a storage service for storing data on the cloud.
[0566] "Viewing" refers to the act of a user playing and watching a video.
[0567] "Sharing" refers to the act of a user sharing video or data with other people.
[0568] This invention is a system that automatically edits soccer match footage taken with fixed cameras and provides videos with commentary, eliminating the need for users to record and edit the footage while watching the match.
[0569] System configuration
[0570] The system includes the following main components:
[0571] 1. Fixed camera
[0572] Users can set up fixed cameras on the sidelines or behind the goals and use wide-angle lenses to capture the entire game.
[0573] 2. Cloud Upload
[0574] Users upload the footage they have taken to the cloud platform, and the video files are sent to the cloud server via the internet.
[0575] 3. Reception and Preprocessing
[0576] The server receives the video uploaded to the cloud. After receiving the video, it checks the format and quality of the video and performs any necessary preprocessing (e.g., adjusting the resolution and converting the frame rate). For preprocessing, it uses an open-source video processing library such as OpenCV.
[0577] 4. AI analysis
[0578] The AI analysis system on the server analyzes the video data frame by frame, and uses OpenCV and deep learning models to detect the position of the ball and the positions and movements of players in each frame.
[0579] 5. Extracting important scenes
[0580] The server analyzes the tracking data to detect when specific events (e.g., goals, saves, fouls) occur, and automatically extracts key moments based on these events to generate highlight clips.
[0581] 6. Explanation Generation
[0582] The server uses a generative AI model (e.g., GPT-4) to generate commentary in the style of a specific commentator. It then uses voice synthesis based on the generated commentary to create the commentary audio. For voice synthesis, it uses a service like Amazon Polly.
[0583] 7. Video Editing
[0584] The commentary audio and subtitles are integrated into the video, and video editing libraries such as FFmpeg are used for video editing. This completes the edited highlight video.
[0585] 8. Cloud Storage
[0586] The edited footage is saved to cloud storage, allowing users to view or share it.
[0587] Specific examples
[0588] For example, a user can film their child's soccer game with their smartphone and upload the footage to the cloud. The server receives the footage and performs pre-processing. The AI analysis system detects the movement of the ball and players and extracts goals and other important plays. Based on this, the generative AI model generates commentary such as, "Player A made a decisive pass here, and Player B scored a shot!" and adds it to the video as audio and subtitles. Finally, the edited video is stored in the cloud, and users can share it with family and friends.
[0589] Prompt Sentence Examples
[0590] For example, when generating explanatory text using a generative AI model, the following prompt is used:
[0591] "I uploaded a video of my child's soccer game. Analyze the ball position and player movements to detect goals and important plays. Based on that, generate commentary such as, 'Player A made the decisive pass, and Player B scored the shot!' and add it to the video as audio and subtitles."
[0592] In this way, the present invention is a system that can significantly reduce the effort required for filming and automatically provide professional-quality video.
[0593] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0594] Step 1:
[0595] A user films a soccer match with a fixed camera. The fixed camera uses a wide-angle lens to capture the entire game, and is fixed on the sidelines or behind the goal. The input is the video data captured by the camera, and the output is a video file stored in local storage.
[0596] Step 2:
[0597] Users upload the videos they have taken from their smartphones to the cloud platform. The uploaded videos are then sent to the cloud server via the Internet. The input is the video file stored on the smartphone, and the output is the video data on the cloud storage.
[0598] Step 3:
[0599] The server receives the video uploaded to the cloud and performs preprocessing. This includes adjusting the resolution and converting the frame rate, and uses a video processing library such as OpenCV. The input is the video data stored in the cloud storage, and the output is the preprocessed video data converted into an analyzable format.
[0600] Step 4:
[0601] The AI analysis system in the server analyzes the preprocessed video frame by frame to detect the position and movement of the ball and players. This analysis uses OpenCV and deep learning models. The input is the preprocessed video data, and the output is tracking data including the position information of the ball and players.
[0602] Step 5:
[0603] The server automatically extracts important scenes from the tracking data. Specifically, it detects events such as goals, saves, and fouls and extracts those moments as highlight clips. The input is the tracking data, and the output is a list of important scenes and their clip data.
[0604] Step 6:
[0605] The server uses a generative AI model (e.g., GPT-4) to create automatically generated commentary. The generated commentary is converted into speech using speech synthesis technology (e.g., Amazon Polly). The input is a list of important scenes and prompts, and the output is commentary audio and subtitle data.
[0606] Step 7:
[0607] The server integrates the commentary audio and subtitles into the video to create an edited highlight video. This editing is done using a video editing library such as FFmpeg. The input is the clip data, commentary audio, and subtitle data, and the output is the edited video.
[0608] Step 8:
[0609] The server saves the completed edited video in cloud storage. Users can access the cloud platform to view and share the video. The input is the edited video data, and the output is the video data saved in cloud storage.
[0610] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0611] This invention is a system that automatically edits soccer match footage on the cloud and provides videos with commentary using an emotion engine. This system saves users the trouble of recording and editing the game while watching it. It also adjusts the commentary content based on the user's emotions, providing a more personalized video experience.
[0612] System configuration
[0613] The system includes the following main components:
[0614] 1. Photography Method:
[0615] Users set up fixed cameras to film the entire soccer match. The fixed cameras are fixed on the sidelines or behind the goals, and use wide-angle lenses to capture the entire match.
[0616] 2. Upload method:
[0617] Users upload the footage they have taken to the cloud platform, and the video files are sent to the cloud server via the internet.
[0618] 3. Receiving and preprocessing means:
[0619] The server receives the video uploaded to the cloud, checks the format and quality of the video, and performs any necessary pre-processing (e.g., adjusting the resolution or converting the frame rate).
[0620] 4. Analysis method:
[0621] The AI analysis system on the server analyzes the received video data frame by frame, detecting the ball and players in the video and saving their positional information as tracking data.
[0622] 5. Extraction means:
[0623] The server automatically extracts important scenes (goals, fouls, saves, etc.) based on the tracking data, and compiles the extracted scenes to generate highlight clips.
[0624] 6. Emotion Engine:
[0625] The server is equipped with an emotion engine that recognizes the user's emotions. It analyzes the user's emotion data (e.g., facial expression recognition and biometric data) and grasps the user's emotional state in real time.
[0626] 7. Explanation Generation Method:
[0627] The server uses a generative AI model to generate commentary in the style of a specific commentator, adjusts the commentary content based on the user's emotions recognized by the emotion engine, and performs voice synthesis based on the generated commentary to create a commentary voice.
[0628] 8. Means of provision:
[0629] The server stores the edited video in cloud storage, and users can access the cloud platform to view or download the edited video.
[0630] Program processing and specific examples
[0631] Natural language explanation of the process
[0632] Users film the match with fixed cameras and upload the footage to the cloud after the match ends. The server preprocesses the received footage and uses an AI analysis system to detect the ball and players. An emotion engine then collects and analyzes the user's emotions in real time. Key scenes are then extracted, and a generative AI model generates commentary based on the emotion data and integrates it into the video. Finally, the edited video is stored in the cloud and made available for users to watch or download.
[0633] Specific examples
[0634] For example, suppose a user films their child's soccer game and uploads the footage to the cloud. The server receives the footage and performs preprocessing, while the AI analysis system detects and tracks the movement of the ball and players. At the same time, while the user is watching the game, the emotion engine analyzes the user's facial expressions, heart rate, and other data to collect emotional data. Goals and important plays are detected, and highlights are automatically generated based on them. The generative AI model generates commentary such as, "Player A made a decisive pass here, and Player B scored a shot!" If it determines that the user is happy, the commentary will be written in a positive tone. Finally, the edited video is saved in the cloud, where the user can enjoy it with their family.
[0635] In this way, the present invention not only significantly reduces the effort required for filming and automatically provides professional-quality video, but also provides a more personalized video experience based on the user's emotions.
[0636] The processing flow will be explained below.
[0637] Step 1:
[0638] A user films a soccer match with a fixed camera. The camera is set up and recording begins. The video is set to capture the entire match, using a wide-angle lens, etc.
[0639] Step 2:
[0640] The video footage taken by the user is uploaded to the cloud server. After the match, the user connects the camera's memory card to a computer and uploads the video files to the cloud platform.
[0641] Step 3:
[0642] The server receives the video uploaded to the cloud, checks the format and quality of the received video file, and performs pre-processing such as adjusting the resolution and converting the frame rate as necessary.
[0643] Step 4:
[0644] The AI analysis system on the server analyzes the received video data frame by frame, detecting the ball and players in the video and saving their positional information as tracking data.
[0645] Step 5:
[0646] The server automatically extracts important scenes (goals, fouls, saves, etc.) based on the tracking data, and compiles the extracted scenes to generate a highlight clip.
[0647] Step 6:
[0648] As users watch a game, the emotion engine collects their facial expressions and heart rate in real time to generate emotion data, which is then sent to a cloud server.
[0649] Step 7:
[0650] The server receives and analyzes the emotion data sent from the emotion engine to understand the user's emotional state (happiness, excitement, sadness, etc.).
[0651] Step 8:
[0652] The server uses a generative AI model to generate commentary in the style of a specific commentator, adjusting the commentary based on emotion data, for example, using positive expressions if the user is happy.
[0653] Step 9:
[0654] The server integrates audio commentary and subtitles into the video, and edits and completes the highlight video with commentary.
[0655] Step 10:
[0656] The server saves the edited video to cloud storage and notifies the user that the video is ready.
[0657] Step 11:
[0658] Users log in to the cloud platform to view or download the edited footage, and share it with family and team members.
[0659] Example 2
[0660] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0661] Conventional soccer match filming and editing systems require lengthy editing work after filming, and adding commentary is labor-intensive. Furthermore, they are unable to reflect the user's emotional state while watching the video, making it difficult to provide a personalized viewing experience. There is a need for a system that can solve these problems and easily provide high-quality, personalized video.
[0662] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0663] In this invention, the server includes means for uploading video captured by a fixed camera to the cloud, means for receiving the uploaded video and pre-processing the video, means for detecting the ball and players from the received video, means for automatically extracting important scenes based on the detected information, means for collecting user emotion data, means for analyzing the collected emotion data and generating commentary using a generative AI model, means for adding the generated commentary to the video, and means for providing the edited video to the user, thereby enabling users to watch high-quality video containing match highlights and personalized commentary without any hassle.
[0664] A "fixed camera" is a camera that is fixed at a specific location and captures a wide range of images using a wide-angle lens.
[0665] "Cloud" is a general term for an online platform that provides data storage and computing resources by accessing remote servers via the Internet.
[0666] "Uploading means" refers to the method or technology used to transfer digital data from a local device to a remote server, such as cloud storage.
[0667] "Means of receiving" refers to the method or technology by which the remote server obtains data uploaded to the cloud.
[0668] "Preprocessing means" refers to processing carried out in advance, such as converting data format or adjusting quality, to ensure smooth subsequent analysis and editing processes.
[0669] "Detection means" refers to methods or technologies that automatically recognize specific objects (e.g., balls or players) from video data and track their location information.
[0670] "Means for automatically extracting important scenes" refers to methods and technologies that use detected information to identify specific events such as goals and fouls and extract them as highlights from the video.
[0671] "Means for collecting emotional data" refers to methods and technologies for acquiring a user's facial expressions and biometric data in real time and analyzing their emotional state.
[0672] A "generative AI model" is a technology that uses a pre-trained artificial intelligence model to generate explanatory text in natural language from given data and prompts.
[0673] "Means for adding commentary" refers to methods or techniques for integrating the generated commentary text and audio into the video data and providing it as the final edited video.
[0674] "Means for providing edited video" refers to the methods and technologies for storing the final edited video in cloud storage and allowing users to access, view, or download it.
[0675] This system automatically edits soccer match footage on the cloud and provides videos with commentary using an emotion engine. This allows users to concentrate on watching the game, while automating the editing and commentary tasks. It also provides a personalized video experience based on the user's emotions.
[0676] First, the user sets up a fixed camera to film a soccer match. The fixed camera uses a wide-angle lens and is installed on the sidelines or behind the goal. After the match ends, the user uploads the filmed video files to a cloud platform. This can be done using cloud storage services such as Google Drive or Dropbox.
[0677] The server then receives the uploaded video on the cloud, checks the format and quality of the video file, and performs any necessary pre-processing, such as converting AVI video to MP4, reducing the resolution from 1080p to 720p, and converting the frame rate from 30fps to 24fps.
[0678] Once preprocessing is complete, the server uses an AI analysis system to analyze the received video data frame by frame. The AI analysis system uses deep learning models and libraries such as OpenCV to detect the ball and players in the video and save their location information as tracking data. This allows the progress of the game and the movements of the players to be understood.
[0679] The server then automatically extracts important moments (e.g., goals, fouls, saves, etc.) from the tracking data. In this step, AI identifies specific events and cuts them out into highlight videos.
[0680] Furthermore, the server uses an emotion engine to collect user emotional data, including biometric data such as facial expressions and heart rate, and analyzes it in real time, allowing the server to understand the user's emotional state during a match.
[0681] The server then uses a generative AI model to generate commentary in the style of a particular commentator. The generative AI model is given the following prompt:
[0682] "Generate commentary on goal scenes."
[0683] "Please provide an inspiring commentary based on the emotional data."
[0684] "Describe a highlight scene of a key play."
[0685] The generated commentary text is used to synthesize speech, creating a commentary voice, which adds a commentary tailored to the user's emotions to the video.
[0686] Finally, the server integrates the generated commentary audio into the highlight video. Specifically, it uses video editing software such as FFmpeg to combine the highlight footage and commentary audio into a single file. The edited video file is then stored in cloud storage, and users can access the cloud platform to watch or download it.
[0687] For example, consider the case where a user films a child's soccer game and uploads the video to the cloud. The server receives the video and performs preprocessing, and the AI analysis system detects and tracks the movement of the ball and players. Then, while the user is watching the game, the emotion engine analyzes the user's facial expressions, heart rate, etc., and collects emotional data. Based on this data, the generative AI model generates commentary, providing commentary in positive terms, allowing the user to enjoy watching the highlights of the game.
[0688] In this way, the present invention not only significantly reduces the effort required for filming and editing, and automatically provides professional-quality video, but also enables a personally tailored video experience that reflects the user's emotions.
[0689] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0690] Step 1:
[0691] The user sets up a fixed camera to film a soccer match. The fixed camera used for recording uses a wide-angle lens and is fixed to the sideline or behind the goal. After the match ends, the user transfers the filmed footage to a computer. The input data is the filmed video file, and the output data is the video file saved on the computer.
[0692] Step 2:
[0693] Users upload video files stored on their computers to a cloud platform. Specifically, they use cloud storage services such as Google Drive and Dropbox to upload files. The input data is the video file stored on their computer, and the output data is the video file stored on the cloud.
[0694] Step 3:
[0695] The server receives the video uploaded to the cloud. After receiving it, it checks the format and quality of the video file and performs any necessary pre-processing. Pre-processing includes converting AVI format video to MP4 format, reducing the resolution from 1080p to 720p, and converting the frame rate from 30fps to 24fps. The input data is the video file on the cloud, and the output data is the video file after pre-processing. Specific operations involve the use of video conversion software such as FFmpeg.
[0696] Step 4:
[0697] The server uses an AI analysis system to analyze the pre-processed video data frame by frame. The AI analysis system uses a deep learning model and libraries such as OpenCV to detect the ball and players in the video. It then saves this position information as tracking data. The input data is the pre-processed video file, and the output data is the analysis results, including the tracking data.
[0698] Step 5:
[0699] The server automatically extracts key moments (e.g., goals, fouls, saves, etc.) from the tracking data. This involves AI identifying specific events and extracting them into highlight videos. The input data is the tracking data, and the output data is video clips of key moments.
[0700] Step 6:
[0701] The server uses an emotion engine to collect user emotion data in real time. It also uses devices such as webcams and smartwatches to acquire biometric data such as the user's facial expressions and heart rate. The input data is the user's biometric data, and the output data is the analyzed emotion data.
[0702] Step 7:
[0703] The server uses a generative AI model to generate commentary in the style of a specific commentator. The generative AI model is given the following prompt:
[0704] "Generate commentary on goal scenes."
[0705] "Please provide an inspiring commentary based on the emotional data."
[0706] "Describe a highlight scene of a key play."
[0707] The input data is emotion data and prompt sentences, and the output data is the generated commentary. Specifically, a generative AI model such as GPT-3 is used.
[0708] Step 8:
[0709] The server synthesizes the commentary based on the generated commentary text and creates the commentary audio. The input data is the generated commentary text, and the output data is the commentary audio file. Specifically, speech synthesis software is used.
[0710] Step 9:
[0711] The server then integrates the generated commentary audio into the highlight video. Specifically, it uses video editing software such as FFmpeg to combine the highlight video and commentary audio into a single file. The input data is the commentary audio file and the video clips of key scenes, and the output data is the edited video file.
[0712] Step 10:
[0713] The server stores the edited video files in cloud storage, allowing users to access and watch or download them. The input data is the edited video files, and the output data is the video files stored on the cloud. Specifically, a cloud storage platform is used.
[0714] (Application example 2)
[0715] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0716] Conventional video editing systems require manual editing after shooting footage, which requires a great deal of time and effort. Furthermore, the generated commentary is standardized, making it difficult to provide a personalized experience that reflects each user's individual emotions and circumstances. This requires a lot of effort for users to record and edit events, and the viewing experience is not personalized enough.
[0717] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for uploading video captured by a fixed camera to the cloud, means for receiving the uploaded video and detecting objects and people, means for automatically extracting important scenes based on the detected information, generation AI model means for analyzing user emotion data and generating customized commentary based on the analysis results, means for adding the generated commentary to the video, and means for providing the edited video to the user. This eliminates the need for manual editing of the video and makes it possible to provide a personalized viewing experience based on the user's emotions.
[0718] A "fixed camera" is a camera that continuously captures a specific area from a fixed position.
[0719] The "cloud" refers to computing resources and data storage provided via the Internet.
[0720] "Uploading" is the act of sending local data to a remote server over the Internet or a network.
[0721] "Object" is a general term that refers to any concrete thing that can be recognized in an image.
[0722] "Person" is a general term that refers to a person recognized in the video.
[0723] "Detection" is the act of using specific algorithms to identify and track objects or people in a video.
[0724] "Key scenes" refer to moments or events that are particularly noteworthy in footage of sporting events or everyday life.
[0725] "Extraction" is the act of extracting a specific part from the whole data.
[0726] "Emotional data" refers to the emotional state obtained from the user's facial expressions, biometric data, etc.
[0727] A "generative AI model" is an algorithm that uses machine learning and natural language processing to automatically generate explanatory text and other content.
[0728] "Commentary" refers to text or audio that adds explanations or comments to the content of the video.
[0729] "Customization" is the act of tailoring content to a particular user or situation.
[0730] "Edited footage" refers to the final video file in which processing and modification have been applied to the original video data.
[0731] "Providing" refers to the act of making a video available for viewing or acquisition by a user through the cloud or other media.
[0732] The following system is used as an embodiment of the present invention: The main components of the system and their respective roles will be described below.
[0733] System Components
[0734] 1. Fixed camera:
[0735] Used by users to capture events, fixed cameras are fixed in specific locations and use wide-angle lenses to capture a wide range of footage.
[0736] 2. Cloud Platform:
[0737] Users upload footage taken with fixed cameras, and the cloud platform receives the video files via the internet and performs the analysis and editing processes described below.
[0738] 3. Receiving and Preprocessing Server:
[0739] The server receives the uploaded video data. It checks the format and quality of the video and performs any necessary preprocessing (adjusting the resolution and converting the frame rate).
[0740] 4. AI analysis system:
[0741] Video data is analyzed frame by frame, objects and people are detected in the video, and their location information is saved as tracking data.
[0742] 5. Sentiment Analysis System:
[0743] Analyze the user's emotional state. Collect user emotional data (facial recognition and biometric data) using sensors on smartphones and smart glasses, and analyze it in real time.
[0744] 6. Generative AI Models:
[0745] Based on the collected emotional data, a commentary in the style of a specific commentator is generated. Speech synthesis is performed based on the generated commentary to create a commentary voice. The software used here includes a natural language processing library and a speech synthesis module.
[0746] 7. Editing System:
[0747] Based on the analysis results and audio commentary, the video is edited to generate highlight clips. This editing system uses video editing libraries (e.g., OpenCV and MoviePy).
[0748] 8. Delivery system:
[0749] The edited video is stored in cloud storage, and users can view or download the edited video from the cloud platform.
[0750] A description of what the program does
[0751] In this system, a user first captures an event with a fixed camera and uploads the footage to a cloud platform. The cloud server receives the video data, checks its format and quality, and performs any necessary preprocessing. The AI analysis system then analyzes the received video data frame by frame to obtain information for detecting and tracking objects and people.
[0752] At the same time, devices such as smartphones and smart glasses are used to collect the user's emotional data (e.g., facial expressions, heart rate, etc.), which is then analyzed in real time by an emotion analysis system to determine the user's emotional state.
[0753] Based on the results of the sentiment analysis and AI analysis, the generative AI model generates commentary in the style of a specific commentator. For example, if it determines that the user is happy, a positive commentary will be selected. The generated commentary is converted into an audio file using text-to-speech software. These commentary voices are then integrated into existing footage, and the editing system generates a highlight clip.
[0754] Finally, the edited video is stored in cloud storage. Users can access the cloud platform and watch or download the generated personalized video. This process allows users to easily obtain high-quality edited video and provides a personalized experience based on their emotions.
[0755] Specific examples
[0756] For example, if a user films a family birthday party, the app uploads the video to the cloud, and the AI performs emotion analysis and video editing as follows: When the emotion analysis system determines that the user is happy at the climax of the party, the generative AI model generates a caption such as, "Here, the whole family has fun and cut the cake!"
[0757] Prompt Sentence Examples
[0758] Here is a video of a family birthday party. Highlight the most exciting scenes and generate commentary based on the emotion data.
[0759] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0760] Step 1:
[0761] A user captures an event with a fixed camera or a mobile information terminal. As input, the user obtains video data. As output, a video file is generated.
[0762] Step 2:
[0763] The user uploads the video they have taken to the cloud platform. The input is the video file they have taken, which the cloud server receives. The output is the video data stored on the cloud.
[0764] Step 3:
[0765] The server receives the uploaded video data and checks the format and quality. The input is the video data stored on the cloud server, and the server performs preprocessing such as adjusting the resolution and converting the frame rate. The output is the preprocessed video data.
[0766] Step 4:
[0767] The server's AI analysis system analyzes the preprocessed video data frame by frame to detect objects and people. The input is the preprocessed video data, and the server applies detection algorithms to identify objects and people, saving their location information as tracking data. The output is tracking data.
[0768] Step 5:
[0769] The device (smartphone or smart glasses) collects emotion data. The input is sensor data such as the user's facial expression and heart rate, and the device sends this data to the emotion analysis system. The output is the user's emotion data as an analysis result.
[0770] Step 6:
[0771] The server uses a generative AI model to generate an explanatory text based on the user's emotional data and the results of AI analysis. The inputs are emotional data and tracking data, and the generative AI model generates a prompt based on this data and then generates an explanatory text (e.g., "That was a great goal!" if the emotion is positive). The output is an explanatory text.
[0772] Step 7:
[0773] The server converts the generated commentary into an audio file using a speech synthesis module. The commentary is input, the server synthesizes the speech, and obtains the audio file. The audio file is obtained as output.
[0774] Step 8:
[0775] The server's editing system integrates the video and audio commentary to generate a highlight clip. The inputs are the pre-processed video data, tracking data, and audio commentary files, and the server uses a video editing library to integrate them. The output is an edited highlight video.
[0776] Step 9:
[0777] The server stores the edited footage in cloud storage for users to access. As input, there is the edited highlight footage, which the server stores in cloud storage. As output, users can watch or download the footage from the cloud platform.
[0778] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0779] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0780] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0781] [Third embodiment]
[0782] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0783] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0784] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0785] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0786] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0787] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0788] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0789] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0790] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0791] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0792] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0793] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0794] This invention is a system that automatically edits soccer match footage on the cloud and provides videos with commentary, eliminating the need for users to record and edit the footage while watching the match.
[0795] System configuration
[0796] The system includes the following main components:
[0797] 1. Photography Method:
[0798] Users set up fixed cameras to film the entire soccer match. The fixed cameras are fixed on the sidelines or behind the goals, and use wide-angle lenses to capture the entire match.
[0799] 2. Upload method:
[0800] Users upload the footage they have taken to the cloud platform, and the video files are sent to the cloud server via the internet.
[0801] 3. Receiving and preprocessing means:
[0802] The server receives the video uploaded to the cloud, checks the format and quality of the video, and performs any necessary pre-processing (e.g., adjusting the resolution or converting the frame rate).
[0803] 4. Analysis method:
[0804] The AI analysis system on the server analyzes the video data frame by frame. The AI uses image processing algorithms to detect the position and movement of the ball and players in each frame. Based on this information, the system tracks the ball and players in real time.
[0805] 5. Extraction means:
[0806] The server analyzes the tracking data to detect when specific events (e.g., goals, saves, fouls) occur, and automatically extracts key moments based on these events to generate highlight clips.
[0807] 6. Explanation Generation Method:
[0808] The server uses a generative AI model to generate commentary in the style of a specific commentator, synthesizes the generated commentary to create a commentary audio, and integrates the commentary audio and subtitles into the video to create an edited highlight video.
[0809] 7. Means of provision:
[0810] The server stores the edited video in cloud storage, and users can access the cloud platform to view or download the edited video.
[0811] Program processing and specific examples
[0812] Natural language explanation of the process
[0813] Users film the match with fixed cameras and upload the footage to the cloud after the match ends. The server preprocesses the received footage and uses an AI analysis system to detect the ball and players. It then extracts key scenes, generates commentary using a generative AI model, and adds commentary audio and subtitles to the video. The edited footage is then stored in the cloud and made available for users to watch or download.
[0814] Specific examples
[0815] For example, suppose a user films their child's soccer game and uploads the footage to the cloud. The server receives the footage and performs preprocessing, and the AI analysis system detects and tracks the movement of the ball and players. Goals and important plays are detected, and highlights are automatically generated based on them. The generative AI model generates commentary such as, "Player A made a decisive pass here, and Player B scored a shot!" and adds it to the video as audio and subtitles. Finally, the edited footage is stored in the cloud, where the user can enjoy it with their family.
[0816] In this way, the present invention is a system that can significantly reduce the effort required for filming and automatically provide professional-quality video.
[0817] The processing flow will be explained below.
[0818] Step 1:
[0819] A user films a soccer match with a fixed camera. The camera is set up and recording begins. The video is set to capture the entire match, using a wide-angle lens, etc.
[0820] Step 2:
[0821] The video footage taken by the user is uploaded to the cloud server. After the match, the user connects the camera's memory card to a computer and uploads the video files to the cloud platform.
[0822] Step 3:
[0823] The server receives the video uploaded to the cloud, checks the format and quality of the received video file, and performs pre-processing such as adjusting the resolution and converting the frame rate as necessary.
[0824] Step 4:
[0825] The AI analysis system on the server analyzes the received video data frame by frame, detecting the ball and players in the video and saving their positional information as tracking data.
[0826] Step 5:
[0827] The server automatically extracts important scenes (goals, fouls, saves, etc.) based on the tracking data, and compiles the extracted scenes to generate a highlight clip.
[0828] Step 6:
[0829] The server uses the generative AI model to generate commentary in the style of a specific commentator, and then performs voice synthesis based on the generated commentary to create the commentary audio.
[0830] Step 7:
[0831] The server integrates audio commentary and subtitles into the video, and edits and completes the highlight video with commentary.
[0832] Step 8:
[0833] The server saves the edited video to cloud storage and notifies the user that the video is ready.
[0834] Step 9:
[0835] Users log in to the cloud platform to view or download the edited footage, and share it with family and team members.
[0836] Example 1
[0837] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0838] Previously, shooting, editing, and creating a video with commentary of a soccer game required a great deal of time and effort. Especially for amateur and family games, where professional video editing skills and commentators are not required, there was a demand for a method to easily create high-quality video with commentary, but no system existed that could achieve this. Furthermore, manual editing and the use of specialized software required specialized knowledge, making it a significant burden for average users.
[0839] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0840] In this invention, the server includes a means for uploading video to the cloud, a means for preprocessing the received video, a means for detecting the ball and players using an image processing algorithm, a means for automatically extracting important scenes and generating highlight clips, a means for creating commentary audio using a generative AI model for generating commentary text and speech synthesis technology, and a means for integrating the commentary audio and subtitles into the video, and a means for saving the edited video in cloud storage and providing it to users, thereby enabling users to easily create highlight videos with professional quality commentary and easily watch and share them.
[0841] A "fixed camera" is a camera that is fixed in a certain position and continuously captures a specific range or scene.
[0842] The "cloud" is a service platform for storing and processing data on remote servers on the Internet.
[0843] "Uploading" is the act of transferring data stored on a local device to a remote server over the Internet.
[0844] "Preprocessing" refers to the initial processing performed on received data, primarily to improve data analysis and algorithm performance.
[0845] An "image processing algorithm" is a computational method for extracting and analyzing specific features or information from image data.
[0846] "Ball and player detection" refers to identifying the position of the ball and players in each frame of the video and tracking their movements.
[0847] A "key moment" is a particularly noteworthy event in a soccer match (e.g., a goal, a save, a foul, etc.).
[0848] A "highlight clip" is a shortened, edited video clip that selects particularly important scenes from the entire game.
[0849] A "generative AI model" is an artificial intelligence model that generates natural-looking explanatory text and content based on given data and prompt text.
[0850] "Speech synthesis technology" is a technology that converts text data into actual speech.
[0851] "Commentary audio" is audio data that has been converted using speech synthesis technology from explanatory text created by a generative AI model.
[0852] "Subtitles" are textual information that is visually displayed within a video and complements the audio content.
[0853] "Cloud storage" is a data storage service on the Internet that stores data on a remote server and allows it to be accessed as needed.
[0854] "Providing to the user" means supplying the edited video in a form that the user can access.
[0855] This invention is a system that automatically edits soccer match footage on the cloud and provides videos with commentary, eliminating the need for users to record and edit the footage while watching the match.
[0856] System configuration
[0857] The system includes the following main components:
[0858] 1. Photography Method:
[0859] Users set up fixed cameras to film the entire soccer match. The fixed cameras are fixed on the sidelines or behind the goals, and use wide-angle lenses to capture the entire match.
[0860] 2. Upload method:
[0861] Users upload the footage they have taken to a cloud platform. The video files are sent to a cloud server via the Internet. The cloud platform uses commonly used storage services (e.g., Google Drive or Dropbox).
[0862] 3. Receiving and preprocessing means:
[0863] The server receives the video uploaded to the cloud. After receiving it, it checks the format and quality of the video and performs any necessary pre-processing (e.g., adjusting the resolution or converting the frame rate) using a video processing library (e.g., FFmpeg).
[0864] 4. Analysis method:
[0865] The AI analysis system on the server analyzes the video data frame by frame. The AI uses an image processing algorithm (e.g., YOLO (You Only Look Once)) to detect the position and movement of the ball and players in each frame. Based on this detected information, the ball and players are tracked in real time.
[0866] 5. Extraction means:
[0867] The server analyzes the tracking data to detect when specific events (e.g., goals, saves, fouls) occur, and automatically extracts key moments based on these events to generate highlight clips.
[0868] 6. Explanation Generation Method:
[0869] The server uses a generative AI model to generate commentary in the style of a specific commentator. Using a prompt as input, the generative AI model generates appropriate commentary. For example, a prompt such as "Player A made a pass and Player B scored a goal" is used. Based on the generated commentary, a voice synthesis technology (e.g., Google Text-to-Speech API) is used to create a commentary voice.
[0870] 7. Integration Methods:
[0871] The server integrates the commentary audio and subtitles into the video to create an edited highlight video. A video editing library (e.g., OpenCV or MoviePy) is used to integrate the audio file with the video and overlay the subtitle text.
[0872] 8. Means of provision:
[0873] The server stores the edited video in cloud storage (e.g., AWS S3), and users can access the cloud platform to view or download the edited video.
[0874] Specific examples
[0875] For example, suppose a user films their child's soccer game and uploads the footage to the cloud. The server receives the footage and performs preprocessing, and the AI analysis system detects and tracks the movement of the ball and players. Goals and important plays are detected, and highlights are automatically generated based on them. The generative AI model generates commentary such as, "Player A made a decisive pass here, and Player B scored a shot!" and adds it to the video as audio and subtitles. Finally, the edited footage is stored in the cloud, where the user can enjoy it with their family.
[0876] In this way, the present invention is a system that can significantly reduce the effort required for filming and automatically provide professional-quality video.
[0877] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0878] Step 1:
[0879] The user sets up a fixed camera and captures footage of the match.
[0880] Specifically, the user fixes a fixed camera on the sideline or behind the goal and adjusts the camera settings, for example, using a wide-angle lens to capture the entire game in the frame.
[0881] Input: The physical camera equipment and the soccer match to be filmed.
[0882] Output: Footage files of the captured match.
[0883] Step 2:
[0884] The user uploads the footage they have taken to a cloud platform.
[0885] Users use their home computer or smartphone to access the designated upload form on the cloud platform, select and submit the video file.
[0886] Input: The captured video file.
[0887] Output: Video files uploaded to the cloud platform.
[0888] Step 3:
[0889] The server receives the video uploaded to the cloud platform and performs pre-processing.
[0890] The server uses a video processing library (e.g., FFmpeg) to check the format and quality of the received video file and make any necessary resolution adjustments and frame rate conversions, for example, converting 4K video to 1080p and changing the frame rate from 60 fps to 30 fps.
[0891] Input: Video files on the cloud.
[0892] Output: Pre-processed video files.
[0893] Step 4:
[0894] The server uses an AI analysis system to analyze the pre-processed video data frame by frame.
[0895] The server uses an image processing algorithm (e.g., YOLO) to detect the positions of the ball and players in each frame. Specifically, it obtains the coordinate data of the bounding boxes of the ball and players included in each frame.
[0896] Input: Preprocessed video files.
[0897] Output: Ball and player position data.
[0898] Step 5:
[0899] The server analyzes the tracking data to detect when a particular event (e.g., goal, save, foul) occurs.
[0900] The server records the timestamps of the frames corresponding to these events and automatically extracts important scenes, such as five seconds before and after a goal.
[0901] Input: Ball and player position data.
[0902] Output: timestamps and footage clips of key scenes.
[0903] Step 6:
[0904] The server uses a generative AI model to generate commentary in the style of a specific commentator.
[0905] The server inputs a prompt, and the generative AI model generates an appropriate commentary. For example, a prompt such as "Player A passes the ball, and Player B scores a goal" is used.
[0906] Input: prompt text and video clips of key scenes.
[0907] Output: The generated description.
[0908] Step 7:
[0909] The server creates a commentary audio based on the generated commentary text using speech synthesis technology (e.g., Google Text-to-Speech API).
[0910] The commentary is converted into an audio file and is also prepared as subtitle text. In detail, the server sends the commentary to the speech synthesis API and obtains the audio file.
[0911] Input: The generated description.
[0912] Output: Descriptive audio file and subtitle text.
[0913] Step 8:
[0914] The server integrates audio commentary and subtitles into the video clips to create an edited highlight video.
[0915] Use a video editing library (e.g., OpenCV or MoviePy) to integrate the audio file and subtitle text into the video clip. Specifically, the audio description is overlaid on the video and the subtitle text is displayed in the appropriate position.
[0916] Input: Descriptive audio files, subtitle text, and video clips of key scenes.
[0917] Output: Edited highlight video.
[0918] Step 9:
[0919] The server saves the completed highlight video to cloud storage (e.g. AWS S3).
[0920] Generate a link that allows users to access the saved video and provide that link to the user. For example, generate a URL like "https: / / cloudstorage.example.com / highlight_video.mp4".
[0921] Input: edited highlight video.
[0922] Output: Highlight video URL on cloud storage.
[0923] Step 10:
[0924] The user accesses the cloud platform and clicks on the provided link to view the edited footage.
[0925] Users can use a browser or a dedicated app to play or download videos from cloud storage.
[0926] Input: Highlight video URL on cloud storage.
[0927] Output: Highlights viewed or downloaded.
[0928] Through the above process, the system automatically edits soccer match footage shot by the user and provides a professional highlight video with commentary.
[0929] (Application example 1)
[0930] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0931] Shooting soccer game footage, editing the footage, and creating a highlight video with commentary has traditionally required a great deal of effort. Editing the footage and creating the commentary, in particular, requires time and effort, making it difficult to easily achieve professional-quality results. The objective of this invention is to solve this problem and provide a system that enables users to easily create high-quality highlight videos with commentary.
[0932] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0933] In this invention, the server includes means for uploading video captured by a fixed camera to the cloud, means for receiving the uploaded video and detecting the ball and players, means for automatically extracting key scenes based on the detected information, means for adding automatically generated commentary to the video using a generative AI model, means for integrating commentary audio and subtitles to complete the edited video, and means for saving the edited video in cloud storage and enabling users to view or share it. This allows users to easily generate, view, and share professional-quality highlight videos with commentary simply by uploading their game video to the cloud.
[0934] A "fixed camera" is a camera that captures video from a fixed position.
[0935] The "cloud" is a collection of servers for storing, managing, and processing data over the Internet.
[0936] "Uploading" is the act of sending data from a local device to a cloud server.
[0937] "Reception" refers to the act of the cloud server receiving data sent from the user.
[0938] "Detection" is the process of identifying a specific object (the ball or a player) through video analysis.
[0939] A "Key Moment" is a specific key moment during a match (e.g., a goal, save, foul).
[0940] A "generative AI model" is an artificial intelligence model that uses machine learning algorithms to generate content such as text and images.
[0941] "Explanation" is written or audio that explains a particular event.
[0942] The "explanatory voice" is the generated explanatory text output as voice through voice synthesis.
[0943] "Subtitles" are textual information displayed on video.
[0944] "Editing" is the process of adding audio commentary and subtitles to the footage to complete the final video.
[0945] "Cloud storage" is a storage service for storing data on the cloud.
[0946] "Viewing" refers to the act of a user playing and watching a video.
[0947] "Sharing" refers to the act of a user sharing video or data with other people.
[0948] This invention is a system that automatically edits soccer match footage taken with fixed cameras and provides videos with commentary, eliminating the need for users to record and edit the footage while watching the match.
[0949] System configuration
[0950] The system includes the following main components:
[0951] 1. Fixed camera
[0952] Users can set up fixed cameras on the sidelines or behind the goals and use wide-angle lenses to capture the entire game.
[0953] 2. Cloud Upload
[0954] Users upload the footage they have taken to the cloud platform, and the video files are sent to the cloud server via the internet.
[0955] 3. Reception and Preprocessing
[0956] The server receives the video uploaded to the cloud. After receiving the video, it checks the format and quality of the video and performs any necessary preprocessing (e.g., adjusting the resolution and converting the frame rate). For preprocessing, it uses an open-source video processing library such as OpenCV.
[0957] 4. AI analysis
[0958] The AI analysis system on the server analyzes the video data frame by frame, and uses OpenCV and deep learning models to detect the position of the ball and the positions and movements of players in each frame.
[0959] 5. Extracting important scenes
[0960] The server analyzes the tracking data to detect when specific events (e.g., goals, saves, fouls) occur, and automatically extracts key moments based on these events to generate highlight clips.
[0961] 6. Explanation Generation
[0962] The server uses a generative AI model (e.g., GPT-4) to generate commentary in the style of a specific commentator. It then uses voice synthesis based on the generated commentary to create the commentary audio. For voice synthesis, it uses a service like Amazon Polly.
[0963] 7. Video Editing
[0964] The commentary audio and subtitles are integrated into the video, and video editing libraries such as FFmpeg are used for video editing. This completes the edited highlight video.
[0965] 8. Cloud Storage
[0966] The edited footage is saved to cloud storage, allowing users to view or share it.
[0967] Specific examples
[0968] For example, a user can film their child's soccer game with their smartphone and upload the footage to the cloud. The server receives the footage and performs pre-processing. The AI analysis system detects the movement of the ball and players and extracts goals and other important plays. Based on this, the generative AI model generates commentary such as, "Player A made a decisive pass here, and Player B scored a shot!" and adds it to the video as audio and subtitles. Finally, the edited video is stored in the cloud, and users can share it with family and friends.
[0969] Prompt Sentence Examples
[0970] For example, when generating explanatory text using a generative AI model, the following prompt is used:
[0971] "I uploaded a video of my child's soccer game. Analyze the ball position and player movements to detect goals and important plays. Based on that, generate commentary such as, 'Player A made the decisive pass, and Player B scored the shot!' and add it to the video as audio and subtitles."
[0972] In this way, the present invention is a system that can significantly reduce the effort required for filming and automatically provide professional-quality video.
[0973] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0974] Step 1:
[0975] A user films a soccer match with a fixed camera. The fixed camera uses a wide-angle lens to capture the entire game, and is fixed on the sidelines or behind the goal. The input is the video data captured by the camera, and the output is a video file stored in local storage.
[0976] Step 2:
[0977] Users upload the videos they have taken from their smartphones to the cloud platform. The uploaded videos are then sent to the cloud server via the Internet. The input is the video file stored on the smartphone, and the output is the video data on the cloud storage.
[0978] Step 3:
[0979] The server receives the video uploaded to the cloud and performs preprocessing. This includes adjusting the resolution and converting the frame rate, and uses a video processing library such as OpenCV. The input is the video data stored in the cloud storage, and the output is the preprocessed video data converted into an analyzable format.
[0980] Step 4:
[0981] The AI analysis system in the server analyzes the preprocessed video frame by frame to detect the position and movement of the ball and players. This analysis uses OpenCV and deep learning models. The input is the preprocessed video data, and the output is tracking data including the position information of the ball and players.
[0982] Step 5:
[0983] The server automatically extracts important scenes from the tracking data. Specifically, it detects events such as goals, saves, and fouls and extracts those moments as highlight clips. The input is the tracking data, and the output is a list of important scenes and their clip data.
[0984] Step 6:
[0985] The server uses a generative AI model (e.g., GPT-4) to create automatically generated commentary. The generated commentary is converted into speech using speech synthesis technology (e.g., Amazon Polly). The input is a list of important scenes and prompts, and the output is commentary audio and subtitle data.
[0986] Step 7:
[0987] The server integrates the commentary audio and subtitles into the video to create an edited highlight video. This editing is done using a video editing library such as FFmpeg. The input is the clip data, commentary audio, and subtitle data, and the output is the edited video.
[0988] Step 8:
[0989] The server saves the completed edited video in cloud storage. Users can access the cloud platform to view and share the video. The input is the edited video data, and the output is the video data saved in cloud storage.
[0990] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0991] This invention is a system that automatically edits soccer match footage on the cloud and provides videos with commentary using an emotion engine. This system saves users the trouble of recording and editing the game while watching it. It also adjusts the commentary content based on the user's emotions, providing a more personalized video experience.
[0992] System configuration
[0993] The system includes the following main components:
[0994] 1. Photography Method:
[0995] Users set up fixed cameras to film the entire soccer match. The fixed cameras are fixed on the sidelines or behind the goals, and use wide-angle lenses to capture the entire match.
[0996] 2. Upload method:
[0997] Users upload the footage they have taken to the cloud platform, and the video files are sent to the cloud server via the internet.
[0998] 3. Receiving and preprocessing means:
[0999] The server receives the video uploaded to the cloud, checks the format and quality of the video, and performs any necessary pre-processing (e.g., adjusting the resolution or converting the frame rate).
[1000] 4. Analysis method:
[1001] The AI analysis system on the server analyzes the received video data frame by frame, detecting the ball and players in the video and saving their positional information as tracking data.
[1002] 5. Extraction means:
[1003] The server automatically extracts important scenes (goals, fouls, saves, etc.) based on the tracking data, and compiles the extracted scenes to generate highlight clips.
[1004] 6. Emotion Engine:
[1005] The server is equipped with an emotion engine that recognizes the user's emotions. It analyzes the user's emotion data (e.g., facial expression recognition and biometric data) and grasps the user's emotional state in real time.
[1006] 7. Explanation Generation Method:
[1007] The server uses a generative AI model to generate commentary in the style of a specific commentator, adjusts the commentary content based on the user's emotions recognized by the emotion engine, and performs voice synthesis based on the generated commentary to create a commentary voice.
[1008] 8. Means of provision:
[1009] The server stores the edited video in cloud storage, and users can access the cloud platform to view or download the edited video.
[1010] Program processing and specific examples
[1011] Natural language explanation of the process
[1012] Users film the match with fixed cameras and upload the footage to the cloud after the match ends. The server preprocesses the received footage and uses an AI analysis system to detect the ball and players. An emotion engine then collects and analyzes the user's emotions in real time. Key scenes are then extracted, and a generative AI model generates commentary based on the emotion data and integrates it into the video. Finally, the edited video is stored in the cloud and made available for users to watch or download.
[1013] Specific examples
[1014] For example, suppose a user films their child's soccer game and uploads the footage to the cloud. The server receives the footage and performs preprocessing, while the AI analysis system detects and tracks the movement of the ball and players. At the same time, while the user is watching the game, the emotion engine analyzes the user's facial expressions, heart rate, and other data to collect emotional data. Goals and important plays are detected, and highlights are automatically generated based on them. The generative AI model generates commentary such as, "Player A made a decisive pass here, and Player B scored a shot!" If it determines that the user is happy, the commentary will be written in a positive tone. Finally, the edited video is saved in the cloud, where the user can enjoy it with their family.
[1015] In this way, the present invention not only significantly reduces the effort required for filming and automatically provides professional-quality video, but also provides a more personalized video experience based on the user's emotions.
[1016] The processing flow will be explained below.
[1017] Step 1:
[1018] A user films a soccer match with a fixed camera. The camera is set up and recording begins. The video is set to capture the entire match, using a wide-angle lens, etc.
[1019] Step 2:
[1020] The video footage taken by the user is uploaded to the cloud server. After the match, the user connects the camera's memory card to a computer and uploads the video files to the cloud platform.
[1021] Step 3:
[1022] The server receives the video uploaded to the cloud, checks the format and quality of the received video file, and performs pre-processing such as adjusting the resolution and converting the frame rate as necessary.
[1023] Step 4:
[1024] The AI analysis system on the server analyzes the received video data frame by frame, detecting the ball and players in the video and saving their positional information as tracking data.
[1025] Step 5:
[1026] The server automatically extracts important scenes (goals, fouls, saves, etc.) based on the tracking data, and compiles the extracted scenes to generate a highlight clip.
[1027] Step 6:
[1028] As users watch a game, the emotion engine collects their facial expressions and heart rate in real time to generate emotion data, which is then sent to a cloud server.
[1029] Step 7:
[1030] The server receives and analyzes the emotion data sent from the emotion engine to understand the user's emotional state (happiness, excitement, sadness, etc.).
[1031] Step 8:
[1032] The server uses a generative AI model to generate commentary in the style of a specific commentator, adjusting the commentary based on emotion data, for example, using positive expressions if the user is happy.
[1033] Step 9:
[1034] The server integrates audio commentary and subtitles into the video, and edits and completes the highlight video with commentary.
[1035] Step 10:
[1036] The server saves the edited video to cloud storage and notifies the user that the video is ready.
[1037] Step 11:
[1038] Users log in to the cloud platform to view or download the edited footage, and share it with family and team members.
[1039] Example 2
[1040] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1041] Conventional soccer match filming and editing systems require lengthy editing work after filming, and adding commentary is labor-intensive. Furthermore, they are unable to reflect the user's emotional state while watching the video, making it difficult to provide a personalized viewing experience. There is a need for a system that can solve these problems and easily provide high-quality, personalized video.
[1042] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1043] In this invention, the server includes means for uploading video captured by a fixed camera to the cloud, means for receiving the uploaded video and pre-processing the video, means for detecting the ball and players from the received video, means for automatically extracting important scenes based on the detected information, means for collecting user emotion data, means for analyzing the collected emotion data and generating commentary using a generative AI model, means for adding the generated commentary to the video, and means for providing the edited video to the user, thereby enabling users to watch high-quality video containing match highlights and personalized commentary without any hassle.
[1044] A "fixed camera" is a camera that is fixed at a specific location and captures a wide range of images using a wide-angle lens.
[1045] "Cloud" is a general term for an online platform that provides data storage and computing resources by accessing remote servers via the Internet.
[1046] "Uploading means" refers to the method or technology used to transfer digital data from a local device to a remote server, such as cloud storage.
[1047] "Means of receiving" refers to the method or technology by which the remote server obtains data uploaded to the cloud.
[1048] "Preprocessing means" refers to processing carried out in advance, such as converting data format or adjusting quality, to ensure smooth subsequent analysis and editing processes.
[1049] "Detection means" refers to methods or technologies that automatically recognize specific objects (e.g., balls or players) from video data and track their location information.
[1050] "Means for automatically extracting important scenes" refers to methods and technologies that use detected information to identify specific events such as goals and fouls and extract them as highlights from the video.
[1051] "Means for collecting emotional data" refers to methods and technologies for acquiring a user's facial expressions and biometric data in real time and analyzing their emotional state.
[1052] A "generative AI model" is a technology that uses a pre-trained artificial intelligence model to generate explanatory text in natural language from given data and prompts.
[1053] "Means for adding commentary" refers to methods or techniques for integrating the generated commentary text and audio into the video data and providing it as the final edited video.
[1054] "Means for providing edited video" refers to the methods and technologies for storing the final edited video in cloud storage and allowing users to access, view, or download it.
[1055] This system automatically edits soccer match footage on the cloud and provides videos with commentary using an emotion engine. This allows users to concentrate on watching the game, while automating the editing and commentary tasks. It also provides a personalized video experience based on the user's emotions.
[1056] First, the user sets up a fixed camera to film a soccer match. The fixed camera uses a wide-angle lens and is installed on the sidelines or behind the goal. After the match ends, the user uploads the filmed video files to a cloud platform. This can be done using cloud storage services such as Google Drive or Dropbox.
[1057] The server then receives the uploaded video on the cloud, checks the format and quality of the video file, and performs any necessary pre-processing, such as converting AVI video to MP4, reducing the resolution from 1080p to 720p, and converting the frame rate from 30fps to 24fps.
[1058] Once preprocessing is complete, the server uses an AI analysis system to analyze the received video data frame by frame. The AI analysis system uses deep learning models and libraries such as OpenCV to detect the ball and players in the video and save their location information as tracking data. This allows the progress of the game and the movements of the players to be understood.
[1059] The server then automatically extracts important moments (e.g., goals, fouls, saves, etc.) from the tracking data. In this step, AI identifies specific events and cuts them out into highlight videos.
[1060] Furthermore, the server uses an emotion engine to collect user emotional data, including biometric data such as facial expressions and heart rate, and analyzes it in real time, allowing the server to understand the user's emotional state during a match.
[1061] The server then uses a generative AI model to generate commentary in the style of a particular commentator. The generative AI model is given the following prompt:
[1062] "Generate commentary on goal scenes."
[1063] "Please provide an inspiring commentary based on the emotional data."
[1064] "Describe a highlight scene of a key play."
[1065] The generated commentary text is used to synthesize speech, creating a commentary voice, which adds a commentary tailored to the user's emotions to the video.
[1066] Finally, the server integrates the generated commentary audio into the highlight video. Specifically, it uses video editing software such as FFmpeg to combine the highlight footage and commentary audio into a single file. The edited video file is then stored in cloud storage, and users can access the cloud platform to watch or download it.
[1067] For example, consider the case where a user films a child's soccer game and uploads the video to the cloud. The server receives the video and performs preprocessing, and the AI analysis system detects and tracks the movement of the ball and players. Then, while the user is watching the game, the emotion engine analyzes the user's facial expressions, heart rate, etc., and collects emotional data. Based on this data, the generative AI model generates commentary, providing commentary in positive terms, allowing the user to enjoy watching the highlights of the game.
[1068] In this way, the present invention not only significantly reduces the effort required for filming and editing, and automatically provides professional-quality video, but also enables a personally tailored video experience that reflects the user's emotions.
[1069] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1070] Step 1:
[1071] The user sets up a fixed camera to film a soccer match. The fixed camera used for recording uses a wide-angle lens and is fixed to the sideline or behind the goal. After the match ends, the user transfers the filmed footage to a computer. The input data is the filmed video file, and the output data is the video file saved on the computer.
[1072] Step 2:
[1073] Users upload video files stored on their computers to a cloud platform. Specifically, they use cloud storage services such as Google Drive and Dropbox to upload files. The input data is the video file stored on their computer, and the output data is the video file stored on the cloud.
[1074] Step 3:
[1075] The server receives the video uploaded to the cloud. After receiving it, it checks the format and quality of the video file and performs any necessary pre-processing. Pre-processing includes converting AVI format video to MP4 format, reducing the resolution from 1080p to 720p, and converting the frame rate from 30fps to 24fps. The input data is the video file on the cloud, and the output data is the video file after pre-processing. Specific operations involve the use of video conversion software such as FFmpeg.
[1076] Step 4:
[1077] The server uses an AI analysis system to analyze the pre-processed video data frame by frame. The AI analysis system uses a deep learning model and libraries such as OpenCV to detect the ball and players in the video. It then saves this position information as tracking data. The input data is the pre-processed video file, and the output data is the analysis results, including the tracking data.
[1078] Step 5:
[1079] The server automatically extracts key moments (e.g., goals, fouls, saves, etc.) from the tracking data. This involves AI identifying specific events and extracting them into highlight videos. The input data is the tracking data, and the output data is video clips of key moments.
[1080] Step 6:
[1081] The server uses an emotion engine to collect user emotion data in real time. It also uses devices such as webcams and smartwatches to acquire biometric data such as the user's facial expressions and heart rate. The input data is the user's biometric data, and the output data is the analyzed emotion data.
[1082] Step 7:
[1083] The server uses a generative AI model to generate commentary in the style of a specific commentator. The generative AI model is given the following prompt:
[1084] "Generate commentary on goal scenes."
[1085] "Please provide an inspiring commentary based on the emotional data."
[1086] "Describe a highlight scene of a key play."
[1087] The input data is emotion data and prompt sentences, and the output data is the generated commentary. Specifically, a generative AI model such as GPT-3 is used.
[1088] Step 8:
[1089] The server synthesizes the commentary based on the generated commentary text and creates the commentary audio. The input data is the generated commentary text, and the output data is the commentary audio file. Specifically, speech synthesis software is used.
[1090] Step 9:
[1091] The server then integrates the generated commentary audio into the highlight video. Specifically, it uses video editing software such as FFmpeg to combine the highlight video and commentary audio into a single file. The input data is the commentary audio file and the video clips of key scenes, and the output data is the edited video file.
[1092] Step 10:
[1093] The server stores the edited video files in cloud storage, allowing users to access and watch or download them. The input data is the edited video files, and the output data is the video files stored on the cloud. Specifically, a cloud storage platform is used.
[1094] (Application example 2)
[1095] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1096] Conventional video editing systems require manual editing after shooting footage, which requires a great deal of time and effort. Furthermore, the generated commentary is standardized, making it difficult to provide a personalized experience that reflects each user's individual emotions and circumstances. This requires a lot of effort for users to record and edit events, and the viewing experience is not personalized enough.
[1097] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for uploading video captured by a fixed camera to the cloud, means for receiving the uploaded video and detecting objects and people, means for automatically extracting important scenes based on the detected information, generation AI model means for analyzing user emotion data and generating customized commentary based on the analysis results, means for adding the generated commentary to the video, and means for providing the edited video to the user. This eliminates the need for manual editing of the video and makes it possible to provide a personalized viewing experience based on the user's emotions.
[1098] A "fixed camera" is a camera that continuously captures a specific area from a fixed position.
[1099] The "cloud" refers to computing resources and data storage provided via the Internet.
[1100] "Uploading" is the act of sending local data to a remote server over the Internet or a network.
[1101] "Object" is a general term that refers to any concrete thing that can be recognized in an image.
[1102] "Person" is a general term that refers to a person recognized in the video.
[1103] "Detection" is the act of using specific algorithms to identify and track objects or people in a video.
[1104] "Key scenes" refer to moments or events that are particularly noteworthy in footage of sporting events or everyday life.
[1105] "Extraction" is the act of extracting a specific part from the whole data.
[1106] "Emotional data" refers to the emotional state obtained from the user's facial expressions, biometric data, etc.
[1107] A "generative AI model" is an algorithm that uses machine learning and natural language processing to automatically generate explanatory text and other content.
[1108] "Commentary" refers to text or audio that adds explanations or comments to the content of the video.
[1109] "Customization" is the act of tailoring content to a particular user or situation.
[1110] "Edited footage" refers to the final video file in which processing and modification have been applied to the original video data.
[1111] "Providing" refers to the act of making a video available for viewing or acquisition by a user through the cloud or other media.
[1112] The following system is used as an embodiment of the present invention: The main components of the system and their respective roles will be described below.
[1113] System Components
[1114] 1. Fixed camera:
[1115] Used by users to capture events, fixed cameras are fixed in specific locations and use wide-angle lenses to capture a wide range of footage.
[1116] 2. Cloud Platform:
[1117] Users upload footage taken with fixed cameras, and the cloud platform receives the video files via the internet and performs the analysis and editing processes described below.
[1118] 3. Receiving and Preprocessing Server:
[1119] The server receives the uploaded video data. It checks the format and quality of the video and performs any necessary preprocessing (adjusting the resolution and converting the frame rate).
[1120] 4. AI analysis system:
[1121] Video data is analyzed frame by frame, objects and people are detected in the video, and their location information is saved as tracking data.
[1122] 5. Sentiment Analysis System:
[1123] Analyze the user's emotional state. Collect user emotional data (facial recognition and biometric data) using sensors on smartphones and smart glasses, and analyze it in real time.
[1124] 6. Generative AI Models:
[1125] Based on the collected emotional data, a commentary in the style of a specific commentator is generated. Speech synthesis is performed based on the generated commentary to create a commentary voice. The software used here includes a natural language processing library and a speech synthesis module.
[1126] 7. Editing System:
[1127] Based on the analysis results and audio commentary, the video is edited to generate highlight clips. This editing system uses video editing libraries (e.g., OpenCV and MoviePy).
[1128] 8. Delivery system:
[1129] The edited video is stored in cloud storage, and users can view or download the edited video from the cloud platform.
[1130] A description of what the program does
[1131] In this system, a user first captures an event with a fixed camera and uploads the footage to a cloud platform. The cloud server receives the video data, checks its format and quality, and performs any necessary preprocessing. The AI analysis system then analyzes the received video data frame by frame to obtain information for detecting and tracking objects and people.
[1132] At the same time, devices such as smartphones and smart glasses are used to collect the user's emotional data (e.g., facial expressions, heart rate, etc.), which is then analyzed in real time by an emotion analysis system to determine the user's emotional state.
[1133] Based on the results of the sentiment analysis and AI analysis, the generative AI model generates commentary in the style of a specific commentator. For example, if it determines that the user is happy, a positive commentary will be selected. The generated commentary is converted into an audio file using text-to-speech software. These commentary voices are then integrated into existing footage, and the editing system generates a highlight clip.
[1134] Finally, the edited video is stored in cloud storage. Users can access the cloud platform and watch or download the generated personalized video. This process allows users to easily obtain high-quality edited video and provides a personalized experience based on their emotions.
[1135] Specific examples
[1136] For example, if a user films a family birthday party, the app uploads the video to the cloud, and the AI performs emotion analysis and video editing as follows: When the emotion analysis system determines that the user is happy at the climax of the party, the generative AI model generates a caption such as, "Here, the whole family has fun and cut the cake!"
[1137] Prompt Sentence Examples
[1138] Here is a video of a family birthday party. Highlight the most exciting scenes and generate commentary based on the emotion data.
[1139] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1140] Step 1:
[1141] A user captures an event with a fixed camera or a mobile information terminal. As input, the user obtains video data. As output, a video file is generated.
[1142] Step 2:
[1143] The user uploads the video they have taken to the cloud platform. The input is the video file they have taken, which the cloud server receives. The output is the video data stored on the cloud.
[1144] Step 3:
[1145] The server receives the uploaded video data and checks the format and quality. The input is the video data stored on the cloud server, and the server performs preprocessing such as adjusting the resolution and converting the frame rate. The output is the preprocessed video data.
[1146] Step 4:
[1147] The server's AI analysis system analyzes the preprocessed video data frame by frame to detect objects and people. The input is the preprocessed video data, and the server applies detection algorithms to identify objects and people, saving their location information as tracking data. The output is tracking data.
[1148] Step 5:
[1149] The device (smartphone or smart glasses) collects emotion data. The input is sensor data such as the user's facial expression and heart rate, and the device sends this data to the emotion analysis system. The output is the user's emotion data as an analysis result.
[1150] Step 6:
[1151] The server uses a generative AI model to generate an explanatory text based on the user's emotional data and the results of AI analysis. The inputs are emotional data and tracking data, and the generative AI model generates a prompt based on this data and then generates an explanatory text (e.g., "That was a great goal!" if the emotion is positive). The output is an explanatory text.
[1152] Step 7:
[1153] The server converts the generated commentary into an audio file using a speech synthesis module. The commentary is input, the server synthesizes the speech, and obtains the audio file. The audio file is obtained as output.
[1154] Step 8:
[1155] The server's editing system integrates the video and audio commentary to generate a highlight clip. The inputs are the pre-processed video data, tracking data, and audio commentary files, and the server uses a video editing library to integrate them. The output is an edited highlight video.
[1156] Step 9:
[1157] The server stores the edited footage in cloud storage for users to access. As input, there is the edited highlight footage, which the server stores in cloud storage. As output, users can watch or download the footage from the cloud platform.
[1158] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1159] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1160] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1161] [Fourth embodiment]
[1162] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1163] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1164] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1165] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1166] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1167] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1168] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1169] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1170] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1171] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1172] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1173] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1174] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1175] This invention is a system that automatically edits soccer match footage on the cloud and provides videos with commentary, eliminating the need for users to record and edit the footage while watching the match.
[1176] System configuration
[1177] The system includes the following main components:
[1178] 1. Photography Method:
[1179] Users set up fixed cameras to film the entire soccer match. The fixed cameras are fixed on the sidelines or behind the goals, and use wide-angle lenses to capture the entire match.
[1180] 2. Upload method:
[1181] Users upload the footage they have taken to the cloud platform, and the video files are sent to the cloud server via the internet.
[1182] 3. Receiving and preprocessing means:
[1183] The server receives the video uploaded to the cloud, checks the format and quality of the video, and performs any necessary pre-processing (e.g., adjusting the resolution or converting the frame rate).
[1184] 4. Analysis method:
[1185] The AI analysis system on the server analyzes the video data frame by frame. The AI uses image processing algorithms to detect the position and movement of the ball and players in each frame. Based on this information, the system tracks the ball and players in real time.
[1186] 5. Extraction means:
[1187] The server analyzes the tracking data to detect when specific events (e.g., goals, saves, fouls) occur, and automatically extracts key moments based on these events to generate highlight clips.
[1188] 6. Explanation Generation Method:
[1189] The server uses a generative AI model to generate commentary in the style of a specific commentator, synthesizes the generated commentary to create a commentary audio, and integrates the commentary audio and subtitles into the video to create an edited highlight video.
[1190] 7. Means of provision:
[1191] The server stores the edited video in cloud storage, and users can access the cloud platform to view or download the edited video.
[1192] Program processing and specific examples
[1193] Natural language explanation of the process
[1194] Users film the match with fixed cameras and upload the footage to the cloud after the match ends. The server preprocesses the received footage and uses an AI analysis system to detect the ball and players. It then extracts key scenes, generates commentary using a generative AI model, and adds commentary audio and subtitles to the video. The edited footage is then stored in the cloud and made available for users to watch or download.
[1195] Specific examples
[1196] For example, suppose a user films their child's soccer game and uploads the footage to the cloud. The server receives the footage and performs preprocessing, and the AI analysis system detects and tracks the movement of the ball and players. Goals and important plays are detected, and highlights are automatically generated based on them. The generative AI model generates commentary such as, "Player A made a decisive pass here, and Player B scored a shot!" and adds it to the video as audio and subtitles. Finally, the edited footage is stored in the cloud, where the user can enjoy it with their family.
[1197] In this way, the present invention is a system that can significantly reduce the effort required for filming and automatically provide professional-quality video.
[1198] The processing flow will be explained below.
[1199] Step 1:
[1200] A user films a soccer match with a fixed camera. The camera is set up and recording begins. The video is set to capture the entire match, using a wide-angle lens, etc.
[1201] Step 2:
[1202] The video footage taken by the user is uploaded to the cloud server. After the match, the user connects the camera's memory card to a computer and uploads the video files to the cloud platform.
[1203] Step 3:
[1204] The server receives the video uploaded to the cloud, checks the format and quality of the received video file, and performs pre-processing such as adjusting the resolution and converting the frame rate as necessary.
[1205] Step 4:
[1206] The AI analysis system on the server analyzes the received video data frame by frame, detecting the ball and players in the video and saving their positional information as tracking data.
[1207] Step 5:
[1208] The server automatically extracts important scenes (goals, fouls, saves, etc.) based on the tracking data, and compiles the extracted scenes to generate a highlight clip.
[1209] Step 6:
[1210] The server uses the generative AI model to generate commentary in the style of a specific commentator, and then performs voice synthesis based on the generated commentary to create the commentary audio.
[1211] Step 7:
[1212] The server integrates audio commentary and subtitles into the video, and edits and completes the highlight video with commentary.
[1213] Step 8:
[1214] The server saves the edited video to cloud storage and notifies the user that the video is ready.
[1215] Step 9:
[1216] Users log in to the cloud platform to view or download the edited footage, and share it with family and team members.
[1217] Example 1
[1218] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1219] Previously, shooting, editing, and creating a video with commentary of a soccer game required a great deal of time and effort. Especially for amateur and family games, where professional video editing skills and commentators are not required, there was a demand for a method to easily create high-quality video with commentary, but no system existed that could achieve this. Furthermore, manual editing and the use of specialized software required specialized knowledge, making it a significant burden for average users.
[1220] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1221] In this invention, the server includes a means for uploading video to the cloud, a means for preprocessing the received video, a means for detecting the ball and players using an image processing algorithm, a means for automatically extracting important scenes and generating highlight clips, a means for creating commentary audio using a generative AI model for generating commentary text and speech synthesis technology, and a means for integrating the commentary audio and subtitles into the video, and a means for saving the edited video in cloud storage and providing it to users, thereby enabling users to easily create highlight videos with professional quality commentary and easily watch and share them.
[1222] A "fixed camera" is a camera that is fixed in a certain position and continuously captures a specific range or scene.
[1223] The "cloud" is a service platform for storing and processing data on remote servers on the Internet.
[1224] "Uploading" is the act of transferring data stored on a local device to a remote server over the Internet.
[1225] "Preprocessing" refers to the initial processing performed on received data, primarily to improve data analysis and algorithm performance.
[1226] An "image processing algorithm" is a computational method for extracting and analyzing specific features or information from image data.
[1227] "Ball and player detection" refers to identifying the position of the ball and players in each frame of the video and tracking their movements.
[1228] A "key moment" is a particularly noteworthy event in a soccer match (e.g., a goal, a save, a foul, etc.).
[1229] A "highlight clip" is a shortened, edited video clip that selects particularly important scenes from the entire game.
[1230] A "generative AI model" is an artificial intelligence model that generates natural-looking explanatory text and content based on given data and prompt text.
[1231] "Speech synthesis technology" is a technology that converts text data into actual speech.
[1232] "Commentary audio" is audio data that has been converted using speech synthesis technology from explanatory text created by a generative AI model.
[1233] "Subtitles" are textual information that is visually displayed within a video and complements the audio content.
[1234] "Cloud storage" is a data storage service on the Internet that stores data on a remote server and allows it to be accessed as needed.
[1235] "Providing to the user" means supplying the edited video in a form that the user can access.
[1236] This invention is a system that automatically edits soccer match footage on the cloud and provides videos with commentary, eliminating the need for users to record and edit the footage while watching the match.
[1237] System configuration
[1238] The system includes the following main components:
[1239] 1. Photography Method:
[1240] Users set up fixed cameras to film the entire soccer match. The fixed cameras are fixed on the sidelines or behind the goals, and use wide-angle lenses to capture the entire match.
[1241] 2. Upload method:
[1242] Users upload the footage they have taken to a cloud platform. The video files are sent to a cloud server via the Internet. The cloud platform uses commonly used storage services (e.g., Google Drive or Dropbox).
[1243] 3. Receiving and preprocessing means:
[1244] The server receives the video uploaded to the cloud. After receiving it, it checks the format and quality of the video and performs any necessary pre-processing (e.g., adjusting the resolution or converting the frame rate) using a video processing library (e.g., FFmpeg).
[1245] 4. Analysis method:
[1246] The AI analysis system on the server analyzes the video data frame by frame. The AI uses an image processing algorithm (e.g., YOLO (You Only Look Once)) to detect the position and movement of the ball and players in each frame. Based on this detected information, the ball and players are tracked in real time.
[1247] 5. Extraction means:
[1248] The server analyzes the tracking data to detect when specific events (e.g., goals, saves, fouls) occur, and automatically extracts key moments based on these events to generate highlight clips.
[1249] 6. Explanation Generation Method:
[1250] The server uses a generative AI model to generate commentary in the style of a specific commentator. Using a prompt as input, the generative AI model generates appropriate commentary. For example, a prompt such as "Player A made a pass and Player B scored a goal" is used. Based on the generated commentary, a voice synthesis technology (e.g., Google Text-to-Speech API) is used to create a commentary voice.
[1251] 7. Integration Methods:
[1252] The server integrates the commentary audio and subtitles into the video to create an edited highlight video. A video editing library (e.g., OpenCV or MoviePy) is used to integrate the audio file with the video and overlay the subtitle text.
[1253] 8. Means of provision:
[1254] The server stores the edited video in cloud storage (e.g., AWS S3), and users can access the cloud platform to view or download the edited video.
[1255] Specific examples
[1256] For example, suppose a user films their child's soccer game and uploads the footage to the cloud. The server receives the footage and performs preprocessing, and the AI analysis system detects and tracks the movement of the ball and players. Goals and important plays are detected, and highlights are automatically generated based on them. The generative AI model generates commentary such as, "Player A made a decisive pass here, and Player B scored a shot!" and adds it to the video as audio and subtitles. Finally, the edited footage is stored in the cloud, where the user can enjoy it with their family.
[1257] In this way, the present invention is a system that can significantly reduce the effort required for filming and automatically provide professional-quality video.
[1258] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1259] Step 1:
[1260] The user sets up a fixed camera and captures footage of the match.
[1261] Specifically, the user fixes a fixed camera on the sideline or behind the goal and adjusts the camera settings, for example, using a wide-angle lens to capture the entire game in the frame.
[1262] Input: The physical camera equipment and the soccer match to be filmed.
[1263] Output: Footage files of the captured match.
[1264] Step 2:
[1265] The user uploads the footage they have taken to a cloud platform.
[1266] Users use their home computer or smartphone to access the designated upload form on the cloud platform, select and submit the video file.
[1267] Input: The captured video file.
[1268] Output: Video files uploaded to the cloud platform.
[1269] Step 3:
[1270] The server receives the video uploaded to the cloud platform and performs pre-processing.
[1271] The server uses a video processing library (e.g., FFmpeg) to check the format and quality of the received video file and make any necessary resolution adjustments and frame rate conversions, for example, converting 4K video to 1080p and changing the frame rate from 60 fps to 30 fps.
[1272] Input: Video files on the cloud.
[1273] Output: Pre-processed video files.
[1274] Step 4:
[1275] The server uses an AI analysis system to analyze the pre-processed video data frame by frame.
[1276] The server uses an image processing algorithm (e.g., YOLO) to detect the positions of the ball and players in each frame. Specifically, it obtains the coordinate data of the bounding boxes of the ball and players included in each frame.
[1277] Input: Preprocessed video files.
[1278] Output: Ball and player position data.
[1279] Step 5:
[1280] The server analyzes the tracking data to detect when a particular event (e.g., goal, save, foul) occurs.
[1281] The server records the timestamps of the frames corresponding to these events and automatically extracts important scenes, such as five seconds before and after a goal.
[1282] Input: Ball and player position data.
[1283] Output: timestamps and footage clips of key scenes.
[1284] Step 6:
[1285] The server uses a generative AI model to generate commentary in the style of a specific commentator.
[1286] The server inputs a prompt, and the generative AI model generates an appropriate commentary. For example, a prompt such as "Player A passes the ball, and Player B scores a goal" is used.
[1287] Input: prompt text and video clips of key scenes.
[1288] Output: The generated description.
[1289] Step 7:
[1290] The server creates a commentary audio based on the generated commentary text using speech synthesis technology (e.g., Google Text-to-Speech API).
[1291] The commentary is converted into an audio file and is also prepared as subtitle text. In detail, the server sends the commentary to the speech synthesis API and obtains the audio file.
[1292] Input: The generated description.
[1293] Output: Descriptive audio file and subtitle text.
[1294] Step 8:
[1295] The server integrates audio commentary and subtitles into the video clips to create an edited highlight video.
[1296] Use a video editing library (e.g., OpenCV or MoviePy) to integrate the audio file and subtitle text into the video clip. Specifically, the audio description is overlaid on the video and the subtitle text is displayed in the appropriate position.
[1297] Input: Descriptive audio files, subtitle text, and video clips of key scenes.
[1298] Output: Edited highlight video.
[1299] Step 9:
[1300] The server saves the completed highlight video to cloud storage (e.g. AWS S3).
[1301] Generate a link that allows users to access the saved video and provide that link to the user. For example, generate a URL like "https: / / cloudstorage.example.com / highlight_video.mp4".
[1302] Input: edited highlight video.
[1303] Output: Highlight video URL on cloud storage.
[1304] Step 10:
[1305] The user accesses the cloud platform and clicks on the provided link to view the edited footage.
[1306] Users can use a browser or a dedicated app to play or download videos from cloud storage.
[1307] Input: Highlight video URL on cloud storage.
[1308] Output: Highlights viewed or downloaded.
[1309] Through the above process, the system automatically edits soccer match footage shot by the user and provides a professional highlight video with commentary.
[1310] (Application example 1)
[1311] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1312] Shooting soccer game footage, editing the footage, and creating a highlight video with commentary has traditionally required a great deal of effort. Editing the footage and creating the commentary, in particular, requires time and effort, making it difficult to easily achieve professional-quality results. The objective of this invention is to solve this problem and provide a system that enables users to easily create high-quality highlight videos with commentary.
[1313] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1314] In this invention, the server includes means for uploading video captured by a fixed camera to the cloud, means for receiving the uploaded video and detecting the ball and players, means for automatically extracting key scenes based on the detected information, means for adding automatically generated commentary to the video using a generative AI model, means for integrating commentary audio and subtitles to complete the edited video, and means for saving the edited video in cloud storage and enabling users to view or share it. This allows users to easily generate, view, and share professional-quality highlight videos with commentary simply by uploading their game video to the cloud.
[1315] A "fixed camera" is a camera that captures video from a fixed position.
[1316] The "cloud" is a collection of servers for storing, managing, and processing data over the Internet.
[1317] "Uploading" is the act of sending data from a local device to a cloud server.
[1318] "Reception" refers to the act of the cloud server receiving data sent from the user.
[1319] "Detection" is the process of identifying a specific object (the ball or a player) through video analysis.
[1320] A "Key Moment" is a specific key moment during a match (e.g., a goal, save, foul).
[1321] A "generative AI model" is an artificial intelligence model that uses machine learning algorithms to generate content such as text and images.
[1322] "Explanation" is written or audio that explains a particular event.
[1323] The "explanatory voice" is the generated explanatory text output as voice through voice synthesis.
[1324] "Subtitles" are textual information displayed on video.
[1325] "Editing" is the process of adding audio commentary and subtitles to the footage to complete the final video.
[1326] "Cloud storage" is a storage service for storing data on the cloud.
[1327] "Viewing" refers to the act of a user playing and watching a video.
[1328] "Sharing" refers to the act of a user sharing video or data with other people.
[1329] This invention is a system that automatically edits soccer match footage taken with fixed cameras and provides videos with commentary, eliminating the need for users to record and edit the footage while watching the match.
[1330] System configuration
[1331] The system includes the following main components:
[1332] 1. Fixed camera
[1333] Users can set up fixed cameras on the sidelines or behind the goals and use wide-angle lenses to capture the entire game.
[1334] 2. Cloud Upload
[1335] Users upload the footage they have taken to the cloud platform, and the video files are sent to the cloud server via the internet.
[1336] 3. Reception and Preprocessing
[1337] The server receives the video uploaded to the cloud. After receiving the video, it checks the format and quality of the video and performs any necessary preprocessing (e.g., adjusting the resolution and converting the frame rate). For preprocessing, it uses an open-source video processing library such as OpenCV.
[1338] 4. AI analysis
[1339] The AI analysis system on the server analyzes the video data frame by frame, and uses OpenCV and deep learning models to detect the position of the ball and the positions and movements of players in each frame.
[1340] 5. Extracting important scenes
[1341] The server analyzes the tracking data to detect when specific events (e.g., goals, saves, fouls) occur, and automatically extracts key moments based on these events to generate highlight clips.
[1342] 6. Explanation Generation
[1343] The server uses a generative AI model (e.g., GPT-4) to generate commentary in the style of a specific commentator. It then uses voice synthesis based on the generated commentary to create the commentary audio. For voice synthesis, it uses a service like Amazon Polly.
[1344] 7. Video Editing
[1345] The commentary audio and subtitles are integrated into the video, and video editing libraries such as FFmpeg are used for video editing. This completes the edited highlight video.
[1346] 8. Cloud Storage
[1347] The edited footage is saved to cloud storage, allowing users to view or share it.
[1348] Specific examples
[1349] For example, a user can film their child's soccer game with their smartphone and upload the footage to the cloud. The server receives the footage and performs pre-processing. The AI analysis system detects the movement of the ball and players and extracts goals and other important plays. Based on this, the generative AI model generates commentary such as, "Player A made a decisive pass here, and Player B scored a shot!" and adds it to the video as audio and subtitles. Finally, the edited video is stored in the cloud, and users can share it with family and friends.
[1350] Prompt Sentence Examples
[1351] For example, when generating explanatory text using a generative AI model, the following prompt is used:
[1352] "I uploaded a video of my child's soccer game. Analyze the ball position and player movements to detect goals and important plays. Based on that, generate commentary such as, 'Player A made the decisive pass, and Player B scored the shot!' and add it to the video as audio and subtitles."
[1353] In this way, the present invention is a system that can significantly reduce the effort required for filming and automatically provide professional-quality video.
[1354] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1355] Step 1:
[1356] A user films a soccer match with a fixed camera. The fixed camera uses a wide-angle lens to capture the entire game, and is fixed on the sidelines or behind the goal. The input is the video data captured by the camera, and the output is a video file stored in local storage.
[1357] Step 2:
[1358] Users upload the videos they have taken from their smartphones to the cloud platform. The uploaded videos are then sent to the cloud server via the Internet. The input is the video file stored on the smartphone, and the output is the video data on the cloud storage.
[1359] Step 3:
[1360] The server receives the video uploaded to the cloud and performs preprocessing. This includes adjusting the resolution and converting the frame rate, and uses a video processing library such as OpenCV. The input is the video data stored in the cloud storage, and the output is the preprocessed video data converted into an analyzable format.
[1361] Step 4:
[1362] The AI analysis system in the server analyzes the preprocessed video frame by frame to detect the position and movement of the ball and players. This analysis uses OpenCV and deep learning models. The input is the preprocessed video data, and the output is tracking data including the position information of the ball and players.
[1363] Step 5:
[1364] The server automatically extracts important scenes from the tracking data. Specifically, it detects events such as goals, saves, and fouls and extracts those moments as highlight clips. The input is the tracking data, and the output is a list of important scenes and their clip data.
[1365] Step 6:
[1366] The server uses a generative AI model (e.g., GPT-4) to create automatically generated commentary. The generated commentary is converted into speech using speech synthesis technology (e.g., Amazon Polly). The input is a list of important scenes and prompts, and the output is commentary audio and subtitle data.
[1367] Step 7:
[1368] The server integrates the commentary audio and subtitles into the video to create an edited highlight video. This editing is done using a video editing library such as FFmpeg. The input is the clip data, commentary audio, and subtitle data, and the output is the edited video.
[1369] Step 8:
[1370] The server saves the completed edited video in cloud storage. Users can access the cloud platform to view and share the video. The input is the edited video data, and the output is the video data saved in cloud storage.
[1371] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1372] This invention is a system that automatically edits soccer match footage on the cloud and provides videos with commentary using an emotion engine. This system saves users the trouble of recording and editing the game while watching it. It also adjusts the commentary content based on the user's emotions, providing a more personalized video experience.
[1373] System configuration
[1374] The system includes the following main components:
[1375] 1. Photography Method:
[1376] Users set up fixed cameras to film the entire soccer match. The fixed cameras are fixed on the sidelines or behind the goals, and use wide-angle lenses to capture the entire match.
[1377] 2. Upload method:
[1378] Users upload the footage they have taken to the cloud platform, and the video files are sent to the cloud server via the internet.
[1379] 3. Receiving and preprocessing means:
[1380] The server receives the video uploaded to the cloud, checks the format and quality of the video, and performs any necessary pre-processing (e.g., adjusting the resolution or converting the frame rate).
[1381] 4. Analysis method:
[1382] The AI analysis system on the server analyzes the received video data frame by frame, detecting the ball and players in the video and saving their positional information as tracking data.
[1383] 5. Extraction means:
[1384] The server automatically extracts important scenes (goals, fouls, saves, etc.) based on the tracking data, and compiles the extracted scenes to generate highlight clips.
[1385] 6. Emotion Engine:
[1386] The server is equipped with an emotion engine that recognizes the user's emotions. It analyzes the user's emotion data (e.g., facial expression recognition and biometric data) and grasps the user's emotional state in real time.
[1387] 7. Explanation Generation Method:
[1388] The server uses a generative AI model to generate commentary in the style of a specific commentator, adjusts the commentary content based on the user's emotions recognized by the emotion engine, and performs voice synthesis based on the generated commentary to create a commentary voice.
[1389] 8. Means of provision:
[1390] The server stores the edited video in cloud storage, and users can access the cloud platform to view or download the edited video.
[1391] Program processing and specific examples
[1392] Natural language explanation of the process
[1393] Users film the match with fixed cameras and upload the footage to the cloud after the match ends. The server preprocesses the received footage and uses an AI analysis system to detect the ball and players. An emotion engine then collects and analyzes the user's emotions in real time. Key scenes are then extracted, and a generative AI model generates commentary based on the emotion data and integrates it into the video. Finally, the edited video is stored in the cloud and made available for users to watch or download.
[1394] Specific examples
[1395] For example, suppose a user films their child's soccer game and uploads the footage to the cloud. The server receives the footage and performs preprocessing, while the AI analysis system detects and tracks the movement of the ball and players. At the same time, while the user is watching the game, the emotion engine analyzes the user's facial expressions, heart rate, and other data to collect emotional data. Goals and important plays are detected, and highlights are automatically generated based on them. The generative AI model generates commentary such as, "Player A made a decisive pass here, and Player B scored a shot!" If it determines that the user is happy, the commentary will be written in a positive tone. Finally, the edited video is saved in the cloud, where the user can enjoy it with their family.
[1396] In this way, the present invention not only significantly reduces the effort required for filming and automatically provides professional-quality video, but also provides a more personalized video experience based on the user's emotions.
[1397] The processing flow will be explained below.
[1398] Step 1:
[1399] A user films a soccer match with a fixed camera. The camera is set up and recording begins. The video is set to capture the entire match, using a wide-angle lens, etc.
[1400] Step 2:
[1401] The video footage taken by the user is uploaded to the cloud server. After the match, the user connects the camera's memory card to a computer and uploads the video files to the cloud platform.
[1402] Step 3:
[1403] The server receives the video uploaded to the cloud, checks the format and quality of the received video file, and performs pre-processing such as adjusting the resolution and converting the frame rate as necessary.
[1404] Step 4:
[1405] The AI analysis system on the server analyzes the received video data frame by frame, detecting the ball and players in the video and saving their positional information as tracking data.
[1406] Step 5:
[1407] The server automatically extracts important scenes (goals, fouls, saves, etc.) based on the tracking data, and compiles the extracted scenes to generate a highlight clip.
[1408] Step 6:
[1409] As users watch a game, the emotion engine collects their facial expressions and heart rate in real time to generate emotion data, which is then sent to a cloud server.
[1410] Step 7:
[1411] The server receives and analyzes the emotion data sent from the emotion engine to understand the user's emotional state (happiness, excitement, sadness, etc.).
[1412] Step 8:
[1413] The server uses a generative AI model to generate commentary in the style of a specific commentator, adjusting the commentary based on emotion data, for example, using positive expressions if the user is happy.
[1414] Step 9:
[1415] The server integrates audio commentary and subtitles into the video, and edits and completes the highlight video with commentary.
[1416] Step 10:
[1417] The server saves the edited video to cloud storage and notifies the user that the video is ready.
[1418] Step 11:
[1419] Users log in to the cloud platform to view or download the edited footage, and share it with family and team members.
[1420] Example 2
[1421] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1422] Conventional soccer match filming and editing systems require lengthy editing work after filming, and adding commentary is labor-intensive. Furthermore, they are unable to reflect the user's emotional state while watching the video, making it difficult to provide a personalized viewing experience. There is a need for a system that can solve these problems and easily provide high-quality, personalized video.
[1423] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1424] In this invention, the server includes means for uploading video captured by a fixed camera to the cloud, means for receiving the uploaded video and pre-processing the video, means for detecting the ball and players from the received video, means for automatically extracting important scenes based on the detected information, means for collecting user emotion data, means for analyzing the collected emotion data and generating commentary using a generative AI model, means for adding the generated commentary to the video, and means for providing the edited video to the user, thereby enabling users to watch high-quality video containing match highlights and personalized commentary without any hassle.
[1425] A "fixed camera" is a camera that is fixed at a specific location and captures a wide range of images using a wide-angle lens.
[1426] "Cloud" is a general term for an online platform that provides data storage and computing resources by accessing remote servers via the Internet.
[1427] "Uploading means" refers to the method or technology used to transfer digital data from a local device to a remote server, such as cloud storage.
[1428] "Means of receiving" refers to the method or technology by which the remote server obtains data uploaded to the cloud.
[1429] "Preprocessing means" refers to processing carried out in advance, such as converting data format or adjusting quality, to ensure smooth subsequent analysis and editing processes.
[1430] "Detection means" refers to methods or technologies that automatically recognize specific objects (e.g., balls or players) from video data and track their location information.
[1431] "Means for automatically extracting important scenes" refers to methods and technologies that use detected information to identify specific events such as goals and fouls and extract them as highlights from the video.
[1432] "Means for collecting emotional data" refers to methods and technologies for acquiring a user's facial expressions and biometric data in real time and analyzing their emotional state.
[1433] A "generative AI model" is a technology that uses a pre-trained artificial intelligence model to generate explanatory text in natural language from given data and prompts.
[1434] "Means for adding commentary" refers to methods or techniques for integrating the generated commentary text and audio into the video data and providing it as the final edited video.
[1435] "Means for providing edited video" refers to the methods and technologies for storing the final edited video in cloud storage and allowing users to access, view, or download it.
[1436] This system automatically edits soccer match footage on the cloud and provides videos with commentary using an emotion engine. This allows users to concentrate on watching the game, while automating the editing and commentary tasks. It also provides a personalized video experience based on the user's emotions.
[1437] First, the user sets up a fixed camera to film a soccer match. The fixed camera uses a wide-angle lens and is installed on the sidelines or behind the goal. After the match ends, the user uploads the filmed video files to a cloud platform. This can be done using cloud storage services such as Google Drive or Dropbox.
[1438] The server then receives the uploaded video on the cloud, checks the format and quality of the video file, and performs any necessary pre-processing, such as converting AVI video to MP4, reducing the resolution from 1080p to 720p, and converting the frame rate from 30fps to 24fps.
[1439] Once preprocessing is complete, the server uses an AI analysis system to analyze the received video data frame by frame. The AI analysis system uses deep learning models and libraries such as OpenCV to detect the ball and players in the video and save their location information as tracking data. This allows the progress of the game and the movements of the players to be understood.
[1440] The server then automatically extracts important moments (e.g., goals, fouls, saves, etc.) from the tracking data. In this step, AI identifies specific events and cuts them out into highlight videos.
[1441] Furthermore, the server uses an emotion engine to collect user emotional data, including biometric data such as facial expressions and heart rate, and analyzes it in real time, allowing the server to understand the user's emotional state during a match.
[1442] The server then uses a generative AI model to generate commentary in the style of a particular commentator. The generative AI model is given the following prompt:
[1443] "Generate commentary on goal scenes."
[1444] "Please provide an inspiring commentary based on the emotional data."
[1445] "Describe a highlight scene of a key play."
[1446] The generated commentary text is used to synthesize speech, creating a commentary voice, which adds a commentary tailored to the user's emotions to the video.
[1447] Finally, the server integrates the generated commentary audio into the highlight video. Specifically, it uses video editing software such as FFmpeg to combine the highlight footage and commentary audio into a single file. The edited video file is then stored in cloud storage, and users can access the cloud platform to watch or download it.
[1448] For example, consider the case where a user films a child's soccer game and uploads the video to the cloud. The server receives the video and performs preprocessing, and the AI analysis system detects and tracks the movement of the ball and players. Then, while the user is watching the game, the emotion engine analyzes the user's facial expressions, heart rate, etc., and collects emotional data. Based on this data, the generative AI model generates commentary, providing commentary in positive terms, allowing the user to enjoy watching the highlights of the game.
[1449] In this way, the present invention not only significantly reduces the effort required for filming and editing, and automatically provides professional-quality video, but also enables a personally tailored video experience that reflects the user's emotions.
[1450] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1451] Step 1:
[1452] The user sets up a fixed camera to film a soccer match. The fixed camera used for recording uses a wide-angle lens and is fixed to the sideline or behind the goal. After the match ends, the user transfers the filmed footage to a computer. The input data is the filmed video file, and the output data is the video file saved on the computer.
[1453] Step 2:
[1454] Users upload video files stored on their computers to a cloud platform. Specifically, they use cloud storage services such as Google Drive and Dropbox to upload files. The input data is the video file stored on their computer, and the output data is the video file stored on the cloud.
[1455] Step 3:
[1456] The server receives the video uploaded to the cloud. After receiving it, it checks the format and quality of the video file and performs any necessary pre-processing. Pre-processing includes converting AVI format video to MP4 format, reducing the resolution from 1080p to 720p, and converting the frame rate from 30fps to 24fps. The input data is the video file on the cloud, and the output data is the video file after pre-processing. Specific operations involve the use of video conversion software such as FFmpeg.
[1457] Step 4:
[1458] The server uses an AI analysis system to analyze the pre-processed video data frame by frame. The AI analysis system uses a deep learning model and libraries such as OpenCV to detect the ball and players in the video. It then saves this position information as tracking data. The input data is the pre-processed video file, and the output data is the analysis results, including the tracking data.
[1459] Step 5:
[1460] The server automatically extracts key moments (e.g., goals, fouls, saves, etc.) from the tracking data. This involves AI identifying specific events and extracting them into highlight videos. The input data is the tracking data, and the output data is video clips of key moments.
[1461] Step 6:
[1462] The server uses an emotion engine to collect user emotion data in real time. It also uses devices such as webcams and smartwatches to acquire biometric data such as the user's facial expressions and heart rate. The input data is the user's biometric data, and the output data is the analyzed emotion data.
[1463] Step 7:
[1464] The server uses a generative AI model to generate commentary in the style of a specific commentator. The generative AI model is given the following prompt:
[1465] "Generate commentary on goal scenes."
[1466] "Please provide an inspiring commentary based on the emotional data."
[1467] "Describe a highlight scene of a key play."
[1468] The input data is emotion data and prompt sentences, and the output data is the generated commentary. Specifically, a generative AI model such as GPT-3 is used.
[1469] Step 8:
[1470] The server synthesizes the commentary based on the generated commentary text and creates the commentary audio. The input data is the generated commentary text, and the output data is the commentary audio file. Specifically, speech synthesis software is used.
[1471] Step 9:
[1472] The server then integrates the generated commentary audio into the highlight video. Specifically, it uses video editing software such as FFmpeg to combine the highlight video and commentary audio into a single file. The input data is the commentary audio file and the video clips of key scenes, and the output data is the edited video file.
[1473] Step 10:
[1474] The server stores the edited video files in cloud storage, allowing users to access and watch or download them. The input data is the edited video files, and the output data is the video files stored on the cloud. Specifically, a cloud storage platform is used.
[1475] (Application example 2)
[1476] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1477] Conventional video editing systems require manual editing after shooting footage, which requires a great deal of time and effort. Furthermore, the generated commentary is standardized, making it difficult to provide a personalized experience that reflects each user's individual emotions and circumstances. This requires a lot of effort for users to record and edit events, and the viewing experience is not personalized enough.
[1478] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for uploading video captured by a fixed camera to the cloud, means for receiving the uploaded video and detecting objects and people, means for automatically extracting important scenes based on the detected information, generation AI model means for analyzing user emotion data and generating customized commentary based on the analysis results, means for adding the generated commentary to the video, and means for providing the edited video to the user. This eliminates the need for manual editing of the video and makes it possible to provide a personalized viewing experience based on the user's emotions.
[1479] A "fixed camera" is a camera that continuously captures a specific area from a fixed position.
[1480] The "cloud" refers to computing resources and data storage provided via the Internet.
[1481] "Uploading" is the act of sending local data to a remote server over the Internet or a network.
[1482] "Object" is a general term that refers to any concrete thing that can be recognized in an image.
[1483] "Person" is a general term that refers to a person recognized in the video.
[1484] "Detection" is the act of using specific algorithms to identify and track objects or people in a video.
[1485] "Key scenes" refer to moments or events that are particularly noteworthy in footage of sporting events or everyday life.
[1486] "Extraction" is the act of extracting a specific part from the whole data.
[1487] "Emotional data" refers to the emotional state obtained from the user's facial expressions, biometric data, etc.
[1488] A "generative AI model" is an algorithm that uses machine learning and natural language processing to automatically generate explanatory text and other content.
[1489] "Commentary" refers to text or audio that adds explanations or comments to the content of the video.
[1490] "Customization" is the act of tailoring content to a particular user or situation.
[1491] "Edited footage" refers to the final video file in which processing and modification have been applied to the original video data.
[1492] "Providing" refers to the act of making a video available for viewing or acquisition by a user through the cloud or other media.
[1493] The following system is used as an embodiment of the present invention: The main components of the system and their respective roles will be described below.
[1494] System Components
[1495] 1. Fixed camera:
[1496] Used by users to capture events, fixed cameras are fixed in specific locations and use wide-angle lenses to capture a wide range of footage.
[1497] 2. Cloud Platform:
[1498] Users upload footage taken with fixed cameras, and the cloud platform receives the video files via the internet and performs the analysis and editing processes described below.
[1499] 3. Receiving and Preprocessing Server:
[1500] The server receives the uploaded video data. It checks the format and quality of the video and performs any necessary preprocessing (adjusting the resolution and converting the frame rate).
[1501] 4. AI analysis system:
[1502] Video data is analyzed frame by frame, objects and people are detected in the video, and their location information is saved as tracking data.
[1503] 5. Sentiment Analysis System:
[1504] Analyze the user's emotional state. Collect user emotional data (facial recognition and biometric data) using sensors on smartphones and smart glasses, and analyze it in real time.
[1505] 6. Generative AI Models:
[1506] Based on the collected emotional data, a commentary in the style of a specific commentator is generated. Speech synthesis is performed based on the generated commentary to create a commentary voice. The software used here includes a natural language processing library and a speech synthesis module.
[1507] 7. Editing System:
[1508] Based on the analysis results and audio commentary, the video is edited to generate highlight clips. This editing system uses video editing libraries (e.g., OpenCV and MoviePy).
[1509] 8. Delivery system:
[1510] The edited video is stored in cloud storage, and users can view or download the edited video from the cloud platform.
[1511] A description of what the program does
[1512] In this system, a user first captures an event with a fixed camera and uploads the footage to a cloud platform. The cloud server receives the video data, checks its format and quality, and performs any necessary preprocessing. The AI analysis system then analyzes the received video data frame by frame to obtain information for detecting and tracking objects and people.
[1513] At the same time, devices such as smartphones and smart glasses are used to collect the user's emotional data (e.g., facial expressions, heart rate, etc.), which is then analyzed in real time by an emotion analysis system to determine the user's emotional state.
[1514] Based on the results of the sentiment analysis and AI analysis, the generative AI model generates commentary in the style of a specific commentator. For example, if it determines that the user is happy, a positive commentary will be selected. The generated commentary is converted into an audio file using text-to-speech software. These commentary voices are then integrated into existing footage, and the editing system generates a highlight clip.
[1515] Finally, the edited video is stored in cloud storage. Users can access the cloud platform and watch or download the generated personalized video. This process allows users to easily obtain high-quality edited video and provides a personalized experience based on their emotions.
[1516] Specific examples
[1517] For example, if a user films a family birthday party, the app uploads the video to the cloud, and the AI performs emotion analysis and video editing as follows: When the emotion analysis system determines that the user is happy at the climax of the party, the generative AI model generates a caption such as, "Here, the whole family has fun and cut the cake!"
[1518] Prompt Sentence Examples
[1519] Here is a video of a family birthday party. Highlight the most exciting scenes and generate commentary based on the emotion data.
[1520] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1521] Step 1:
[1522] A user captures an event with a fixed camera or a mobile information terminal. As input, the user obtains video data. As output, a video file is generated.
[1523] Step 2:
[1524] The user uploads the video they have taken to the cloud platform. The input is the video file they have taken, which the cloud server receives. The output is the video data stored on the cloud.
[1525] Step 3:
[1526] The server receives the uploaded video data and checks the format and quality. The input is the video data stored on the cloud server, and the server performs preprocessing such as adjusting the resolution and converting the frame rate. The output is the preprocessed video data.
[1527] Step 4:
[1528] The server's AI analysis system analyzes the preprocessed video data frame by frame to detect objects and people. The input is the preprocessed video data, and the server applies detection algorithms to identify objects and people, saving their location information as tracking data. The output is tracking data.
[1529] Step 5:
[1530] The device (smartphone or smart glasses) collects emotion data. The input is sensor data such as the user's facial expression and heart rate, and the device sends this data to the emotion analysis system. The output is the user's emotion data as an analysis result.
[1531] Step 6:
[1532] The server uses a generative AI model to generate an explanatory text based on the user's emotional data and the results of AI analysis. The inputs are emotional data and tracking data, and the generative AI model generates a prompt based on this data and then generates an explanatory text (e.g., "That was a great goal!" if the emotion is positive). The output is an explanatory text.
[1533] Step 7:
[1534] The server converts the generated commentary into an audio file using a speech synthesis module. The commentary is input, the server synthesizes the speech, and obtains the audio file. The audio file is obtained as output.
[1535] Step 8:
[1536] The server's editing system integrates the video and audio commentary to generate a highlight clip. The inputs are the pre-processed video data, tracking data, and audio commentary files, and the server uses a video editing library to integrate them. The output is an edited highlight video.
[1537] Step 9:
[1538] The server stores the edited footage in cloud storage for users to access. As input, there is the edited highlight footage, which the server stores in cloud storage. As output, users can watch or download the footage from the cloud platform.
[1539] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1540] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1541] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1542] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1543] FIG. 9 illustrates an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and behaviors arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1544] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1545] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1546] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1547] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1548] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1549] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1550] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1551] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1552] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1553] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1554] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1555] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1556] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1557] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1558] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1559] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1560] The following is further disclosed regarding the above embodiment.
[1561] (Claim 1)
[1562] A means to upload footage taken by fixed cameras to the cloud,
[1563] means for receiving the uploaded footage and detecting the ball and players;
[1564] A means for automatically extracting important scenes based on the detected information;
[1565] a means for adding the generated commentary to the video;
[1566] a means for providing the edited video to a user;
[1567] A system including:
[1568] (Claim 2)
[1569] The system of claim 1, wherein a user uploads video captured by a fixed camera to the cloud.
[1570] (Claim 3)
[1571] 10. The system of claim 1, wherein important scenes are extracted as highlights.
[1572] "Example 1"
[1573] (Claim 1)
[1574] A means to upload footage taken by fixed cameras to the cloud,
[1575] a pre-processing means for receiving the uploaded video and adjusting the resolution and converting the frame rate;
[1576] means for detecting the ball and players from the video using image processing algorithms;
[1577] A means for automatically extracting important scenes based on the detected tracking data and generating highlight clips;
[1578] A means for creating a commentary audio using a generative AI model and speech synthesis technology that generates commentary using a prompt sentence as input, and integrating the commentary audio and subtitles into the video;
[1579] A means for storing the edited video in cloud storage and providing it to users;
[1580] A system including:
[1581] (Claim 2)
[1582] The system of claim 1, wherein a user uploads video captured by a fixed camera to the cloud.
[1583] (Claim 3)
[1584] The system of claim 1 extracts important scenes as highlights and generates explanatory text using a generative AI model.
[1585] "Application Example 1"
[1586] (Claim 1)
[1587] A means to upload footage taken by fixed cameras to the cloud,
[1588] means for receiving the uploaded footage and detecting the ball and players;
[1589] A means for automatically extracting important scenes based on the detected information;
[1590] A way to add automatically generated commentary to videos using generative AI models;
[1591] a means of integrating audio commentary and subtitles to complete the edited footage;
[1592] A means for storing the edited video in cloud storage and enabling users to view or share the video;
[1593] A system including:
[1594] (Claim 2)
[1595] The system of claim 1, wherein a user uploads video captured by a fixed camera to the cloud.
[1596] (Claim 3)
[1597] The system of claim 1 extracts important scenes as highlights and integrates commentary audio and subtitles to complete the edited video.
[1598] "Example 2: Combining Emotion Engines"
[1599] (Claim 1)
[1600] A means to upload footage taken by fixed cameras to the cloud,
[1601] means for receiving the uploaded video and pre-processing the video;
[1602] means for detecting a ball and players from the received video;
[1603] A means for automatically extracting important scenes based on the detected information;
[1604] means for collecting user emotion data;
[1605] A means for analyzing the collected emotion data and generating explanatory text using a generative AI model;
[1606] a means for adding the generated commentary to the video;
[1607] a means for providing the edited video to a user;
[1608] A system including:
[1609] (Claim 2)
[1610] The system of claim 1, wherein a user uploads video captured by a fixed camera to the cloud.
[1611] (Claim 3)
[1612] The system of claim 1, wherein important scenes are extracted as highlights and commentary is generated based on the collected emotion data.
[1613] "Application example 2 when combining emotion engines"
[1614] (Claim 1)
[1615] A means to upload footage taken by fixed cameras to the cloud,
[1616] means for receiving the uploaded video and detecting objects and people;
[1617] A means for automatically extracting important scenes based on the detected information;
[1618] A generative AI model means for analyzing the user's emotional data and generating a customized commentary based on the analysis result;
[1619] a means for adding the generated commentary to the video;
[1620] a means for providing the edited video to a user;
[1621] A system including:
[1622] (Claim 2)
[1623] The system according to claim 1, wherein a user uploads video taken by a fixed camera or a mobile information terminal to the cloud.
[1624] (Claim 3)
[1625] 10. The system of claim 1, wherein important scenes are extracted as highlights and the video is customized based on user sentiment analysis. [Explanation of symbols]
[1626] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. A means to upload footage taken by fixed cameras to the cloud, means for receiving the uploaded footage and detecting the ball and players; A means for automatically extracting important scenes based on the detected information; a means for adding the generated commentary to the video; a means for providing the edited video to a user; A system including:
2. The system according to claim 1, wherein a user uploads video captured by a fixed camera to the cloud.
3. The system of claim 1, wherein important scenes are extracted as highlights.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A