System

A system using a wide-angle fixed camera, cloud-based AI analysis, and generative AI commentary addresses the challenge of manual editing in sporting events, enabling users to easily produce high-quality videos with customizable commentary.

JP2026023987APending Publication Date: 2026-02-13SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024126308
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-01
Publication Date
2026-02-13

AI Technical Summary

Technical Problem

Conventional filming and editing of sporting events, particularly by non-professionals, is labor-intensive and requires expert knowledge for creating high-quality videos with commentary, making it difficult for average users to easily produce such content.

Method used

A system utilizing a wide-angle fixed camera, cloud-based video data uploading, AI analysis for scene identification, and generative AI for commentary generation, allowing users to automatically edit and add commentary to soccer match videos.

Benefits of technology

Enables average users to efficiently create high-quality videos with commentary, focusing on game analysis and enjoyment, and supports various sports events with user-friendly commentary styles.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026023987000001_ABST
    Figure 2026023987000001_ABST
Patent Text Reader

Abstract

To provide a system for automatically generating a moving image with high-quality moving image editing and explanation.SOLUTION: The specific processing unit 290 of the information processing device 12 in the system receives and stores the video captured by the wide-angle fixed-point camera for capturing videos in the storage, analyzes the stored video by the AI module, tracks the positions of the ball and the players, automatically extracts important scenes and performs zoom editing based on the analysis result, generates commentary by the generative AI model based on the commentary style selected by the user, adds the generated commentary, stores the final edited video in the cloud storage, and provides access links.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] Conventional filming and editing of sporting events such as soccer matches requires manual camera operation and video editing, which requires a great deal of time and effort. In particular, for children's, student, and adult matches, where there are no professional cameramen, parents or coaches must operate the camera themselves. Furthermore, creating videos with commentary requires expert knowledge, making it difficult for average users to easily achieve this.

[0005] In this situation, there is a need for a system that can automatically generate high-quality video editing and commentary videos, allowing users to focus on the game itself while enjoying the videos later, and analyzing and reflecting on them. [Means for solving the problem]

[0006] This invention solves the above problem by using AI to automatically edit and add commentary to soccer match videos uploaded to the cloud. Specifically, we provide a system that includes the following means.

[0007] 1. Wide-angle fixed camera method: Video is shot using a wide-angle fixed camera that covers the entire game.

[0008] 2. Video data uploading method: Upload the captured video data to the cloud server.

[0009] 3. Video data receiving means: The cloud server receives the video data and stores it in storage.

[0010] 4. AI module: Analyzes the stored video data and tracks the positions of the ball and players, thereby identifying important scenes.

[0011] 5. Zoom editing method: Based on the analysis results, goal scenes and important plays are automatically extracted and zoom edited.

[0012] 6. Explanation generation method: A generative AI model generates text and audio explanations based on the explanation style selected by the user.

[0013] 7. Commentary integration method: The generated commentary is integrated into the video, and the final edited video is saved in cloud storage and an access link is provided to the user.

[0014] The app also adds a means to register player names and uniform numbers, and supports sports events other than soccer, improving versatility and usability. In this way, users can easily generate high-quality videos that can be used for memorization and analysis.

[0015] The "wide-angle fixed camera means" is a camera device that takes pictures from a fixed position using a wide-angle lens that covers the entire game or event.

[0016] "Video data uploading means" refers to the function or process of sending captured video data to a cloud server via an internet connection.

[0017] The "video data receiving means" is a mechanism or process for receiving video data uploaded to the cloud server and storing it in storage.

[0018] "AI module means" refers to an artificial intelligence algorithm or program that analyzes received video data and tracks the positions of the ball and players.

[0019] The "zoom editing means" is a function or process for zoom editing important scenes or events based on the analysis results of the AI ​​module means.

[0020] "Description generation means" means a function or process that generates text and audio commentary using a generated AI model based on a user-selected commentary style.

[0021] The "commentary integration means" is a function or process that integrates the generated commentary into the video, saves the final edited video in cloud storage, and provides the user with an access link.

[0022] The "means for registering player names and uniform numbers" is a function or process that allows a user to input the names and uniform numbers of players participating in a match or event and record them in the system.

[0023] "Cloud storage" is an online storage service for storing and managing data over the Internet.

[0024] "Access Link" refers to the URL or other identifying information that allows a user to access a video stored in cloud storage via the Internet.

[0025] "Commentary Style" is a setting item that defines the type and tone of commentary that the user can select, to provide commentary in different commentary styles. [Brief explanation of the drawings]

[0026] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0027] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0028] First, the terms used in the following description will be explained.

[0029] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0030] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0031] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0032] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0033] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0034] [First embodiment]

[0035] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0036] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0037] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0038] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0039] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0040] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0041] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0042] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0043] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0044] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0045] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0046] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0047] This system automatically edits footage of soccer matches and other sporting events into high-quality videos and adds user-selected commentary. The system uploads video data captured with a wide-angle fixed camera to the cloud, analyzes and edits it using an AI module, and finally provides the user with a video that integrates the generated commentary.

[0048] System configuration and program processing

[0049] Recording and uploading videos

[0050] The device (wide-angle fixed camera) is equipped with a wide-angle lens that can capture the entire game, and records the game footage from a fixed position. The user sets up the camera device and presses the start button to start recording. After recording is complete, the device uploads the captured video data to a cloud server via Wi-Fi or a wired connection.

[0051] Receiving and storing video data

[0052] The server receives the uploaded video data and stores it in a storage service. In addition, the information on the player names and uniform numbers entered by the user in advance is also stored in a database.

[0053] Video Analysis and Editing

[0054] The server then passes the video data stored in storage to the AI ​​module, which then begins video analysis. The AI ​​module analyzes the ball position and player movements frame by frame to identify goal scenes and important plays. Based on the analysis results, the server extracts specific scenes, automatically zooms in, and generates clips.

[0055] Explanation generation and integration

[0056] The user selects a commentary style (e.g., "in the style of famous commentator A" or "in the style of famous commentator B") and sends that information to the cloud server. Based on the selected commentary style, the server invokes a generative AI model to generate text and audio commentary. The generated commentary is integrated into the video clip and saved in cloud storage as the final edited video.

[0057] Video distribution and viewing

[0058] The server then notifies the user of the download link for the final video. The notification is sent via email or an in-system notification. The user can then view or download the generated high-quality video with commentary via the provided link.

[0059] Specific examples

[0060] Editing of goal scenes and adding commentary

[0061] 1. Before the start of the game, the user registers the player's name and uniform number and sets up the device (camera device).

[0062] 2. The device records the match and uploads the video data to a cloud server after the match ends.

[0063] 3. The server stores the received video in a database and analyzes it using an AI module.

[0064] 4. The AI ​​module analyzes the ball's movement and identifies goal opportunities.

[0065] 5. The server zooms in on the goal scene and creates a clip.

[0066] 6. The user selects the commentary style "Famous Commentator A Style."

[0067] 7. The generative AI model generates commentary text and audio, for example, "Goal! What a great shot!"

[0068] 8. The server integrates the generated commentary to produce the final goal clip.

[0069] 9. The user then watches the final video with commentary via the provided link.

[0070] The system can also be used for a variety of events, including sports events other than soccer and athletic meets, and helps parents and coaches easily create high-quality footage for analysis and preservation of memories.

[0071] The processing flow will be explained below.

[0072] Step 1:

[0073] A user logs in to the system and registers a soccer match as a new event. The user enters the player names and uniform numbers through the interface and sends the registration information to the cloud server.

[0074] Step 2:

[0075] The device (a wide-angle fixed camera) captures the entire game with a wide angle. The user sets up the camera device and presses the start recording button to start recording. The device continues to record video data throughout the entire game.

[0076] Step 3:

[0077] After the game ends, the device uploads the captured video data to a cloud server via Wi-Fi or a wired connection.

[0078] Step 4:

[0079] The server receives the video data uploaded to the cloud. The received video data is saved in the storage service. The player names and uniform numbers registered by the user in advance are also saved in the database.

[0080] Step 5:

[0081] The server calls the video analysis AI module and analyzes the stored video data. The AI ​​module analyzes the ball position and player movements for each frame and identifies important scenes (goals, important plays, etc.).

[0082] Step 6:

[0083] The server extracts specific scenes based on the analysis results from the AI ​​module, and then runs an automatic zoom editing algorithm to generate a zoomed-in clip of important scenes or events.

[0084] Step 7:

[0085] After registering a player, the user selects a commentary style and sends that information to the cloud server. For example, the user can select a commentary style similar to that of famous commentator A.

[0086] Step 8:

[0087] Based on the selected commentary style, the server calls a generative AI model to generate commentary text and audio data. The generative AI model generates appropriate commentary for each scene based on static text and pre-recorded audio.

[0088] Step 9:

[0089] The server then integrates the generated commentary text and audio into the video clip, thereby generating the final edited video with commentary.

[0090] Step 10:

[0091] The server saves the final edited video to cloud storage and generates an access link for the user. Once saving is complete, the access link is notified to the user.

[0092] Step 11:

[0093] The user watches or downloads the generated high-quality commentary video via the provided link. If the user downloads the video, it is saved locally on the user's device.

[0094] In this way, a system is realized that provides users with high-quality videos with commentary that have been automatically analyzed and edited through a series of processing steps.

[0095] Example 1

[0096] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0097] It is difficult to efficiently edit video of sports games and events with high quality and add visually easy-to-understand commentary. In particular, there is a lack of systems that can efficiently extract only important scenes without watching the entire game and quickly provide edited videos with commentary. There is also a need for a method to easily generate videos that adapt to the user's preferred commentary style.

[0098] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0099] In this invention, the server includes means for uploading video data to a cloud server, means for receiving the video data in the cloud server and saving it in storage, artificial intelligence module means for analyzing the saved video data and tracking the positions of the ball and players, means for automatically extracting key scenes based on the analysis results and performing zoom editing, means for generating commentary using a generative artificial intelligence model based on a commentary style selected by the user, means for the user to select a commentary style and send that information to the cloud server, means for integrating the generated commentary into a video clip, means for notifying the user of a download link for the generated video with commentary, and means for saving the final video with commentary in cloud storage. This allows users to easily access high-quality video with commentary and efficiently enjoy key scenes from games and events.

[0100] A "wide-angle fixed camera" is a camera device that can capture a wide range of images from a fixed position.

[0101] A "cloud server" is a remote server that stores and processes data over the Internet.

[0102] "Storage" is a place or device for storing digital data.

[0103] "Artificial Intelligence Module" means a software module that contains artificial intelligence algorithms designed to perform data analysis or automated processing.

[0104] "Key Scenes" refers to scenes or moments that are particularly noteworthy in a sporting event or sporting event.

[0105] "Zoom editing" is an editing method that emphasizes important scenes by enlarging specific parts of the video.

[0106] "Commentary Style" is a setting that expresses the particular commentary style or tone desired by the user.

[0107] A "generative artificial intelligence model" is an artificial intelligence model for generating text or audio commentary based on input from a user.

[0108] A "video clip" is a short video segment extracted from a longer video.

[0109] A "download link" is a URL that allows a user to download a specific file via the Internet.

[0110] An "explanatory video" is a video that integrates explanatory audio and text with the video.

[0111] "Cloud storage" is an online storage service that allows you to store data via the Internet.

[0112] This invention is a system that automatically edits footage of sports games and events into high-quality videos with commentary. Specifically, it uses a wide-angle fixed camera to shoot video, and includes a process for analyzing and editing the video data on a cloud server. The details are as follows.

[0113] Video recording and uploading

[0114] Before the start of a game or event, users set up a wide-angle fixed camera in an appropriate position. When the camera presses the start button, it starts recording video using the wide-angle lens. When filming is finished, the device (camera device) uploads the video data to a cloud server via Wi-Fi or a wired connection.

[0115] Receiving and storing video data

[0116] The cloud server receives the uploaded video data and stores it in storage. Information such as player names and uniform numbers entered by the user before the game is also stored in the database. This data is used during analysis.

[0117] Video analysis and key scene extraction

[0118] The server then passes the stored video data to an AI module for video analysis. The AI ​​module tracks the ball position and player movements frame by frame to detect goals and other important plays. Based on the detected scenes, the server automatically performs zoom editing, extracts key scenes, and generates clips.

[0119] Selection of commentary style and generation of commentary

[0120] Users select their desired commentary style through an interface and send that information to a cloud server, which then invokes a generative artificial intelligence model to generate text and audio commentary.

[0121] Specifically, for example, a prompt sentence in the style of "famous commentator A" is sent to the generative AI model, and an audio file corresponding to the commentary, such as "Goal! What a great play!", is generated. The generative AI model used in this process utilizes advanced natural language processing technology.

[0122] Explanation and video integration

[0123] The server then integrates the generated commentary into the video clip, resulting in a video clip with commentary, which is then finally stored in cloud storage and made accessible to users.

[0124] Distribution and viewing of the final video

[0125] Finally, the server notifies the user of the download link for the generated commentary video. The notification method is email or in-system notification. The user can watch or download the high-quality commentary video through the provided link.

[0126] Examples of prompt statements

[0127] "Generate commentary for when a goal occurs immediately after the start of a match"

[0128] "When Player A scores a goal, please generate a commentary in the style of famous commentator B."

[0129] "Generate commentary on the most impressive plays in this game"

[0130] The system can be used to record not only sports matches, but also athletic meets and other events, helping parents and coaches easily create high-quality footage for analysis and preservation of memories.

[0131] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0132] Step 1: Camera setup and shooting

[0133] The user sets up a wide-angle fixed camera in an appropriate position and begins recording a game or event. When the user presses the start button, the camera records video using the wide-angle lens. The input includes the camera settings and information about when the start button was pressed. The output generates the captured video data. This video data is temporarily stored on an SD card or internal memory.

[0134] Step 2: Upload video data

[0135] Once the recording is complete, the device (camera device) uploads the video data to a cloud server via Wi-Fi or a wired connection. The input includes the video data stored in the camera. The data is compressed and transmitted over the network. The output is video data temporarily stored on the cloud server.

[0136] Step 3: Receiving and saving video data

[0137] The server receives video data uploaded from the device. The input includes video data sent via the network. After receiving it, the server stores it in a storage service. Information such as player names and uniform numbers registered by the user before the match is also stored in the database. As output, the video data stored in cloud storage and player information stored in the database are generated.

[0138] Step 4: Video analysis and key scene extraction

[0139] The server passes the stored video data to an artificial intelligence module to begin video analysis. The input includes the video data retrieved from storage. The artificial intelligence module tracks the ball position and player movements frame by frame to detect goals and other important plays. Based on the analysis results, the server automatically extracts key scenes and performs zoom editing. The output is a zoom-edited video clip containing key scenes.

[0140] Step 5: Choose your commentary style

[0141] The user selects the desired commentary style through the interface and sends the information to the cloud server. The input includes the user's selected commentary style. The output is saved in the server.

[0142] Step 6: Generate a description

[0143] The server invokes the generative AI model based on the commentary style selected by the user. The input includes the user's commentary style selection information and a prompt sentence for a specific scene. The generative AI model generates text and audio commentary. The output is the text commentary and audio file generated by the generative AI model.

[0144] Example prompt sentences:

[0145] "Generate commentary for when a goal occurs immediately after the start of a match"

[0146] "When Player A scores a goal, please generate a commentary in the style of famous commentator B."

[0147] "Generate commentary on the most impressive plays in this game"

[0148] Step 7: Integrating the explanation and video

[0149] The server integrates the generated description into the video clip. The input includes the generated text and audio description and the edited video clip. Using video editing software (e.g., FFmpeg), the description audio is merged into a specific timeline of the video clip. The output is a video clip with the description.

[0150] Step 8: Stream and view your final video

[0151] The server notifies the user of a download link for the generated annotated video. The notification method is via email or in-system notification. The input includes the annotated video clip stored in cloud storage. The output is a download link notified to the user. The user can watch or download the generated high-quality annotated video through the provided link.

[0152] Through these steps, users can generate and easily use high-quality videos with commentary that allow them to efficiently enjoy important scenes from games and events.

[0153] (Application example 1)

[0154] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0155] Conventional video editing systems for sporting events require manual editing, which is labor-intensive and time-consuming. Furthermore, specialized knowledge is required to add commentary, making it difficult for average users to create high-quality videos with commentary. Furthermore, managing and distributing edited videos is complicated. To solve these issues, it is necessary to provide a system that automatically analyzes and edits videos shot with a wide-angle fixed camera and adds commentary.

[0156] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0157] In this invention, the server includes a wide-angle fixed camera means for shooting videos, a means for uploading the shot video data to a cloud server, a means for receiving the video data in the cloud server and saving it in storage, an AI module means for analyzing the saved video data and tracking the positions of objects and people, a means for automatically extracting important scenes based on the analysis results and performing zoom editing, a means for generating commentary using a generative AI model based on a commentary style selected by the user, a means for adding the generated commentary, saving the final edited video in cloud storage and providing an access link, and a means for managing video data through an application installed on a smartphone. This allows users to easily create and manage high-quality videos with commentary.

[0158] The "wide-angle fixed-point camera means" is a camera device that uses a wide-angle lens to capture an overall image from a fixed position.

[0159] A "cloud server" is a remote server that stores, analyzes, and processes data over a network.

[0160] "Means of storing data in storage" refers to a method of storing data for the long term inside or outside the cloud server.

[0161] An "AI module means" is a software component that realizes artificial intelligence technology that mimics human intelligence and performs data analysis and predictions.

[0162] The "means for automatically extracting important scenes based on the analysis results and performing zoom editing" is a method for automatically selecting specific scenes using the analysis data generated by the AI ​​module means and focusing on a portion of the video for zoom editing.

[0163] "Means for generating commentary using a generative AI model based on the commentary style selected by the user" refers to a method for generating commentary with specific writing style and voice characteristics using AI technology in accordance with the user's selection.

[0164] "Means for adding generated commentary, saving the final edited video in cloud storage, and providing an access link" refers to a method for integrating commentary created by a generative AI model into an edited video and generating a URL that can be accessed by users after saving it in the cloud.

[0165] An "application installed on a smartphone" is software that is downloaded to and executed on a mobile device, and is a means for a user to manage video data.

[0166] An "identification number" is a number assigned to uniquely identify a particular player or object.

[0167] "Event" is a general term that refers to a series of activities or programs that take place at a specific time and place.

[0168] The system for implementing this invention uses a wide-angle fixed camera, a cloud server, storage, an AI module, and a generative AI model. The role of each component and the operation of the entire system are described in detail below.

[0169] A wide-angle fixed camera is used to capture the entire game or event. This camera captures footage from a fixed position using a wide-angle lens. A user sets up the camera device and begins capturing footage of the game or event.

[0170] The video data captured by the device is uploaded to a cloud server via Wi-Fi or a wired connection. The cloud server receives the uploaded video data and saves it in storage. In addition, information such as player names and identification numbers entered by the user in advance is also saved in a database.

[0171] The cloud server then passes the stored video data to the AI ​​module, which then begins video analysis. The AI ​​module uses technology that mimics human intelligence to analyze the movement of the ball and players frame by frame and identify key scenes. This analysis uses image processing libraries such as OpenCV.

[0172] Based on the analysis results, the server automatically extracts specific scenes and performs zoom editing as necessary. Users select a commentary style through an application installed on their smartphone. For example, they can choose a style such as "famous commentator A style." This information is sent to the cloud server.

[0173] The cloud server then invokes a generative AI model based on the user's selection to generate commentary with a specific writing style and voice characteristics. Using natural language processing, the generative AI model might create a commentary such as, "Goal! What a great shot!" The generated commentary is then integrated into the edited video to create the final video file.

[0174] The server stores the final video file in cloud storage and provides a download link to the user, who can then view or download the final video with commentary via the provided link.

[0175] Specific examples

[0176] For example, suppose a user films a soccer match and connects the camera device to a cloud server. The user uses an application to register player names and identification numbers, and uploads the video data after the match. The cloud server receives the video data and analyzes it using an AI module. Based on the analysis results, important scenes are automatically extracted and zoomed in for editing. The user selects "Famous Commentator A" as the commentary style, and as a result, the generative AI model generates a commentary such as "That was an excellent shot!"

[0177] Prompt Sentence Examples

[0178] Analyze a soccer match video provided by the user and identify goal scenes. The commentary style selected is "Famous Commentator A Style." Based on this style, generate commentary such as "Goal! What a great shot!"

[0179] The above is an embodiment of the present invention. By using this system, users can easily create and manage high-quality videos with commentary.

[0180] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0181] Step 1:

[0182] A user sets up a wide-angle fixed camera and begins capturing footage of a game or event.

[0183] Input: Camera device setup and recording start command

[0184] Output: Recorded video data (saved in local storage)

[0185] Step 2:

[0186] After the device finishes shooting, it uploads the captured video data to a cloud server.

[0187] Input: Video data stored in local storage

[0188] Data Processing: Transferring video data via Wi-Fi or wired connection

[0189] Output: Video data uploaded to the cloud server

[0190] Step 3:

[0191] The server receives the video data uploaded to the cloud server and stores it in storage.

[0192] Input: Video data uploaded to the cloud server

[0193] Data processing: Writing to cloud storage

[0194] Output: Video data stored in cloud storage

[0195] Step 4:

[0196] The server passes the stored video data to the AI ​​module, which then begins video analysis.

[0197] Input: Video data stored in cloud storage

[0198] Data calculation: AI module analyzes the movement of objects and people (ball and players) for each frame

[0199] Output: Analysis result data (object location information, movement patterns)

[0200] Step 5:

[0201] Based on the analysis results, the server automatically extracts important scenes and performs zoom editing.

[0202] Input: Analysis result data of AI module

[0203] Data processing: Extraction of important scenes and video processing by zoom editing

[0204] Output: Edited clip data

[0205] Step 6:

[0206] The user selects a commentary style via a smartphone app. For example, they select "Famous Commentator A Style."

[0207] Input: User's commentary style selection in the app

[0208] Output: Selected commentary style information (sent to cloud server)

[0209] Step 7:

[0210] The server generates commentary using a generative AI model based on the selected commentary style.

[0211] Input: User selected commentary style information

[0212] Data Computing: Generating Explanatory Text and Audio with Generative AI Models

[0213] Output: Generated commentary data (e.g., "Goal! What a great shot!" audio and text)

[0214] Step 8:

[0215] The server adds the generated commentary and saves the final edited video to cloud storage.

[0216] Input: Edited clip data and generated commentary data

[0217] Data processing: Integrating audio commentary and text into video

[0218] Output: Final edited video data

[0219] Step 9:

[0220] The server then generates a download link for the final video and notifies the user.

[0221] Input: Final edited video data

[0222] Data processing: Download link generation and URL generation

[0223] Output: Notification to user (email or in-system notification)

[0224] This allows users to view or download the final video with commentary via the provided link.

[0225] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0226] This system automatically edits footage of soccer matches and other sporting events into high-quality videos and adds user-selected commentary. Furthermore, the system recognizes the user's emotions and automatically selects and refines the commentary style based on those emotions. The system uploads video data captured with a wide-angle fixed camera to the cloud, where it is analyzed and edited by an AI module, and finally provides the user with a video that integrates the generated commentary.

[0227] System configuration and program processing

[0228] Recording and uploading videos

[0229] The device (wide-angle fixed camera) is equipped with a wide-angle lens that can capture the entire game, and records the game footage from a fixed position. The user sets up the camera device and presses the start button to start recording. After recording is complete, the device uploads the captured video data to a cloud server via Wi-Fi or a wired connection.

[0230] Receiving and storing video data

[0231] The server receives the uploaded video data and stores it in a storage service. In addition, the information on the player names and uniform numbers entered by the user in advance is also stored in a database.

[0232] Video Analysis and Editing

[0233] The server then passes the video data stored in storage to the AI ​​module, which then begins video analysis. The AI ​​module analyzes the ball position and player movements frame by frame to identify goal scenes and important plays. Based on the analysis results, the server extracts specific scenes, automatically zooms in, and generates clips.

[0234] Recognizing user emotions with an emotion engine

[0235] When a user accesses the system through a device, the emotion engine recognizes the user's emotions, obtains emotional data from the user's facial expressions and voice, and adjusts the commentary style accordingly.

[0236] Explanation generation and integration

[0237] The user manually selects the commentary style, or it is automatically selected by the emotion engine. Based on the selected commentary style, the server invokes a generative AI model to generate commentary text and audio data. The generated commentary is integrated into the video clip and saved in cloud storage as the final edited video.

[0238] Video distribution and viewing

[0239] The server then notifies the user of the download link for the final video. The notification is sent via email or an in-system notification. The user can then view or download the generated high-quality video with commentary via the provided link.

[0240] Specific examples

[0241] Editing of goal scenes and adding commentary

[0242] 1. Before the start of the game, the user registers the player's name and uniform number and sets up the device (camera device).

[0243] 2. The device records the match and uploads the video data to a cloud server after the match ends.

[0244] 3. The server stores the received video in a database and analyzes it using an AI module.

[0245] 4. The AI ​​module analyzes the ball's movement and identifies goal opportunities.

[0246] 5. The server zooms in on the goal scene and creates a clip.

[0247] 6. When the user manually or automatically selects a commentary style, the emotion engine recognizes the user's excitement and selects a more energetic commentary style, such as "Goal! What a great shot!"

[0248] 7. The generative AI model generates explanatory text and audio.

[0249] 8. The server integrates the generated commentary to produce the final goal clip.

[0250] 9. The user then watches the final video with commentary via the provided link.

[0251] The introduction of an emotion engine makes it possible to provide customized commentary that corresponds to the user's specific emotional state, further enhancing the viewing experience.The system can also be used for various events other than soccer, such as sports events and athletic meets, and helps families and coaches easily create high-quality footage for analysis and recording memories.

[0252] The processing flow will be explained below.

[0253] Step 1:

[0254] A user logs in to the system and registers a soccer match as a new event. The user enters the player names and uniform numbers through the interface and sends the registration information to the cloud server.

[0255] Step 2:

[0256] The device (wide-angle fixed camera) captures the entire game with a wide angle. Once the setup is complete, the user presses the start button on the camera device to begin recording the game. The device continues to record video data throughout the entire game.

[0257] Step 3:

[0258] After the game ends, the device uploads the captured video data to a cloud server via Wi-Fi or a wired connection. The video data is sent directly to the cloud server.

[0259] Step 4:

[0260] The server receives the video data uploaded to the cloud and stores it in the storage service. At the same time, the player names and uniform numbers registered by the user in advance are also stored in the database.

[0261] Step 5:

[0262] The server then calls up the video analysis AI module, which analyzes the stored video data. The AI ​​module analyzes the ball position and player movements for each frame to identify goal scenes and key plays.

[0263] Step 6:

[0264] The server extracts specific scenes based on the analysis results from the AI ​​module, runs an automatic zoom editing algorithm, zooms in on important scenes and events, and generates clips.

[0265] Step 7:

[0266] After registering a player, the user manually selects a commentary style. In some cases, the emotion engine recognizes the user's emotions and automatically selects an appropriate commentary style. This allows the system to provide commentary that reflects the user's emotions, such as joy or excitement.

[0267] Step 8:

[0268] The emotion engine analyzes the user's emotions from their facial expressions and voice, obtains emotion data in real time, and provides feedback to the selection of commentary style.

[0269] Step 9:

[0270] The server calls a generative AI model to generate commentary text and audio data based on the selected commentary style. For example, if the user is excited, an energetic commentary will be generated.

[0271] Step 10:

[0272] The server integrates the generated commentary text and audio into the video clip, generates a final edited video with commentary, and stores the video in cloud storage.

[0273] Step 11:

[0274] The server will notify the user with a download link for the final edited video. Notifications will be sent via email or in-system notifications. Once the link is provided, the user will have immediate access.

[0275] Step 12:

[0276] The user can then view or download the generated high-quality commentary video via the provided link, which can be saved locally on the user's device.

[0277] Through these processing steps, high-quality videos with annotations that match the user's emotions are automatically generated and provided, allowing the user to enjoy videos with annotations that reflect specific emotions when played back, resulting in a deeper viewing experience.

[0278] Example 2

[0279] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0280] There is a demand for high-quality editing of sporting event videos and the automatic addition of appropriate commentary based on user emotions. However, conventional technologies have difficulty automating video editing and providing emotion-based commentary, requiring a high level of specialized knowledge and manual work. As a result, it is difficult to generate timely and personalized video content.

[0281] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0282] In this invention, the server includes a wide-angle fixed-point camera means for shooting video, a means for uploading the shot video data to an information processing device, and a means for receiving the video data in the information processing device and storing it in a storage device. This enables highly accurate analysis and editing of the video data. The server also includes an artificial intelligence module means for analyzing the saved video data and tracking the positions of the sphere and people, a means for automatically extracting important scenes based on the analysis results and performing zoom editing, a means for generating commentary using a generative artificial intelligence model based on a commentary style selected by the user, a means for adding the generated commentary, storing the final edited video in a storage device, and providing an access link, and an emotion recognition engine means for recognizing the user's emotions and automatically selecting a commentary style based on the emotions. This enables automatic and efficient generation of high-quality video with commentary that corresponds to the user's emotions.

[0283] A "wide-angle fixed-point camera" is a device that captures images from a fixed position with a wide field of view, and is primarily used to record the overall picture of sporting events and the like.

[0284] An "information processing device" is a device that processes, analyzes, stores, and transmits data, and includes a cloud server, a local server, and the like.

[0285] A "storage device" is a hardware or software means for storing digital data, including cloud storage and hard disk drives.

[0286] An "artificial intelligence module" is a software module that uses machine learning and deep learning to analyze data and recognize specific patterns and features.

[0287] "Tracking the position of spheres and people" refers to the process of detecting specific objects, such as balls or players, within video frames and continuously tracking their position information.

[0288] "Automatic extraction of important scenes" is the process of identifying specific events or actions within a video based on the analysis results of an artificial intelligence module and cutting out only the relevant parts.

[0289] "Zoom editing" is an editing technique that enlarges and visually enhances portions of the original footage, and is used to provide viewers with important scene details.

[0290] A "generative AI model" is a model that uses AI technology to generate text and speech based on user selections and emotional data.

[0291] An "emotion recognition engine" is a software module that analyzes and recognizes a user's emotional state from their facial expressions and voice, acquiring emotional data and reflecting it in the analysis results.

[0292] MODE FOR CARRYING OUT THE INVENTION

[0293] The present invention is a system that automatically edits video data from sporting events into high-quality videos and adds user-selected commentary. It also provides the ability to recognize user emotions and automatically select and refine commentary styles based on those emotions. The system begins by capturing game footage using a fixed-point camera with a wide field of view and uploading the raw data to a cloud server. Specific embodiments of the present invention are described below.

[0294] System Configuration

[0295] 1. Wide-angle fixed-point shooting device

[0296] The wide-angle fixed-point camera is a dedicated camera that records the entire sporting event over a wide area. Users set up the camera and start recording at the start of the game. The video data captured by the camera is uploaded to a cloud server via Wi-Fi or a wired connection.

[0297] 2. Cloud Server

[0298] The cloud server receives the video data and stores it securely in a storage device. The cloud server also stores information such as player names and uniform numbers that users have registered in advance. The video data is then analyzed by an AI module.

[0299] 3. Artificial Intelligence Module

[0300] An AI module in the cloud server analyzes the video data frame by frame, tracking the positions of balls and people. This analysis identifies goals and other key moments. The AI ​​module extracts key scenes and performs zoom editing to generate clips.

[0301] 4. Emotion Recognition Engine

[0302] When a user accesses the system, the emotion recognition engine acquires emotional data from the user's facial expressions and voice, which is used as an important factor in determining the commentary style.

[0303] 5. Generative AI Models

[0304] The generative AI model generates commentary text and audio data based on the user's emotion recognition results and selected commentary style, which is then integrated with the video clip to create the final edited video.

[0305] 6. Video streaming and viewing

[0306] The cloud server then sends the user a download link for the final video, allowing the user to view or download the high-quality video with commentary.

[0307] Specific examples

[0308] 1. Editing goal scenes and adding commentary

[0309] Before the game, the user registers the player names and uniform numbers and sets up the wide-angle fixed-point camera.

[0310] The device records the match and uploads the video data to a cloud server after the match ends.

[0311] The video data received by the server is stored in a storage device and analyzed by an AI module.

[0312] The AI ​​module analyzes the movement of the ball and identifies goal opportunities.

[0313] The server zooms in on the goal scene and creates a clip.

[0314] The user can select a commentary style, or an emotion recognition engine can automatically select one based on the emotion. In this case, an energetic commentary such as "Goal! What a great shot!" is generated.

[0315] A generative AI model generates explanatory text and audio.

[0316] The server then integrates the generated commentary to produce the final goal clip.

[0317] The user then watches the final video with commentary via the provided link.

[0318] Prompt Sentence Examples

[0319] "The entire match should be filmed with a wide-angle camera, and specific scenes should be automatically edited and commentary should be added. The commentary should be selected and generated based on the user's emotions."

[0320] The above is a specific embodiment for carrying out the present invention. This system can be applied to various sporting events other than soccer, and efficiently generates high-quality video.

[0321] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0322] Step 1: Record your video and upload it to the cloud

[0323] The terminal (wide-angle fixed-point camera) is equipped with a wide-angle lens that can capture the entire game. The user sets up the camera and presses the start button to begin recording. The input is live footage of the game, and the output is the captured video data. This video data is uploaded to a cloud server via Wi-Fi or a wired connection after the game ends. Specific operations include the user adjusting the camera position and instructing the game to start.

[0324] Step 2: Receiving and saving video data

[0325] The server receives video data from the cloud and stores it in a storage device. The input is the video data uploaded from the device, and the output is the data stored in a storage device on the cloud. Specific operations include the server confirming receipt of the data and storing it in the appropriate folder. At the same time, the player name and uniform number information previously entered by the user is also stored in the database.

[0326] Step 3: Video Analysis and Editing

[0327] The server passes the video data stored in the storage device to an AI module and begins data analysis. The input is the stored video data, and the output is the analysis results identifying key scenes. The AI ​​module tracks the ball position and player movements for each frame to identify goals and important plays. Specific operations include extracting the identified scenes and performing zoom editing to generate clips.

[0328] Step 4: Recognizing user emotions with the emotion engine

[0329] When a user accesses the system, the emotion recognition engine analyzes the user's emotions. The input is the user's facial expression and voice data, and the output is recognized emotion data. Specific operations include the process in which the user provides emotion data using a webcam or microphone, and the emotion engine analyzes that data.

[0330] Step 5: Explanation generation and integration

[0331] The user manually selects the commentary style, or the emotion recognition engine automatically selects it. The input is the recognized emotion data and the selected commentary style, and the output is the generated commentary text and audio data. The server uses a generative AI model to generate commentary and integrate it into the video clip. Specific operations include the server calling the generative AI model and automatically generating and integrating commentary.

[0332] Step 6: Stream and view your video

[0333] The server notifies the user of the download link for the final video. The input is the final video data, and the output is a link that the user can access. Specific operations include the server sending a notification email or an in-system notification and providing the user with a viewing link. The user can then view or download the high-quality video with commentary via the link.

[0334] (Application example 2)

[0335] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0336] Conventional video editing systems for sporting events make it difficult for users to add their own commentary and lack the ability to dynamically change the commentary style based on the user's emotions. As a result, viewers' viewing experience is consistent and they are unable to respond to the needs of individual users. Furthermore, there is a need for similar high-quality editing functions for a variety of sporting events other than soccer.

[0337] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0338] In this invention, the server includes a wide-angle fixed camera means for shooting videos, a means for uploading the shot video data to a cloud server, a means for receiving the video data in the cloud server and storing it in storage, an AI module means for analyzing the stored video data and tracking the positions of the ball and players, a means for automatically extracting important scenes based on the analysis results and performing zoom editing, a means for generating commentary using a generative AI model based on a commentary style selected by the user, an emotion engine means for recognizing the user's emotions and adjusting the commentary style based on the emotions, and a means for adding the generated commentary, storing the final edited video in cloud storage, and providing an access link. This makes it possible to provide high-quality videos with commentary that are adapted to the user's emotions.

[0339] The "wide-angle fixed-point camera means" is a device equipped with a wide-angle lens that captures moving images from a fixed position.

[0340] A "cloud server" is a computer system that stores and processes data and is accessible remotely over the Internet.

[0341] "Storage" refers to a storage device or service for saving data.

[0342] "AI module means" is software that uses artificial intelligence to analyze video and track the movements of the ball and players.

[0343] "Zoom editing" is an editing technique that visually emphasizes a particular scene by enlarging it.

[0344] A "generative AI model" is an artificial intelligence model that automatically generates natural language based on user input and data.

[0345] The "emotion engine means" is software for analyzing the user's facial expressions and voice and recognizing their emotions.

[0346] An "access link" is a URL or hyperlink that allows a user to access specific content via the Internet.

[0347] "Means for registering player names and uniform numbers" is a function for entering player names and uniform numbers into the system for video analysis.

[0348] The "emotion-based commentary style automatic adjustment function" is a function that dynamically changes the style and tone of commentary using the user's emotional data.

[0349] In this invention, a user shoots a sporting event with a wide-angle fixed camera and uploads the video data to a cloud server. The cloud server receives the video data and saves it in storage. The saved video data is analyzed by an AI module, and the positions of the ball and players are tracked. Based on the analysis results, important scenes are automatically extracted and zoomed in for editing.

[0350] If the user selects a commentary style, a commentary is generated using the generative AI model. If the user's emotion is recognized, an emotion engine means analyzes the user's facial expressions and voice and adjusts the commentary style. The generated commentary is added to the video, and the final edited video is saved in cloud storage, and an access link is provided to the user.

[0351] A specific system configuration operates as follows.

[0352] 1. Hardware and Software Description

[0353] Wide-angle fixed camera means: A device equipped with a wide-angle lens that shoots video from a fixed position.

[0354] Cloud Server: A computer system that stores and processes data and is accessible remotely over the Internet.

[0355] Storage: A storage device or service for storing data (e.g., Amazon S3, Google Cloud Storage).

[0356] AI module means: Software (e.g. OpenCV) that uses artificial intelligence to analyze video and track ball and player movements.

[0357] Generative AI models: Artificial intelligence models that generate commentary based on user input and data (e.g., GPT-3).

[0358] Emotion engine means: Software that analyzes the user's facial expressions and voice and recognizes their emotions (e.g., OpenFace).

[0359] 2. Description of Data Processing and Data Calculations

[0360] Uploading and saving video data: The video data taken by the user is uploaded to the cloud server and saved in storage. The saved data is passed to the AI ​​module.

[0361] Video analysis: The AI ​​module analyzes the video data and tracks the position of the ball and players frame by frame. Key scenes are extracted and zoomed in for editing.

[0362] Emotion Recognition: Emotion data is acquired from the user's facial expressions and voice. The emotion engine analyzes the user's emotions and adjusts the commentary style accordingly.

[0363] Commentary generation and integration: A generative AI model generates commentary text and audio based on the emotion data and integrates it into the video. The final edited video is saved in cloud storage and an access link is provided to the user.

[0364] 3. Examples of concrete examples and prompts

[0365] For example, a user can film a soccer game using their smartphone and upload the video to the cloud. The cloud server then analyzes the video data and adds energetic commentary based on the user's emotional data. Finally, a high-quality video with commentary is generated, and the user receives a link to watch it.

[0366] Example prompts for generative AI models

[0367] If the user is excited while watching:

[0368] "It was an exciting goal scene! The moment the shot shook the net, the whole stadium erupted in cheers!"

[0369] If the user is relaxing while watching:

[0370] "Here, a relaxed ball movement resulted in a fantastic goal. This play is a great example of the teamwork."

[0371] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0372] Step 1:

[0373] The user takes a video

[0374] A user shoots video of a sporting event using a wide-angle fixed camera, and video data is generated. The inputs are the user's operations and the sporting event being shot, and the output is a video file of the shot video.

[0375] Step 2:

[0376] Upload videos to a cloud server

[0377] After the user has finished shooting, they upload the video data to the cloud server. The input for the upload is a video file, and the output is the video data saved on the cloud server. The user is also notified of the upload status.

[0378] Step 3:

[0379] The cloud server receives and stores the data.

[0380] The cloud server receives the uploaded video data and saves it to storage. The input is the uploaded video data, and the output is the storage of the data in the destination storage. A notification of successful saving is also generated.

[0381] Step 4:

[0382] Video analysis and key scene extraction

[0383] The cloud server uses an AI module to analyze the stored video data and track the positions of the ball and players. The input is the stored video data, and the output is the position data of the ball and players, as well as identification information for important scenes. Specifically, the system analyzes each frame and records the position data.

[0384] Step 5:

[0385] Performing Zoom Edits

[0386] The cloud server automatically extracts important scenes and performs zoom editing. The input is the identification information and video data of the important scenes, and the output is a zoom-edited clip of the scene. Specifically, editing software is used to enlarge and display the important scenes.

[0387] Step 6:

[0388] Recognize user emotions with an emotion engine

[0389] When a user watches a video, the emotion engine means analyzes the user's facial expressions and voice to recognize emotions. The input is the user's facial expression data and voice data, and the output is recognized emotion data. Specifically, data is collected using a camera and microphone, and an emotion analysis algorithm is applied.

[0390] Step 7:

[0391] Generating and synthesizing explanations

[0392] The user selects a commentary style, or the generative AI model generates commentary based on recognized emotion data. The input is emotion data and clips of key scenes, and the output is a video clip with commentary. Specifically, a prompt is generated for the generative AI model, and commentary text and audio data are generated.

[0393] Step 8:

[0394] Save the final edited video and provide an access link

[0395] The cloud server integrates the generated commentary into the video and saves the final edited video in cloud storage. The input is a video clip with commentary, and the output is saving the final edited video and generating an access link. Specifically, the video editing software integrates the commentary, uploads it to the cloud, and generates an access link that is notified to the user.

[0396] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0397] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0398] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0399] [Second embodiment]

[0400] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0401] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0402] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0403] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0404] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0405] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0406] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0407] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0408] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0409] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0410] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0411] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0412] This system automatically edits footage of soccer matches and other sporting events into high-quality videos and adds user-selected commentary. The system uploads video data captured with a wide-angle fixed camera to the cloud, analyzes and edits it using an AI module, and finally provides the user with a video that integrates the generated commentary.

[0413] System configuration and program processing

[0414] Recording and uploading videos

[0415] The device (wide-angle fixed camera) is equipped with a wide-angle lens that can capture the entire game, and records the game footage from a fixed position. The user sets up the camera device and presses the start button to start recording. After recording is complete, the device uploads the captured video data to a cloud server via Wi-Fi or a wired connection.

[0416] Receiving and storing video data

[0417] The server receives the uploaded video data and stores it in a storage service. In addition, the information on the player names and uniform numbers entered by the user in advance is also stored in a database.

[0418] Video Analysis and Editing

[0419] The server then passes the video data stored in storage to the AI ​​module, which then begins video analysis. The AI ​​module analyzes the ball position and player movements frame by frame to identify goal scenes and important plays. Based on the analysis results, the server extracts specific scenes, automatically zooms in, and generates clips.

[0420] Explanation generation and integration

[0421] The user selects a commentary style (e.g., "in the style of famous commentator A" or "in the style of famous commentator B") and sends that information to the cloud server. Based on the selected commentary style, the server invokes a generative AI model to generate text and audio commentary. The generated commentary is integrated into the video clip and saved in cloud storage as the final edited video.

[0422] Video distribution and viewing

[0423] The server then notifies the user of the download link for the final video. The notification is sent via email or an in-system notification. The user can then view or download the generated high-quality video with commentary via the provided link.

[0424] Specific examples

[0425] Editing of goal scenes and adding commentary

[0426] 1. Before the start of the game, the user registers the player's name and uniform number and sets up the device (camera device).

[0427] 2. The device records the match and uploads the video data to a cloud server after the match ends.

[0428] 3. The server stores the received video in a database and analyzes it using an AI module.

[0429] 4. The AI ​​module analyzes the ball's movement and identifies goal opportunities.

[0430] 5. The server zooms in on the goal scene and creates a clip.

[0431] 6. The user selects the commentary style "Famous Commentator A Style."

[0432] 7. The generative AI model generates commentary text and audio, for example, "Goal! What a great shot!"

[0433] 8. The server integrates the generated commentary to produce the final goal clip.

[0434] 9. The user then watches the final video with commentary via the provided link.

[0435] The system can also be used for a variety of events, including sports events other than soccer and athletic meets, and helps parents and coaches easily create high-quality footage for analysis and preservation of memories.

[0436] The processing flow will be explained below.

[0437] Step 1:

[0438] A user logs in to the system and registers a soccer match as a new event. The user enters the player names and uniform numbers through the interface and sends the registration information to the cloud server.

[0439] Step 2:

[0440] The device (a wide-angle fixed camera) captures the entire game with a wide angle. The user sets up the camera device and presses the start recording button to start recording. The device continues to record video data throughout the entire game.

[0441] Step 3:

[0442] After the game ends, the device uploads the captured video data to a cloud server via Wi-Fi or a wired connection.

[0443] Step 4:

[0444] The server receives the video data uploaded to the cloud. The received video data is saved in the storage service. The player names and uniform numbers registered by the user in advance are also saved in the database.

[0445] Step 5:

[0446] The server calls the video analysis AI module and analyzes the stored video data. The AI ​​module analyzes the ball position and player movements for each frame and identifies important scenes (goals, important plays, etc.).

[0447] Step 6:

[0448] The server extracts specific scenes based on the analysis results from the AI ​​module, and then runs an automatic zoom editing algorithm to generate a zoomed-in clip of important scenes or events.

[0449] Step 7:

[0450] After registering a player, the user selects a commentary style and sends that information to the cloud server. For example, the user can select a commentary style similar to that of famous commentator A.

[0451] Step 8:

[0452] Based on the selected commentary style, the server calls a generative AI model to generate commentary text and audio data. The generative AI model generates appropriate commentary for each scene based on static text and pre-recorded audio.

[0453] Step 9:

[0454] The server then integrates the generated commentary text and audio into the video clip, thereby generating the final edited video with commentary.

[0455] Step 10:

[0456] The server saves the final edited video to cloud storage and generates an access link for the user. Once saving is complete, the access link is notified to the user.

[0457] Step 11:

[0458] The user watches or downloads the generated high-quality commentary video via the provided link. If the user downloads the video, it is saved locally on the user's device.

[0459] In this way, a system is realized that provides users with high-quality videos with commentary that have been automatically analyzed and edited through a series of processing steps.

[0460] Example 1

[0461] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0462] It is difficult to efficiently edit video of sports games and events with high quality and add visually easy-to-understand commentary. In particular, there is a lack of systems that can efficiently extract only important scenes without watching the entire game and quickly provide edited videos with commentary. There is also a need for a method to easily generate videos that adapt to the user's preferred commentary style.

[0463] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0464] In this invention, the server includes means for uploading video data to a cloud server, means for receiving the video data in the cloud server and saving it in storage, artificial intelligence module means for analyzing the saved video data and tracking the positions of the ball and players, means for automatically extracting key scenes based on the analysis results and performing zoom editing, means for generating commentary using a generative artificial intelligence model based on a commentary style selected by the user, means for the user to select a commentary style and send that information to the cloud server, means for integrating the generated commentary into a video clip, means for notifying the user of a download link for the generated video with commentary, and means for saving the final video with commentary in cloud storage. This allows users to easily access high-quality video with commentary and efficiently enjoy key scenes from games and events.

[0465] A "wide-angle fixed camera" is a camera device that can capture a wide range of images from a fixed position.

[0466] A "cloud server" is a remote server that stores and processes data over the Internet.

[0467] "Storage" is a place or device for storing digital data.

[0468] "Artificial Intelligence Module" means a software module that contains artificial intelligence algorithms designed to perform data analysis or automated processing.

[0469] "Key Scenes" refers to scenes or moments that are particularly noteworthy in a sporting event or sporting event.

[0470] "Zoom editing" is an editing method that emphasizes important scenes by enlarging specific parts of the video.

[0471] "Commentary Style" is a setting that expresses the particular commentary style or tone desired by the user.

[0472] A "generative artificial intelligence model" is an artificial intelligence model for generating text or audio commentary based on input from a user.

[0473] A "video clip" is a short video segment extracted from a longer video.

[0474] A "download link" is a URL that allows a user to download a specific file via the Internet.

[0475] An "explanatory video" is a video that integrates explanatory audio and text with the video.

[0476] "Cloud storage" is an online storage service that allows you to store data via the Internet.

[0477] This invention is a system that automatically edits footage of sports games and events into high-quality videos with commentary. Specifically, it uses a wide-angle fixed camera to shoot video, and includes a process for analyzing and editing the video data on a cloud server. The details are as follows.

[0478] Video recording and uploading

[0479] Before the start of a game or event, users set up a wide-angle fixed camera in an appropriate position. When the camera presses the start button, it starts recording video using the wide-angle lens. When filming is finished, the device (camera device) uploads the video data to a cloud server via Wi-Fi or a wired connection.

[0480] Receiving and storing video data

[0481] The cloud server receives the uploaded video data and stores it in storage. Information such as player names and uniform numbers entered by the user before the game is also stored in the database. This data is used during analysis.

[0482] Video analysis and key scene extraction

[0483] The server then passes the stored video data to an AI module for video analysis. The AI ​​module tracks the ball position and player movements frame by frame to detect goals and other important plays. Based on the detected scenes, the server automatically performs zoom editing, extracts key scenes, and generates clips.

[0484] Selection of commentary style and generation of commentary

[0485] Users select their desired commentary style through an interface and send that information to a cloud server, which then invokes a generative artificial intelligence model to generate text and audio commentary.

[0486] Specifically, for example, a prompt sentence in the style of "famous commentator A" is sent to the generative AI model, and an audio file corresponding to the commentary, such as "Goal! What a great play!", is generated. The generative AI model used in this process utilizes advanced natural language processing technology.

[0487] Explanation and video integration

[0488] The server then integrates the generated commentary into the video clip, resulting in a video clip with commentary, which is then finally stored in cloud storage and made accessible to users.

[0489] Distribution and viewing of the final video

[0490] Finally, the server notifies the user of the download link for the generated commentary video. The notification method is email or in-system notification. The user can watch or download the high-quality commentary video through the provided link.

[0491] Examples of prompt statements

[0492] "Generate commentary for when a goal occurs immediately after the start of a match"

[0493] "When Player A scores a goal, please generate a commentary in the style of famous commentator B."

[0494] "Generate commentary on the most impressive plays in this game"

[0495] The system can be used to record not only sports matches, but also athletic meets and other events, helping parents and coaches easily create high-quality footage for analysis and preservation of memories.

[0496] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0497] Step 1: Camera setup and shooting

[0498] The user sets up a wide-angle fixed camera in an appropriate position and begins recording a game or event. When the user presses the start button, the camera records video using the wide-angle lens. The input includes the camera settings and information about when the start button was pressed. The output generates the captured video data. This video data is temporarily stored on an SD card or internal memory.

[0499] Step 2: Upload video data

[0500] Once the recording is complete, the device (camera device) uploads the video data to a cloud server via Wi-Fi or a wired connection. The input includes the video data stored in the camera. The data is compressed and transmitted over the network. The output is video data temporarily stored on the cloud server.

[0501] Step 3: Receiving and saving video data

[0502] The server receives video data uploaded from the device. The input includes video data sent via the network. After receiving it, the server stores it in a storage service. Information such as player names and uniform numbers registered by the user before the match is also stored in the database. As output, the video data stored in cloud storage and player information stored in the database are generated.

[0503] Step 4: Video analysis and key scene extraction

[0504] The server passes the stored video data to an artificial intelligence module to begin video analysis. The input includes the video data retrieved from storage. The artificial intelligence module tracks the ball position and player movements frame by frame to detect goals and other important plays. Based on the analysis results, the server automatically extracts key scenes and performs zoom editing. The output is a zoom-edited video clip containing key scenes.

[0505] Step 5: Choose your commentary style

[0506] The user selects the desired commentary style through the interface and sends the information to the cloud server. The input includes the user's selected commentary style. The output is saved in the server.

[0507] Step 6: Generate a description

[0508] The server invokes the generative AI model based on the commentary style selected by the user. The input includes the user's commentary style selection information and a prompt sentence for a specific scene. The generative AI model generates text and audio commentary. The output is the text commentary and audio file generated by the generative AI model.

[0509] Example prompt sentences:

[0510] "Generate commentary for when a goal occurs immediately after the start of a match"

[0511] "When Player A scores a goal, please generate a commentary in the style of famous commentator B."

[0512] "Generate commentary on the most impressive plays in this game"

[0513] Step 7: Integrating the explanation and video

[0514] The server integrates the generated description into the video clip. The input includes the generated text and audio description and the edited video clip. Using video editing software (e.g., FFmpeg), the description audio is merged into a specific timeline of the video clip. The output is a video clip with the description.

[0515] Step 8: Stream and view your final video

[0516] The server notifies the user of a download link for the generated annotated video. The notification method is via email or in-system notification. The input includes the annotated video clip stored in cloud storage. The output is a download link notified to the user. The user can watch or download the generated high-quality annotated video through the provided link.

[0517] Through these steps, users can generate and easily use high-quality videos with commentary that allow them to efficiently enjoy important scenes from games and events.

[0518] (Application example 1)

[0519] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0520] Conventional video editing systems for sporting events require manual editing, which is labor-intensive and time-consuming. Furthermore, specialized knowledge is required to add commentary, making it difficult for average users to create high-quality videos with commentary. Furthermore, managing and distributing edited videos is complicated. To solve these issues, it is necessary to provide a system that automatically analyzes and edits videos shot with a wide-angle fixed camera and adds commentary.

[0521] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0522] In this invention, the server includes a wide-angle fixed camera means for shooting videos, a means for uploading the shot video data to a cloud server, a means for receiving the video data in the cloud server and saving it in storage, an AI module means for analyzing the saved video data and tracking the positions of objects and people, a means for automatically extracting important scenes based on the analysis results and performing zoom editing, a means for generating commentary using a generative AI model based on a commentary style selected by the user, a means for adding the generated commentary, saving the final edited video in cloud storage and providing an access link, and a means for managing video data through an application installed on a smartphone. This allows users to easily create and manage high-quality videos with commentary.

[0523] The "wide-angle fixed-point camera means" is a camera device that uses a wide-angle lens to capture an overall image from a fixed position.

[0524] A "cloud server" is a remote server that stores, analyzes, and processes data over a network.

[0525] "Means of storing data in storage" refers to a method of storing data for the long term inside or outside the cloud server.

[0526] An "AI module means" is a software component that realizes artificial intelligence technology that mimics human intelligence and performs data analysis and predictions.

[0527] The "means for automatically extracting important scenes based on the analysis results and performing zoom editing" is a method for automatically selecting specific scenes using the analysis data generated by the AI ​​module means and focusing on a portion of the video for zoom editing.

[0528] "Means for generating commentary using a generative AI model based on the commentary style selected by the user" refers to a method for generating commentary with specific writing style and voice characteristics using AI technology in accordance with the user's selection.

[0529] "Means for adding generated commentary, saving the final edited video in cloud storage, and providing an access link" refers to a method for integrating commentary created by a generative AI model into an edited video and generating a URL that can be accessed by users after saving it in the cloud.

[0530] An "application installed on a smartphone" is software that is downloaded to and executed on a mobile device, and is a means for a user to manage video data.

[0531] An "identification number" is a number assigned to uniquely identify a particular player or object.

[0532] "Event" is a general term that refers to a series of activities or programs that take place at a specific time and place.

[0533] The system for implementing this invention uses a wide-angle fixed camera, a cloud server, storage, an AI module, and a generative AI model. The role of each component and the operation of the entire system are described in detail below.

[0534] A wide-angle fixed camera is used to capture the entire game or event. This camera captures footage from a fixed position using a wide-angle lens. A user sets up the camera device and begins capturing footage of the game or event.

[0535] The video data captured by the device is uploaded to a cloud server via Wi-Fi or a wired connection. The cloud server receives the uploaded video data and saves it in storage. In addition, information such as player names and identification numbers entered by the user in advance is also saved in a database.

[0536] The cloud server then passes the stored video data to the AI ​​module, which then begins video analysis. The AI ​​module uses technology that mimics human intelligence to analyze the movement of the ball and players frame by frame and identify key scenes. This analysis uses image processing libraries such as OpenCV.

[0537] Based on the analysis results, the server automatically extracts specific scenes and performs zoom editing as necessary. Users select a commentary style through an application installed on their smartphone. For example, they can choose a style such as "famous commentator A style." This information is sent to the cloud server.

[0538] The cloud server then invokes a generative AI model based on the user's selection to generate commentary with a specific writing style and voice characteristics. Using natural language processing, the generative AI model might create a commentary such as, "Goal! What a great shot!" The generated commentary is then integrated into the edited video to create the final video file.

[0539] The server stores the final video file in cloud storage and provides a download link to the user, who can then view or download the final video with commentary via the provided link.

[0540] Specific examples

[0541] For example, suppose a user films a soccer match and connects the camera device to a cloud server. The user uses an application to register player names and identification numbers, and uploads the video data after the match. The cloud server receives the video data and analyzes it using an AI module. Based on the analysis results, important scenes are automatically extracted and zoomed in for editing. The user selects "Famous Commentator A" as the commentary style, and as a result, the generative AI model generates a commentary such as "That was an excellent shot!"

[0542] Prompt Sentence Examples

[0543] Analyze a soccer match video provided by the user and identify goal scenes. The commentary style selected is "Famous Commentator A Style." Based on this style, generate commentary such as "Goal! What a great shot!"

[0544] The above is an embodiment of the present invention. By using this system, users can easily create and manage high-quality videos with commentary.

[0545] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0546] Step 1:

[0547] A user sets up a wide-angle fixed camera and begins capturing footage of a game or event.

[0548] Input: Camera device setup and recording start command

[0549] Output: Recorded video data (saved in local storage)

[0550] Step 2:

[0551] After the device finishes shooting, it uploads the captured video data to a cloud server.

[0552] Input: Video data stored in local storage

[0553] Data Processing: Transferring video data via Wi-Fi or wired connection

[0554] Output: Video data uploaded to the cloud server

[0555] Step 3:

[0556] The server receives the video data uploaded to the cloud server and stores it in storage.

[0557] Input: Video data uploaded to the cloud server

[0558] Data processing: Writing to cloud storage

[0559] Output: Video data stored in cloud storage

[0560] Step 4:

[0561] The server passes the stored video data to the AI ​​module, which then begins video analysis.

[0562] Input: Video data stored in cloud storage

[0563] Data calculation: AI module analyzes the movement of objects and people (ball and players) for each frame

[0564] Output: Analysis result data (object location information, movement patterns)

[0565] Step 5:

[0566] Based on the analysis results, the server automatically extracts important scenes and performs zoom editing.

[0567] Input: Analysis result data of AI module

[0568] Data processing: Extraction of important scenes and video processing by zoom editing

[0569] Output: Edited clip data

[0570] Step 6:

[0571] The user selects a commentary style via a smartphone app. For example, they select "Famous Commentator A Style."

[0572] Input: User's commentary style selection in the app

[0573] Output: Selected commentary style information (sent to cloud server)

[0574] Step 7:

[0575] The server generates commentary using a generative AI model based on the selected commentary style.

[0576] Input: User selected commentary style information

[0577] Data Computing: Generating Explanatory Text and Audio with Generative AI Models

[0578] Output: Generated commentary data (e.g., "Goal! What a great shot!" audio and text)

[0579] Step 8:

[0580] The server adds the generated commentary and saves the final edited video to cloud storage.

[0581] Input: Edited clip data and generated commentary data

[0582] Data processing: Integrating audio commentary and text into video

[0583] Output: Final edited video data

[0584] Step 9:

[0585] The server then generates a download link for the final video and notifies the user.

[0586] Input: Final edited video data

[0587] Data processing: Download link generation and URL generation

[0588] Output: Notification to user (email or in-system notification)

[0589] This allows users to view or download the final video with commentary via the provided link.

[0590] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0591] This system automatically edits footage of soccer matches and other sporting events into high-quality videos and adds user-selected commentary. Furthermore, the system recognizes the user's emotions and automatically selects and refines the commentary style based on those emotions. The system uploads video data captured with a wide-angle fixed camera to the cloud, where it is analyzed and edited by an AI module, and finally provides the user with a video that integrates the generated commentary.

[0592] System configuration and program processing

[0593] Recording and uploading videos

[0594] The device (wide-angle fixed camera) is equipped with a wide-angle lens that can capture the entire game, and records the game footage from a fixed position. The user sets up the camera device and presses the start button to start recording. After recording is complete, the device uploads the captured video data to a cloud server via Wi-Fi or a wired connection.

[0595] Receiving and storing video data

[0596] The server receives the uploaded video data and stores it in a storage service. In addition, the information on the player names and uniform numbers entered by the user in advance is also stored in a database.

[0597] Video Analysis and Editing

[0598] The server then passes the video data stored in storage to the AI ​​module, which then begins video analysis. The AI ​​module analyzes the ball position and player movements frame by frame to identify goal scenes and important plays. Based on the analysis results, the server extracts specific scenes, automatically zooms in, and generates clips.

[0599] Recognizing user emotions with an emotion engine

[0600] When a user accesses the system through a device, the emotion engine recognizes the user's emotions, obtains emotional data from the user's facial expressions and voice, and adjusts the commentary style accordingly.

[0601] Explanation generation and integration

[0602] The user manually selects the commentary style, or it is automatically selected by the emotion engine. Based on the selected commentary style, the server invokes a generative AI model to generate commentary text and audio data. The generated commentary is integrated into the video clip and saved in cloud storage as the final edited video.

[0603] Video distribution and viewing

[0604] The server then notifies the user of the download link for the final video. The notification is sent via email or an in-system notification. The user can then view or download the generated high-quality video with commentary via the provided link.

[0605] Specific examples

[0606] Editing of goal scenes and adding commentary

[0607] 1. Before the start of the game, the user registers the player's name and uniform number and sets up the device (camera device).

[0608] 2. The device records the match and uploads the video data to a cloud server after the match ends.

[0609] 3. The server stores the received video in a database and analyzes it using an AI module.

[0610] 4. The AI ​​module analyzes the ball's movement and identifies goal opportunities.

[0611] 5. The server zooms in on the goal scene and creates a clip.

[0612] 6. When the user manually or automatically selects a commentary style, the emotion engine recognizes the user's excitement and selects a more energetic commentary style, such as "Goal! What a great shot!"

[0613] 7. The generative AI model generates explanatory text and audio.

[0614] 8. The server integrates the generated commentary to produce the final goal clip.

[0615] 9. The user then watches the final video with commentary via the provided link.

[0616] The introduction of an emotion engine makes it possible to provide customized commentary that corresponds to the user's specific emotional state, further enhancing the viewing experience.The system can also be used for various events other than soccer, such as sports events and athletic meets, and helps families and coaches easily create high-quality footage for analysis and recording memories.

[0617] The processing flow will be explained below.

[0618] Step 1:

[0619] A user logs in to the system and registers a soccer match as a new event. The user enters the player names and uniform numbers through the interface and sends the registration information to the cloud server.

[0620] Step 2:

[0621] The device (wide-angle fixed camera) captures the entire game with a wide angle. Once the setup is complete, the user presses the start button on the camera device to begin recording the game. The device continues to record video data throughout the entire game.

[0622] Step 3:

[0623] After the game ends, the device uploads the captured video data to a cloud server via Wi-Fi or a wired connection. The video data is sent directly to the cloud server.

[0624] Step 4:

[0625] The server receives the video data uploaded to the cloud and stores it in the storage service. At the same time, the player names and uniform numbers registered by the user in advance are also stored in the database.

[0626] Step 5:

[0627] The server then calls up the video analysis AI module, which analyzes the stored video data. The AI ​​module analyzes the ball position and player movements for each frame to identify goal scenes and key plays.

[0628] Step 6:

[0629] The server extracts specific scenes based on the analysis results from the AI ​​module, runs an automatic zoom editing algorithm, zooms in on important scenes and events, and generates clips.

[0630] Step 7:

[0631] After registering a player, the user manually selects a commentary style. In some cases, the emotion engine recognizes the user's emotions and automatically selects an appropriate commentary style. This allows the system to provide commentary that reflects the user's emotions, such as joy or excitement.

[0632] Step 8:

[0633] The emotion engine analyzes the user's emotions from their facial expressions and voice, obtains emotion data in real time, and provides feedback to the selection of commentary style.

[0634] Step 9:

[0635] The server calls a generative AI model to generate commentary text and audio data based on the selected commentary style. For example, if the user is excited, an energetic commentary will be generated.

[0636] Step 10:

[0637] The server integrates the generated commentary text and audio into the video clip, generates a final edited video with commentary, and stores the video in cloud storage.

[0638] Step 11:

[0639] The server will notify the user with a download link for the final edited video. Notifications will be sent via email or in-system notifications. Once the link is provided, the user will have immediate access.

[0640] Step 12:

[0641] The user can then view or download the generated high-quality commentary video via the provided link, which can be saved locally on the user's device.

[0642] Through these processing steps, high-quality videos with annotations that match the user's emotions are automatically generated and provided, allowing the user to enjoy videos with annotations that reflect specific emotions when played back, resulting in a deeper viewing experience.

[0643] Example 2

[0644] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0645] There is a demand for high-quality editing of sporting event videos and the automatic addition of appropriate commentary based on user emotions. However, conventional technologies have difficulty automating video editing and providing emotion-based commentary, requiring a high level of specialized knowledge and manual work. As a result, it is difficult to generate timely and personalized video content.

[0646] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0647] In this invention, the server includes a wide-angle fixed-point camera means for shooting video, a means for uploading the shot video data to an information processing device, and a means for receiving the video data in the information processing device and storing it in a storage device. This enables highly accurate analysis and editing of the video data. The server also includes an artificial intelligence module means for analyzing the saved video data and tracking the positions of the sphere and people, a means for automatically extracting important scenes based on the analysis results and performing zoom editing, a means for generating commentary using a generative artificial intelligence model based on a commentary style selected by the user, a means for adding the generated commentary, storing the final edited video in a storage device, and providing an access link, and an emotion recognition engine means for recognizing the user's emotions and automatically selecting a commentary style based on the emotions. This enables automatic and efficient generation of high-quality video with commentary that corresponds to the user's emotions.

[0648] A "wide-angle fixed-point camera" is a device that captures images from a fixed position with a wide field of view, and is primarily used to record the overall picture of sporting events and the like.

[0649] An "information processing device" is a device that processes, analyzes, stores, and transmits data, and includes a cloud server, a local server, and the like.

[0650] A "storage device" is a hardware or software means for storing digital data, including cloud storage and hard disk drives.

[0651] An "artificial intelligence module" is a software module that uses machine learning and deep learning to analyze data and recognize specific patterns and features.

[0652] "Tracking the position of spheres and people" refers to the process of detecting specific objects, such as balls or players, within video frames and continuously tracking their position information.

[0653] "Automatic extraction of important scenes" is the process of identifying specific events or actions within a video based on the analysis results of an artificial intelligence module and cutting out only the relevant parts.

[0654] "Zoom editing" is an editing technique that enlarges and visually enhances portions of the original footage, and is used to provide viewers with important scene details.

[0655] A "generative AI model" is a model that uses AI technology to generate text and speech based on user selections and emotional data.

[0656] An "emotion recognition engine" is a software module that analyzes and recognizes a user's emotional state from their facial expressions and voice, acquiring emotional data and reflecting it in the analysis results.

[0657] MODE FOR CARRYING OUT THE INVENTION

[0658] The present invention is a system that automatically edits video data from sporting events into high-quality videos and adds user-selected commentary. It also provides the ability to recognize user emotions and automatically select and refine commentary styles based on those emotions. The system begins by capturing game footage using a fixed-point camera with a wide field of view and uploading the raw data to a cloud server. Specific embodiments of the present invention are described below.

[0659] System Configuration

[0660] 1. Wide-angle fixed-point shooting device

[0661] The wide-angle fixed-point camera is a dedicated camera that records the entire sporting event over a wide area. Users set up the camera and start recording at the start of the game. The video data captured by the camera is uploaded to a cloud server via Wi-Fi or a wired connection.

[0662] 2. Cloud Server

[0663] The cloud server receives the video data and stores it securely in a storage device. The cloud server also stores information such as player names and uniform numbers that users have registered in advance. The video data is then analyzed by an AI module.

[0664] 3. Artificial Intelligence Module

[0665] An AI module in the cloud server analyzes the video data frame by frame, tracking the positions of balls and people. This analysis identifies goals and other key moments. The AI ​​module extracts key scenes and performs zoom editing to generate clips.

[0666] 4. Emotion Recognition Engine

[0667] When a user accesses the system, the emotion recognition engine acquires emotional data from the user's facial expressions and voice, which is used as an important factor in determining the commentary style.

[0668] 5. Generative AI Models

[0669] The generative AI model generates commentary text and audio data based on the user's emotion recognition results and selected commentary style, which is then integrated with the video clip to create the final edited video.

[0670] 6. Video streaming and viewing

[0671] The cloud server then sends the user a download link for the final video, allowing the user to view or download the high-quality video with commentary.

[0672] Specific examples

[0673] 1. Editing goal scenes and adding commentary

[0674] Before the game, the user registers the player names and uniform numbers and sets up the wide-angle fixed-point camera.

[0675] The device records the match and uploads the video data to a cloud server after the match ends.

[0676] The video data received by the server is stored in a storage device and analyzed by an AI module.

[0677] The AI ​​module analyzes the movement of the ball and identifies goal opportunities.

[0678] The server zooms in on the goal scene and creates a clip.

[0679] The user can select a commentary style, or an emotion recognition engine can automatically select one based on the emotion. In this case, an energetic commentary such as "Goal! What a great shot!" is generated.

[0680] A generative AI model generates explanatory text and audio.

[0681] The server then integrates the generated commentary to produce the final goal clip.

[0682] The user then watches the final video with commentary via the provided link.

[0683] Prompt Sentence Examples

[0684] "The entire match should be filmed with a wide-angle camera, and specific scenes should be automatically edited and commentary should be added. The commentary should be selected and generated based on the user's emotions."

[0685] The above is a specific embodiment for carrying out the present invention. This system can be applied to various sporting events other than soccer, and efficiently generates high-quality video.

[0686] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0687] Step 1: Record your video and upload it to the cloud

[0688] The terminal (wide-angle fixed-point camera) is equipped with a wide-angle lens that can capture the entire game. The user sets up the camera and presses the start button to begin recording. The input is live footage of the game, and the output is the captured video data. This video data is uploaded to a cloud server via Wi-Fi or a wired connection after the game ends. Specific operations include the user adjusting the camera position and instructing the game to start.

[0689] Step 2: Receiving and saving video data

[0690] The server receives video data from the cloud and stores it in a storage device. The input is the video data uploaded from the device, and the output is the data stored in a storage device on the cloud. Specific operations include the server confirming receipt of the data and storing it in the appropriate folder. At the same time, the player name and uniform number information previously entered by the user is also stored in the database.

[0691] Step 3: Video Analysis and Editing

[0692] The server passes the video data stored in the storage device to an AI module and begins data analysis. The input is the stored video data, and the output is the analysis results identifying key scenes. The AI ​​module tracks the ball position and player movements for each frame to identify goals and important plays. Specific operations include extracting the identified scenes and performing zoom editing to generate clips.

[0693] Step 4: Recognizing user emotions with the emotion engine

[0694] When a user accesses the system, the emotion recognition engine analyzes the user's emotions. The input is the user's facial expression and voice data, and the output is recognized emotion data. Specific operations include the process in which the user provides emotion data using a webcam or microphone, and the emotion engine analyzes that data.

[0695] Step 5: Explanation generation and integration

[0696] The user manually selects the commentary style, or the emotion recognition engine automatically selects it. The input is the recognized emotion data and the selected commentary style, and the output is the generated commentary text and audio data. The server uses a generative AI model to generate commentary and integrate it into the video clip. Specific operations include the server calling the generative AI model and automatically generating and integrating commentary.

[0697] Step 6: Stream and view your video

[0698] The server notifies the user of the download link for the final video. The input is the final video data, and the output is a link that the user can access. Specific operations include the server sending a notification email or an in-system notification and providing the user with a viewing link. The user can then view or download the high-quality video with commentary via the link.

[0699] (Application example 2)

[0700] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0701] Conventional video editing systems for sporting events make it difficult for users to add their own commentary and lack the ability to dynamically change the commentary style based on the user's emotions. As a result, viewers' viewing experience is consistent and they are unable to respond to the needs of individual users. Furthermore, there is a need for similar high-quality editing functions for a variety of sporting events other than soccer.

[0702] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0703] In this invention, the server includes a wide-angle fixed camera means for shooting videos, a means for uploading the shot video data to a cloud server, a means for receiving the video data in the cloud server and storing it in storage, an AI module means for analyzing the stored video data and tracking the positions of the ball and players, a means for automatically extracting important scenes based on the analysis results and performing zoom editing, a means for generating commentary using a generative AI model based on a commentary style selected by the user, an emotion engine means for recognizing the user's emotions and adjusting the commentary style based on the emotions, and a means for adding the generated commentary, storing the final edited video in cloud storage, and providing an access link. This makes it possible to provide high-quality videos with commentary that are adapted to the user's emotions.

[0704] The "wide-angle fixed-point camera means" is a device equipped with a wide-angle lens that captures moving images from a fixed position.

[0705] A "cloud server" is a computer system that stores and processes data and is accessible remotely over the Internet.

[0706] "Storage" refers to a storage device or service for saving data.

[0707] "AI module means" is software that uses artificial intelligence to analyze video and track the movements of the ball and players.

[0708] "Zoom editing" is an editing technique that visually emphasizes a particular scene by enlarging it.

[0709] A "generative AI model" is an artificial intelligence model that automatically generates natural language based on user input and data.

[0710] The "emotion engine means" is software for analyzing the user's facial expressions and voice and recognizing their emotions.

[0711] An "access link" is a URL or hyperlink that allows a user to access specific content via the Internet.

[0712] "Means for registering player names and uniform numbers" is a function for entering player names and uniform numbers into the system for video analysis.

[0713] The "emotion-based commentary style automatic adjustment function" is a function that dynamically changes the style and tone of commentary using the user's emotional data.

[0714] In this invention, a user shoots a sporting event with a wide-angle fixed camera and uploads the video data to a cloud server. The cloud server receives the video data and saves it in storage. The saved video data is analyzed by an AI module, and the positions of the ball and players are tracked. Based on the analysis results, important scenes are automatically extracted and zoomed in for editing.

[0715] If the user selects a commentary style, a commentary is generated using the generative AI model. If the user's emotion is recognized, an emotion engine means analyzes the user's facial expressions and voice and adjusts the commentary style. The generated commentary is added to the video, and the final edited video is saved in cloud storage, and an access link is provided to the user.

[0716] A specific system configuration operates as follows.

[0717] 1. Hardware and Software Description

[0718] Wide-angle fixed camera means: A device equipped with a wide-angle lens that shoots video from a fixed position.

[0719] Cloud Server: A computer system that stores and processes data and is accessible remotely over the Internet.

[0720] Storage: A storage device or service for storing data (e.g., Amazon S3, Google Cloud Storage).

[0721] AI module means: Software (e.g. OpenCV) that uses artificial intelligence to analyze video and track ball and player movements.

[0722] Generative AI models: Artificial intelligence models that generate commentary based on user input and data (e.g., GPT-3).

[0723] Emotion engine means: Software that analyzes the user's facial expressions and voice and recognizes their emotions (e.g., OpenFace).

[0724] 2. Description of Data Processing and Data Calculations

[0725] Uploading and saving video data: The video data taken by the user is uploaded to the cloud server and saved in storage. The saved data is passed to the AI ​​module.

[0726] Video analysis: The AI ​​module analyzes the video data and tracks the position of the ball and players frame by frame. Key scenes are extracted and zoomed in for editing.

[0727] Emotion Recognition: Emotion data is acquired from the user's facial expressions and voice. The emotion engine analyzes the user's emotions and adjusts the commentary style accordingly.

[0728] Commentary generation and integration: A generative AI model generates commentary text and audio based on the emotion data and integrates it into the video. The final edited video is saved in cloud storage and an access link is provided to the user.

[0729] 3. Examples of concrete examples and prompts

[0730] For example, a user can film a soccer game using their smartphone and upload the video to the cloud. The cloud server then analyzes the video data and adds energetic commentary based on the user's emotional data. Finally, a high-quality video with commentary is generated, and the user receives a link to watch it.

[0731] Example prompts for generative AI models

[0732] If the user is excited while watching:

[0733] "It was an exciting goal scene! The moment the shot shook the net, the whole stadium erupted in cheers!"

[0734] If the user is relaxing while watching:

[0735] "Here, a relaxed ball movement resulted in a fantastic goal. This play is a great example of the teamwork."

[0736] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0737] Step 1:

[0738] The user takes a video

[0739] A user shoots video of a sporting event using a wide-angle fixed camera, and video data is generated. The inputs are the user's operations and the sporting event being shot, and the output is a video file of the shot video.

[0740] Step 2:

[0741] Upload videos to a cloud server

[0742] After the user has finished shooting, they upload the video data to the cloud server. The input for the upload is a video file, and the output is the video data saved on the cloud server. The user is also notified of the upload status.

[0743] Step 3:

[0744] The cloud server receives and stores the data.

[0745] The cloud server receives the uploaded video data and saves it to storage. The input is the uploaded video data, and the output is the storage of the data in the destination storage. A notification of successful saving is also generated.

[0746] Step 4:

[0747] Video analysis and key scene extraction

[0748] The cloud server uses an AI module to analyze the stored video data and track the positions of the ball and players. The input is the stored video data, and the output is the position data of the ball and players, as well as identification information for important scenes. Specifically, the system analyzes each frame and records the position data.

[0749] Step 5:

[0750] Performing Zoom Edits

[0751] The cloud server automatically extracts important scenes and performs zoom editing. The input is the identification information and video data of the important scenes, and the output is a zoom-edited clip of the scene. Specifically, editing software is used to enlarge and display the important scenes.

[0752] Step 6:

[0753] Recognize user emotions with an emotion engine

[0754] When a user watches a video, the emotion engine means analyzes the user's facial expressions and voice to recognize emotions. The input is the user's facial expression data and voice data, and the output is recognized emotion data. Specifically, data is collected using a camera and microphone, and an emotion analysis algorithm is applied.

[0755] Step 7:

[0756] Generating and synthesizing explanations

[0757] The user selects a commentary style, or the generative AI model generates commentary based on recognized emotion data. The input is emotion data and clips of key scenes, and the output is a video clip with commentary. Specifically, a prompt is generated for the generative AI model, and commentary text and audio data are generated.

[0758] Step 8:

[0759] Save the final edited video and provide an access link

[0760] The cloud server integrates the generated commentary into the video and saves the final edited video in cloud storage. The input is a video clip with commentary, and the output is saving the final edited video and generating an access link. Specifically, the video editing software integrates the commentary, uploads it to the cloud, and generates an access link that is notified to the user.

[0761] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0762] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0763] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0764] [Third embodiment]

[0765] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0766] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[0767] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0768] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0769] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0770] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0771] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0772] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0773] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0774] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0775] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0776] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0777] This system automatically edits footage of soccer matches and other sporting events into high-quality videos and adds user-selected commentary. The system uploads video data captured with a wide-angle fixed camera to the cloud, analyzes and edits it using an AI module, and finally provides the user with a video that integrates the generated commentary.

[0778] System configuration and program processing

[0779] Recording and uploading videos

[0780] The device (wide-angle fixed camera) is equipped with a wide-angle lens that can capture the entire game, and records the game footage from a fixed position. The user sets up the camera device and presses the start button to start recording. After recording is complete, the device uploads the captured video data to a cloud server via Wi-Fi or a wired connection.

[0781] Receiving and storing video data

[0782] The server receives the uploaded video data and stores it in a storage service. In addition, the information on the player names and uniform numbers entered by the user in advance is also stored in a database.

[0783] Video Analysis and Editing

[0784] The server then passes the video data stored in storage to the AI ​​module, which then begins video analysis. The AI ​​module analyzes the ball position and player movements frame by frame to identify goal scenes and important plays. Based on the analysis results, the server extracts specific scenes, automatically zooms in, and generates clips.

[0785] Explanation generation and integration

[0786] The user selects a commentary style (e.g., "in the style of famous commentator A" or "in the style of famous commentator B") and sends that information to the cloud server. Based on the selected commentary style, the server invokes a generative AI model to generate text and audio commentary. The generated commentary is integrated into the video clip and saved in cloud storage as the final edited video.

[0787] Video distribution and viewing

[0788] The server then notifies the user of the download link for the final video. The notification is sent via email or an in-system notification. The user can then view or download the generated high-quality video with commentary via the provided link.

[0789] Specific examples

[0790] Editing of goal scenes and adding commentary

[0791] 1. Before the start of the game, the user registers the player's name and uniform number and sets up the device (camera device).

[0792] 2. The device records the match and uploads the video data to a cloud server after the match ends.

[0793] 3. The server stores the received video in a database and analyzes it using an AI module.

[0794] 4. The AI ​​module analyzes the ball's movement and identifies goal opportunities.

[0795] 5. The server zooms in on the goal scene and creates a clip.

[0796] 6. The user selects the commentary style "Famous Commentator A Style."

[0797] 7. The generative AI model generates commentary text and audio, for example, "Goal! What a great shot!"

[0798] 8. The server integrates the generated commentary to produce the final goal clip.

[0799] 9. The user then watches the final video with commentary via the provided link.

[0800] The system can also be used for a variety of events, including sports events other than soccer and athletic meets, and helps parents and coaches easily create high-quality footage for analysis and preservation of memories.

[0801] The processing flow will be explained below.

[0802] Step 1:

[0803] A user logs in to the system and registers a soccer match as a new event. The user enters the player names and uniform numbers through the interface and sends the registration information to the cloud server.

[0804] Step 2:

[0805] The device (a wide-angle fixed camera) captures the entire game with a wide angle. The user sets up the camera device and presses the start recording button to start recording. The device continues to record video data throughout the entire game.

[0806] Step 3:

[0807] After the game ends, the device uploads the captured video data to a cloud server via Wi-Fi or a wired connection.

[0808] Step 4:

[0809] The server receives the video data uploaded to the cloud. The received video data is saved in the storage service. The player names and uniform numbers registered by the user in advance are also saved in the database.

[0810] Step 5:

[0811] The server calls the video analysis AI module and analyzes the stored video data. The AI ​​module analyzes the ball position and player movements for each frame and identifies important scenes (goals, important plays, etc.).

[0812] Step 6:

[0813] The server extracts specific scenes based on the analysis results from the AI ​​module, and then runs an automatic zoom editing algorithm to generate a zoomed-in clip of important scenes or events.

[0814] Step 7:

[0815] After registering a player, the user selects a commentary style and sends that information to the cloud server. For example, the user can select a commentary style similar to that of famous commentator A.

[0816] Step 8:

[0817] Based on the selected commentary style, the server calls a generative AI model to generate commentary text and audio data. The generative AI model generates appropriate commentary for each scene based on static text and pre-recorded audio.

[0818] Step 9:

[0819] The server then integrates the generated commentary text and audio into the video clip, thereby generating the final edited video with commentary.

[0820] Step 10:

[0821] The server saves the final edited video to cloud storage and generates an access link for the user. Once saving is complete, the access link is notified to the user.

[0822] Step 11:

[0823] The user watches or downloads the generated high-quality commentary video via the provided link. If the user downloads the video, it is saved locally on the user's device.

[0824] In this way, a system is realized that provides users with high-quality videos with commentary that have been automatically analyzed and edited through a series of processing steps.

[0825] Example 1

[0826] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0827] It is difficult to efficiently edit video of sports games and events with high quality and add visually easy-to-understand commentary. In particular, there is a lack of systems that can efficiently extract only important scenes without watching the entire game and quickly provide edited videos with commentary. There is also a need for a method to easily generate videos that adapt to the user's preferred commentary style.

[0828] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0829] In this invention, the server includes means for uploading video data to a cloud server, means for receiving the video data in the cloud server and saving it in storage, artificial intelligence module means for analyzing the saved video data and tracking the positions of the ball and players, means for automatically extracting key scenes based on the analysis results and performing zoom editing, means for generating commentary using a generative artificial intelligence model based on a commentary style selected by the user, means for the user to select a commentary style and send that information to the cloud server, means for integrating the generated commentary into a video clip, means for notifying the user of a download link for the generated video with commentary, and means for saving the final video with commentary in cloud storage. This allows users to easily access high-quality video with commentary and efficiently enjoy key scenes from games and events.

[0830] A "wide-angle fixed camera" is a camera device that can capture a wide range of images from a fixed position.

[0831] A "cloud server" is a remote server that stores and processes data over the Internet.

[0832] "Storage" is a place or device for storing digital data.

[0833] "Artificial Intelligence Module" means a software module that contains artificial intelligence algorithms designed to perform data analysis or automated processing.

[0834] "Key Scenes" refers to scenes or moments that are particularly noteworthy in a sporting event or sporting event.

[0835] "Zoom editing" is an editing method that emphasizes important scenes by enlarging specific parts of the video.

[0836] "Commentary Style" is a setting that expresses the particular commentary style or tone desired by the user.

[0837] A "generative artificial intelligence model" is an artificial intelligence model for generating text or audio commentary based on input from a user.

[0838] A "video clip" is a short video segment extracted from a longer video.

[0839] A "download link" is a URL that allows a user to download a specific file via the Internet.

[0840] An "explanatory video" is a video that integrates explanatory audio and text with the video.

[0841] "Cloud storage" is an online storage service that allows you to store data via the Internet.

[0842] This invention is a system that automatically edits footage of sports games and events into high-quality videos with commentary. Specifically, it uses a wide-angle fixed camera to shoot video, and includes a process for analyzing and editing the video data on a cloud server. The details are as follows.

[0843] Video recording and uploading

[0844] Before the start of a game or event, users set up a wide-angle fixed camera in an appropriate position. When the camera presses the start button, it starts recording video using the wide-angle lens. When filming is finished, the device (camera device) uploads the video data to a cloud server via Wi-Fi or a wired connection.

[0845] Receiving and storing video data

[0846] The cloud server receives the uploaded video data and stores it in storage. Information such as player names and uniform numbers entered by the user before the game is also stored in the database. This data is used during analysis.

[0847] Video analysis and key scene extraction

[0848] The server then passes the stored video data to an AI module for video analysis. The AI ​​module tracks the ball position and player movements frame by frame to detect goals and other important plays. Based on the detected scenes, the server automatically performs zoom editing, extracts key scenes, and generates clips.

[0849] Selection of commentary style and generation of commentary

[0850] Users select their desired commentary style through an interface and send that information to a cloud server, which then invokes a generative artificial intelligence model to generate text and audio commentary.

[0851] Specifically, for example, a prompt sentence in the style of "famous commentator A" is sent to the generative AI model, and an audio file corresponding to the commentary, such as "Goal! What a great play!", is generated. The generative AI model used in this process utilizes advanced natural language processing technology.

[0852] Explanation and video integration

[0853] The server then integrates the generated commentary into the video clip, resulting in a video clip with commentary, which is then finally stored in cloud storage and made accessible to users.

[0854] Distribution and viewing of the final video

[0855] Finally, the server notifies the user of the download link for the generated commentary video. The notification method is email or in-system notification. The user can watch or download the high-quality commentary video through the provided link.

[0856] Examples of prompt statements

[0857] "Generate commentary for when a goal occurs immediately after the start of a match"

[0858] "When Player A scores a goal, please generate a commentary in the style of famous commentator B."

[0859] "Generate commentary on the most impressive plays in this game"

[0860] The system can be used to record not only sports matches, but also athletic meets and other events, helping parents and coaches easily create high-quality footage for analysis and preservation of memories.

[0861] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0862] Step 1: Camera setup and shooting

[0863] The user sets up a wide-angle fixed camera in an appropriate position and begins recording a game or event. When the user presses the start button, the camera records video using the wide-angle lens. The input includes the camera settings and information about when the start button was pressed. The output generates the captured video data. This video data is temporarily stored on an SD card or internal memory.

[0864] Step 2: Upload video data

[0865] Once the recording is complete, the device (camera device) uploads the video data to a cloud server via Wi-Fi or a wired connection. The input includes the video data stored in the camera. The data is compressed and transmitted over the network. The output is video data temporarily stored on the cloud server.

[0866] Step 3: Receiving and saving video data

[0867] The server receives video data uploaded from the device. The input includes video data sent via the network. After receiving it, the server stores it in a storage service. Information such as player names and uniform numbers registered by the user before the match is also stored in the database. As output, the video data stored in cloud storage and player information stored in the database are generated.

[0868] Step 4: Video analysis and key scene extraction

[0869] The server passes the stored video data to an artificial intelligence module to begin video analysis. The input includes the video data retrieved from storage. The artificial intelligence module tracks the ball position and player movements frame by frame to detect goals and other important plays. Based on the analysis results, the server automatically extracts key scenes and performs zoom editing. The output is a zoom-edited video clip containing key scenes.

[0870] Step 5: Choose your commentary style

[0871] The user selects the desired commentary style through the interface and sends the information to the cloud server. The input includes the user's selected commentary style. The output is saved in the server.

[0872] Step 6: Generate a description

[0873] The server invokes the generative AI model based on the commentary style selected by the user. The input includes the user's commentary style selection information and a prompt sentence for a specific scene. The generative AI model generates text and audio commentary. The output is the text commentary and audio file generated by the generative AI model.

[0874] Example prompt sentences:

[0875] "Generate commentary for when a goal occurs immediately after the start of a match"

[0876] "When Player A scores a goal, please generate a commentary in the style of famous commentator B."

[0877] "Generate commentary on the most impressive plays in this game"

[0878] Step 7: Integrating the explanation and video

[0879] The server integrates the generated description into the video clip. The input includes the generated text and audio description and the edited video clip. Using video editing software (e.g., FFmpeg), the description audio is merged into a specific timeline of the video clip. The output is a video clip with the description.

[0880] Step 8: Stream and view your final video

[0881] The server notifies the user of a download link for the generated annotated video. The notification method is via email or in-system notification. The input includes the annotated video clip stored in cloud storage. The output is a download link notified to the user. The user can watch or download the generated high-quality annotated video through the provided link.

[0882] Through these steps, users can generate and easily use high-quality videos with commentary that allow them to efficiently enjoy important scenes from games and events.

[0883] (Application example 1)

[0884] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0885] Conventional video editing systems for sporting events require manual editing, which is labor-intensive and time-consuming. Furthermore, specialized knowledge is required to add commentary, making it difficult for average users to create high-quality videos with commentary. Furthermore, managing and distributing edited videos is complicated. To solve these issues, it is necessary to provide a system that automatically analyzes and edits videos shot with a wide-angle fixed camera and adds commentary.

[0886] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0887] In this invention, the server includes a wide-angle fixed camera means for shooting videos, a means for uploading the shot video data to a cloud server, a means for receiving the video data in the cloud server and saving it in storage, an AI module means for analyzing the saved video data and tracking the positions of objects and people, a means for automatically extracting important scenes based on the analysis results and performing zoom editing, a means for generating commentary using a generative AI model based on a commentary style selected by the user, a means for adding the generated commentary, saving the final edited video in cloud storage and providing an access link, and a means for managing video data through an application installed on a smartphone. This allows users to easily create and manage high-quality videos with commentary.

[0888] The "wide-angle fixed-point camera means" is a camera device that uses a wide-angle lens to capture an overall image from a fixed position.

[0889] A "cloud server" is a remote server that stores, analyzes, and processes data over a network.

[0890] "Means of storing data in storage" refers to a method of storing data for the long term inside or outside the cloud server.

[0891] An "AI module means" is a software component that realizes artificial intelligence technology that mimics human intelligence and performs data analysis and predictions.

[0892] The "means for automatically extracting important scenes based on the analysis results and performing zoom editing" is a method for automatically selecting specific scenes using the analysis data generated by the AI ​​module means and focusing on a portion of the video for zoom editing.

[0893] "Means for generating commentary using a generative AI model based on the commentary style selected by the user" refers to a method for generating commentary with specific writing style and voice characteristics using AI technology in accordance with the user's selection.

[0894] "Means for adding generated commentary, saving the final edited video in cloud storage, and providing an access link" refers to a method for integrating commentary created by a generative AI model into an edited video and generating a URL that can be accessed by users after saving it in the cloud.

[0895] An "application installed on a smartphone" is software that is downloaded to and executed on a mobile device, and is a means for a user to manage video data.

[0896] An "identification number" is a number assigned to uniquely identify a particular player or object.

[0897] "Event" is a general term that refers to a series of activities or programs that take place at a specific time and place.

[0898] The system for implementing this invention uses a wide-angle fixed camera, a cloud server, storage, an AI module, and a generative AI model. The role of each component and the operation of the entire system are described in detail below.

[0899] A wide-angle fixed camera is used to capture the entire game or event. This camera captures footage from a fixed position using a wide-angle lens. A user sets up the camera device and begins capturing footage of the game or event.

[0900] The video data captured by the device is uploaded to a cloud server via Wi-Fi or a wired connection. The cloud server receives the uploaded video data and saves it in storage. In addition, information such as player names and identification numbers entered by the user in advance is also saved in a database.

[0901] The cloud server then passes the stored video data to the AI ​​module, which then begins video analysis. The AI ​​module uses technology that mimics human intelligence to analyze the movement of the ball and players frame by frame and identify key scenes. This analysis uses image processing libraries such as OpenCV.

[0902] Based on the analysis results, the server automatically extracts specific scenes and performs zoom editing as necessary. Users select a commentary style through an application installed on their smartphone. For example, they can choose a style such as "famous commentator A style." This information is sent to the cloud server.

[0903] The cloud server then invokes a generative AI model based on the user's selection to generate commentary with a specific writing style and voice characteristics. Using natural language processing, the generative AI model might create a commentary such as, "Goal! What a great shot!" The generated commentary is then integrated into the edited video to create the final video file.

[0904] The server stores the final video file in cloud storage and provides a download link to the user, who can then view or download the final video with commentary via the provided link.

[0905] Specific examples

[0906] For example, suppose a user films a soccer match and connects the camera device to a cloud server. The user uses an application to register player names and identification numbers, and uploads the video data after the match. The cloud server receives the video data and analyzes it using an AI module. Based on the analysis results, important scenes are automatically extracted and zoomed in for editing. The user selects "Famous Commentator A" as the commentary style, and as a result, the generative AI model generates a commentary such as "That was an excellent shot!"

[0907] Prompt Sentence Examples

[0908] Analyze a soccer match video provided by the user and identify goal scenes. The commentary style selected is "Famous Commentator A Style." Based on this style, generate commentary such as "Goal! What a great shot!"

[0909] The above is an embodiment of the present invention. By using this system, users can easily create and manage high-quality videos with commentary.

[0910] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0911] Step 1:

[0912] A user sets up a wide-angle fixed camera and begins capturing footage of a game or event.

[0913] Input: Camera device setup and recording start command

[0914] Output: Recorded video data (saved in local storage)

[0915] Step 2:

[0916] After the device finishes shooting, it uploads the captured video data to a cloud server.

[0917] Input: Video data stored in local storage

[0918] Data Processing: Transferring video data via Wi-Fi or wired connection

[0919] Output: Video data uploaded to the cloud server

[0920] Step 3:

[0921] The server receives the video data uploaded to the cloud server and stores it in storage.

[0922] Input: Video data uploaded to the cloud server

[0923] Data processing: Writing to cloud storage

[0924] Output: Video data stored in cloud storage

[0925] Step 4:

[0926] The server passes the stored video data to the AI ​​module, which then begins video analysis.

[0927] Input: Video data stored in cloud storage

[0928] Data calculation: AI module analyzes the movement of objects and people (ball and players) for each frame

[0929] Output: Analysis result data (object location information, movement patterns)

[0930] Step 5:

[0931] Based on the analysis results, the server automatically extracts important scenes and performs zoom editing.

[0932] Input: Analysis result data of AI module

[0933] Data processing: Extraction of important scenes and video processing by zoom editing

[0934] Output: Edited clip data

[0935] Step 6:

[0936] The user selects a commentary style via a smartphone app. For example, they select "Famous Commentator A Style."

[0937] Input: User's commentary style selection in the app

[0938] Output: Selected commentary style information (sent to cloud server)

[0939] Step 7:

[0940] The server generates commentary using a generative AI model based on the selected commentary style.

[0941] Input: User selected commentary style information

[0942] Data Computing: Generating Explanatory Text and Audio with Generative AI Models

[0943] Output: Generated commentary data (e.g., "Goal! What a great shot!" audio and text)

[0944] Step 8:

[0945] The server adds the generated commentary and saves the final edited video to cloud storage.

[0946] Input: Edited clip data and generated commentary data

[0947] Data processing: Integrating audio commentary and text into video

[0948] Output: Final edited video data

[0949] Step 9:

[0950] The server then generates a download link for the final video and notifies the user.

[0951] Input: Final edited video data

[0952] Data processing: Download link generation and URL generation

[0953] Output: Notification to user (email or in-system notification)

[0954] This allows users to view or download the final video with commentary via the provided link.

[0955] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0956] This system automatically edits footage of soccer matches and other sporting events into high-quality videos and adds user-selected commentary. Furthermore, the system recognizes the user's emotions and automatically selects and refines the commentary style based on those emotions. The system uploads video data captured with a wide-angle fixed camera to the cloud, where it is analyzed and edited by an AI module, and finally provides the user with a video that integrates the generated commentary.

[0957] System configuration and program processing

[0958] Recording and uploading videos

[0959] The device (wide-angle fixed camera) is equipped with a wide-angle lens that can capture the entire game, and records the game footage from a fixed position. The user sets up the camera device and presses the start button to start recording. After recording is complete, the device uploads the captured video data to a cloud server via Wi-Fi or a wired connection.

[0960] Receiving and storing video data

[0961] The server receives the uploaded video data and stores it in a storage service. In addition, the information on the player names and uniform numbers entered by the user in advance is also stored in a database.

[0962] Video Analysis and Editing

[0963] The server then passes the video data stored in storage to the AI ​​module, which then begins video analysis. The AI ​​module analyzes the ball position and player movements frame by frame to identify goal scenes and important plays. Based on the analysis results, the server extracts specific scenes, automatically zooms in, and generates clips.

[0964] Recognizing user emotions with an emotion engine

[0965] When a user accesses the system through a device, the emotion engine recognizes the user's emotions, obtains emotional data from the user's facial expressions and voice, and adjusts the commentary style accordingly.

[0966] Explanation generation and integration

[0967] The user manually selects the commentary style, or it is automatically selected by the emotion engine. Based on the selected commentary style, the server invokes a generative AI model to generate commentary text and audio data. The generated commentary is integrated into the video clip and saved in cloud storage as the final edited video.

[0968] Video distribution and viewing

[0969] The server then notifies the user of the download link for the final video. The notification is sent via email or an in-system notification. The user can then view or download the generated high-quality video with commentary via the provided link.

[0970] Specific examples

[0971] Editing of goal scenes and adding commentary

[0972] 1. Before the start of the game, the user registers the player's name and uniform number and sets up the device (camera device).

[0973] 2. The device records the match and uploads the video data to a cloud server after the match ends.

[0974] 3. The server stores the received video in a database and analyzes it using an AI module.

[0975] 4. The AI ​​module analyzes the ball's movement and identifies goal opportunities.

[0976] 5. The server zooms in on the goal scene and creates a clip.

[0977] 6. When the user manually or automatically selects a commentary style, the emotion engine recognizes the user's excitement and selects a more energetic commentary style, such as "Goal! What a great shot!"

[0978] 7. The generative AI model generates explanatory text and audio.

[0979] 8. The server integrates the generated commentary to produce the final goal clip.

[0980] 9. The user then watches the final video with commentary via the provided link.

[0981] The introduction of an emotion engine makes it possible to provide customized commentary that corresponds to the user's specific emotional state, further enhancing the viewing experience.The system can also be used for various events other than soccer, such as sports events and athletic meets, and helps families and coaches easily create high-quality footage for analysis and recording memories.

[0982] The processing flow will be explained below.

[0983] Step 1:

[0984] A user logs in to the system and registers a soccer match as a new event. The user enters the player names and uniform numbers through the interface and sends the registration information to the cloud server.

[0985] Step 2:

[0986] The device (wide-angle fixed camera) captures the entire game with a wide angle. Once the setup is complete, the user presses the start button on the camera device to begin recording the game. The device continues to record video data throughout the entire game.

[0987] Step 3:

[0988] After the game ends, the device uploads the captured video data to a cloud server via Wi-Fi or a wired connection. The video data is sent directly to the cloud server.

[0989] Step 4:

[0990] The server receives the video data uploaded to the cloud and stores it in the storage service. At the same time, the player names and uniform numbers registered by the user in advance are also stored in the database.

[0991] Step 5:

[0992] The server then calls up the video analysis AI module, which analyzes the stored video data. The AI ​​module analyzes the ball position and player movements for each frame to identify goal scenes and key plays.

[0993] Step 6:

[0994] The server extracts specific scenes based on the analysis results from the AI ​​module, runs an automatic zoom editing algorithm, zooms in on important scenes and events, and generates clips.

[0995] Step 7:

[0996] After registering a player, the user manually selects a commentary style. In some cases, the emotion engine recognizes the user's emotions and automatically selects an appropriate commentary style. This allows the system to provide commentary that reflects the user's emotions, such as joy or excitement.

[0997] Step 8:

[0998] The emotion engine analyzes the user's emotions from their facial expressions and voice, obtains emotion data in real time, and provides feedback to the selection of commentary style.

[0999] Step 9:

[1000] The server calls a generative AI model to generate commentary text and audio data based on the selected commentary style. For example, if the user is excited, an energetic commentary will be generated.

[1001] Step 10:

[1002] The server integrates the generated commentary text and audio into the video clip, generates a final edited video with commentary, and stores the video in cloud storage.

[1003] Step 11:

[1004] The server will notify the user with a download link for the final edited video. Notifications will be sent via email or in-system notifications. Once the link is provided, the user will have immediate access.

[1005] Step 12:

[1006] The user can then view or download the generated high-quality commentary video via the provided link, which can be saved locally on the user's device.

[1007] Through these processing steps, high-quality videos with annotations that match the user's emotions are automatically generated and provided, allowing the user to enjoy videos with annotations that reflect specific emotions when played back, resulting in a deeper viewing experience.

[1008] Example 2

[1009] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1010] There is a demand for high-quality editing of sporting event videos and the automatic addition of appropriate commentary based on user emotions. However, conventional technologies have difficulty automating video editing and providing emotion-based commentary, requiring a high level of specialized knowledge and manual work. As a result, it is difficult to generate timely and personalized video content.

[1011] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1012] In this invention, the server includes a wide-angle fixed-point camera means for shooting video, a means for uploading the shot video data to an information processing device, and a means for receiving the video data in the information processing device and storing it in a storage device. This enables highly accurate analysis and editing of the video data. The server also includes an artificial intelligence module means for analyzing the saved video data and tracking the positions of the sphere and people, a means for automatically extracting important scenes based on the analysis results and performing zoom editing, a means for generating commentary using a generative artificial intelligence model based on a commentary style selected by the user, a means for adding the generated commentary, storing the final edited video in a storage device, and providing an access link, and an emotion recognition engine means for recognizing the user's emotions and automatically selecting a commentary style based on the emotions. This enables automatic and efficient generation of high-quality video with commentary that corresponds to the user's emotions.

[1013] A "wide-angle fixed-point camera" is a device that captures images from a fixed position with a wide field of view, and is primarily used to record the overall picture of sporting events and the like.

[1014] An "information processing device" is a device that processes, analyzes, stores, and transmits data, and includes a cloud server, a local server, and the like.

[1015] A "storage device" is a hardware or software means for storing digital data, including cloud storage and hard disk drives.

[1016] An "artificial intelligence module" is a software module that uses machine learning and deep learning to analyze data and recognize specific patterns and features.

[1017] "Tracking the position of spheres and people" refers to the process of detecting specific objects, such as balls or players, within video frames and continuously tracking their position information.

[1018] "Automatic extraction of important scenes" is the process of identifying specific events or actions within a video based on the analysis results of an artificial intelligence module and cutting out only the relevant parts.

[1019] "Zoom editing" is an editing technique that enlarges and visually enhances portions of the original footage, and is used to provide viewers with important scene details.

[1020] A "generative AI model" is a model that uses AI technology to generate text and speech based on user selections and emotional data.

[1021] An "emotion recognition engine" is a software module that analyzes and recognizes a user's emotional state from their facial expressions and voice, acquiring emotional data and reflecting it in the analysis results.

[1022] MODE FOR CARRYING OUT THE INVENTION

[1023] The present invention is a system that automatically edits video data from sporting events into high-quality videos and adds user-selected commentary. It also provides the ability to recognize user emotions and automatically select and refine commentary styles based on those emotions. The system begins by capturing game footage using a fixed-point camera with a wide field of view and uploading the raw data to a cloud server. Specific embodiments of the present invention are described below.

[1024] System Configuration

[1025] 1. Wide-angle fixed-point shooting device

[1026] The wide-angle fixed-point camera is a dedicated camera that records the entire sporting event over a wide area. Users set up the camera and start recording at the start of the game. The video data captured by the camera is uploaded to a cloud server via Wi-Fi or a wired connection.

[1027] 2. Cloud Server

[1028] The cloud server receives the video data and stores it securely in a storage device. The cloud server also stores information such as player names and uniform numbers that users have registered in advance. The video data is then analyzed by an AI module.

[1029] 3. Artificial Intelligence Module

[1030] An AI module in the cloud server analyzes the video data frame by frame, tracking the positions of balls and people. This analysis identifies goals and other key moments. The AI ​​module extracts key scenes and performs zoom editing to generate clips.

[1031] 4. Emotion Recognition Engine

[1032] When a user accesses the system, the emotion recognition engine acquires emotional data from the user's facial expressions and voice, which is used as an important factor in determining the commentary style.

[1033] 5. Generative AI Models

[1034] The generative AI model generates commentary text and audio data based on the user's emotion recognition results and selected commentary style, which is then integrated with the video clip to create the final edited video.

[1035] 6. Video streaming and viewing

[1036] The cloud server then sends the user a download link for the final video, allowing the user to view or download the high-quality video with commentary.

[1037] Specific examples

[1038] 1. Editing goal scenes and adding commentary

[1039] Before the game, the user registers the player names and uniform numbers and sets up the wide-angle fixed-point camera.

[1040] The device records the match and uploads the video data to a cloud server after the match ends.

[1041] The video data received by the server is stored in a storage device and analyzed by an AI module.

[1042] The AI ​​module analyzes the movement of the ball and identifies goal opportunities.

[1043] The server zooms in on the goal scene and creates a clip.

[1044] The user can select a commentary style, or an emotion recognition engine can automatically select one based on the emotion. In this case, an energetic commentary such as "Goal! What a great shot!" is generated.

[1045] A generative AI model generates explanatory text and audio.

[1046] The server then integrates the generated commentary to produce the final goal clip.

[1047] The user then watches the final video with commentary via the provided link.

[1048] Prompt Sentence Examples

[1049] "The entire match should be filmed with a wide-angle camera, and specific scenes should be automatically edited and commentary should be added. The commentary should be selected and generated based on the user's emotions."

[1050] The above is a specific embodiment for carrying out the present invention. This system can be applied to various sporting events other than soccer, and efficiently generates high-quality video.

[1051] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1052] Step 1: Record your video and upload it to the cloud

[1053] The terminal (wide-angle fixed-point camera) is equipped with a wide-angle lens that can capture the entire game. The user sets up the camera and presses the start button to begin recording. The input is live footage of the game, and the output is the captured video data. This video data is uploaded to a cloud server via Wi-Fi or a wired connection after the game ends. Specific operations include the user adjusting the camera position and instructing the game to start.

[1054] Step 2: Receiving and saving video data

[1055] The server receives video data from the cloud and stores it in a storage device. The input is the video data uploaded from the device, and the output is the data stored in a storage device on the cloud. Specific operations include the server confirming receipt of the data and storing it in the appropriate folder. At the same time, the player name and uniform number information previously entered by the user is also stored in the database.

[1056] Step 3: Video Analysis and Editing

[1057] The server passes the video data stored in the storage device to an AI module and begins data analysis. The input is the stored video data, and the output is the analysis results identifying key scenes. The AI ​​module tracks the ball position and player movements for each frame to identify goals and important plays. Specific operations include extracting the identified scenes and performing zoom editing to generate clips.

[1058] Step 4: Recognizing user emotions with the emotion engine

[1059] When a user accesses the system, the emotion recognition engine analyzes the user's emotions. The input is the user's facial expression and voice data, and the output is recognized emotion data. Specific operations include the process in which the user provides emotion data using a webcam or microphone, and the emotion engine analyzes that data.

[1060] Step 5: Explanation generation and integration

[1061] The user manually selects the commentary style, or the emotion recognition engine automatically selects it. The input is the recognized emotion data and the selected commentary style, and the output is the generated commentary text and audio data. The server uses a generative AI model to generate commentary and integrate it into the video clip. Specific operations include the server calling the generative AI model and automatically generating and integrating commentary.

[1062] Step 6: Stream and view your video

[1063] The server notifies the user of the download link for the final video. The input is the final video data, and the output is a link that the user can access. Specific operations include the server sending a notification email or an in-system notification and providing the user with a viewing link. The user can then view or download the high-quality video with commentary via the link.

[1064] (Application example 2)

[1065] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1066] Conventional video editing systems for sporting events make it difficult for users to add their own commentary and lack the ability to dynamically change the commentary style based on the user's emotions. As a result, viewers' viewing experience is consistent and they are unable to respond to the needs of individual users. Furthermore, there is a need for similar high-quality editing functions for a variety of sporting events other than soccer.

[1067] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1068] In this invention, the server includes a wide-angle fixed camera means for shooting videos, a means for uploading the shot video data to a cloud server, a means for receiving the video data in the cloud server and storing it in storage, an AI module means for analyzing the stored video data and tracking the positions of the ball and players, a means for automatically extracting important scenes based on the analysis results and performing zoom editing, a means for generating commentary using a generative AI model based on a commentary style selected by the user, an emotion engine means for recognizing the user's emotions and adjusting the commentary style based on the emotions, and a means for adding the generated commentary, storing the final edited video in cloud storage, and providing an access link. This makes it possible to provide high-quality videos with commentary that are adapted to the user's emotions.

[1069] The "wide-angle fixed-point camera means" is a device equipped with a wide-angle lens that captures moving images from a fixed position.

[1070] A "cloud server" is a computer system that stores and processes data and is accessible remotely over the Internet.

[1071] "Storage" refers to a storage device or service for saving data.

[1072] "AI module means" is software that uses artificial intelligence to analyze video and track the movements of the ball and players.

[1073] "Zoom editing" is an editing technique that visually emphasizes a particular scene by enlarging it.

[1074] A "generative AI model" is an artificial intelligence model that automatically generates natural language based on user input and data.

[1075] The "emotion engine means" is software for analyzing the user's facial expressions and voice and recognizing their emotions.

[1076] An "access link" is a URL or hyperlink that allows a user to access specific content via the Internet.

[1077] "Means for registering player names and uniform numbers" is a function for entering player names and uniform numbers into the system for video analysis.

[1078] The "emotion-based commentary style automatic adjustment function" is a function that dynamically changes the style and tone of commentary using the user's emotional data.

[1079] In this invention, a user shoots a sporting event with a wide-angle fixed camera and uploads the video data to a cloud server. The cloud server receives the video data and saves it in storage. The saved video data is analyzed by an AI module, and the positions of the ball and players are tracked. Based on the analysis results, important scenes are automatically extracted and zoomed in for editing.

[1080] If the user selects a commentary style, a commentary is generated using the generative AI model. If the user's emotion is recognized, an emotion engine means analyzes the user's facial expressions and voice and adjusts the commentary style. The generated commentary is added to the video, and the final edited video is saved in cloud storage, and an access link is provided to the user.

[1081] A specific system configuration operates as follows.

[1082] 1. Hardware and Software Description

[1083] Wide-angle fixed camera means: A device equipped with a wide-angle lens that shoots video from a fixed position.

[1084] Cloud Server: A computer system that stores and processes data and is accessible remotely over the Internet.

[1085] Storage: A storage device or service for storing data (e.g., Amazon S3, Google Cloud Storage).

[1086] AI module means: Software (e.g. OpenCV) that uses artificial intelligence to analyze video and track ball and player movements.

[1087] Generative AI models: Artificial intelligence models that generate commentary based on user input and data (e.g., GPT-3).

[1088] Emotion engine means: Software that analyzes the user's facial expressions and voice and recognizes their emotions (e.g., OpenFace).

[1089] 2. Description of Data Processing and Data Calculations

[1090] Uploading and saving video data: The video data taken by the user is uploaded to the cloud server and saved in storage. The saved data is passed to the AI ​​module.

[1091] Video analysis: The AI ​​module analyzes the video data and tracks the position of the ball and players frame by frame. Key scenes are extracted and zoomed in for editing.

[1092] Emotion Recognition: Emotion data is acquired from the user's facial expressions and voice. The emotion engine analyzes the user's emotions and adjusts the commentary style accordingly.

[1093] Commentary generation and integration: A generative AI model generates commentary text and audio based on the emotion data and integrates it into the video. The final edited video is saved in cloud storage and an access link is provided to the user.

[1094] 3. Examples of concrete examples and prompts

[1095] For example, a user can film a soccer game using their smartphone and upload the video to the cloud. The cloud server then analyzes the video data and adds energetic commentary based on the user's emotional data. Finally, a high-quality video with commentary is generated, and the user receives a link to watch it.

[1096] Example prompts for generative AI models

[1097] If the user is excited while watching:

[1098] "It was an exciting goal scene! The moment the shot shook the net, the whole stadium erupted in cheers!"

[1099] If the user is relaxing while watching:

[1100] "Here, a relaxed ball movement resulted in a fantastic goal. This play is a great example of the teamwork."

[1101] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1102] Step 1:

[1103] The user takes a video

[1104] A user shoots video of a sporting event using a wide-angle fixed camera, and video data is generated. The inputs are the user's operations and the sporting event being shot, and the output is a video file of the shot video.

[1105] Step 2:

[1106] Upload videos to a cloud server

[1107] After the user has finished shooting, they upload the video data to the cloud server. The input for the upload is a video file, and the output is the video data saved on the cloud server. The user is also notified of the upload status.

[1108] Step 3:

[1109] The cloud server receives and stores the data.

[1110] The cloud server receives the uploaded video data and saves it to storage. The input is the uploaded video data, and the output is the storage of the data in the destination storage. A notification of successful saving is also generated.

[1111] Step 4:

[1112] Video analysis and key scene extraction

[1113] The cloud server uses an AI module to analyze the stored video data and track the positions of the ball and players. The input is the stored video data, and the output is the position data of the ball and players, as well as identification information for important scenes. Specifically, the system analyzes each frame and records the position data.

[1114] Step 5:

[1115] Performing Zoom Edits

[1116] The cloud server automatically extracts important scenes and performs zoom editing. The input is the identification information and video data of the important scenes, and the output is a zoom-edited clip of the scene. Specifically, editing software is used to enlarge and display the important scenes.

[1117] Step 6:

[1118] Recognize user emotions with an emotion engine

[1119] When a user watches a video, the emotion engine means analyzes the user's facial expressions and voice to recognize emotions. The input is the user's facial expression data and voice data, and the output is recognized emotion data. Specifically, data is collected using a camera and microphone, and an emotion analysis algorithm is applied.

[1120] Step 7:

[1121] Generating and synthesizing explanations

[1122] The user selects a commentary style, or the generative AI model generates commentary based on recognized emotion data. The input is emotion data and clips of key scenes, and the output is a video clip with commentary. Specifically, a prompt is generated for the generative AI model, and commentary text and audio data are generated.

[1123] Step 8:

[1124] Save the final edited video and provide an access link

[1125] The cloud server integrates the generated commentary into the video and saves the final edited video in cloud storage. The input is a video clip with commentary, and the output is saving the final edited video and generating an access link. Specifically, the video editing software integrates the commentary, uploads it to the cloud, and generates an access link that is notified to the user.

[1126] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1127] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1128] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1129] [Fourth embodiment]

[1130] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1131] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1132] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1133] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1134] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1135] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1136] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1137] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1138] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1139] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1140] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1141] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1142] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1143] This system automatically edits footage of soccer matches and other sporting events into high-quality videos and adds user-selected commentary. The system uploads video data captured with a wide-angle fixed camera to the cloud, analyzes and edits it using an AI module, and finally provides the user with a video that integrates the generated commentary.

[1144] System configuration and program processing

[1145] Recording and uploading videos

[1146] The device (wide-angle fixed camera) is equipped with a wide-angle lens that can capture the entire game, and records the game footage from a fixed position. The user sets up the camera device and presses the start button to start recording. After recording is complete, the device uploads the captured video data to a cloud server via Wi-Fi or a wired connection.

[1147] Receiving and storing video data

[1148] The server receives the uploaded video data and stores it in a storage service. In addition, the information on the player names and uniform numbers entered by the user in advance is also stored in a database.

[1149] Video Analysis and Editing

[1150] The server then passes the video data stored in storage to the AI ​​module, which then begins video analysis. The AI ​​module analyzes the ball position and player movements frame by frame to identify goal scenes and important plays. Based on the analysis results, the server extracts specific scenes, automatically zooms in, and generates clips.

[1151] Explanation generation and integration

[1152] The user selects a commentary style (e.g., "in the style of famous commentator A" or "in the style of famous commentator B") and sends that information to the cloud server. Based on the selected commentary style, the server invokes a generative AI model to generate text and audio commentary. The generated commentary is integrated into the video clip and saved in cloud storage as the final edited video.

[1153] Video distribution and viewing

[1154] The server then notifies the user of the download link for the final video. The notification is sent via email or an in-system notification. The user can then view or download the generated high-quality video with commentary via the provided link.

[1155] Specific examples

[1156] Editing of goal scenes and adding commentary

[1157] 1. Before the start of the game, the user registers the player's name and uniform number and sets up the device (camera device).

[1158] 2. The device records the match and uploads the video data to a cloud server after the match ends.

[1159] 3. The server stores the received video in a database and analyzes it using an AI module.

[1160] 4. The AI ​​module analyzes the ball's movement and identifies goal opportunities.

[1161] 5. The server zooms in on the goal scene and creates a clip.

[1162] 6. The user selects the commentary style "Famous Commentator A Style."

[1163] 7. The generative AI model generates commentary text and audio, for example, "Goal! What a great shot!"

[1164] 8. The server integrates the generated commentary to produce the final goal clip.

[1165] 9. The user then watches the final video with commentary via the provided link.

[1166] The system can also be used for a variety of events, including sports events other than soccer and athletic meets, and helps parents and coaches easily create high-quality footage for analysis and preservation of memories.

[1167] The processing flow will be explained below.

[1168] Step 1:

[1169] A user logs in to the system and registers a soccer match as a new event. The user enters the player names and uniform numbers through the interface and sends the registration information to the cloud server.

[1170] Step 2:

[1171] The device (a wide-angle fixed camera) captures the entire game with a wide angle. The user sets up the camera device and presses the start recording button to start recording. The device continues to record video data throughout the entire game.

[1172] Step 3:

[1173] After the game ends, the device uploads the captured video data to a cloud server via Wi-Fi or a wired connection.

[1174] Step 4:

[1175] The server receives the video data uploaded to the cloud. The received video data is saved in the storage service. The player names and uniform numbers registered by the user in advance are also saved in the database.

[1176] Step 5:

[1177] The server calls the video analysis AI module and analyzes the stored video data. The AI ​​module analyzes the ball position and player movements for each frame and identifies important scenes (goals, important plays, etc.).

[1178] Step 6:

[1179] The server extracts specific scenes based on the analysis results from the AI ​​module, and then runs an automatic zoom editing algorithm to generate a zoomed-in clip of important scenes or events.

[1180] Step 7:

[1181] After registering a player, the user selects a commentary style and sends that information to the cloud server. For example, the user can select a commentary style similar to that of famous commentator A.

[1182] Step 8:

[1183] Based on the selected commentary style, the server calls a generative AI model to generate commentary text and audio data. The generative AI model generates appropriate commentary for each scene based on static text and pre-recorded audio.

[1184] Step 9:

[1185] The server then integrates the generated commentary text and audio into the video clip, thereby generating the final edited video with commentary.

[1186] Step 10:

[1187] The server saves the final edited video to cloud storage and generates an access link for the user. Once saving is complete, the access link is notified to the user.

[1188] Step 11:

[1189] The user watches or downloads the generated high-quality commentary video via the provided link. If the user downloads the video, it is saved locally on the user's device.

[1190] In this way, a system is realized that provides users with high-quality videos with commentary that have been automatically analyzed and edited through a series of processing steps.

[1191] Example 1

[1192] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1193] It is difficult to efficiently edit video of sports games and events with high quality and add visually easy-to-understand commentary. In particular, there is a lack of systems that can efficiently extract only important scenes without watching the entire game and quickly provide edited videos with commentary. There is also a need for a method to easily generate videos that adapt to the user's preferred commentary style.

[1194] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1195] In this invention, the server includes means for uploading video data to a cloud server, means for receiving the video data in the cloud server and saving it in storage, artificial intelligence module means for analyzing the saved video data and tracking the positions of the ball and players, means for automatically extracting key scenes based on the analysis results and performing zoom editing, means for generating commentary using a generative artificial intelligence model based on a commentary style selected by the user, means for the user to select a commentary style and send that information to the cloud server, means for integrating the generated commentary into a video clip, means for notifying the user of a download link for the generated video with commentary, and means for saving the final video with commentary in cloud storage. This allows users to easily access high-quality video with commentary and efficiently enjoy key scenes from games and events.

[1196] A "wide-angle fixed camera" is a camera device that can capture a wide range of images from a fixed position.

[1197] A "cloud server" is a remote server that stores and processes data over the Internet.

[1198] "Storage" is a place or device for storing digital data.

[1199] "Artificial Intelligence Module" means a software module that contains artificial intelligence algorithms designed to perform data analysis or automated processing.

[1200] "Key Scenes" refers to scenes or moments that are particularly noteworthy in a sporting event or sporting event.

[1201] "Zoom editing" is an editing method that emphasizes important scenes by enlarging specific parts of the video.

[1202] "Commentary Style" is a setting that expresses the particular commentary style or tone desired by the user.

[1203] A "generative artificial intelligence model" is an artificial intelligence model for generating text or audio commentary based on input from a user.

[1204] A "video clip" is a short video segment extracted from a longer video.

[1205] A "download link" is a URL that allows a user to download a specific file via the Internet.

[1206] An "explanatory video" is a video that integrates explanatory audio and text with the video.

[1207] "Cloud storage" is an online storage service that allows you to store data via the Internet.

[1208] This invention is a system that automatically edits footage of sports games and events into high-quality videos with commentary. Specifically, it uses a wide-angle fixed camera to shoot video, and includes a process for analyzing and editing the video data on a cloud server. The details are as follows.

[1209] Video recording and uploading

[1210] Before the start of a game or event, users set up a wide-angle fixed camera in an appropriate position. When the camera presses the start button, it starts recording video using the wide-angle lens. When filming is finished, the device (camera device) uploads the video data to a cloud server via Wi-Fi or a wired connection.

[1211] Receiving and storing video data

[1212] The cloud server receives the uploaded video data and stores it in storage. Information such as player names and uniform numbers entered by the user before the game is also stored in the database. This data is used during analysis.

[1213] Video analysis and key scene extraction

[1214] The server then passes the stored video data to an AI module for video analysis. The AI ​​module tracks the ball position and player movements frame by frame to detect goals and other important plays. Based on the detected scenes, the server automatically performs zoom editing, extracts key scenes, and generates clips.

[1215] Selection of commentary style and generation of commentary

[1216] Users select their desired commentary style through an interface and send that information to a cloud server, which then invokes a generative artificial intelligence model to generate text and audio commentary.

[1217] Specifically, for example, a prompt sentence in the style of "famous commentator A" is sent to the generative AI model, and an audio file corresponding to the commentary, such as "Goal! What a great play!", is generated. The generative AI model used in this process utilizes advanced natural language processing technology.

[1218] Explanation and video integration

[1219] The server then integrates the generated commentary into the video clip, resulting in a video clip with commentary, which is then finally stored in cloud storage and made accessible to users.

[1220] Distribution and viewing of the final video

[1221] Finally, the server notifies the user of the download link for the generated commentary video. The notification method is email or in-system notification. The user can watch or download the high-quality commentary video through the provided link.

[1222] Examples of prompt statements

[1223] "Generate commentary for when a goal occurs immediately after the start of a match"

[1224] "When Player A scores a goal, please generate a commentary in the style of famous commentator B."

[1225] "Generate commentary on the most impressive plays in this game"

[1226] The system can be used to record not only sports matches, but also athletic meets and other events, helping parents and coaches easily create high-quality footage for analysis and preservation of memories.

[1227] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1228] Step 1: Camera setup and shooting

[1229] The user sets up a wide-angle fixed camera in an appropriate position and begins recording a game or event. When the user presses the start button, the camera records video using the wide-angle lens. The input includes the camera settings and information about when the start button was pressed. The output generates the captured video data. This video data is temporarily stored on an SD card or internal memory.

[1230] Step 2: Upload video data

[1231] Once the recording is complete, the device (camera device) uploads the video data to a cloud server via Wi-Fi or a wired connection. The input includes the video data stored in the camera. The data is compressed and transmitted over the network. The output is video data temporarily stored on the cloud server.

[1232] Step 3: Receiving and saving video data

[1233] The server receives video data uploaded from the device. The input includes video data sent via the network. After receiving it, the server stores it in a storage service. Information such as player names and uniform numbers registered by the user before the match is also stored in the database. As output, the video data stored in cloud storage and player information stored in the database are generated.

[1234] Step 4: Video analysis and key scene extraction

[1235] The server passes the stored video data to an artificial intelligence module to begin video analysis. The input includes the video data retrieved from storage. The artificial intelligence module tracks the ball position and player movements frame by frame to detect goals and other important plays. Based on the analysis results, the server automatically extracts key scenes and performs zoom editing. The output is a zoom-edited video clip containing key scenes.

[1236] Step 5: Choose your commentary style

[1237] The user selects the desired commentary style through the interface and sends the information to the cloud server. The input includes the user's selected commentary style. The output is saved in the server.

[1238] Step 6: Generate a description

[1239] The server invokes the generative AI model based on the commentary style selected by the user. The input includes the user's commentary style selection information and a prompt sentence for a specific scene. The generative AI model generates text and audio commentary. The output is the text commentary and audio file generated by the generative AI model.

[1240] Example prompt sentences:

[1241] "Generate commentary for when a goal occurs immediately after the start of a match"

[1242] "When Player A scores a goal, please generate a commentary in the style of famous commentator B."

[1243] "Generate commentary on the most impressive plays in this game"

[1244] Step 7: Integrating the explanation and video

[1245] The server integrates the generated description into the video clip. The input includes the generated text and audio description and the edited video clip. Using video editing software (e.g., FFmpeg), the description audio is merged into a specific timeline of the video clip. The output is a video clip with the description.

[1246] Step 8: Stream and view your final video

[1247] The server notifies the user of a download link for the generated annotated video. The notification method is via email or in-system notification. The input includes the annotated video clip stored in cloud storage. The output is a download link notified to the user. The user can watch or download the generated high-quality annotated video through the provided link.

[1248] Through these steps, users can generate and easily use high-quality videos with commentary that allow them to efficiently enjoy important scenes from games and events.

[1249] (Application example 1)

[1250] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1251] Conventional video editing systems for sporting events require manual editing, which is labor-intensive and time-consuming. Furthermore, specialized knowledge is required to add commentary, making it difficult for average users to create high-quality videos with commentary. Furthermore, managing and distributing edited videos is complicated. To solve these issues, it is necessary to provide a system that automatically analyzes and edits videos shot with a wide-angle fixed camera and adds commentary.

[1252] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1253] In this invention, the server includes a wide-angle fixed camera means for shooting videos, a means for uploading the shot video data to a cloud server, a means for receiving the video data in the cloud server and saving it in storage, an AI module means for analyzing the saved video data and tracking the positions of objects and people, a means for automatically extracting important scenes based on the analysis results and performing zoom editing, a means for generating commentary using a generative AI model based on a commentary style selected by the user, a means for adding the generated commentary, saving the final edited video in cloud storage and providing an access link, and a means for managing video data through an application installed on a smartphone. This allows users to easily create and manage high-quality videos with commentary.

[1254] The "wide-angle fixed-point camera means" is a camera device that uses a wide-angle lens to capture an overall image from a fixed position.

[1255] A "cloud server" is a remote server that stores, analyzes, and processes data over a network.

[1256] "Means of storing data in storage" refers to a method of storing data for the long term inside or outside the cloud server.

[1257] An "AI module means" is a software component that realizes artificial intelligence technology that mimics human intelligence and performs data analysis and predictions.

[1258] The "means for automatically extracting important scenes based on the analysis results and performing zoom editing" is a method for automatically selecting specific scenes using the analysis data generated by the AI ​​module means and focusing on a portion of the video for zoom editing.

[1259] "Means for generating commentary using a generative AI model based on the commentary style selected by the user" refers to a method for generating commentary with specific writing style and voice characteristics using AI technology in accordance with the user's selection.

[1260] "Means for adding generated commentary, saving the final edited video in cloud storage, and providing an access link" refers to a method for integrating commentary created by a generative AI model into an edited video and generating a URL that can be accessed by users after saving it in the cloud.

[1261] An "application installed on a smartphone" is software that is downloaded to and executed on a mobile device, and is a means for a user to manage video data.

[1262] An "identification number" is a number assigned to uniquely identify a particular player or object.

[1263] "Event" is a general term that refers to a series of activities or programs that take place at a specific time and place.

[1264] The system for implementing this invention uses a wide-angle fixed camera, a cloud server, storage, an AI module, and a generative AI model. The role of each component and the operation of the entire system are described in detail below.

[1265] A wide-angle fixed camera is used to capture the entire game or event. This camera captures footage from a fixed position using a wide-angle lens. A user sets up the camera device and begins capturing footage of the game or event.

[1266] The video data captured by the device is uploaded to a cloud server via Wi-Fi or a wired connection. The cloud server receives the uploaded video data and saves it in storage. In addition, information such as player names and identification numbers entered by the user in advance is also saved in a database.

[1267] The cloud server then passes the stored video data to the AI ​​module, which then begins video analysis. The AI ​​module uses technology that mimics human intelligence to analyze the movement of the ball and players frame by frame and identify key scenes. This analysis uses image processing libraries such as OpenCV.

[1268] Based on the analysis results, the server automatically extracts specific scenes and performs zoom editing as necessary. Users select a commentary style through an application installed on their smartphone. For example, they can choose a style such as "famous commentator A style." This information is sent to the cloud server.

[1269] The cloud server then invokes a generative AI model based on the user's selection to generate commentary with a specific writing style and voice characteristics. Using natural language processing, the generative AI model might create a commentary such as, "Goal! What a great shot!" The generated commentary is then integrated into the edited video to create the final video file.

[1270] The server stores the final video file in cloud storage and provides a download link to the user, who can then view or download the final video with commentary via the provided link.

[1271] Specific examples

[1272] For example, suppose a user films a soccer match and connects the camera device to a cloud server. The user uses an application to register player names and identification numbers, and uploads the video data after the match. The cloud server receives the video data and analyzes it using an AI module. Based on the analysis results, important scenes are automatically extracted and zoomed in for editing. The user selects "Famous Commentator A" as the commentary style, and as a result, the generative AI model generates a commentary such as "That was an excellent shot!"

[1273] Prompt Sentence Examples

[1274] Analyze a soccer match video provided by the user and identify goal scenes. The commentary style selected is "Famous Commentator A Style." Based on this style, generate commentary such as "Goal! What a great shot!"

[1275] The above is an embodiment of the present invention. By using this system, users can easily create and manage high-quality videos with commentary.

[1276] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1277] Step 1:

[1278] A user sets up a wide-angle fixed camera and begins capturing footage of a game or event.

[1279] Input: Camera device setup and recording start command

[1280] Output: Recorded video data (saved in local storage)

[1281] Step 2:

[1282] After the device finishes shooting, it uploads the captured video data to a cloud server.

[1283] Input: Video data stored in local storage

[1284] Data Processing: Transferring video data via Wi-Fi or wired connection

[1285] Output: Video data uploaded to the cloud server

[1286] Step 3:

[1287] The server receives the video data uploaded to the cloud server and stores it in storage.

[1288] Input: Video data uploaded to the cloud server

[1289] Data processing: Writing to cloud storage

[1290] Output: Video data stored in cloud storage

[1291] Step 4:

[1292] The server passes the stored video data to the AI ​​module, which then begins video analysis.

[1293] Input: Video data stored in cloud storage

[1294] Data calculation: AI module analyzes the movement of objects and people (ball and players) for each frame

[1295] Output: Analysis result data (object location information, movement patterns)

[1296] Step 5:

[1297] Based on the analysis results, the server automatically extracts important scenes and performs zoom editing.

[1298] Input: Analysis result data of AI module

[1299] Data processing: Extraction of important scenes and video processing by zoom editing

[1300] Output: Edited clip data

[1301] Step 6:

[1302] The user selects a commentary style via a smartphone app. For example, they select "Famous Commentator A Style."

[1303] Input: User's commentary style selection in the app

[1304] Output: Selected commentary style information (sent to cloud server)

[1305] Step 7:

[1306] The server generates commentary using a generative AI model based on the selected commentary style.

[1307] Input: User selected commentary style information

[1308] Data Computing: Generating Explanatory Text and Audio with Generative AI Models

[1309] Output: Generated commentary data (e.g., "Goal! What a great shot!" audio and text)

[1310] Step 8:

[1311] The server adds the generated commentary and saves the final edited video to cloud storage.

[1312] Input: Edited clip data and generated commentary data

[1313] Data processing: Integrating audio commentary and text into video

[1314] Output: Final edited video data

[1315] Step 9:

[1316] The server then generates a download link for the final video and notifies the user.

[1317] Input: Final edited video data

[1318] Data processing: Download link generation and URL generation

[1319] Output: Notification to user (email or in-system notification)

[1320] This allows users to view or download the final video with commentary via the provided link.

[1321] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1322] This system automatically edits footage of soccer matches and other sporting events into high-quality videos and adds user-selected commentary. Furthermore, the system recognizes the user's emotions and automatically selects and refines the commentary style based on those emotions. The system uploads video data captured with a wide-angle fixed camera to the cloud, where it is analyzed and edited by an AI module, and finally provides the user with a video that integrates the generated commentary.

[1323] System configuration and program processing

[1324] Recording and uploading videos

[1325] The device (wide-angle fixed camera) is equipped with a wide-angle lens that can capture the entire game, and records the game footage from a fixed position. The user sets up the camera device and presses the start button to start recording. After recording is complete, the device uploads the captured video data to a cloud server via Wi-Fi or a wired connection.

[1326] Receiving and storing video data

[1327] The server receives the uploaded video data and stores it in a storage service. In addition, the information on the player names and uniform numbers entered by the user in advance is also stored in a database.

[1328] Video Analysis and Editing

[1329] The server then passes the video data stored in storage to the AI ​​module, which then begins video analysis. The AI ​​module analyzes the ball position and player movements frame by frame to identify goal scenes and important plays. Based on the analysis results, the server extracts specific scenes, automatically zooms in, and generates clips.

[1330] Recognizing user emotions with an emotion engine

[1331] When a user accesses the system through a device, the emotion engine recognizes the user's emotions, obtains emotional data from the user's facial expressions and voice, and adjusts the commentary style accordingly.

[1332] Explanation generation and integration

[1333] The user manually selects the commentary style, or it is automatically selected by the emotion engine. Based on the selected commentary style, the server invokes a generative AI model to generate commentary text and audio data. The generated commentary is integrated into the video clip and saved in cloud storage as the final edited video.

[1334] Video distribution and viewing

[1335] The server then notifies the user of the download link for the final video. The notification is sent via email or an in-system notification. The user can then view or download the generated high-quality video with commentary via the provided link.

[1336] Specific examples

[1337] Editing of goal scenes and adding commentary

[1338] 1. Before the start of the game, the user registers the player's name and uniform number and sets up the device (camera device).

[1339] 2. The device records the match and uploads the video data to a cloud server after the match ends.

[1340] 3. The server stores the received video in a database and analyzes it using an AI module.

[1341] 4. The AI ​​module analyzes the ball's movement and identifies goal opportunities.

[1342] 5. The server zooms in on the goal scene and creates a clip.

[1343] 6. When the user manually or automatically selects a commentary style, the emotion engine recognizes the user's excitement and selects a more energetic commentary style, such as "Goal! What a great shot!"

[1344] 7. The generative AI model generates explanatory text and audio.

[1345] 8. The server integrates the generated commentary to produce the final goal clip.

[1346] 9. The user then watches the final video with commentary via the provided link.

[1347] The introduction of an emotion engine makes it possible to provide customized commentary that corresponds to the user's specific emotional state, further enhancing the viewing experience.The system can also be used for various events other than soccer, such as sports events and athletic meets, and helps families and coaches easily create high-quality footage for analysis and recording memories.

[1348] The processing flow will be explained below.

[1349] Step 1:

[1350] A user logs in to the system and registers a soccer match as a new event. The user enters the player names and uniform numbers through the interface and sends the registration information to the cloud server.

[1351] Step 2:

[1352] The device (wide-angle fixed camera) captures the entire game with a wide angle. Once the setup is complete, the user presses the start button on the camera device to begin recording the game. The device continues to record video data throughout the entire game.

[1353] Step 3:

[1354] After the game ends, the device uploads the captured video data to a cloud server via Wi-Fi or a wired connection. The video data is sent directly to the cloud server.

[1355] Step 4:

[1356] The server receives the video data uploaded to the cloud and stores it in the storage service. At the same time, the player names and uniform numbers registered by the user in advance are also stored in the database.

[1357] Step 5:

[1358] The server then calls up the video analysis AI module, which analyzes the stored video data. The AI ​​module analyzes the ball position and player movements for each frame to identify goal scenes and key plays.

[1359] Step 6:

[1360] The server extracts specific scenes based on the analysis results from the AI ​​module, runs an automatic zoom editing algorithm, zooms in on important scenes and events, and generates clips.

[1361] Step 7:

[1362] After registering a player, the user manually selects a commentary style. In some cases, the emotion engine recognizes the user's emotions and automatically selects an appropriate commentary style. This allows the system to provide commentary that reflects the user's emotions, such as joy or excitement.

[1363] Step 8:

[1364] The emotion engine analyzes the user's emotions from their facial expressions and voice, obtains emotion data in real time, and provides feedback to the selection of commentary style.

[1365] Step 9:

[1366] The server calls a generative AI model to generate commentary text and audio data based on the selected commentary style. For example, if the user is excited, an energetic commentary will be generated.

[1367] Step 10:

[1368] The server integrates the generated commentary text and audio into the video clip, generates a final edited video with commentary, and stores the video in cloud storage.

[1369] Step 11:

[1370] The server will notify the user with a download link for the final edited video. Notifications will be sent via email or in-system notifications. Once the link is provided, the user will have immediate access.

[1371] Step 12:

[1372] The user can then view or download the generated high-quality commentary video via the provided link, which can be saved locally on the user's device.

[1373] Through these processing steps, high-quality videos with annotations that match the user's emotions are automatically generated and provided, allowing the user to enjoy videos with annotations that reflect specific emotions when played back, resulting in a deeper viewing experience.

[1374] Example 2

[1375] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1376] There is a demand for high-quality editing of sporting event videos and the automatic addition of appropriate commentary based on user emotions. However, conventional technologies have difficulty automating video editing and providing emotion-based commentary, requiring a high level of specialized knowledge and manual work. As a result, it is difficult to generate timely and personalized video content.

[1377] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1378] In this invention, the server includes a wide-angle fixed-point camera means for shooting video, a means for uploading the shot video data to an information processing device, and a means for receiving the video data in the information processing device and storing it in a storage device. This enables highly accurate analysis and editing of the video data. The server also includes an artificial intelligence module means for analyzing the saved video data and tracking the positions of the sphere and people, a means for automatically extracting important scenes based on the analysis results and performing zoom editing, a means for generating commentary using a generative artificial intelligence model based on a commentary style selected by the user, a means for adding the generated commentary, storing the final edited video in a storage device, and providing an access link, and an emotion recognition engine means for recognizing the user's emotions and automatically selecting a commentary style based on the emotions. This enables automatic and efficient generation of high-quality video with commentary that corresponds to the user's emotions.

[1379] A "wide-angle fixed-point camera" is a device that captures images from a fixed position with a wide field of view, and is primarily used to record the overall picture of sporting events and the like.

[1380] An "information processing device" is a device that processes, analyzes, stores, and transmits data, and includes a cloud server, a local server, and the like.

[1381] A "storage device" is a hardware or software means for storing digital data, including cloud storage and hard disk drives.

[1382] An "artificial intelligence module" is a software module that uses machine learning and deep learning to analyze data and recognize specific patterns and features.

[1383] "Tracking the position of spheres and people" refers to the process of detecting specific objects, such as balls or players, within video frames and continuously tracking their position information.

[1384] "Automatic extraction of important scenes" is the process of identifying specific events or actions within a video based on the analysis results of an artificial intelligence module and cutting out only the relevant parts.

[1385] "Zoom editing" is an editing technique that enlarges and visually enhances portions of the original footage, and is used to provide viewers with important scene details.

[1386] A "generative AI model" is a model that uses AI technology to generate text and speech based on user selections and emotional data.

[1387] An "emotion recognition engine" is a software module that analyzes and recognizes a user's emotional state from their facial expressions and voice, acquiring emotional data and reflecting it in the analysis results.

[1388] MODE FOR CARRYING OUT THE INVENTION

[1389] The present invention is a system that automatically edits video data from sporting events into high-quality videos and adds user-selected commentary. It also provides the ability to recognize user emotions and automatically select and refine commentary styles based on those emotions. The system begins by capturing game footage using a fixed-point camera with a wide field of view and uploading the raw data to a cloud server. Specific embodiments of the present invention are described below.

[1390] System Configuration

[1391] 1. Wide-angle fixed-point shooting device

[1392] The wide-angle fixed-point camera is a dedicated camera that records the entire sporting event over a wide area. Users set up the camera and start recording at the start of the game. The video data captured by the camera is uploaded to a cloud server via Wi-Fi or a wired connection.

[1393] 2. Cloud Server

[1394] The cloud server receives the video data and stores it securely in a storage device. The cloud server also stores information such as player names and uniform numbers that users have registered in advance. The video data is then analyzed by an AI module.

[1395] 3. Artificial Intelligence Module

[1396] An AI module in the cloud server analyzes the video data frame by frame, tracking the positions of balls and people. This analysis identifies goals and other key moments. The AI ​​module extracts key scenes and performs zoom editing to generate clips.

[1397] 4. Emotion Recognition Engine

[1398] When a user accesses the system, the emotion recognition engine acquires emotional data from the user's facial expressions and voice, which is used as an important factor in determining the commentary style.

[1399] 5. Generative AI Models

[1400] The generative AI model generates commentary text and audio data based on the user's emotion recognition results and selected commentary style, which is then integrated with the video clip to create the final edited video.

[1401] 6. Video streaming and viewing

[1402] The cloud server then sends the user a download link for the final video, allowing the user to view or download the high-quality video with commentary.

[1403] Specific examples

[1404] 1. Editing goal scenes and adding commentary

[1405] Before the game, the user registers the player names and uniform numbers and sets up the wide-angle fixed-point camera.

[1406] The device records the match and uploads the video data to a cloud server after the match ends.

[1407] The video data received by the server is stored in a storage device and analyzed by an AI module.

[1408] The AI ​​module analyzes the movement of the ball and identifies goal opportunities.

[1409] The server zooms in on the goal scene and creates a clip.

[1410] The user can select a commentary style, or an emotion recognition engine can automatically select one based on the emotion. In this case, an energetic commentary such as "Goal! What a great shot!" is generated.

[1411] A generative AI model generates explanatory text and audio.

[1412] The server then integrates the generated commentary to produce the final goal clip.

[1413] The user then watches the final video with commentary via the provided link.

[1414] Prompt Sentence Examples

[1415] "The entire match should be filmed with a wide-angle camera, and specific scenes should be automatically edited and commentary should be added. The commentary should be selected and generated based on the user's emotions."

[1416] The above is a specific embodiment for carrying out the present invention. This system can be applied to various sporting events other than soccer, and efficiently generates high-quality video.

[1417] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1418] Step 1: Record your video and upload it to the cloud

[1419] The terminal (wide-angle fixed-point camera) is equipped with a wide-angle lens that can capture the entire game. The user sets up the camera and presses the start button to begin recording. The input is live footage of the game, and the output is the captured video data. This video data is uploaded to a cloud server via Wi-Fi or a wired connection after the game ends. Specific operations include the user adjusting the camera position and instructing the game to start.

[1420] Step 2: Receiving and saving video data

[1421] The server receives video data from the cloud and stores it in a storage device. The input is the video data uploaded from the device, and the output is the data stored in a storage device on the cloud. Specific operations include the server confirming receipt of the data and storing it in the appropriate folder. At the same time, the player name and uniform number information previously entered by the user is also stored in the database.

[1422] Step 3: Video Analysis and Editing

[1423] The server passes the video data stored in the storage device to an AI module and begins data analysis. The input is the stored video data, and the output is the analysis results identifying key scenes. The AI ​​module tracks the ball position and player movements for each frame to identify goals and important plays. Specific operations include extracting the identified scenes and performing zoom editing to generate clips.

[1424] Step 4: Recognizing user emotions with the emotion engine

[1425] When a user accesses the system, the emotion recognition engine analyzes the user's emotions. The input is the user's facial expression and voice data, and the output is recognized emotion data. Specific operations include the process in which the user provides emotion data using a webcam or microphone, and the emotion engine analyzes that data.

[1426] Step 5: Explanation generation and integration

[1427] The user manually selects the commentary style, or the emotion recognition engine automatically selects it. The input is the recognized emotion data and the selected commentary style, and the output is the generated commentary text and audio data. The server uses a generative AI model to generate commentary and integrate it into the video clip. Specific operations include the server calling the generative AI model and automatically generating and integrating commentary.

[1428] Step 6: Stream and view your video

[1429] The server notifies the user of the download link for the final video. The input is the final video data, and the output is a link that the user can access. Specific operations include the server sending a notification email or an in-system notification and providing the user with a viewing link. The user can then view or download the high-quality video with commentary via the link.

[1430] (Application example 2)

[1431] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1432] Conventional video editing systems for sporting events make it difficult for users to add their own commentary and lack the ability to dynamically change the commentary style based on the user's emotions. As a result, viewers' viewing experience is consistent and they are unable to respond to the needs of individual users. Furthermore, there is a need for similar high-quality editing functions for a variety of sporting events other than soccer.

[1433] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1434] In this invention, the server includes a wide-angle fixed camera means for shooting videos, a means for uploading the shot video data to a cloud server, a means for receiving the video data in the cloud server and storing it in storage, an AI module means for analyzing the stored video data and tracking the positions of the ball and players, a means for automatically extracting important scenes based on the analysis results and performing zoom editing, a means for generating commentary using a generative AI model based on a commentary style selected by the user, an emotion engine means for recognizing the user's emotions and adjusting the commentary style based on the emotions, and a means for adding the generated commentary, storing the final edited video in cloud storage, and providing an access link. This makes it possible to provide high-quality videos with commentary that are adapted to the user's emotions.

[1435] The "wide-angle fixed-point camera means" is a device equipped with a wide-angle lens that captures moving images from a fixed position.

[1436] A "cloud server" is a computer system that stores and processes data and is accessible remotely over the Internet.

[1437] "Storage" refers to a storage device or service for saving data.

[1438] "AI module means" is software that uses artificial intelligence to analyze video and track the movements of the ball and players.

[1439] "Zoom editing" is an editing technique that visually emphasizes a particular scene by enlarging it.

[1440] A "generative AI model" is an artificial intelligence model that automatically generates natural language based on user input and data.

[1441] The "emotion engine means" is software for analyzing the user's facial expressions and voice and recognizing their emotions.

[1442] An "access link" is a URL or hyperlink that allows a user to access specific content via the Internet.

[1443] "Means for registering player names and uniform numbers" is a function for entering player names and uniform numbers into the system for video analysis.

[1444] The "emotion-based commentary style automatic adjustment function" is a function that dynamically changes the style and tone of commentary using the user's emotional data.

[1445] In this invention, a user shoots a sporting event with a wide-angle fixed camera and uploads the video data to a cloud server. The cloud server receives the video data and saves it in storage. The saved video data is analyzed by an AI module, and the positions of the ball and players are tracked. Based on the analysis results, important scenes are automatically extracted and zoomed in for editing.

[1446] If the user selects a commentary style, a commentary is generated using the generative AI model. If the user's emotion is recognized, an emotion engine means analyzes the user's facial expressions and voice and adjusts the commentary style. The generated commentary is added to the video, and the final edited video is saved in cloud storage, and an access link is provided to the user.

[1447] A specific system configuration operates as follows.

[1448] 1. Hardware and Software Description

[1449] Wide-angle fixed camera means: A device equipped with a wide-angle lens that shoots video from a fixed position.

[1450] Cloud Server: A computer system that stores and processes data and is accessible remotely over the Internet.

[1451] Storage: A storage device or service for storing data (e.g., Amazon S3, Google Cloud Storage).

[1452] AI module means: Software (e.g. OpenCV) that uses artificial intelligence to analyze video and track ball and player movements.

[1453] Generative AI models: Artificial intelligence models that generate commentary based on user input and data (e.g., GPT-3).

[1454] Emotion engine means: Software that analyzes the user's facial expressions and voice and recognizes their emotions (e.g., OpenFace).

[1455] 2. Description of Data Processing and Data Calculations

[1456] Uploading and saving video data: The video data taken by the user is uploaded to the cloud server and saved in storage. The saved data is passed to the AI ​​module.

[1457] Video analysis: The AI ​​module analyzes the video data and tracks the position of the ball and players frame by frame. Key scenes are extracted and zoomed in for editing.

[1458] Emotion Recognition: Emotion data is acquired from the user's facial expressions and voice. The emotion engine analyzes the user's emotions and adjusts the commentary style accordingly.

[1459] Commentary generation and integration: A generative AI model generates commentary text and audio based on the emotion data and integrates it into the video. The final edited video is saved in cloud storage and an access link is provided to the user.

[1460] 3. Examples of concrete examples and prompts

[1461] For example, a user can film a soccer game using their smartphone and upload the video to the cloud. The cloud server then analyzes the video data and adds energetic commentary based on the user's emotional data. Finally, a high-quality video with commentary is generated, and the user receives a link to watch it.

[1462] Example prompts for generative AI models

[1463] If the user is excited while watching:

[1464] "It was an exciting goal scene! The moment the shot shook the net, the whole stadium erupted in cheers!"

[1465] If the user is relaxing while watching:

[1466] "Here, a relaxed ball movement resulted in a fantastic goal. This play is a great example of the teamwork."

[1467] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1468] Step 1:

[1469] The user takes a video

[1470] A user shoots video of a sporting event using a wide-angle fixed camera, and video data is generated. The inputs are the user's operations and the sporting event being shot, and the output is a video file of the shot video.

[1471] Step 2:

[1472] Upload videos to a cloud server

[1473] After the user has finished shooting, they upload the video data to the cloud server. The input for the upload is a video file, and the output is the video data saved on the cloud server. The user is also notified of the upload status.

[1474] Step 3:

[1475] The cloud server receives and stores the data.

[1476] The cloud server receives the uploaded video data and saves it to storage. The input is the uploaded video data, and the output is the storage of the data in the destination storage. A notification of successful saving is also generated.

[1477] Step 4:

[1478] Video analysis and key scene extraction

[1479] The cloud server uses an AI module to analyze the stored video data and track the positions of the ball and players. The input is the stored video data, and the output is the position data of the ball and players, as well as identification information for important scenes. Specifically, the system analyzes each frame and records the position data.

[1480] Step 5:

[1481] Performing Zoom Edits

[1482] The cloud server automatically extracts important scenes and performs zoom editing. The input is the identification information and video data of the important scenes, and the output is a zoom-edited clip of the scene. Specifically, editing software is used to enlarge and display the important scenes.

[1483] Step 6:

[1484] Recognize user emotions with an emotion engine

[1485] When a user watches a video, the emotion engine means analyzes the user's facial expressions and voice to recognize emotions. The input is the user's facial expression data and voice data, and the output is recognized emotion data. Specifically, data is collected using a camera and microphone, and an emotion analysis algorithm is applied.

[1486] Step 7:

[1487] Generating and synthesizing explanations

[1488] The user selects a commentary style, or the generative AI model generates commentary based on recognized emotion data. The input is emotion data and clips of key scenes, and the output is a video clip with commentary. Specifically, a prompt is generated for the generative AI model, and commentary text and audio data are generated.

[1489] Step 8:

[1490] Save the final edited video and provide an access link

[1491] The cloud server integrates the generated commentary into the video and saves the final edited video in cloud storage. The input is a video clip with commentary, and the output is saving the final edited video and generating an access link. Specifically, the video editing software integrates the commentary, uploads it to the cloud, and generates an access link that is notified to the user.

[1492] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1493] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1494] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1495] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1496] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1497] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1498] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1499] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1500] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1501] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1502] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1503] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1504] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1505] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1506] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1507] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1508] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1509] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1510] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1511] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1512] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1513] The following is further disclosed regarding the above embodiment.

[1514] (Claim 1)

[1515] a wide-angle fixed camera means for capturing video;

[1516] A means for uploading the captured video data to a cloud server;

[1517] A means for receiving and storing video data in a cloud server;

[1518] an AI module means for analyzing the stored video data and tracking the positions of the ball and players;

[1519] A means for automatically extracting important scenes based on the analysis results and performing zoom editing;

[1520] A means for generating commentary using a generative AI model based on a commentary style selected by a user;

[1521] A means to add the generated commentary and save the final edited video to cloud storage and provide an access link;

[1522] A system including:

[1523] (Claim 2)

[1524] 2. The system according to claim 1, further comprising means for registering player names and uniform numbers.

[1525] (Claim 3)

[1526] 10. The system of claim 1, further comprising means for supporting video editing of sporting events and events other than soccer.

[1527] "Example 1"

[1528] (Claim 1)

[1529] a wide-angle fixed camera means for capturing video;

[1530] A means for uploading the captured video data to a cloud server;

[1531] A means for receiving and storing video data in a cloud server;

[1532] an artificial intelligence module means for analyzing the stored video data and tracking the positions of the ball and players;

[1533] A means for automatically extracting important scenes based on the analysis results and performing zoom editing;

[1534] A means for generating commentary using a generative artificial intelligence model based on a commentary style selected by a user;

[1535] A means for a user to select an explanation style and transmit the information to a cloud server;

[1536] a means for integrating the generated commentary into the video clip;

[1537] A means for notifying a user of a download link for the generated video with commentary;

[1538] A means to save the final video with commentary in cloud storage,

[1539] A system including:

[1540] (Claim 2)

[1541] 2. The system according to claim 1, further comprising means for registering player names and uniform numbers.

[1542] (Claim 3)

[1543] 10. The system of claim 1, further comprising means for supporting video editing of sporting events and entertainment.

[1544] "Application Example 1"

[1545] (Claim 1)

[1546] a wide-angle fixed camera means for capturing video;

[1547] A means for uploading the captured video data to a cloud server;

[1548] A means for receiving and storing video data in a cloud server;

[1549] An AI module means for analyzing the stored video data and tracking the positions of objects and people;

[1550] A means for automatically extracting important scenes based on the analysis results and performing zoom editing;

[1551] A means for generating commentary using a generative AI model based on a commentary style selected by a user;

[1552] A means to add the generated commentary and save the final edited video to cloud storage and provide an access link;

[1553] A system including a means for managing video data through an application installed on a smartphone.

[1554] (Claim 2)

[1555] 10. The system of claim 1, further comprising means for registering player names and identification numbers.

[1556] (Claim 3)

[1557] 10. The system of claim 1, further comprising means for supporting video editing of the event.

[1558] "Example 2: Combining Emotion Engines"

[1559] (Claim 1)

[1560] a wide-angle fixed-point photographing device means for photographing a moving image;

[1561] means for uploading the captured video data to an information processing device;

[1562] a means for receiving video data in an information processing device and storing the video data in a storage device;

[1563] an artificial intelligence module means for analyzing the stored video data and tracking the positions of the sphere and the person;

[1564] A means for automatically extracting important scenes based on the analysis results and performing zoom editing;

[1565] A means for generating commentary using a generative artificial intelligence model based on a commentary style selected by a user;

[1566] means for adding the generated commentary and saving the final edited video to a storage device and providing an access link;

[1567] an emotion recognition engine means for recognizing an emotion of a user and automatically selecting an explanation style based on the emotion;

[1568] A system including:

[1569] (Claim 2)

[1570] 2. The system according to claim 1, further comprising means for registering player names and uniform numbers.

[1571] (Claim 3)

[1572] 10. The system of claim 1, further comprising means for supporting video editing of various sporting events and other occasions.

[1573] "Application example 2 when combining emotion engines"

[1574] (Claim 1)

[1575] a wide-angle fixed camera means for capturing video;

[1576] A means for uploading the captured video data to a cloud server;

[1577] A means for receiving and storing video data in a cloud server;

[1578] an AI module means for analyzing the stored video data and tracking the positions of the ball and players;

[1579] A means for automatically extracting important scenes based on the analysis results and performing zoom editing;

[1580] A means for generating commentary using a generative AI model based on a commentary style selected by a user;

[1581] an emotion engine means for recognizing a user's emotion and adjusting a commentary style based thereon;

[1582] A means to add the generated commentary and save the final edited video to cloud storage and provide an access link;

[1583] A system including:

[1584] (Claim 2)

[1585] 2. The system according to claim 1, further comprising means for registering player names and uniform numbers, and means for recognizing user emotions.

[1586] (Claim 3)

[1587] The system of claim 1, which also supports video editing of sporting events and events other than soccer, and further includes a function for automatically adjusting commentary style based on user emotions. [Explanation of symbols]

[1588] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. a wide-angle fixed camera means for capturing video; A means for uploading the captured video data to a cloud server; A means for receiving and storing video data in a cloud server; an AI module means for analyzing the stored video data and tracking the positions of the ball and players; A means for automatically extracting important scenes based on the analysis results and performing zoom editing; A means for generating commentary using a generative AI model based on a commentary style selected by a user; A means to add the generated commentary and save the final edited video to cloud storage and provide an access link; A system including:

2. 2. The system according to claim 1, further comprising means for registering player names and uniform numbers.

3. 10. The system of claim 1, further comprising means for supporting video editing of sporting events and events other than soccer.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A