System

The system uses a generative AI model to create highlight videos from user-selected players, addressing the issue of missed scenes in live broadcasts, enhancing viewer engagement.

JP2026018089APending Publication Date: 2026-02-05SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024119150
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-24
Publication Date
2026-02-05

AI Technical Summary

Technical Problem

Viewers starting to watch games or live broadcasts midway often miss important scenes and performances of specific players or team members, leading to a diminished viewing experience.

Method used

A system that generates highlight videos based on user customization settings, using a generative AI model to analyze video data, filter scenes with specific players or members, and edit them into a highlight video for real-time viewing.

Benefits of technology

Enables users to efficiently watch key moments of their favorite players or members without missing important scenes, improving the viewing experience and increasing content engagement.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026018089000001_ABST
    Figure 2026018089000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system for generating a highlight video from a game or a live video based on a customization setting of a user, the system comprising: means for acquiring data of the game or the live video; means for receiving information of a specific player or member set by the user; means for analyzing the acquired video data and detecting an important scene or a characteristic event; means for filtering a scene including the specific player or member from the detected scene based on the setting of the user; means for generating the highlight video by editing the filtered scene; and means for providing the generated highlight video to the user.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] When viewers start watching a game or live broadcast midway through, they lose track of what has happened up to that point, which reduces the enjoyment of the viewing experience. Also, if you are rooting for a particular player or team member, it is difficult to fully see their performance in regular broadcasts or highlights. Technological solutions are needed to improve the viewing experience and meet individual needs. [Means for solving the problem]

[0005] The present invention is a system that generates a highlight video from a match or live video based on a user's customization settings, and solves these problems by including: means for acquiring data on the match or live video; means for accepting information on specific players or members set by the user; means for analyzing the acquired video data to detect important scenes or characteristic events; means for filtering scenes that include specific players or members from the detected scenes based on the user's settings; means for editing the filtered scenes to generate a highlight video; and means for providing the generated highlight video to the user.

[0006] "Games and live footage" refers to real-time or recorded video data such as live sports broadcasts and live music performances.

[0007] "User" refers to an individual or organization that uses the system to watch matches and live footage and set specific players or members.

[0008] A "highlight video" is a short video that is edited and contains excerpts of particularly important scenes or notable events from a match or live footage.

[0009] "Means of obtaining data" refers to the function of collecting data on matches and live footage from servers and streaming services.

[0010] "Means for accepting information on specific players or members" refers to the function that accepts settings for the players or members that the user supports.

[0011] "Means for analyzing video data" refers to the function of analyzing acquired video data to detect important scenes and characteristic events.

[0012] "Means for filtering scenes" refers to the function of selecting scenes that include players or members set by the user from the analysis results.

[0013] "Means for generating a highlight video" refers to a function for editing filtered scenes to create a highlight video for the user.

[0014] "Means for providing to users" refers to the function of distributing or notifying users of the generated highlight video in a format that allows them to view it. [Brief explanation of the drawings]

[0015] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13]FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0016] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0017] First, the terms used in the following description will be explained.

[0018] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0019] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0020] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0021] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0022] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0023] [First embodiment]

[0024] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0025] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0026] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0027] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0028] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0029] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0030] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0031] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0032] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0033] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0034] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0035] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0036] The present invention relates to a system that automatically generates highlight videos of matches and live footage based on user-specified players and team members. The system includes a user device, a server, and video analysis technology using a generative AI model.

[0037] Program processing and specific operations

[0038] Accepting user settings

[0039] 1. The user launches the application and logs in. On the settings screen, the user selects the specific player or member they want to support. For example, they can select soccer player Lionel Messi or music group member Alice.

[0040] 2. The device sends the selected information to the server, which stores it in a database as the user's customized settings.

[0041] Start watching content

[0042] 1. The user selects the sports broadcast or live video they want to watch. The device sends this selection information and setting information to the server.

[0043] 2. The server works with the streaming service to set up the start of viewing and prepares to obtain video data in real time.

[0044] AI model analysis

[0045] 1. The server inputs the acquired video data into a generative AI model and automatically analyzes important scenes and distinctive events. For example, it identifies goal scenes and interviews in a soccer game, and performance scenes and MC talk in a music concert.

[0046] 2. The analysis results are temporarily stored in a database.

[0047] Extracting your favorite scenes

[0048] 1. The server filters out relevant scenes from the analysis results based on the information of specific players or members set by the user. For example, it can select scenes where Lionel Messi shoots or Alice sings.

[0049] 2. Organize these selected scenes in chronological order and review them as necessary.

[0050] Highlight video generation

[0051] 1. The server edits the filtered scenes to generate a smooth highlight video, for example by adding transition effects, background music, and editing clips including text overlays.

[0052] 2. Render the edited video, convert it into the optimal viewing format, and save it in the database.

[0053] Highlights viewing available

[0054] 1. The server sends a link to the generated highlight video to the user's device. The notification includes a URL for starting viewing.

[0055] 2. The user receives a notification and plays the highlight video. The device streams or downloads the video from the server and plays it.

[0056] For example, if a user starts watching a soccer match halfway through, a highlight video will be automatically generated, including key moments (e.g., goals and assists) of the player selected by the user, such as Lionel Messi. Similarly, if a user starts watching a live stream halfway through, a highlight video focusing on the performance of their favorite member will be provided.

[0057] This allows users to enjoy watching their favorite players and members' performances without missing important scenes, even when they join a match or live show. This system also improves the viewing experience, and is expected to increase content viewing and service users.

[0058] The processing flow will be explained below.

[0059] Step 1:

[0060] The user launches the application and logs in. On the settings screen, the user selects the specific player or member they want to support.

[0061] Step 2:

[0062] The device temporarily stores the information of the specific players or members selected by the user and displays a confirmation dialog. When the user clicks Confirm, the setting information is sent to the server.

[0063] Step 3:

[0064] The server stores the received configuration information in a user database, thereby establishing the user's customized settings.

[0065] Step 4:

[0066] The user selects the sports broadcast or live video they want to watch on the application screen, and this selection information is also sent from the device to the server.

[0067] Step 5:

[0068] The server receives the instruction to start viewing and prepares to obtain video data in real time in cooperation with the streaming service.

[0069] Step 6:

[0070] The server inputs video data obtained from the streaming service into a generative AI model, which analyzes the video data and detects important scenes and distinctive events.

[0071] Step 7:

[0072] The server stores the analysis results of the AI ​​model in temporary storage, which preserves important scene and event data.

[0073] Step 8:

[0074] The server filters out relevant scenes from the analysis results based on specific player and team member information set by the user, such as Lionel Messi's shooting scenes or Alice's singing scenes.

[0075] Step 9:

[0076] The server then organizes the filtered video clips into chronological order, rechecking and fine-tuning them as needed.

[0077] Step 10:

[0078] The server then edits the organized scene clips, adds transitions, background music, and text overlays to create a smooth highlight video.

[0079] Step 11:

[0080] The server converts the rendered highlight video into the appropriate viewing format, ensuring the video format is playable on the user's device.

[0081] Step 12:

[0082] The server stores the generated highlight video in a database and prepares to notify the user's device of the link to the video.

[0083] Step 13:

[0084] The device receives the link to the highlight video sent from the server and displays a notification to the user. When the user clicks on the notification, the highlight video starts playing.

[0085] Step 14:

[0086] Users can play highlight videos and enjoy key moments of specific players or members.

[0087] These steps allow users to enjoy customized highlight videos of specific players or members, even from the middle of the video.

[0088] Example 1

[0089] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0090] Conventional highlight video generation systems make it difficult for users to efficiently watch the highlight scenes of specific players or team members. Furthermore, content customization based on user settings is insufficient, leading to users often missing important or interesting scenes. Furthermore, real-time highlight video generation and viewing was not possible, limiting the viewing experience.

[0091] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0092] In this invention, the server includes means for acquiring game and live video data, means for accepting information on specific players and members set by the user, means for using a generative AI model to analyze the acquired video data and detect important scenes and characteristic events, means for filtering scenes containing specific players and members from the detected scenes based on the user's settings, means for using an editing tool to edit the filtered scenes to generate a highlight video, and means for notifying and providing the generated highlight video to the user's terminal. This enables users to efficiently watch the performances of specific players and members, and the generation and viewing of highlight videos is realized in real time, significantly improving the viewing experience.

[0093] "Means for acquiring data on games and live video" refers to a combination of hardware and software for capturing game and live video in real time and transmitting the data to a server.

[0094] The "means for accepting information on specific players or members set by the user" refers to an interface and protocol for allowing the user to input information on the players or members selected by the user and transmitting that information to the server.

[0095] "Means of using a generative AI model to analyze acquired video data and detect significant scenes and distinctive events" refers to a function that uses AI technology to analyze video data and identify specific events, such as goal scenes and performance scenes.

[0096] "Means for filtering scenes that include specific players or members from the detected scenes based on user settings" refers to software processing for selecting and filtering scenes related to players or members from the analyzed data.

[0097] The "means of using an editing tool to edit the filtered scenes to generate a highlight video" refers to a video editing function for combining the selected scenes and adding transitions, background music, etc. to generate a highlight video.

[0098] The "means of notifying and providing the generated highlight video to the user's device" is a function of sending a viewing link for the generated highlight video to the user's device by means of a push notification or the like.

[0099] This invention relates to a system that allows users to set specific players or members and automatically generates highlight videos of matches and live footage based on that information. The system includes a user's device, a server, and video analysis technology using a generative AI model.

[0100] Accepting user settings

[0101] First, the user launches the application and logs in. The login process involves the user entering their email address and password and pressing the "Login" button. The user then selects the specific player or member they want to support on the settings screen. For example, if a user supports Lionel Messi, they can enter his name in the search bar and select him from the list.

[0102] The device sends this selected information to a server, which stores the information as the user's customized settings in a database, for example, using a relational database management system such as MySQL.

[0103] Start watching content

[0104] Next, the user selects the sports or live video they want to watch, for example, the "2023 Champions League Final." The device then sends this selection and configuration information to the server.

[0105] The server interacts with streaming service APIs (e.g., YouTube API, Twitch API) to set up the start of viewing. The server also obtains video data in real time and creates a viewing session.

[0106] AI model analysis

[0107] The server inputs the acquired video data into a generative AI model (e.g., OpenAI's GPT-4). The video is first broken down into frames, and the generative AI model analyzes the content of each frame and tags important scenes and distinctive events. For example, goals and important plays in a soccer game, or the start and end of a performance in a live event.

[0108] The analysis results are temporarily stored in a database and include timestamp and tag information.

[0109] Extracting your favorite scenes

[0110] The server filters the database based on specific players and team members set by the user. For example, it extracts scenes of Lionel Messi's shots. Filtering is done using SQL queries.

[0111] The filtered scenes are then sorted chronologically and subjected to a video review process, where automated checking algorithms are used to verify video continuity and content.

[0112] Highlight video generation

[0113] The server then edits the filtered scenes using an editing tool (e.g., FFmpeg), adding transition effects (fade in / out), adding background music, overlaying text, etc. The edited video is then rendered and converted to the optimal format (e.g., MP4, 1080p).

[0114] The generated video is stored in the server's storage system.

[0115] Highlights viewing available

[0116] Finally, the server sends a link to the generated highlight video to the user's device via a notification service (e.g., Firebase Cloud Messaging). The notification includes a URL for starting the video.

[0117] Users can view the video by clicking the URL provided and checking the notification on their device. The device will then stream or download the video from the server and play it. Buffering is performed during this process to prevent interruptions to viewing.

[0118] Examples and prompts

[0119] For example, when a user watches a soccer match, a highlight video including key scenes of the player selected, Lionel Messi, will be automatically generated. Similarly, when a user starts watching live footage, a highlight video focusing on the performance of their favorite member will be provided.

[0120] An example of a prompt might be "Generate highlights of Lionel Messi from the following soccer game."

[0121] This will allow users to efficiently watch the performances of specific players or team members without missing any important scenes. The viewing experience will be significantly improved, which is expected to lead to an increase in content viewing and an expansion of the service user base.

[0122] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0123] Step 1: Accepting User Settings

[0124] The user launches the application and logs in. The user selects the specific player or member they want to support on the settings screen. The input is an email address and password, and the output is a message indicating whether the login was successful or failed. The user then selects a specific player or member, and the device sends this information to the server. The input is a list of players or members, and the output is information about the selected player or member.

[0125] Step 2: Save your user settings information

[0126] The server receives the user's selection information sent from the device and stores it in a database. The input is the player and member information selected by the user, and the output is the customized settings stored in the database. Specifically, the server analyzes the received data and issues SQL queries to a relational database such as MySQL.

[0127] Step 3: Select content to watch

[0128] The user selects the sports broadcast or live video they want to watch within the application. The input is information about the content they want to watch, and the output is the identification information of the selected content. The device sends this information to the server and associates it with the settings information. The server calls the streaming service API to set up viewing. The input is the user's settings information and content identification information, and the output is the streaming URL and session ID.

[0129] Step 4: Acquiring real-time video data

[0130] The server prepares to acquire video data in real time from the streaming service. The input is the streaming URL and session ID, and the output is the acquired video data. Specifically, the server starts the real-time stream using the streaming service API.

[0131] Step 5: Analysis by AI model

[0132] The server inputs the acquired video data into a generative AI model. The input is real-time video data, and the output is analyzed scene information (tagged data). The generative AI model (e.g., GPT-4) analyzes the video frame by frame and tags important scenes and events. Specifically, the video data is divided into frames and each frame is input into the model.

[0133] Step 6: Save the analysis results

[0134] The server temporarily stores the analyzed scene information in a database. The input is tagged scene information, and the output is the analysis results stored in the database. The server issues SQL queries and stores data including timestamps and tag information.

[0135] Step 7: Filtering your favorite scenes

[0136] The server filters the saved scene information based on specific players and members set by the user. The input is the user's setting information and analysis results, and the output is the filtered scene information. Specifically, the server extracts relevant scenes from the database using SQL queries.

[0137] Step 8: Edit your highlight video

[0138] The server edits the filtered scenes using an editing tool (e.g., FFmpeg). The input is the extracted scene information, and the output is an edited highlight video. Specific operations include adding transition effects, background music, and text overlays, and rendering the video.

[0139] Step 9: Save your highlight video

[0140] The server renders the edited highlight video, converts it to the optimal format, and saves it. The input is the edited video file, and the output is the saved video file. Specifically, the server sets the video encoding settings and saves it in the server's storage system.

[0141] Step 10: Notification of highlight videos and provision of viewing

[0142] The server sends a link to the generated highlight video to the user's device via a notification service (e.g., Firebase Cloud Messaging). The input is the URL of the saved video file, and the output is a notification message. The user receives the notification and clicks the provided URL to play the video. The device streams or downloads the video from the server and plays it. The input is the video URL from the server, and the output is the video that is played.

[0143] This allows users to efficiently enjoy the performances of specific players and team members without missing anything. It also enables highlight videos to be generated and viewed in real time, improving the viewing experience.

[0144] (Application example 1)

[0145] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0146] In modern sports and music bars, customers often want to watch highlight videos of specific athletes or artists in real time, but there is a lack of an efficient system to make this possible. There is also a lack of a way to easily project the highlight videos that customers individually want. This has led to a demand for improved customer satisfaction and increased customer attraction.

[0147] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0148] In this invention, the server includes means for acquiring game and live video data, means for accepting information on specific players or members set by the user, means for analyzing the acquired video data to detect important scenes and characteristic events, means for filtering scenes including specific players or members from the detected scenes based on the user's settings, means for editing the filtered scenes to generate a highlight video, means for providing the generated highlight video to the user, and connection means for projecting the video from the user's terminal onto a display device in the store, thereby enabling users to easily view highlight videos of specific players or artists in real time in the store.

[0149] "Game and live footage" refers to video data recorded from live performances such as sports games and music concerts.

[0150] "User Customization Settings" refers to the ability for users to select specific players or artists and individually customize content based on those settings.

[0151] A "highlight video" is a shortened version of a video that is edited to include particularly important and distinctive scenes from a game or live footage.

[0152] "Specific player or member" refers to a specific individual or member of a group that the user supports.

[0153] "Server" refers to a computer system for storing and processing data.

[0154] "User's device" refers to the electronic device used by the User, such as a smartphone, tablet, or PC.

[0155] "In-store display devices" refers to displays and projectors installed in physical stores such as sports bars and music bars.

[0156] "Connection means" refers to the technology for connecting the user's terminal with the display device in the store and sending and receiving data.

[0157] "Means for analyzing video data" refers to the function of analyzing acquired video data using technologies such as AI models to detect important scenes and characteristic events.

[0158] "Filtering means" refers to the function of selecting only scenes that include specific players or members from video data.

[0159] "Editing means" refers to the function of combining filtered scenes, adding transition effects and background music, and completing the highlight video.

[0160] "Means of providing" refers to the technology used to deliver the generated highlight video to users.

[0161] This invention relates to a system that automatically generates highlight videos of matches and live footage based on user-specified players and team members. The system includes a user terminal, a server, and video analysis technology using a generative AI model.

[0162] Accepting user settings

[0163] Users launch the smartphone application and log in. They select the specific player or artist they want to support on the settings screen and enter their information. This selection information is sent from the device to the server, which then stores it in a database as the user's customized settings.

[0164] Start watching content

[0165] The user selects the game or live video they want to watch. The device sends this selection information to the server, which then works with the streaming service to prepare to obtain the video data in real time, taking the user's settings into consideration.

[0166] AI model analysis

[0167] The server inputs the acquired video data into a generative AI model and automatically analyzes important scenes and distinctive events. For example, it identifies goal scenes in a game, or performance scenes and MC talk in a live performance. The results of this analysis are temporarily stored in a database.

[0168] Extracting your favorite scenes

[0169] The server filters out relevant scenes from the analysis results based on the information of specific players or members set by the user. This allows for the extraction of only scenes featuring a specific player's shot or an artist's performance. The extracted scenes are organized in chronological order and can be reviewed.

[0170] Highlight video generation

[0171] The server then edits the filtered scenes to create a smooth highlight video, adding transition effects, background music, and clip editing including text overlays. The edited video is then rendered, converted into the optimal viewing format, and stored in a database.

[0172] Highlights viewing available

[0173] The server notifies the user's device of a link to the generated highlight video. The user receives the notification and clicks the link to play the highlight video. The device streams or downloads the video from the server and plays it. A connection means is also provided for the user's device to project the video onto a display device in the store. This allows users to watch specific videos on a large screen in physical stores such as sports bars and music bars.

[0174] Specific examples

[0175] Consider the case where a customer at a sports bar requests a highlight video of a specific soccer player. When the customer requests a Lionel Messi goal scene through a smartphone app, the server retrieves the video data, inputs it into a generative AI model for analysis, and then extracts only Messi's goal scenes from the analysis results, edits them, and generates a highlight video. A link to this video is sent to the customer's device, and the video is projected onto a large screen display in the bar. The customer can then use the app to enjoy Messi's goal scene on the big screen.

[0176] Prompt Sentence Examples

[0177] "Generate a highlight video that includes a specific player's goal."

[0178] This allows users to efficiently watch only the scenes they particularly want to support without missing any of their favorite athletes or artists.

[0179] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0180] Step 1:

[0181] The user launches the smartphone app and logs in. The user selects a specific player or artist on the settings screen. This selection information is sent from the device to the server.

[0182] Input: Selected information about players and artists that users enter into the smartphone app

[0183] Data processing: Send the selected information to the server as an HTTP request

[0184] Output: Player and artist selection information is saved in the server database.

[0185] Specific operation: The user opens the "Settings" screen of the smartphone app and selects the player or artist they want to support. The selected information is sent to the server by pressing the send button.

[0186] Step 2:

[0187] Users select the game or live video they want to watch and send the selection information from their device to the server, which then works in conjunction with the streaming service to obtain the video data in real time.

[0188] Input: Information about the game or live video you want to watch that you enter into the smartphone app

[0189] Data processing: Send desired viewing information to the server as an HTTP request

[0190] Output: The server connects to the streaming service and acquires video data in real time.

[0191] Specific operation: The user opens the "Watch" screen on the smartphone app and selects the game or live event they want to watch. The selection information is sent to the server by pressing the send button.

[0192] Step 3:

[0193] The server inputs the acquired video data into the generative AI model, which analyzes important scenes and distinctive events. For example, it identifies goal scenes in a soccer match, or performance scenes and MC talk in a live concert.

[0194] Input: Match and live video data acquired by the server

[0195] Data processing: Analyzing video data using generative AI models

[0196] Output: A list of important scenes and notable events

[0197] Specific operation: Run a generative AI model (e.g., Google Cloud Video Intelligence API) to identify important scenes in the video data and generate a scene list.

[0198] Step 4:

[0199] The server filters relevant scenes from the analysis results based on the user's settings, extracting only scenes that include specific players or artists.

[0200] Input: A list of important scenes and events output from the generative AI model, and user settings.

[0201] Data processing: Filtering the scene according to user settings

[0202] Output: A list of scenes that contain a specific player or artist

[0203] Specific operation: The server references the user's settings information and filters scenes that feature athletes or artists based on the analysis results.

[0204] Step 5:

[0205] The server then edits the filtered scenes to generate a highlight video, adding transition effects, background music, and text overlays.

[0206] Input: A filtered list of scenes

[0207] Data processing: Generate highlight videos using video editing software (e.g., FFmpeg)

[0208] Output: Highlight video file

[0209] How it works: The server uses video editing software to stitch the scenes together in order, adding transition effects and background music as needed.

[0210] Step 6:

[0211] The server provides the generated highlight video to the user's terminal, which receives the notification and plays the video. Furthermore, a connection means is provided for projecting the video from the user's terminal onto a display device within the store.

[0212] Input: Generated highlight video file

[0213] Data processing: Notifications and video streaming settings on user devices

[0214] Output: A link that can be played on the user's device

[0215] Specific operation: The server saves the highlight video in a database and sends a link to the user's device. The user clicks the notification to play the video and, if necessary, project the image on a display device in the store.

[0216] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0217] The present invention relates to a system for generating highlight videos from game or live footage based on user customization settings, and further adjusting the content of the videos by recognizing the user's emotions. The system includes a user terminal, a server, and an emotion engine.

[0218] Program processing and specific operations

[0219] Accepting user settings

[0220] 1. The user launches the application and logs in. The user selects the specific player or member they want to support and authorizes the use of the emotion engine.

[0221] 2. The device sends the user's selection information to the server and saves the setting information and permission to use the emotion engine on the server.

[0222] Start watching content

[0223] 1. The user selects the sports broadcast or live footage they want to watch.

[0224] 2. The terminal sends content selection information and setting information to the server.

[0225] Running AI models and emotion engines

[0226] 1. The server connects to the streaming service and prepares to acquire video data. The streaming data is input into the generative AI model, which begins analyzing the video data and detecting key scenes.

[0227] 2. The server starts an emotion engine that analyzes the user's facial expressions and voice in real time. The emotion engine collects emotional data using the user's camera and microphone.

[0228] Integration of favorite scenes and emotional data

[0229] 1. The server detects important scenes from the acquired video data and filters the scenes based on specific players or members set by the user.

[0230] 2. The server analyzes the user's emotional data collected in real time to detect their current emotional state (e.g., joy, excitement, sadness, etc.).

[0231] Highlight video generation and adjustment

[0232] 1. The server edits the filtered scenes, adds transitions, background music, and text overlays to generate a smooth highlight video.

[0233] 2. The server adjusts the content and order of the highlight video based on the user's emotional data. For example, if the user is excited, it will prioritize displaying more dynamic scenes.

[0234] Highlight video provided

[0235] 1. The server converts the generated highlight video into the optimal viewing format, generates a link, and notifies the user's device of the link to the highlight video stored in the database.

[0236] 2. The device receives the highlight video link sent from the server and displays a notification to the user. When the user clicks the notification, the highlight video begins playing.

[0237] Specific examples

[0238] For example, a user can watch a soccer match from the middle of the match and a highlight video will be automatically generated, including key moments (e.g., goals and assists) of Lionel Messi that the user has selected. While watching, the emotion engine analyzes the user's facial expressions, and if it detects excitement, the highlight video will be edited to include more dynamic play scenes and cheers.

[0239] Users can also watch live music concerts from the middle of the show and be provided with a highlight video of their favorite band's performance. If the emotion engine detects the user's smile, it will add more moving visual effects to that performance scene.

[0240] As described above, this system enhances the viewing experience by providing a customized highlight video that corresponds to the user's individual emotional state.

[0241] The processing flow will be explained below.

[0242] Step 1:

[0243] The user launches the application and logs in. They select the specific player or member they want to support, such as Lionel Messi or Alice, and authorize the use of the emotion engine.

[0244] Step 2:

[0245] The terminal receives the user's selection information and permission to use the emotion engine, and transmits the information to the server.

[0246] Step 3:

[0247] The server stores the received configuration information in a user database, establishing the user's customized settings.

[0248] Step 4:

[0249] The user selects the sports broadcast or live video they want to watch, and the device sends the selected information to the server.

[0250] Step 5:

[0251] The server receives the instruction to start viewing and prepares to obtain video data in real time in cooperation with the streaming service.

[0252] Step 6:

[0253] The server inputs the acquired video data into the generative AI model and begins video analysis, which then detects important scenes and characteristic events.

[0254] Step 7:

[0255] The server stores the analysis results in temporary storage and uses the data for subsequent analysis and editing.

[0256] Step 8:

[0257] The server collects the user's facial expressions and voice in real time through the user's camera and microphone, and inputs them into the emotion engine, which analyzes the emotion data and determines the user's current emotional state.

[0258] Step 9:

[0259] The server filters the scenes detected by the AI ​​model based on specific player and team information provided by the user, such as Lionel Messi's goal or Alice's performance.

[0260] Step 10:

[0261] The server organizes the video clips in chronological order and feeds them into editing software, which includes adding transitions, background music, and text overlays.

[0262] Step 11:

[0263] The server adjusts the highlight video composition based on the user's emotional data collected by the emotion engine, for example, prioritizing more action-packed scenes if the user is excited.

[0264] Step 12:

[0265] The server renders the final highlight video and converts it into the appropriate format, ensuring smooth playback on the user's device.

[0266] Step 13:

[0267] The server stores the link of the generated highlight video in a database and notifies the link information to the user's terminal.

[0268] Step 14:

[0269] The device notifies the user of the link to the highlight video received from the server. When the user clicks on the notification, the device starts streaming or downloading the highlight video and plays it.

[0270] These steps allow users to enjoy customized highlight videos that include scenes of specific players or team members. The introduction of the emotion engine further enhances the individual viewing experience, providing dynamic video content tailored to emotions.

[0271] Example 2

[0272] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0273] In modern games and live video viewing, it is difficult to quickly find specific scenes that users are interested in from vast amounts of video data. Furthermore, to improve the user's viewing experience, it is necessary not only to provide video content but also to customize it to match the user's real-time emotional state. However, conventional systems have had difficulty in recognizing emotions in real time and dynamically adjusting content based on those emotions.

[0274] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0275] In this invention, the server includes means for acquiring game and live video data, means for accepting information on specific players and members set by the user, means for analyzing the acquired video data to detect important scenes and characteristic events, means for filtering scenes including specific players and members from the detected scenes based on the user's settings, means for analyzing the collected user emotional data to detect the user's current emotional state, means for dynamically adjusting the generated highlight video based on the user's emotional state, means for generating a highlight video by editing the filtered scenes, and means for providing the generated highlight video to the user. This makes it possible to provide a customized highlight video that corresponds to the individual emotional state of the user.

[0276] "Means for acquiring game or live video data" refers to a function including communication means and interfaces for acquiring game or live event video data in real time or from storage.

[0277] "Means for accepting information on specific players and members set by the user" refers to an input means by which the user specifies the players they want to support or members they are interested in on the system, and a function for acquiring and saving that setting information.

[0278] "Means for analyzing acquired video data to detect important scenes and characteristic events" refers to a function that uses video analysis technology to automatically recognize characteristic actions and important scenes (such as goal scenes and performance scenes) during a match or event.

[0279] "Means for filtering scenes that include specific players or members from detected scenes based on user settings" is a filtering function for selecting only scenes related to specific players or members set by the user.

[0280] "Means for detecting the user's current emotional state by analyzing collected user emotional data" refers to a function for identifying the user's emotional state (e.g., joy, excitement, sadness, etc.) by analyzing data collected from the user's facial expressions and voice.

[0281] The "means for dynamically adjusting the generated highlight video based on the user's emotional state" is an editing function for changing the content and order of the highlight video in real time according to the analyzed emotional data of the user.

[0282] The "means for editing filtered scenes to generate a highlight video" is a function for editing multiple selected scenes into a single video, and then adding visual effects and music to create a visually consistent highlight video.

[0283] The "means for providing the generated highlight video to the user" is a function including a communication means and an interface for transmitting the highlight video to the user terminal so that the user can view it.

[0284] The present invention relates to a system for generating highlight videos from game or live footage based on user customization settings, and further adjusting the content of the videos by recognizing the user's emotions. The system includes a user terminal, a server, and an emotion engine.

[0285] Program processing and specific operations

[0286] Accepting user settings

[0287] Users launch a dedicated application on their smartphone or PC and log in to their account. Through the application, users can select the specific player or member they want to support and authorize the use of the emotion engine. The user's device sends this setting information to the server, which stores it in a database. For example, it is possible to set up the system so that users can "select their favorite basketball player and obtain highlight footage of that player."

[0288] Start watching content

[0289] When a user selects a game or live video they want to watch, the user's device sends the selection information and user settings information to the server. The server then works with the streaming service to prepare to acquire the video data. For example, a scenario could be that a user selects an artist's live performance and watches it in high definition.

[0290] Running AI models and emotion engines

[0291] The server inputs the streaming data into a generative AI model, which begins analyzing the video data and detecting key scenes. This AI model uses, for example, OpenAI's video analysis model. The server also activates an emotion engine, which collects data in real time from the user's camera and microphone. The emotion engine analyzes the user's facial expressions and voice to detect their current emotional state. For example, "if the user shows excitement during a game, that state is recorded."

[0292] Integration of favorite scenes and emotional data

[0293] The server detects important scenes from the analyzed video data and filters them based on specific players or team members set by the user. At the same time, it analyzes the collected user emotional data to detect the user's current emotional state. For example, if the user is "excited about a goal," the scene will be highlighted by filtering.

[0294] Highlight video generation and adjustment

[0295] The server then edits the filtered scenes, adding transitions, background music, and text overlays to create a smooth highlight video. This process uses video editing software such as Adobe Premiere Pro. The server also adjusts the content and order of the video based on the user's emotional data. For example, if the user is sad, it might add encouraging scenes.

[0296] Highlight video provided

[0297] The server encodes the generated highlight video into the optimal viewing format and uploads it to cloud storage. The user's device is then notified of the link to the generated video, which then receives the link and displays a notification to the user. When the user clicks on the notification, playback of the highlight video begins. For example, a scenario could be that a user plays a highlight video of their favorite idol during their morning commute.

[0298] Specific examples

[0299] For example, if a user starts watching a soccer match partway through, a highlight video will be automatically generated, including a goal scored by a specific player. If the emotion engine analyzes the user's facial expressions while watching and detects an excited state, the highlight video will be edited to include more dynamic playing scenes and cheers. Similarly, if a user starts watching a music concert partway through, a highlight video will be provided that includes performance scenes of their favorite band members. If the emotion engine detects the user's smile, moving visual effects will be added to the performance scenes.

[0300] Examples of prompt statements

[0301] "Please explain how you can automatically detect goals scored by specific players in real time while a user is watching a soccer game, and dynamically generate a highlight video based on the user's emotional state."

[0302] The system enhances the viewing experience by providing a customized highlight video that responds to the user's individual emotional state.

[0303] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0304] Step 1:

[0305] Users launch a dedicated application on their smartphone or computer and log in to their account. The information they enter is authentication information such as their user ID and password. The output is the authentication result indicating whether the user successfully logged in.

[0306] Step 2:

[0307] Within the application, users can select the specific player or member they want to support and authorize the use of the emotion engine. The information input is the information about the player or member they are supporting, and whether or not the emotion engine can be used. The output is the setting information and the confirmation result of permission to use the emotion engine. The device sends this setting information to the server, which stores it in a database. Specifically, the device's app screen displays a list of players and a checkbox to allow the emotion engine.

[0308] Step 3:

[0309] The user selects the game or live video they want to watch. The input information is the ID or URL of the game or live video they want to watch. The output is information about the selected content. The device sends this information and user setting information to the server.

[0310] Step 4:

[0311] The server calls the streaming service API and prepares to obtain game or live video data. The input information is the ID and URL of the game or live video to be viewed. The output is streaming data. Here, the server receives the video data in real time via the API.

[0312] Step 5:

[0313] The server inputs the acquired streaming data into the generative AI model and begins analyzing the video data and detecting key scenes. The input information is the streaming data. The output is a list of detected key scenes. The AI ​​model used could be, for example, OpenAI's video analysis model. Specifically, the server identifies goal scenes and performance scenes in real time.

[0314] Step 6:

[0315] The server runs the emotion engine to collect data in real time from the user's camera and microphone. The input information is the user's facial expressions and voice. The output is the user's emotional state (e.g., joy, excitement, sadness, etc.). Specifically, the engine captures a picture of the user's face with a webcam and analyzes their voice with a microphone.

[0316] Step 7:

[0317] The server filters out scenes related to specific players or members set by the user from among the important scenes detected by the generative AI model. The input information is a list of detected important scenes and user settings. The output is a list of filtered scenes. Specifically, the scenes in the list are organized based on the user's settings.

[0318] Step 8:

[0319] The server then analyzes and filters the video data, adjusts the content and order of the highlight video based on the user's emotional data, and begins editing. The input information is the list of filtered scenes and the user's emotional state. The output is an edited highlight video. Specifically, the server adds transitions, background music, and text overlays using video editing tools such as Adobe Premiere Pro API.

[0320] Step 9:

[0321] The server encodes the generated highlight video into an optimal viewing format and uploads it to cloud storage. The input information is the edited highlight video. The output is a link for viewing the video. Specifically, the encoded video is saved in cloud storage and a link is generated.

[0322] Step 10:

[0323] The device displays a link to view the highlight video to the user as a new arrival notification. The input information is the link to view the highlight video. The output is a notification that is displayed to the user. Specifically, the notification is displayed on the user's smartphone or computer, and when the user clicks the link, the video begins playing.

[0324] The above are the specific processing steps and operations of the program of this system.

[0325] (Application example 2)

[0326] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0327] Conventional highlight video generation systems were capable of generating highlight videos based on user customization settings, but it was difficult to reflect the user's emotions in real time. As a result, the viewing experience could not adequately respond to individual emotions, and visual satisfaction and excitement could not be maximized. The present invention aims to solve this problem and provide a more personalized and moving viewing experience by appropriately adjusting the content and order of highlight videos according to the user's emotions.

[0328] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0329] In this invention, the server includes means for acquiring game and live video data, means for accepting information on specific players and members set by the user, means for analyzing the acquired video data to detect important scenes and characteristic events, means for filtering scenes including specific players and members from the detected scenes based on the user's settings, means for editing the filtered scenes to generate a highlight video, means for providing the generated highlight video to the user, means for analyzing the user's face and voice in real time to detect their emotional state, and means for adjusting the content and order of the highlight video based on the detected emotional state, thereby making it possible to generate a customized highlight video that corresponds to the individual emotional state of the user.

[0330] "Means for acquiring game or live video data" refers to devices or software that include interfaces or protocols for acquiring frame data and audio data of games or live video in streaming or file format.

[0331] "Means for accepting information on specific players or members set by the user" refers to an interface or application that allows the user to select the specific players or members they want to support and register that information on a server or terminal.

[0332] "Means for analyzing acquired video data to detect important scenes and characteristic events" refers to hardware or software that analyzes video data using image analysis technology and machine learning algorithms to automatically detect important moments in matches and highlights from live broadcasts.

[0333] "Means for filtering scenes that include specific players or members from scenes detected based on user settings" refers to algorithms or logic for extracting scenes that feature specific players or members set by the user, and for identifying and filtering those scenes.

[0334] The "means for editing the filtered scenes to generate a highlight video" refers to editing software or an engine for editing the filtered scenes in chronological order, adding transitions and background music, overlaying text, and so on, to generate the final highlight video.

[0335] The "means for providing the generated highlight video to the user" refers to an interface or a server for distributing the generated highlight video in a streaming format or for providing a download link.

[0336] "Means for analyzing a user's face and voice in real time to detect their emotional state" refers to an emotion recognition engine or software that captures a user's facial expressions and voice tone in real time and analyzes that data to determine the user's emotional state (for example, joy, excitement, sadness, etc.).

[0337] The "means for adjusting the content and sequence of the highlight video based on the detected emotional state" refers to an algorithm or editing software for changing the selection and arrangement of scenes in the highlight video and adjusting visual effects and background music in accordance with the user's emotional state.

[0338] The present invention provides a system for generating a highlight video from a game or live video based on a user's customized settings, and the system adjusts the content of the video by recognizing the user's emotions. Detailed embodiments of the present invention will be described below.

[0339] The system's main components include a user device, a server, and an emotion engine. The server acquires game and live video data, accepts user-specified player and team member information, and analyzes the acquired video data. It then detects important scenes and distinctive events, filters them based on the user's settings, and generates a highlight video. The server then provides the generated highlight video to the user, adjusting the content and order of the video based on the user's emotional state.

[0340] The server uses a streaming interface or file acquisition protocol to acquire video data. User devices transmit information about specific players or team members using input methods implemented in the interface or application. The server then analyzes the video data using an AI model to detect key moments. This AI model utilizes image recognition technology and machine learning algorithms.

[0341] The emotion engine collects the user's facial and voice data in real time and uses emotion recognition APIs (such as the Azure Emotion API) to determine their emotional state. This emotional data is used to adjust the content and order of the highlight video. For example, if the user is excited, more dynamic scenes can be prioritized.

[0342] The filtered scenes are then edited using editing software such as MoviePy to add transitions, background music, text overlays, etc. The final highlight video is then distributed in streaming format or provided as a download link.

[0343] For example, a user can watch a soccer match and filter out key moments (e.g., goals and assists) of a specific player. While watching, the emotion engine analyzes the user's facial expressions and, if it detects excitement, edits the highlight video to include more dynamic play scenes and cheers.

[0344] Examples of prompts for generative AI models include:

[0345] "Generate dynamic highlight videos for excited users, including specific player goals and assists."

[0346] Examples include:

[0347] As mentioned above, by providing a customized highlight video that reflects the user's emotions, a more personal and moving viewing experience is possible.

[0348] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0349] Step 1: Accepting User Settings

[0350] The user launches the application and logs in. The device selects the specific player or member the user wants to support and accepts permission to use the emotion engine as input. This input information is sent to the server, which then stores it in a database.

[0351] Step 2: Start watching content

[0352] Users select the sports broadcast or live video they want to watch from the application interface. The device receives the selection information as input and sends it along with the user's settings information to the server. Based on this information, the server works with the streaming service to prepare for the acquisition of the video data.

[0353] Step 3: Running the AI ​​model and emotion engine

[0354] The server receives video data from the streaming service and passes it as input to the generative AI model. The generative AI model analyzes the video data to detect important scenes and distinctive events. At the same time, it activates an emotion engine, collecting the user's face and voice as input in real time. The emotion engine then analyzes the user's emotional state from this data and returns the results to the server as output.

[0355] Step 4: Integrating the favorite scenes and emotion data

[0356] The server filters the key scene data obtained from the generative AI model based on specific players and team members selected by the user. This filtered scene data is then combined with the analyzed emotional data to select appropriate scenes based on the user's emotional state at that time.

[0357] Step 5: Generate and adjust the highlight video

[0358] The server then passes the filtered scenes to editing software (such as MoviePy) to add transitions, background music, and text overlays. This editing process produces a smooth highlight video. The server then adjusts the content and sequence of scenes in the highlight video based on the user's emotional data. For example, if the user is excited, more dynamic scenes will be added and edited.

[0359] Step 6: Submit a highlight video

[0360] The server converts the generated highlight video into the optimal viewing format and generates a viewing link. The link stored in the database is sent to the user's device as a notification. The device receives the notification and displays it to the user. When the user clicks the notification, the highlight video is played.

[0361] The above are the specific processing steps for implementing the present invention.

[0362] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0363] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0364] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0365] [Second embodiment]

[0366] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0367] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0368] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0369] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0370] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0371] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0372] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0373] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0374] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0375] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0376] In the smart glasses 214, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0377] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0378] The present invention relates to a system that automatically generates highlight videos of matches and live footage based on user-specified players and team members. The system includes a user device, a server, and video analysis technology using a generative AI model.

[0379] Program processing and specific operations

[0380] Accepting user settings

[0381] 1. The user launches the application and logs in. On the settings screen, the user selects the specific player or member they want to support. For example, they can select soccer player Lionel Messi or music group member Alice.

[0382] 2. The device sends the selected information to the server, which stores it in a database as the user's customized settings.

[0383] Start watching content

[0384] 1. The user selects the sports broadcast or live video they want to watch. The device sends this selection information and setting information to the server.

[0385] 2. The server works with the streaming service to set up the start of viewing and prepares to obtain video data in real time.

[0386] AI model analysis

[0387] 1. The server inputs the acquired video data into a generative AI model and automatically analyzes important scenes and distinctive events. For example, it identifies goal scenes and interviews in a soccer game, and performance scenes and MC talk in a music concert.

[0388] 2. The analysis results are temporarily stored in a database.

[0389] Extracting your favorite scenes

[0390] 1. The server filters out relevant scenes from the analysis results based on the information of specific players or members set by the user. For example, it can select scenes where Lionel Messi shoots or Alice sings.

[0391] 2. Organize these selected scenes in chronological order and review them as necessary.

[0392] Highlight video generation

[0393] 1. The server edits the filtered scenes to generate a smooth highlight video, for example by adding transition effects, background music, and editing clips including text overlays.

[0394] 2. Render the edited video, convert it into the optimal viewing format, and save it in the database.

[0395] Highlights viewing available

[0396] 1. The server sends a link to the generated highlight video to the user's device. The notification includes a URL for starting viewing.

[0397] 2. The user receives a notification and plays the highlight video. The device streams or downloads the video from the server and plays it.

[0398] For example, if a user starts watching a soccer match halfway through, a highlight video will be automatically generated, including key moments (e.g., goals and assists) of the player selected by the user, such as Lionel Messi. Similarly, if a user starts watching a live stream halfway through, a highlight video focusing on the performance of their favorite member will be provided.

[0399] This allows users to enjoy watching their favorite players and members' performances without missing important scenes, even when they join a match or live show. This system also improves the viewing experience, and is expected to increase content viewing and service users.

[0400] The processing flow will be explained below.

[0401] Step 1:

[0402] The user launches the application and logs in. On the settings screen, the user selects the specific player or member they want to support.

[0403] Step 2:

[0404] The device temporarily stores the information of the specific players or members selected by the user and displays a confirmation dialog. When the user clicks Confirm, the setting information is sent to the server.

[0405] Step 3:

[0406] The server stores the received configuration information in a user database, thereby establishing the user's customized settings.

[0407] Step 4:

[0408] The user selects the sports broadcast or live video they want to watch on the application screen, and this selection information is also sent from the device to the server.

[0409] Step 5:

[0410] The server receives the instruction to start viewing and prepares to obtain video data in real time in cooperation with the streaming service.

[0411] Step 6:

[0412] The server inputs video data obtained from the streaming service into a generative AI model, which analyzes the video data and detects important scenes and distinctive events.

[0413] Step 7:

[0414] The server stores the analysis results of the AI ​​model in temporary storage, which preserves important scene and event data.

[0415] Step 8:

[0416] The server filters out relevant scenes from the analysis results based on specific player and team member information set by the user, such as Lionel Messi's shooting scenes or Alice's singing scenes.

[0417] Step 9:

[0418] The server then organizes the filtered video clips into chronological order, rechecking and fine-tuning them as needed.

[0419] Step 10:

[0420] The server then edits the organized scene clips, adds transitions, background music, and text overlays to create a smooth highlight video.

[0421] Step 11:

[0422] The server converts the rendered highlight video into the appropriate viewing format, ensuring the video format is playable on the user's device.

[0423] Step 12:

[0424] The server stores the generated highlight video in a database and prepares to notify the user's device of the link to the video.

[0425] Step 13:

[0426] The device receives the link to the highlight video sent from the server and displays a notification to the user. When the user clicks on the notification, the highlight video starts playing.

[0427] Step 14:

[0428] Users can play highlight videos and enjoy key moments of specific players or members.

[0429] These steps allow users to enjoy customized highlight videos of specific players or members, even from the middle of the video.

[0430] Example 1

[0431] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0432] Conventional highlight video generation systems make it difficult for users to efficiently watch the highlight scenes of specific players or team members. Furthermore, content customization based on user settings is insufficient, leading to users often missing important or interesting scenes. Furthermore, real-time highlight video generation and viewing was not possible, limiting the viewing experience.

[0433] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0434] In this invention, the server includes means for acquiring game and live video data, means for accepting information on specific players and members set by the user, means for using a generative AI model to analyze the acquired video data and detect important scenes and characteristic events, means for filtering scenes containing specific players and members from the detected scenes based on the user's settings, means for using an editing tool to edit the filtered scenes to generate a highlight video, and means for notifying and providing the generated highlight video to the user's terminal. This enables users to efficiently watch the performances of specific players and members, and the generation and viewing of highlight videos is realized in real time, significantly improving the viewing experience.

[0435] "Means for acquiring data on games and live video" refers to a combination of hardware and software for capturing game and live video in real time and transmitting the data to a server.

[0436] The "means for accepting information on specific players or members set by the user" refers to an interface and protocol for allowing the user to input information on the players or members selected by the user and transmitting that information to the server.

[0437] "Means of using a generative AI model to analyze acquired video data and detect significant scenes and distinctive events" refers to a function that uses AI technology to analyze video data and identify specific events, such as goal scenes and performance scenes.

[0438] "Means for filtering scenes that include specific players or members from the detected scenes based on user settings" refers to software processing for selecting and filtering scenes related to players or members from the analyzed data.

[0439] The "means of using an editing tool to edit the filtered scenes to generate a highlight video" refers to a video editing function for combining the selected scenes and adding transitions, background music, etc. to generate a highlight video.

[0440] The "means of notifying and providing the generated highlight video to the user's device" is a function of sending a viewing link for the generated highlight video to the user's device by means of a push notification or the like.

[0441] This invention relates to a system that allows users to set specific players or members and automatically generates highlight videos of matches and live footage based on that information. The system includes a user's device, a server, and video analysis technology using a generative AI model.

[0442] Accepting user settings

[0443] First, the user launches the application and logs in. The login process involves the user entering their email address and password and pressing the "Login" button. The user then selects the specific player or member they want to support on the settings screen. For example, if a user supports Lionel Messi, they can enter his name in the search bar and select him from the list.

[0444] The device sends this selected information to a server, which stores the information as the user's customized settings in a database, for example, using a relational database management system such as MySQL.

[0445] Start watching content

[0446] Next, the user selects the sports or live video they want to watch, for example, the "2023 Champions League Final." The device then sends this selection and configuration information to the server.

[0447] The server interacts with streaming service APIs (e.g., YouTube API, Twitch API) to set up the start of viewing. The server also obtains video data in real time and creates a viewing session.

[0448] AI model analysis

[0449] The server inputs the acquired video data into a generative AI model (e.g., OpenAI's GPT-4). The video is first broken down into frames, and the generative AI model analyzes the content of each frame and tags important scenes and distinctive events. For example, goals and important plays in a soccer game, or the start and end of a performance in a live event.

[0450] The analysis results are temporarily stored in a database and include timestamp and tag information.

[0451] Extracting your favorite scenes

[0452] The server filters the database based on specific players and team members set by the user. For example, it extracts scenes of Lionel Messi's shots. Filtering is done using SQL queries.

[0453] The filtered scenes are then sorted chronologically and subjected to a video review process, where automated checking algorithms are used to verify video continuity and content.

[0454] Highlight video generation

[0455] The server then edits the filtered scenes using an editing tool (e.g., FFmpeg), adding transition effects (fade in / out), adding background music, overlaying text, etc. The edited video is then rendered and converted to the optimal format (e.g., MP4, 1080p).

[0456] The generated video is stored in the server's storage system.

[0457] Highlights viewing available

[0458] Finally, the server sends a link to the generated highlight video to the user's device via a notification service (e.g., Firebase Cloud Messaging). The notification includes a URL for starting the video.

[0459] Users can view the video by clicking the URL provided and checking the notification on their device. The device will then stream or download the video from the server and play it. Buffering is performed during this process to prevent interruptions to viewing.

[0460] Examples and prompts

[0461] For example, when a user watches a soccer match, a highlight video including key scenes of the player selected, Lionel Messi, will be automatically generated. Similarly, when a user starts watching live footage, a highlight video focusing on the performance of their favorite member will be provided.

[0462] An example of a prompt might be "Generate highlights of Lionel Messi from the following soccer game."

[0463] This will allow users to efficiently watch the performances of specific players or team members without missing any important scenes. The viewing experience will be significantly improved, which is expected to lead to an increase in content viewing and an expansion of the service user base.

[0464] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0465] Step 1: Accepting User Settings

[0466] The user launches the application and logs in. The user selects the specific player or member they want to support on the settings screen. The input is an email address and password, and the output is a message indicating whether the login was successful or failed. The user then selects a specific player or member, and the device sends this information to the server. The input is a list of players or members, and the output is information about the selected player or member.

[0467] Step 2: Save your user settings information

[0468] The server receives the user's selection information sent from the device and stores it in a database. The input is the player and member information selected by the user, and the output is the customized settings stored in the database. Specifically, the server analyzes the received data and issues SQL queries to a relational database such as MySQL.

[0469] Step 3: Select content to watch

[0470] The user selects the sports broadcast or live video they want to watch within the application. The input is information about the content they want to watch, and the output is the identification information of the selected content. The device sends this information to the server and associates it with the settings information. The server calls the streaming service API to set up viewing. The input is the user's settings information and content identification information, and the output is the streaming URL and session ID.

[0471] Step 4: Acquiring real-time video data

[0472] The server prepares to acquire video data in real time from the streaming service. The input is the streaming URL and session ID, and the output is the acquired video data. Specifically, the server starts the real-time stream using the streaming service API.

[0473] Step 5: Analysis by AI model

[0474] The server inputs the acquired video data into a generative AI model. The input is real-time video data, and the output is analyzed scene information (tagged data). The generative AI model (e.g., GPT-4) analyzes the video frame by frame and tags important scenes and events. Specifically, the video data is divided into frames and each frame is input into the model.

[0475] Step 6: Save the analysis results

[0476] The server temporarily stores the analyzed scene information in a database. The input is tagged scene information, and the output is the analysis results stored in the database. The server issues SQL queries and stores data including timestamps and tag information.

[0477] Step 7: Filtering your favorite scenes

[0478] The server filters the saved scene information based on specific players and members set by the user. The input is the user's setting information and analysis results, and the output is the filtered scene information. Specifically, the server extracts relevant scenes from the database using SQL queries.

[0479] Step 8: Edit your highlight video

[0480] The server edits the filtered scenes using an editing tool (e.g., FFmpeg). The input is the extracted scene information, and the output is an edited highlight video. Specific operations include adding transition effects, background music, and text overlays, and rendering the video.

[0481] Step 9: Save your highlight video

[0482] The server renders the edited highlight video, converts it to the optimal format, and saves it. The input is the edited video file, and the output is the saved video file. Specifically, the server sets the video encoding settings and saves it in the server's storage system.

[0483] Step 10: Notification of highlight videos and provision of viewing

[0484] The server sends a link to the generated highlight video to the user's device via a notification service (e.g., Firebase Cloud Messaging). The input is the URL of the saved video file, and the output is a notification message. The user receives the notification and clicks the provided URL to play the video. The device streams or downloads the video from the server and plays it. The input is the video URL from the server, and the output is the video that is played.

[0485] This allows users to efficiently enjoy the performances of specific players and team members without missing anything. It also enables highlight videos to be generated and viewed in real time, improving the viewing experience.

[0486] (Application example 1)

[0487] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0488] In modern sports and music bars, customers often want to watch highlight videos of specific athletes or artists in real time, but there is a lack of an efficient system to make this possible. There is also a lack of a way to easily project the highlight videos that customers individually want. This has led to a demand for improved customer satisfaction and increased customer attraction.

[0489] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0490] In this invention, the server includes means for acquiring game and live video data, means for accepting information on specific players or members set by the user, means for analyzing the acquired video data to detect important scenes and characteristic events, means for filtering scenes including specific players or members from the detected scenes based on the user's settings, means for editing the filtered scenes to generate a highlight video, means for providing the generated highlight video to the user, and connection means for projecting the video from the user's terminal onto a display device in the store, thereby enabling users to easily view highlight videos of specific players or artists in real time in the store.

[0491] "Game and live footage" refers to video data recorded from live performances such as sports games and music concerts.

[0492] "User Customization Settings" refers to the ability for users to select specific players or artists and individually customize content based on those settings.

[0493] A "highlight video" is a shortened version of a video that is edited to include particularly important and distinctive scenes from a game or live footage.

[0494] "Specific player or member" refers to a specific individual or member of a group that the user supports.

[0495] "Server" refers to a computer system for storing and processing data.

[0496] "User's device" refers to the electronic device used by the User, such as a smartphone, tablet, or PC.

[0497] "In-store display devices" refers to displays and projectors installed in physical stores such as sports bars and music bars.

[0498] "Connection means" refers to the technology for connecting the user's terminal with the display device in the store and sending and receiving data.

[0499] "Means for analyzing video data" refers to the function of analyzing acquired video data using technologies such as AI models to detect important scenes and characteristic events.

[0500] "Filtering means" refers to the function of selecting only scenes that include specific players or members from video data.

[0501] "Editing means" refers to the function of combining filtered scenes, adding transition effects and background music, and completing the highlight video.

[0502] "Means of providing" refers to the technology used to deliver the generated highlight video to users.

[0503] This invention relates to a system that automatically generates highlight videos of matches and live footage based on user-specified players and team members. The system includes a user terminal, a server, and video analysis technology using a generative AI model.

[0504] Accepting user settings

[0505] Users launch the smartphone application and log in. They select the specific player or artist they want to support on the settings screen and enter their information. This selection information is sent from the device to the server, which then stores it in a database as the user's customized settings.

[0506] Start watching content

[0507] The user selects the game or live video they want to watch. The device sends this selection information to the server, which then works with the streaming service to prepare to obtain the video data in real time, taking the user's settings into consideration.

[0508] AI model analysis

[0509] The server inputs the acquired video data into a generative AI model and automatically analyzes important scenes and distinctive events. For example, it identifies goal scenes in a game, or performance scenes and MC talk in a live performance. The results of this analysis are temporarily stored in a database.

[0510] Extracting your favorite scenes

[0511] The server filters out relevant scenes from the analysis results based on the information of specific players or members set by the user. This allows for the extraction of only scenes featuring a specific player's shot or an artist's performance. The extracted scenes are organized in chronological order and can be reviewed.

[0512] Highlight video generation

[0513] The server then edits the filtered scenes to create a smooth highlight video, adding transition effects, background music, and clip editing including text overlays. The edited video is then rendered, converted into the optimal viewing format, and stored in a database.

[0514] Highlights viewing available

[0515] The server notifies the user's device of a link to the generated highlight video. The user receives the notification and clicks the link to play the highlight video. The device streams or downloads the video from the server and plays it. A connection means is also provided for the user's device to project the video onto a display device in the store. This allows users to watch specific videos on a large screen in physical stores such as sports bars and music bars.

[0516] Specific examples

[0517] Consider the case where a customer at a sports bar requests a highlight video of a specific soccer player. When the customer requests a Lionel Messi goal scene through a smartphone app, the server retrieves the video data, inputs it into a generative AI model for analysis, and then extracts only Messi's goal scenes from the analysis results, edits them, and generates a highlight video. A link to this video is sent to the customer's device, and the video is projected onto a large screen display in the bar. The customer can then use the app to enjoy Messi's goal scene on the big screen.

[0518] Prompt Sentence Examples

[0519] "Generate a highlight video that includes a specific player's goal."

[0520] This allows users to efficiently watch only the scenes they particularly want to support without missing any of their favorite athletes or artists.

[0521] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0522] Step 1:

[0523] The user launches the smartphone app and logs in. The user selects a specific player or artist on the settings screen. This selection information is sent from the device to the server.

[0524] Input: Selected information about players and artists that users enter into the smartphone app

[0525] Data processing: Send the selected information to the server as an HTTP request

[0526] Output: Player and artist selection information is saved in the server database.

[0527] Specific operation: The user opens the "Settings" screen of the smartphone app and selects the player or artist they want to support. The selected information is sent to the server by pressing the send button.

[0528] Step 2:

[0529] Users select the game or live video they want to watch and send the selection information from their device to the server, which then works in conjunction with the streaming service to obtain the video data in real time.

[0530] Input: Information about the game or live video you want to watch that you enter into the smartphone app

[0531] Data processing: Send desired viewing information to the server as an HTTP request

[0532] Output: The server connects to the streaming service and acquires video data in real time.

[0533] Specific operation: The user opens the "Watch" screen on the smartphone app and selects the game or live event they want to watch. The selection information is sent to the server by pressing the send button.

[0534] Step 3:

[0535] The server inputs the acquired video data into the generative AI model, which analyzes important scenes and distinctive events. For example, it identifies goal scenes in a soccer match, or performance scenes and MC talk in a live concert.

[0536] Input: Match and live video data acquired by the server

[0537] Data processing: Analyzing video data using generative AI models

[0538] Output: A list of important scenes and notable events

[0539] Specific operation: Run a generative AI model (e.g., Google Cloud Video Intelligence API) to identify important scenes in the video data and generate a scene list.

[0540] Step 4:

[0541] The server filters relevant scenes from the analysis results based on the user's settings, extracting only scenes that include specific players or artists.

[0542] Input: A list of important scenes and events output from the generative AI model, and user settings.

[0543] Data processing: Filtering the scene according to user settings

[0544] Output: A list of scenes that contain a specific player or artist

[0545] Specific operation: The server references the user's settings information and filters scenes that feature athletes or artists based on the analysis results.

[0546] Step 5:

[0547] The server then edits the filtered scenes to generate a highlight video, adding transition effects, background music, and text overlays.

[0548] Input: A filtered list of scenes

[0549] Data processing: Generate highlight videos using video editing software (e.g., FFmpeg)

[0550] Output: Highlight video file

[0551] How it works: The server uses video editing software to stitch the scenes together in order, adding transition effects and background music as needed.

[0552] Step 6:

[0553] The server provides the generated highlight video to the user's terminal, which receives the notification and plays the video. Furthermore, a connection means is provided for projecting the video from the user's terminal onto a display device within the store.

[0554] Input: Generated highlight video file

[0555] Data processing: Notifications and video streaming settings on user devices

[0556] Output: A link that can be played on the user's device

[0557] Specific operation: The server saves the highlight video in a database and sends a link to the user's device. The user clicks the notification to play the video and, if necessary, project the image on a display device in the store.

[0558] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0559] The present invention relates to a system for generating highlight videos from game or live footage based on user customization settings, and further adjusting the content of the videos by recognizing the user's emotions. The system includes a user terminal, a server, and an emotion engine.

[0560] Program processing and specific operations

[0561] Accepting user settings

[0562] 1. The user launches the application and logs in. The user selects the specific player or member they want to support and authorizes the use of the emotion engine.

[0563] 2. The device sends the user's selection information to the server and saves the setting information and permission to use the emotion engine on the server.

[0564] Start watching content

[0565] 1. The user selects the sports broadcast or live footage they want to watch.

[0566] 2. The terminal sends content selection information and setting information to the server.

[0567] Running AI models and emotion engines

[0568] 1. The server connects to the streaming service and prepares to acquire video data. The streaming data is input into the generative AI model, which begins analyzing the video data and detecting key scenes.

[0569] 2. The server starts an emotion engine that analyzes the user's facial expressions and voice in real time. The emotion engine collects emotional data using the user's camera and microphone.

[0570] Integration of favorite scenes and emotional data

[0571] 1. The server detects important scenes from the acquired video data and filters the scenes based on specific players or members set by the user.

[0572] 2. The server analyzes the user's emotional data collected in real time to detect their current emotional state (e.g., joy, excitement, sadness, etc.).

[0573] Highlight video generation and adjustment

[0574] 1. The server edits the filtered scenes, adds transitions, background music, and text overlays to generate a smooth highlight video.

[0575] 2. The server adjusts the content and order of the highlight video based on the user's emotional data. For example, if the user is excited, it will prioritize displaying more dynamic scenes.

[0576] Highlight video provided

[0577] 1. The server converts the generated highlight video into the optimal viewing format, generates a link, and notifies the user's device of the link to the highlight video stored in the database.

[0578] 2. The device receives the highlight video link sent from the server and displays a notification to the user. When the user clicks the notification, the highlight video begins playing.

[0579] Specific examples

[0580] For example, a user can watch a soccer match from the middle of the match and a highlight video will be automatically generated, including key moments (e.g., goals and assists) of Lionel Messi that the user has selected. While watching, the emotion engine analyzes the user's facial expressions, and if it detects excitement, the highlight video will be edited to include more dynamic play scenes and cheers.

[0581] Users can also watch live music concerts from the middle of the show and be provided with a highlight video of their favorite band's performance. If the emotion engine detects the user's smile, it will add more moving visual effects to that performance scene.

[0582] As described above, this system enhances the viewing experience by providing a customized highlight video that corresponds to the user's individual emotional state.

[0583] The processing flow will be explained below.

[0584] Step 1:

[0585] The user launches the application and logs in. They select the specific player or member they want to support, such as Lionel Messi or Alice, and authorize the use of the emotion engine.

[0586] Step 2:

[0587] The terminal receives the user's selection information and permission to use the emotion engine, and transmits the information to the server.

[0588] Step 3:

[0589] The server stores the received configuration information in a user database, establishing the user's customized settings.

[0590] Step 4:

[0591] The user selects the sports broadcast or live video they want to watch, and the device sends the selected information to the server.

[0592] Step 5:

[0593] The server receives the instruction to start viewing and prepares to obtain video data in real time in cooperation with the streaming service.

[0594] Step 6:

[0595] The server inputs the acquired video data into the generative AI model and begins video analysis, which then detects important scenes and characteristic events.

[0596] Step 7:

[0597] The server stores the analysis results in temporary storage and uses the data for subsequent analysis and editing.

[0598] Step 8:

[0599] The server collects the user's facial expressions and voice in real time through the user's camera and microphone, and inputs them into the emotion engine, which analyzes the emotion data and determines the user's current emotional state.

[0600] Step 9:

[0601] The server filters the scenes detected by the AI ​​model based on specific player and team information provided by the user, such as Lionel Messi's goal or Alice's performance.

[0602] Step 10:

[0603] The server organizes the video clips in chronological order and feeds them into editing software, which includes adding transitions, background music, and text overlays.

[0604] Step 11:

[0605] The server adjusts the highlight video composition based on the user's emotional data collected by the emotion engine, for example, prioritizing more action-packed scenes if the user is excited.

[0606] Step 12:

[0607] The server renders the final highlight video and converts it into the appropriate format, ensuring smooth playback on the user's device.

[0608] Step 13:

[0609] The server stores the link of the generated highlight video in a database and notifies the link information to the user's terminal.

[0610] Step 14:

[0611] The device notifies the user of the link to the highlight video received from the server. When the user clicks on the notification, the device starts streaming or downloading the highlight video and plays it.

[0612] These steps allow users to enjoy customized highlight videos that include scenes of specific players or team members. The introduction of the emotion engine further enhances the individual viewing experience, providing dynamic video content tailored to emotions.

[0613] Example 2

[0614] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0615] In modern games and live video viewing, it is difficult to quickly find specific scenes that users are interested in from vast amounts of video data. Furthermore, to improve the user's viewing experience, it is necessary not only to provide video content but also to customize it to match the user's real-time emotional state. However, conventional systems have had difficulty in recognizing emotions in real time and dynamically adjusting content based on those emotions.

[0616] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0617] In this invention, the server includes means for acquiring game and live video data, means for accepting information on specific players and members set by the user, means for analyzing the acquired video data to detect important scenes and characteristic events, means for filtering scenes including specific players and members from the detected scenes based on the user's settings, means for analyzing the collected user emotional data to detect the user's current emotional state, means for dynamically adjusting the generated highlight video based on the user's emotional state, means for generating a highlight video by editing the filtered scenes, and means for providing the generated highlight video to the user. This makes it possible to provide a customized highlight video that corresponds to the individual emotional state of the user.

[0618] "Means for acquiring game or live video data" refers to a function including communication means and interfaces for acquiring game or live event video data in real time or from storage.

[0619] "Means for accepting information on specific players and members set by the user" refers to an input means by which the user specifies the players they want to support or members they are interested in on the system, and a function for acquiring and saving that setting information.

[0620] "Means for analyzing acquired video data to detect important scenes and characteristic events" refers to a function that uses video analysis technology to automatically recognize characteristic actions and important scenes (such as goal scenes and performance scenes) during a match or event.

[0621] "Means for filtering scenes that include specific players or members from detected scenes based on user settings" is a filtering function for selecting only scenes related to specific players or members set by the user.

[0622] "Means for detecting the user's current emotional state by analyzing collected user emotional data" refers to a function for identifying the user's emotional state (e.g., joy, excitement, sadness, etc.) by analyzing data collected from the user's facial expressions and voice.

[0623] The "means for dynamically adjusting the generated highlight video based on the user's emotional state" is an editing function for changing the content and order of the highlight video in real time according to the analyzed emotional data of the user.

[0624] The "means for editing filtered scenes to generate a highlight video" is a function for editing multiple selected scenes into a single video, and then adding visual effects and music to create a visually consistent highlight video.

[0625] The "means for providing the generated highlight video to the user" is a function including a communication means and an interface for transmitting the highlight video to the user terminal so that the user can view it.

[0626] The present invention relates to a system for generating highlight videos from game or live footage based on user customization settings, and further adjusting the content of the videos by recognizing the user's emotions. The system includes a user terminal, a server, and an emotion engine.

[0627] Program processing and specific operations

[0628] Accepting user settings

[0629] Users launch a dedicated application on their smartphone or PC and log in to their account. Through the application, users can select the specific player or member they want to support and authorize the use of the emotion engine. The user's device sends this setting information to the server, which stores it in a database. For example, it is possible to set up the system so that users can "select their favorite basketball player and obtain highlight footage of that player."

[0630] Start watching content

[0631] When a user selects a game or live video they want to watch, the user's device sends the selection information and user settings information to the server. The server then works with the streaming service to prepare to acquire the video data. For example, a scenario could be that a user selects an artist's live performance and watches it in high definition.

[0632] Running AI models and emotion engines

[0633] The server inputs the streaming data into a generative AI model, which begins analyzing the video data and detecting key scenes. This AI model uses, for example, OpenAI's video analysis model. The server also activates an emotion engine, which collects data in real time from the user's camera and microphone. The emotion engine analyzes the user's facial expressions and voice to detect their current emotional state. For example, "if the user shows excitement during a game, that state is recorded."

[0634] Integration of favorite scenes and emotional data

[0635] The server detects important scenes from the analyzed video data and filters them based on specific players or team members set by the user. At the same time, it analyzes the collected user emotional data to detect the user's current emotional state. For example, if the user is "excited about a goal," the scene will be highlighted by filtering.

[0636] Highlight video generation and adjustment

[0637] The server then edits the filtered scenes, adding transitions, background music, and text overlays to create a smooth highlight video. This process uses video editing software such as Adobe Premiere Pro. The server also adjusts the content and order of the video based on the user's emotional data. For example, if the user is sad, it might add encouraging scenes.

[0638] Highlight video provided

[0639] The server encodes the generated highlight video into the optimal viewing format and uploads it to cloud storage. The user's device is then notified of the link to the generated video, which then receives the link and displays a notification to the user. When the user clicks on the notification, playback of the highlight video begins. For example, a scenario could be that a user plays a highlight video of their favorite idol during their morning commute.

[0640] Specific examples

[0641] For example, if a user starts watching a soccer match partway through, a highlight video will be automatically generated, including a goal scored by a specific player. If the emotion engine analyzes the user's facial expressions while watching and detects an excited state, the highlight video will be edited to include more dynamic playing scenes and cheers. Similarly, if a user starts watching a music concert partway through, a highlight video will be provided that includes performance scenes of their favorite band members. If the emotion engine detects the user's smile, moving visual effects will be added to the performance scenes.

[0642] Examples of prompt statements

[0643] "Please explain how you can automatically detect goals scored by specific players in real time while a user is watching a soccer game, and dynamically generate a highlight video based on the user's emotional state."

[0644] The system enhances the viewing experience by providing a customized highlight video that responds to the user's individual emotional state.

[0645] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0646] Step 1:

[0647] Users launch a dedicated application on their smartphone or computer and log in to their account. The information they enter is authentication information such as their user ID and password. The output is the authentication result indicating whether the user successfully logged in.

[0648] Step 2:

[0649] Within the application, users can select the specific player or member they want to support and authorize the use of the emotion engine. The information input is the information about the player or member they are supporting, and whether or not the emotion engine can be used. The output is the setting information and the confirmation result of permission to use the emotion engine. The device sends this setting information to the server, which stores it in a database. Specifically, the device's app screen displays a list of players and a checkbox to allow the emotion engine.

[0650] Step 3:

[0651] The user selects the game or live video they want to watch. The input information is the ID or URL of the game or live video they want to watch. The output is information about the selected content. The device sends this information and user setting information to the server.

[0652] Step 4:

[0653] The server calls the streaming service API and prepares to obtain game or live video data. The input information is the ID and URL of the game or live video to be viewed. The output is streaming data. Here, the server receives the video data in real time via the API.

[0654] Step 5:

[0655] The server inputs the acquired streaming data into the generative AI model and begins analyzing the video data and detecting key scenes. The input information is the streaming data. The output is a list of detected key scenes. The AI ​​model used could be, for example, OpenAI's video analysis model. Specifically, the server identifies goal scenes and performance scenes in real time.

[0656] Step 6:

[0657] The server runs the emotion engine to collect data in real time from the user's camera and microphone. The input information is the user's facial expressions and voice. The output is the user's emotional state (e.g., joy, excitement, sadness, etc.). Specifically, the engine captures a picture of the user's face with a webcam and analyzes their voice with a microphone.

[0658] Step 7:

[0659] The server filters out scenes related to specific players or members set by the user from among the important scenes detected by the generative AI model. The input information is a list of detected important scenes and user settings. The output is a list of filtered scenes. Specifically, the scenes in the list are organized based on the user's settings.

[0660] Step 8:

[0661] The server then analyzes and filters the video data, adjusts the content and order of the highlight video based on the user's emotional data, and begins editing. The input information is the list of filtered scenes and the user's emotional state. The output is an edited highlight video. Specifically, the server adds transitions, background music, and text overlays using video editing tools such as Adobe Premiere Pro API.

[0662] Step 9:

[0663] The server encodes the generated highlight video into an optimal viewing format and uploads it to cloud storage. The input information is the edited highlight video. The output is a link for viewing the video. Specifically, the encoded video is saved in cloud storage and a link is generated.

[0664] Step 10:

[0665] The device displays a link to view the highlight video to the user as a new arrival notification. The input information is the link to view the highlight video. The output is a notification that is displayed to the user. Specifically, the notification is displayed on the user's smartphone or computer, and when the user clicks the link, the video begins playing.

[0666] The above are the specific processing steps and operations of the program of this system.

[0667] (Application example 2)

[0668] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0669] Conventional highlight video generation systems were capable of generating highlight videos based on user customization settings, but it was difficult to reflect the user's emotions in real time. As a result, the viewing experience could not adequately respond to individual emotions, and visual satisfaction and excitement could not be maximized. The present invention aims to solve this problem and provide a more personalized and moving viewing experience by appropriately adjusting the content and order of highlight videos according to the user's emotions.

[0670] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0671] In this invention, the server includes means for acquiring game and live video data, means for accepting information on specific players and members set by the user, means for analyzing the acquired video data to detect important scenes and characteristic events, means for filtering scenes including specific players and members from the detected scenes based on the user's settings, means for editing the filtered scenes to generate a highlight video, means for providing the generated highlight video to the user, means for analyzing the user's face and voice in real time to detect their emotional state, and means for adjusting the content and order of the highlight video based on the detected emotional state, thereby making it possible to generate a customized highlight video that corresponds to the individual emotional state of the user.

[0672] "Means for acquiring game or live video data" refers to devices or software that include interfaces or protocols for acquiring frame data and audio data of games or live video in streaming or file format.

[0673] "Means for accepting information on specific players or members set by the user" refers to an interface or application that allows the user to select the specific players or members they want to support and register that information on a server or terminal.

[0674] "Means for analyzing acquired video data to detect important scenes and characteristic events" refers to hardware or software that analyzes video data using image analysis technology and machine learning algorithms to automatically detect important moments in matches and highlights from live broadcasts.

[0675] "Means for filtering scenes that include specific players or members from scenes detected based on user settings" refers to algorithms or logic for extracting scenes that feature specific players or members set by the user, and for identifying and filtering those scenes.

[0676] The "means for editing the filtered scenes to generate a highlight video" refers to editing software or an engine for editing the filtered scenes in chronological order, adding transitions and background music, overlaying text, and so on, to generate the final highlight video.

[0677] The "means for providing the generated highlight video to the user" refers to an interface or a server for distributing the generated highlight video in a streaming format or for providing a download link.

[0678] "Means for analyzing a user's face and voice in real time to detect their emotional state" refers to an emotion recognition engine or software that captures a user's facial expressions and voice tone in real time and analyzes that data to determine the user's emotional state (for example, joy, excitement, sadness, etc.).

[0679] The "means for adjusting the content and sequence of the highlight video based on the detected emotional state" refers to an algorithm or editing software for changing the selection and arrangement of scenes in the highlight video and adjusting visual effects and background music in accordance with the user's emotional state.

[0680] The present invention provides a system for generating a highlight video from a game or live video based on a user's customized settings, and the system adjusts the content of the video by recognizing the user's emotions. Detailed embodiments of the present invention will be described below.

[0681] The system's main components include a user device, a server, and an emotion engine. The server acquires game and live video data, accepts user-specified player and team member information, and analyzes the acquired video data. It then detects important scenes and distinctive events, filters them based on the user's settings, and generates a highlight video. The server then provides the generated highlight video to the user, adjusting the content and order of the video based on the user's emotional state.

[0682] The server uses a streaming interface or file acquisition protocol to acquire video data. User devices transmit information about specific players or team members using input methods implemented in the interface or application. The server then analyzes the video data using an AI model to detect key moments. This AI model utilizes image recognition technology and machine learning algorithms.

[0683] The emotion engine collects the user's facial and voice data in real time and uses emotion recognition APIs (such as the Azure Emotion API) to determine their emotional state. This emotional data is used to adjust the content and order of the highlight video. For example, if the user is excited, more dynamic scenes can be prioritized.

[0684] The filtered scenes are then edited using editing software such as MoviePy to add transitions, background music, text overlays, etc. The final highlight video is then distributed in streaming format or provided as a download link.

[0685] For example, a user can watch a soccer match and filter out key moments (e.g., goals and assists) of a specific player. While watching, the emotion engine analyzes the user's facial expressions and, if it detects excitement, edits the highlight video to include more dynamic play scenes and cheers.

[0686] Examples of prompts for generative AI models include:

[0687] "Generate dynamic highlight videos for excited users, including specific player goals and assists."

[0688] Examples include:

[0689] As mentioned above, by providing a customized highlight video that reflects the user's emotions, a more personal and moving viewing experience is possible.

[0690] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0691] Step 1: Accepting User Settings

[0692] The user launches the application and logs in. The device selects the specific player or member the user wants to support and accepts permission to use the emotion engine as input. This input information is sent to the server, which then stores it in a database.

[0693] Step 2: Start watching content

[0694] Users select the sports broadcast or live video they want to watch from the application interface. The device receives the selection information as input and sends it along with the user's settings information to the server. Based on this information, the server works with the streaming service to prepare for the acquisition of the video data.

[0695] Step 3: Running the AI ​​model and emotion engine

[0696] The server receives video data from the streaming service and passes it as input to the generative AI model. The generative AI model analyzes the video data to detect important scenes and distinctive events. At the same time, it activates an emotion engine, collecting the user's face and voice as input in real time. The emotion engine then analyzes the user's emotional state from this data and returns the results to the server as output.

[0697] Step 4: Integrating the favorite scenes and emotion data

[0698] The server filters the key scene data obtained from the generative AI model based on specific players and team members selected by the user. This filtered scene data is then combined with the analyzed emotional data to select appropriate scenes based on the user's emotional state at that time.

[0699] Step 5: Generate and adjust the highlight video

[0700] The server then passes the filtered scenes to editing software (such as MoviePy) to add transitions, background music, and text overlays. This editing process produces a smooth highlight video. The server then adjusts the content and sequence of scenes in the highlight video based on the user's emotional data. For example, if the user is excited, more dynamic scenes will be added and edited.

[0701] Step 6: Submit a highlight video

[0702] The server converts the generated highlight video into the optimal viewing format and generates a viewing link. The link stored in the database is sent to the user's device as a notification. The device receives the notification and displays it to the user. When the user clicks the notification, the highlight video is played.

[0703] The above are the specific processing steps for implementing the present invention.

[0704] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0705] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0706] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0707] [Third embodiment]

[0708] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0709] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0710] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0711] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0712] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0713] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0714] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0715] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0716] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0717] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0718] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0719] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0720] The present invention relates to a system that automatically generates highlight videos of matches and live footage based on user-specified players and team members. The system includes a user device, a server, and video analysis technology using a generative AI model.

[0721] Program processing and specific operations

[0722] Accepting user settings

[0723] 1. The user launches the application and logs in. On the settings screen, the user selects the specific player or member they want to support. For example, they can select soccer player Lionel Messi or music group member Alice.

[0724] 2. The device sends the selected information to the server, which stores it in a database as the user's customized settings.

[0725] Start watching content

[0726] 1. The user selects the sports broadcast or live video they want to watch. The device sends this selection information and setting information to the server.

[0727] 2. The server works with the streaming service to set up the start of viewing and prepares to obtain video data in real time.

[0728] AI model analysis

[0729] 1. The server inputs the acquired video data into a generative AI model and automatically analyzes important scenes and distinctive events. For example, it identifies goal scenes and interviews in a soccer game, and performance scenes and MC talk in a music concert.

[0730] 2. The analysis results are temporarily stored in a database.

[0731] Extracting your favorite scenes

[0732] 1. The server filters out relevant scenes from the analysis results based on the information of specific players or members set by the user. For example, it can select scenes where Lionel Messi shoots or Alice sings.

[0733] 2. Organize these selected scenes in chronological order and review them as necessary.

[0734] Highlight video generation

[0735] 1. The server edits the filtered scenes to generate a smooth highlight video, for example by adding transition effects, background music, and editing clips including text overlays.

[0736] 2. Render the edited video, convert it into the optimal viewing format, and save it in the database.

[0737] Highlights viewing available

[0738] 1. The server sends a link to the generated highlight video to the user's device. The notification includes a URL for starting viewing.

[0739] 2. The user receives a notification and plays the highlight video. The device streams or downloads the video from the server and plays it.

[0740] For example, if a user starts watching a soccer match halfway through, a highlight video will be automatically generated, including key moments (e.g., goals and assists) of the player selected by the user, such as Lionel Messi. Similarly, if a user starts watching a live stream halfway through, a highlight video focusing on the performance of their favorite member will be provided.

[0741] This allows users to enjoy watching their favorite players and members' performances without missing important scenes, even when they join a match or live show. This system also improves the viewing experience, and is expected to increase content viewing and service users.

[0742] The processing flow will be explained below.

[0743] Step 1:

[0744] The user launches the application and logs in. On the settings screen, the user selects the specific player or member they want to support.

[0745] Step 2:

[0746] The device temporarily stores the information of the specific players or members selected by the user and displays a confirmation dialog. When the user clicks Confirm, the setting information is sent to the server.

[0747] Step 3:

[0748] The server stores the received configuration information in a user database, thereby establishing the user's customized settings.

[0749] Step 4:

[0750] The user selects the sports broadcast or live video they want to watch on the application screen, and this selection information is also sent from the device to the server.

[0751] Step 5:

[0752] The server receives the instruction to start viewing and prepares to obtain video data in real time in cooperation with the streaming service.

[0753] Step 6:

[0754] The server inputs video data obtained from the streaming service into a generative AI model, which analyzes the video data and detects important scenes and distinctive events.

[0755] Step 7:

[0756] The server stores the analysis results of the AI ​​model in temporary storage, which preserves important scene and event data.

[0757] Step 8:

[0758] The server filters out relevant scenes from the analysis results based on specific player and team member information set by the user, such as Lionel Messi's shooting scenes or Alice's singing scenes.

[0759] Step 9:

[0760] The server then organizes the filtered video clips into chronological order, rechecking and fine-tuning them as needed.

[0761] Step 10:

[0762] The server then edits the organized scene clips, adds transitions, background music, and text overlays to create a smooth highlight video.

[0763] Step 11:

[0764] The server converts the rendered highlight video into the appropriate viewing format, ensuring the video format is playable on the user's device.

[0765] Step 12:

[0766] The server stores the generated highlight video in a database and prepares to notify the user's device of the link to the video.

[0767] Step 13:

[0768] The device receives the link to the highlight video sent from the server and displays a notification to the user. When the user clicks on the notification, the highlight video starts playing.

[0769] Step 14:

[0770] Users can play highlight videos and enjoy key moments of specific players or members.

[0771] These steps allow users to enjoy customized highlight videos of specific players or members, even from the middle of the video.

[0772] Example 1

[0773] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0774] Conventional highlight video generation systems make it difficult for users to efficiently watch the highlight scenes of specific players or team members. Furthermore, content customization based on user settings is insufficient, leading to users often missing important or interesting scenes. Furthermore, real-time highlight video generation and viewing was not possible, limiting the viewing experience.

[0775] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0776] In this invention, the server includes means for acquiring game and live video data, means for accepting information on specific players and members set by the user, means for using a generative AI model to analyze the acquired video data and detect important scenes and characteristic events, means for filtering scenes containing specific players and members from the detected scenes based on the user's settings, means for using an editing tool to edit the filtered scenes to generate a highlight video, and means for notifying and providing the generated highlight video to the user's terminal. This enables users to efficiently watch the performances of specific players and members, and the generation and viewing of highlight videos is realized in real time, significantly improving the viewing experience.

[0777] "Means for acquiring data on games and live video" refers to a combination of hardware and software for capturing game and live video in real time and transmitting the data to a server.

[0778] The "means for accepting information on specific players or members set by the user" refers to an interface and protocol for allowing the user to input information on the players or members selected by the user and transmitting that information to the server.

[0779] "Means of using a generative AI model to analyze acquired video data and detect significant scenes and distinctive events" refers to a function that uses AI technology to analyze video data and identify specific events, such as goal scenes and performance scenes.

[0780] "Means for filtering scenes that include specific players or members from the detected scenes based on user settings" refers to software processing for selecting and filtering scenes related to players or members from the analyzed data.

[0781] The "means of using an editing tool to edit the filtered scenes to generate a highlight video" refers to a video editing function for combining the selected scenes and adding transitions, background music, etc. to generate a highlight video.

[0782] The "means of notifying and providing the generated highlight video to the user's device" is a function of sending a viewing link for the generated highlight video to the user's device by means of a push notification or the like.

[0783] This invention relates to a system that allows users to set specific players or members and automatically generates highlight videos of matches and live footage based on that information. The system includes a user's device, a server, and video analysis technology using a generative AI model.

[0784] Accepting user settings

[0785] First, the user launches the application and logs in. The login process involves the user entering their email address and password and pressing the "Login" button. The user then selects the specific player or member they want to support on the settings screen. For example, if a user supports Lionel Messi, they can enter his name in the search bar and select him from the list.

[0786] The device sends this selected information to a server, which stores the information as the user's customized settings in a database, for example, using a relational database management system such as MySQL.

[0787] Start watching content

[0788] Next, the user selects the sports or live video they want to watch, for example, the "2023 Champions League Final." The device then sends this selection and configuration information to the server.

[0789] The server interacts with streaming service APIs (e.g., YouTube API, Twitch API) to set up the start of viewing. The server also obtains video data in real time and creates a viewing session.

[0790] AI model analysis

[0791] The server inputs the acquired video data into a generative AI model (e.g., OpenAI's GPT-4). The video is first broken down into frames, and the generative AI model analyzes the content of each frame and tags important scenes and distinctive events. For example, goals and important plays in a soccer game, or the start and end of a performance in a live event.

[0792] The analysis results are temporarily stored in a database and include timestamp and tag information.

[0793] Extracting your favorite scenes

[0794] The server filters the database based on specific players and team members set by the user. For example, it extracts scenes of Lionel Messi's shots. Filtering is done using SQL queries.

[0795] The filtered scenes are then sorted chronologically and subjected to a video review process, where automated checking algorithms are used to verify video continuity and content.

[0796] Highlight video generation

[0797] The server then edits the filtered scenes using an editing tool (e.g., FFmpeg), adding transition effects (fade in / out), adding background music, overlaying text, etc. The edited video is then rendered and converted to the optimal format (e.g., MP4, 1080p).

[0798] The generated video is stored in the server's storage system.

[0799] Highlights viewing available

[0800] Finally, the server sends a link to the generated highlight video to the user's device via a notification service (e.g., Firebase Cloud Messaging). The notification includes a URL for starting the video.

[0801] Users can view the video by clicking the URL provided and checking the notification on their device. The device will then stream or download the video from the server and play it. Buffering is performed during this process to prevent interruptions to viewing.

[0802] Examples and prompts

[0803] For example, when a user watches a soccer match, a highlight video including key scenes of the player selected, Lionel Messi, will be automatically generated. Similarly, when a user starts watching live footage, a highlight video focusing on the performance of their favorite member will be provided.

[0804] An example of a prompt might be "Generate highlights of Lionel Messi from the following soccer game."

[0805] This will allow users to efficiently watch the performances of specific players or team members without missing any important scenes. The viewing experience will be significantly improved, which is expected to lead to an increase in content viewing and an expansion of the service user base.

[0806] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0807] Step 1: Accepting User Settings

[0808] The user launches the application and logs in. The user selects the specific player or member they want to support on the settings screen. The input is an email address and password, and the output is a message indicating whether the login was successful or failed. The user then selects a specific player or member, and the device sends this information to the server. The input is a list of players or members, and the output is information about the selected player or member.

[0809] Step 2: Save your user settings information

[0810] The server receives the user's selection information sent from the device and stores it in a database. The input is the player and member information selected by the user, and the output is the customized settings stored in the database. Specifically, the server analyzes the received data and issues SQL queries to a relational database such as MySQL.

[0811] Step 3: Select content to watch

[0812] The user selects the sports broadcast or live video they want to watch within the application. The input is information about the content they want to watch, and the output is the identification information of the selected content. The device sends this information to the server and associates it with the settings information. The server calls the streaming service API to set up viewing. The input is the user's settings information and content identification information, and the output is the streaming URL and session ID.

[0813] Step 4: Acquiring real-time video data

[0814] The server prepares to acquire video data in real time from the streaming service. The input is the streaming URL and session ID, and the output is the acquired video data. Specifically, the server starts the real-time stream using the streaming service API.

[0815] Step 5: Analysis by AI model

[0816] The server inputs the acquired video data into a generative AI model. The input is real-time video data, and the output is analyzed scene information (tagged data). The generative AI model (e.g., GPT-4) analyzes the video frame by frame and tags important scenes and events. Specifically, the video data is divided into frames and each frame is input into the model.

[0817] Step 6: Save the analysis results

[0818] The server temporarily stores the analyzed scene information in a database. The input is tagged scene information, and the output is the analysis results stored in the database. The server issues SQL queries and stores data including timestamps and tag information.

[0819] Step 7: Filtering your favorite scenes

[0820] The server filters the saved scene information based on specific players and members set by the user. The input is the user's setting information and analysis results, and the output is the filtered scene information. Specifically, the server extracts relevant scenes from the database using SQL queries.

[0821] Step 8: Edit your highlight video

[0822] The server edits the filtered scenes using an editing tool (e.g., FFmpeg). The input is the extracted scene information, and the output is an edited highlight video. Specific operations include adding transition effects, background music, and text overlays, and rendering the video.

[0823] Step 9: Save your highlight video

[0824] The server renders the edited highlight video, converts it to the optimal format, and saves it. The input is the edited video file, and the output is the saved video file. Specifically, the server sets the video encoding settings and saves it in the server's storage system.

[0825] Step 10: Notification of highlight videos and provision of viewing

[0826] The server sends a link to the generated highlight video to the user's device via a notification service (e.g., Firebase Cloud Messaging). The input is the URL of the saved video file, and the output is a notification message. The user receives the notification and clicks the provided URL to play the video. The device streams or downloads the video from the server and plays it. The input is the video URL from the server, and the output is the video that is played.

[0827] This allows users to efficiently enjoy the performances of specific players and team members without missing anything. It also enables highlight videos to be generated and viewed in real time, improving the viewing experience.

[0828] (Application example 1)

[0829] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0830] In modern sports and music bars, customers often want to watch highlight videos of specific athletes or artists in real time, but there is a lack of an efficient system to make this possible. There is also a lack of a way to easily project the highlight videos that customers individually want. This has led to a demand for improved customer satisfaction and increased customer attraction.

[0831] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0832] In this invention, the server includes means for acquiring game and live video data, means for accepting information on specific players or members set by the user, means for analyzing the acquired video data to detect important scenes and characteristic events, means for filtering scenes including specific players or members from the detected scenes based on the user's settings, means for editing the filtered scenes to generate a highlight video, means for providing the generated highlight video to the user, and connection means for projecting the video from the user's terminal onto a display device in the store, thereby enabling users to easily view highlight videos of specific players or artists in real time in the store.

[0833] "Game and live footage" refers to video data recorded from live performances such as sports games and music concerts.

[0834] "User Customization Settings" refers to the ability for users to select specific players or artists and individually customize content based on those settings.

[0835] A "highlight video" is a shortened version of a video that is edited to include particularly important and distinctive scenes from a game or live footage.

[0836] "Specific player or member" refers to a specific individual or member of a group that the user supports.

[0837] "Server" refers to a computer system for storing and processing data.

[0838] "User's device" refers to the electronic device used by the User, such as a smartphone, tablet, or PC.

[0839] "In-store display devices" refers to displays and projectors installed in physical stores such as sports bars and music bars.

[0840] "Connection means" refers to the technology for connecting the user's terminal with the display device in the store and sending and receiving data.

[0841] "Means for analyzing video data" refers to the function of analyzing acquired video data using technologies such as AI models to detect important scenes and characteristic events.

[0842] "Filtering means" refers to the function of selecting only scenes that include specific players or members from video data.

[0843] "Editing means" refers to the function of combining filtered scenes, adding transition effects and background music, and completing the highlight video.

[0844] "Means of providing" refers to the technology used to deliver the generated highlight video to users.

[0845] This invention relates to a system that automatically generates highlight videos of matches and live footage based on user-specified players and team members. The system includes a user terminal, a server, and video analysis technology using a generative AI model.

[0846] Accepting user settings

[0847] Users launch the smartphone application and log in. They select the specific player or artist they want to support on the settings screen and enter their information. This selection information is sent from the device to the server, which then stores it in a database as the user's customized settings.

[0848] Start watching content

[0849] The user selects the game or live video they want to watch. The device sends this selection information to the server, which then works with the streaming service to prepare to obtain the video data in real time, taking the user's settings into consideration.

[0850] AI model analysis

[0851] The server inputs the acquired video data into a generative AI model and automatically analyzes important scenes and distinctive events. For example, it identifies goal scenes in a game, or performance scenes and MC talk in a live performance. The results of this analysis are temporarily stored in a database.

[0852] Extracting your favorite scenes

[0853] The server filters out relevant scenes from the analysis results based on the information of specific players or members set by the user. This allows for the extraction of only scenes featuring a specific player's shot or an artist's performance. The extracted scenes are organized in chronological order and can be reviewed.

[0854] Highlight video generation

[0855] The server then edits the filtered scenes to create a smooth highlight video, adding transition effects, background music, and clip editing including text overlays. The edited video is then rendered, converted into the optimal viewing format, and stored in a database.

[0856] Highlights viewing available

[0857] The server notifies the user's device of a link to the generated highlight video. The user receives the notification and clicks the link to play the highlight video. The device streams or downloads the video from the server and plays it. A connection means is also provided for the user's device to project the video onto a display device in the store. This allows users to watch specific videos on a large screen in physical stores such as sports bars and music bars.

[0858] Specific examples

[0859] Consider the case where a customer at a sports bar requests a highlight video of a specific soccer player. When the customer requests a Lionel Messi goal scene through a smartphone app, the server retrieves the video data, inputs it into a generative AI model for analysis, and then extracts only Messi's goal scenes from the analysis results, edits them, and generates a highlight video. A link to this video is sent to the customer's device, and the video is projected onto a large screen display in the bar. The customer can then use the app to enjoy Messi's goal scene on the big screen.

[0860] Prompt Sentence Examples

[0861] "Generate a highlight video that includes a specific player's goal."

[0862] This allows users to efficiently watch only the scenes they particularly want to support without missing any of their favorite athletes or artists.

[0863] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0864] Step 1:

[0865] The user launches the smartphone app and logs in. The user selects a specific player or artist on the settings screen. This selection information is sent from the device to the server.

[0866] Input: Selected information about players and artists that users enter into the smartphone app

[0867] Data processing: Send the selected information to the server as an HTTP request

[0868] Output: Player and artist selection information is saved in the server database.

[0869] Specific operation: The user opens the "Settings" screen of the smartphone app and selects the player or artist they want to support. The selected information is sent to the server by pressing the send button.

[0870] Step 2:

[0871] Users select the game or live video they want to watch and send the selection information from their device to the server, which then works in conjunction with the streaming service to obtain the video data in real time.

[0872] Input: Information about the game or live video you want to watch that you enter into the smartphone app

[0873] Data processing: Send desired viewing information to the server as an HTTP request

[0874] Output: The server connects to the streaming service and acquires video data in real time.

[0875] Specific operation: The user opens the "Watch" screen on the smartphone app and selects the game or live event they want to watch. The selection information is sent to the server by pressing the send button.

[0876] Step 3:

[0877] The server inputs the acquired video data into the generative AI model, which analyzes important scenes and distinctive events. For example, it identifies goal scenes in a soccer match, or performance scenes and MC talk in a live concert.

[0878] Input: Match and live video data acquired by the server

[0879] Data processing: Analyzing video data using generative AI models

[0880] Output: A list of important scenes and notable events

[0881] Specific operation: Run a generative AI model (e.g., Google Cloud Video Intelligence API) to identify important scenes in the video data and generate a scene list.

[0882] Step 4:

[0883] The server filters relevant scenes from the analysis results based on the user's settings, extracting only scenes that include specific players or artists.

[0884] Input: A list of important scenes and events output from the generative AI model, and user settings.

[0885] Data processing: Filtering the scene according to user settings

[0886] Output: A list of scenes that contain a specific player or artist

[0887] Specific operation: The server references the user's settings information and filters scenes that feature athletes or artists based on the analysis results.

[0888] Step 5:

[0889] The server then edits the filtered scenes to generate a highlight video, adding transition effects, background music, and text overlays.

[0890] Input: A filtered list of scenes

[0891] Data processing: Generate highlight videos using video editing software (e.g., FFmpeg)

[0892] Output: Highlight video file

[0893] How it works: The server uses video editing software to stitch the scenes together in order, adding transition effects and background music as needed.

[0894] Step 6:

[0895] The server provides the generated highlight video to the user's terminal, which receives the notification and plays the video. Furthermore, a connection means is provided for projecting the video from the user's terminal onto a display device within the store.

[0896] Input: Generated highlight video file

[0897] Data processing: Notifications and video streaming settings on user devices

[0898] Output: A link that can be played on the user's device

[0899] Specific operation: The server saves the highlight video in a database and sends a link to the user's device. The user clicks the notification to play the video and, if necessary, project the image on a display device in the store.

[0900] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0901] The present invention relates to a system for generating highlight videos from game or live footage based on user customization settings, and further adjusting the content of the videos by recognizing the user's emotions. The system includes a user terminal, a server, and an emotion engine.

[0902] Program processing and specific operations

[0903] Accepting user settings

[0904] 1. The user launches the application and logs in. The user selects the specific player or member they want to support and authorizes the use of the emotion engine.

[0905] 2. The device sends the user's selection information to the server and saves the setting information and permission to use the emotion engine on the server.

[0906] Start watching content

[0907] 1. The user selects the sports broadcast or live footage they want to watch.

[0908] 2. The terminal sends content selection information and setting information to the server.

[0909] Running AI models and emotion engines

[0910] 1. The server connects to the streaming service and prepares to acquire video data. The streaming data is input into the generative AI model, which begins analyzing the video data and detecting key scenes.

[0911] 2. The server starts an emotion engine that analyzes the user's facial expressions and voice in real time. The emotion engine collects emotional data using the user's camera and microphone.

[0912] Integration of favorite scenes and emotional data

[0913] 1. The server detects important scenes from the acquired video data and filters the scenes based on specific players or members set by the user.

[0914] 2. The server analyzes the user's emotional data collected in real time to detect their current emotional state (e.g., joy, excitement, sadness, etc.).

[0915] Highlight video generation and adjustment

[0916] 1. The server edits the filtered scenes, adds transitions, background music, and text overlays to generate a smooth highlight video.

[0917] 2. The server adjusts the content and order of the highlight video based on the user's emotional data. For example, if the user is excited, it will prioritize displaying more dynamic scenes.

[0918] Highlight video provided

[0919] 1. The server converts the generated highlight video into the optimal viewing format, generates a link, and notifies the user's device of the link to the highlight video stored in the database.

[0920] 2. The device receives the highlight video link sent from the server and displays a notification to the user. When the user clicks the notification, the highlight video begins playing.

[0921] Specific examples

[0922] For example, a user can watch a soccer match from the middle of the match and a highlight video will be automatically generated, including key moments (e.g., goals and assists) of Lionel Messi that the user has selected. While watching, the emotion engine analyzes the user's facial expressions, and if it detects excitement, the highlight video will be edited to include more dynamic play scenes and cheers.

[0923] Users can also watch live music concerts from the middle of the show and be provided with a highlight video of their favorite band's performance. If the emotion engine detects the user's smile, it will add more moving visual effects to that performance scene.

[0924] As described above, this system enhances the viewing experience by providing a customized highlight video that corresponds to the user's individual emotional state.

[0925] The processing flow will be explained below.

[0926] Step 1:

[0927] The user launches the application and logs in. They select the specific player or member they want to support, such as Lionel Messi or Alice, and authorize the use of the emotion engine.

[0928] Step 2:

[0929] The terminal receives the user's selection information and permission to use the emotion engine, and transmits the information to the server.

[0930] Step 3:

[0931] The server stores the received configuration information in a user database, establishing the user's customized settings.

[0932] Step 4:

[0933] The user selects the sports broadcast or live video they want to watch, and the device sends the selected information to the server.

[0934] Step 5:

[0935] The server receives the instruction to start viewing and prepares to obtain video data in real time in cooperation with the streaming service.

[0936] Step 6:

[0937] The server inputs the acquired video data into the generative AI model and begins video analysis, which then detects important scenes and characteristic events.

[0938] Step 7:

[0939] The server stores the analysis results in temporary storage and uses the data for subsequent analysis and editing.

[0940] Step 8:

[0941] The server collects the user's facial expressions and voice in real time through the user's camera and microphone, and inputs them into the emotion engine, which analyzes the emotion data and determines the user's current emotional state.

[0942] Step 9:

[0943] The server filters the scenes detected by the AI ​​model based on specific player and team information provided by the user, such as Lionel Messi's goal or Alice's performance.

[0944] Step 10:

[0945] The server organizes the video clips in chronological order and feeds them into editing software, which includes adding transitions, background music, and text overlays.

[0946] Step 11:

[0947] The server adjusts the highlight video composition based on the user's emotional data collected by the emotion engine, for example, prioritizing more action-packed scenes if the user is excited.

[0948] Step 12:

[0949] The server renders the final highlight video and converts it into the appropriate format, ensuring smooth playback on the user's device.

[0950] Step 13:

[0951] The server stores the link of the generated highlight video in a database and notifies the link information to the user's terminal.

[0952] Step 14:

[0953] The device notifies the user of the link to the highlight video received from the server. When the user clicks on the notification, the device starts streaming or downloading the highlight video and plays it.

[0954] These steps allow users to enjoy customized highlight videos that include scenes of specific players or team members. The introduction of the emotion engine further enhances the individual viewing experience, providing dynamic video content tailored to emotions.

[0955] Example 2

[0956] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0957] In modern games and live video viewing, it is difficult to quickly find specific scenes that users are interested in from vast amounts of video data. Furthermore, to improve the user's viewing experience, it is necessary not only to provide video content but also to customize it to match the user's real-time emotional state. However, conventional systems have had difficulty in recognizing emotions in real time and dynamically adjusting content based on those emotions.

[0958] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0959] In this invention, the server includes means for acquiring game and live video data, means for accepting information on specific players and members set by the user, means for analyzing the acquired video data to detect important scenes and characteristic events, means for filtering scenes including specific players and members from the detected scenes based on the user's settings, means for analyzing the collected user emotional data to detect the user's current emotional state, means for dynamically adjusting the generated highlight video based on the user's emotional state, means for generating a highlight video by editing the filtered scenes, and means for providing the generated highlight video to the user. This makes it possible to provide a customized highlight video that corresponds to the individual emotional state of the user.

[0960] "Means for acquiring game or live video data" refers to a function including communication means and interfaces for acquiring game or live event video data in real time or from storage.

[0961] "Means for accepting information on specific players and members set by the user" refers to an input means by which the user specifies the players they want to support or members they are interested in on the system, and a function for acquiring and saving that setting information.

[0962] "Means for analyzing acquired video data to detect important scenes and characteristic events" refers to a function that uses video analysis technology to automatically recognize characteristic actions and important scenes (such as goal scenes and performance scenes) during a match or event.

[0963] "Means for filtering scenes that include specific players or members from detected scenes based on user settings" is a filtering function for selecting only scenes related to specific players or members set by the user.

[0964] "Means for detecting the user's current emotional state by analyzing collected user emotional data" refers to a function for identifying the user's emotional state (e.g., joy, excitement, sadness, etc.) by analyzing data collected from the user's facial expressions and voice.

[0965] The "means for dynamically adjusting the generated highlight video based on the user's emotional state" is an editing function for changing the content and order of the highlight video in real time according to the analyzed emotional data of the user.

[0966] The "means for editing filtered scenes to generate a highlight video" is a function for editing multiple selected scenes into a single video, and then adding visual effects and music to create a visually consistent highlight video.

[0967] The "means for providing the generated highlight video to the user" is a function including a communication means and an interface for transmitting the highlight video to the user terminal so that the user can view it.

[0968] The present invention relates to a system for generating highlight videos from game or live footage based on user customization settings, and further adjusting the content of the videos by recognizing the user's emotions. The system includes a user terminal, a server, and an emotion engine.

[0969] Program processing and specific operations

[0970] Accepting user settings

[0971] Users launch a dedicated application on their smartphone or PC and log in to their account. Through the application, users can select the specific player or member they want to support and authorize the use of the emotion engine. The user's device sends this setting information to the server, which stores it in a database. For example, it is possible to set up the system so that users can "select their favorite basketball player and obtain highlight footage of that player."

[0972] Start watching content

[0973] When a user selects a game or live video they want to watch, the user's device sends the selection information and user settings information to the server. The server then works with the streaming service to prepare to acquire the video data. For example, a scenario could be that a user selects an artist's live performance and watches it in high definition.

[0974] Running AI models and emotion engines

[0975] The server inputs the streaming data into a generative AI model, which begins analyzing the video data and detecting key scenes. This AI model uses, for example, OpenAI's video analysis model. The server also activates an emotion engine, which collects data in real time from the user's camera and microphone. The emotion engine analyzes the user's facial expressions and voice to detect their current emotional state. For example, "if the user shows excitement during a game, that state is recorded."

[0976] Integration of favorite scenes and emotional data

[0977] The server detects important scenes from the analyzed video data and filters them based on specific players or team members set by the user. At the same time, it analyzes the collected user emotional data to detect the user's current emotional state. For example, if the user is "excited about a goal," the scene will be highlighted by filtering.

[0978] Highlight video generation and adjustment

[0979] The server then edits the filtered scenes, adding transitions, background music, and text overlays to create a smooth highlight video. This process uses video editing software such as Adobe Premiere Pro. The server also adjusts the content and order of the video based on the user's emotional data. For example, if the user is sad, it might add encouraging scenes.

[0980] Highlight video provided

[0981] The server encodes the generated highlight video into the optimal viewing format and uploads it to cloud storage. The user's device is then notified of the link to the generated video, which then receives the link and displays a notification to the user. When the user clicks on the notification, playback of the highlight video begins. For example, a scenario could be that a user plays a highlight video of their favorite idol during their morning commute.

[0982] Specific examples

[0983] For example, if a user starts watching a soccer match partway through, a highlight video will be automatically generated, including a goal scored by a specific player. If the emotion engine analyzes the user's facial expressions while watching and detects an excited state, the highlight video will be edited to include more dynamic playing scenes and cheers. Similarly, if a user starts watching a music concert partway through, a highlight video will be provided that includes performance scenes of their favorite band members. If the emotion engine detects the user's smile, moving visual effects will be added to the performance scenes.

[0984] Examples of prompt statements

[0985] "Please explain how you can automatically detect goals scored by specific players in real time while a user is watching a soccer game, and dynamically generate a highlight video based on the user's emotional state."

[0986] The system enhances the viewing experience by providing a customized highlight video that responds to the user's individual emotional state.

[0987] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0988] Step 1:

[0989] Users launch a dedicated application on their smartphone or computer and log in to their account. The information they enter is authentication information such as their user ID and password. The output is the authentication result indicating whether the user successfully logged in.

[0990] Step 2:

[0991] Within the application, users can select the specific player or member they want to support and authorize the use of the emotion engine. The information input is the information about the player or member they are supporting, and whether or not the emotion engine can be used. The output is the setting information and the confirmation result of permission to use the emotion engine. The device sends this setting information to the server, which stores it in a database. Specifically, the device's app screen displays a list of players and a checkbox to allow the emotion engine.

[0992] Step 3:

[0993] The user selects the game or live video they want to watch. The input information is the ID or URL of the game or live video they want to watch. The output is information about the selected content. The device sends this information and user setting information to the server.

[0994] Step 4:

[0995] The server calls the streaming service API and prepares to obtain game or live video data. The input information is the ID and URL of the game or live video to be viewed. The output is streaming data. Here, the server receives the video data in real time via the API.

[0996] Step 5:

[0997] The server inputs the acquired streaming data into the generative AI model and begins analyzing the video data and detecting key scenes. The input information is the streaming data. The output is a list of detected key scenes. The AI ​​model used could be, for example, OpenAI's video analysis model. Specifically, the server identifies goal scenes and performance scenes in real time.

[0998] Step 6:

[0999] The server runs the emotion engine to collect data in real time from the user's camera and microphone. The input information is the user's facial expressions and voice. The output is the user's emotional state (e.g., joy, excitement, sadness, etc.). Specifically, the engine captures a picture of the user's face with a webcam and analyzes their voice with a microphone.

[1000] Step 7:

[1001] The server filters out scenes related to specific players or members set by the user from among the important scenes detected by the generative AI model. The input information is a list of detected important scenes and user settings. The output is a list of filtered scenes. Specifically, the scenes in the list are organized based on the user's settings.

[1002] Step 8:

[1003] The server then analyzes and filters the video data, adjusts the content and order of the highlight video based on the user's emotional data, and begins editing. The input information is the list of filtered scenes and the user's emotional state. The output is an edited highlight video. Specifically, the server adds transitions, background music, and text overlays using video editing tools such as Adobe Premiere Pro API.

[1004] Step 9:

[1005] The server encodes the generated highlight video into an optimal viewing format and uploads it to cloud storage. The input information is the edited highlight video. The output is a link for viewing the video. Specifically, the encoded video is saved in cloud storage and a link is generated.

[1006] Step 10:

[1007] The device displays a link to view the highlight video to the user as a new arrival notification. The input information is the link to view the highlight video. The output is a notification that is displayed to the user. Specifically, the notification is displayed on the user's smartphone or computer, and when the user clicks the link, the video begins playing.

[1008] The above are the specific processing steps and operations of the program of this system.

[1009] (Application example 2)

[1010] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1011] Conventional highlight video generation systems were capable of generating highlight videos based on user customization settings, but it was difficult to reflect the user's emotions in real time. As a result, the viewing experience could not adequately respond to individual emotions, and visual satisfaction and excitement could not be maximized. The present invention aims to solve this problem and provide a more personalized and moving viewing experience by appropriately adjusting the content and order of highlight videos according to the user's emotions.

[1012] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1013] In this invention, the server includes means for acquiring game and live video data, means for accepting information on specific players and members set by the user, means for analyzing the acquired video data to detect important scenes and characteristic events, means for filtering scenes including specific players and members from the detected scenes based on the user's settings, means for editing the filtered scenes to generate a highlight video, means for providing the generated highlight video to the user, means for analyzing the user's face and voice in real time to detect their emotional state, and means for adjusting the content and order of the highlight video based on the detected emotional state, thereby making it possible to generate a customized highlight video that corresponds to the individual emotional state of the user.

[1014] "Means for acquiring game or live video data" refers to devices or software that include interfaces or protocols for acquiring frame data and audio data of games or live video in streaming or file format.

[1015] "Means for accepting information on specific players or members set by the user" refers to an interface or application that allows the user to select the specific players or members they want to support and register that information on a server or terminal.

[1016] "Means for analyzing acquired video data to detect important scenes and characteristic events" refers to hardware or software that analyzes video data using image analysis technology and machine learning algorithms to automatically detect important moments in matches and highlights from live broadcasts.

[1017] "Means for filtering scenes that include specific players or members from scenes detected based on user settings" refers to algorithms or logic for extracting scenes that feature specific players or members set by the user, and for identifying and filtering those scenes.

[1018] The "means for editing the filtered scenes to generate a highlight video" refers to editing software or an engine for editing the filtered scenes in chronological order, adding transitions and background music, overlaying text, and so on, to generate the final highlight video.

[1019] The "means for providing the generated highlight video to the user" refers to an interface or a server for distributing the generated highlight video in a streaming format or for providing a download link.

[1020] "Means for analyzing a user's face and voice in real time to detect their emotional state" refers to an emotion recognition engine or software that captures a user's facial expressions and voice tone in real time and analyzes that data to determine the user's emotional state (for example, joy, excitement, sadness, etc.).

[1021] The "means for adjusting the content and sequence of the highlight video based on the detected emotional state" refers to an algorithm or editing software for changing the selection and arrangement of scenes in the highlight video and adjusting visual effects and background music in accordance with the user's emotional state.

[1022] The present invention provides a system for generating a highlight video from a game or live video based on a user's customized settings, and the system adjusts the content of the video by recognizing the user's emotions. Detailed embodiments of the present invention will be described below.

[1023] The system's main components include a user device, a server, and an emotion engine. The server acquires game and live video data, accepts user-specified player and team member information, and analyzes the acquired video data. It then detects important scenes and distinctive events, filters them based on the user's settings, and generates a highlight video. The server then provides the generated highlight video to the user, adjusting the content and order of the video based on the user's emotional state.

[1024] The server uses a streaming interface or file acquisition protocol to acquire video data. User devices transmit information about specific players or team members using input methods implemented in the interface or application. The server then analyzes the video data using an AI model to detect key moments. This AI model utilizes image recognition technology and machine learning algorithms.

[1025] The emotion engine collects the user's facial and voice data in real time and uses emotion recognition APIs (such as the Azure Emotion API) to determine their emotional state. This emotional data is used to adjust the content and order of the highlight video. For example, if the user is excited, more dynamic scenes can be prioritized.

[1026] The filtered scenes are then edited using editing software such as MoviePy to add transitions, background music, text overlays, etc. The final highlight video is then distributed in streaming format or provided as a download link.

[1027] For example, a user can watch a soccer match and filter out key moments (e.g., goals and assists) of a specific player. While watching, the emotion engine analyzes the user's facial expressions and, if it detects excitement, edits the highlight video to include more dynamic play scenes and cheers.

[1028] Examples of prompts for generative AI models include:

[1029] "Generate dynamic highlight videos for excited users, including specific player goals and assists."

[1030] Examples include:

[1031] As mentioned above, by providing a customized highlight video that reflects the user's emotions, a more personal and moving viewing experience is possible.

[1032] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1033] Step 1: Accepting User Settings

[1034] The user launches the application and logs in. The device selects the specific player or member the user wants to support and accepts permission to use the emotion engine as input. This input information is sent to the server, which then stores it in a database.

[1035] Step 2: Start watching content

[1036] Users select the sports broadcast or live video they want to watch from the application interface. The device receives the selection information as input and sends it along with the user's settings information to the server. Based on this information, the server works with the streaming service to prepare for the acquisition of the video data.

[1037] Step 3: Running the AI ​​model and emotion engine

[1038] The server receives video data from the streaming service and passes it as input to the generative AI model. The generative AI model analyzes the video data to detect important scenes and distinctive events. At the same time, it activates an emotion engine, collecting the user's face and voice as input in real time. The emotion engine then analyzes the user's emotional state from this data and returns the results to the server as output.

[1039] Step 4: Integrating the favorite scenes and emotion data

[1040] The server filters the key scene data obtained from the generative AI model based on specific players and team members selected by the user. This filtered scene data is then combined with the analyzed emotional data to select appropriate scenes based on the user's emotional state at that time.

[1041] Step 5: Generate and adjust the highlight video

[1042] The server then passes the filtered scenes to editing software (such as MoviePy) to add transitions, background music, and text overlays. This editing process produces a smooth highlight video. The server then adjusts the content and sequence of scenes in the highlight video based on the user's emotional data. For example, if the user is excited, more dynamic scenes will be added and edited.

[1043] Step 6: Submit a highlight video

[1044] The server converts the generated highlight video into the optimal viewing format and generates a viewing link. The link stored in the database is sent to the user's device as a notification. The device receives the notification and displays it to the user. When the user clicks the notification, the highlight video is played.

[1045] The above are the specific processing steps for implementing the present invention.

[1046] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1047] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1048] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1049] [Fourth embodiment]

[1050] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1051] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1052] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1053] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1054] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1055] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1056] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1057] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1058] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1059] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1060] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1061] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1062] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1063] The present invention relates to a system that automatically generates highlight videos of matches and live footage based on user-specified players and team members. The system includes a user device, a server, and video analysis technology using a generative AI model.

[1064] Program processing and specific operations

[1065] Accepting user settings

[1066] 1. The user launches the application and logs in. On the settings screen, the user selects the specific player or member they want to support. For example, they can select soccer player Lionel Messi or music group member Alice.

[1067] 2. The device sends the selected information to the server, which stores it in a database as the user's customized settings.

[1068] Start watching content

[1069] 1. The user selects the sports broadcast or live video they want to watch. The device sends this selection information and setting information to the server.

[1070] 2. The server works with the streaming service to set up the start of viewing and prepares to obtain video data in real time.

[1071] AI model analysis

[1072] 1. The server inputs the acquired video data into a generative AI model and automatically analyzes important scenes and distinctive events. For example, it identifies goal scenes and interviews in a soccer game, and performance scenes and MC talk in a music concert.

[1073] 2. The analysis results are temporarily stored in a database.

[1074] Extracting your favorite scenes

[1075] 1. The server filters out relevant scenes from the analysis results based on the information of specific players or members set by the user. For example, it can select scenes where Lionel Messi shoots or Alice sings.

[1076] 2. Organize these selected scenes in chronological order and review them as necessary.

[1077] Highlight video generation

[1078] 1. The server edits the filtered scenes to generate a smooth highlight video, for example by adding transition effects, background music, and editing clips including text overlays.

[1079] 2. Render the edited video, convert it into the optimal viewing format, and save it in the database.

[1080] Highlights viewing available

[1081] 1. The server sends a link to the generated highlight video to the user's device. The notification includes a URL for starting viewing.

[1082] 2. The user receives a notification and plays the highlight video. The device streams or downloads the video from the server and plays it.

[1083] For example, if a user starts watching a soccer match halfway through, a highlight video will be automatically generated, including key moments (e.g., goals and assists) of the player selected by the user, such as Lionel Messi. Similarly, if a user starts watching a live stream halfway through, a highlight video focusing on the performance of their favorite member will be provided.

[1084] This allows users to enjoy watching their favorite players and members' performances without missing important scenes, even when they join a match or live show. This system also improves the viewing experience, and is expected to increase content viewing and service users.

[1085] The processing flow will be explained below.

[1086] Step 1:

[1087] The user launches the application and logs in. On the settings screen, the user selects the specific player or member they want to support.

[1088] Step 2:

[1089] The device temporarily stores the information of the specific players or members selected by the user and displays a confirmation dialog. When the user clicks Confirm, the setting information is sent to the server.

[1090] Step 3:

[1091] The server stores the received configuration information in a user database, thereby establishing the user's customized settings.

[1092] Step 4:

[1093] The user selects the sports broadcast or live video they want to watch on the application screen, and this selection information is also sent from the device to the server.

[1094] Step 5:

[1095] The server receives the instruction to start viewing and prepares to obtain video data in real time in cooperation with the streaming service.

[1096] Step 6:

[1097] The server inputs video data obtained from the streaming service into a generative AI model, which analyzes the video data and detects important scenes and distinctive events.

[1098] Step 7:

[1099] The server stores the analysis results of the AI ​​model in temporary storage, which preserves important scene and event data.

[1100] Step 8:

[1101] The server filters out relevant scenes from the analysis results based on specific player and team member information set by the user, such as Lionel Messi's shooting scenes or Alice's singing scenes.

[1102] Step 9:

[1103] The server then organizes the filtered video clips into chronological order, rechecking and fine-tuning them as needed.

[1104] Step 10:

[1105] The server then edits the organized scene clips, adds transitions, background music, and text overlays to create a smooth highlight video.

[1106] Step 11:

[1107] The server converts the rendered highlight video into the appropriate viewing format, ensuring the video format is playable on the user's device.

[1108] Step 12:

[1109] The server stores the generated highlight video in a database and prepares to notify the user's device of the link to the video.

[1110] Step 13:

[1111] The device receives the link to the highlight video sent from the server and displays a notification to the user. When the user clicks on the notification, the highlight video starts playing.

[1112] Step 14:

[1113] Users can play highlight videos and enjoy key moments of specific players or members.

[1114] These steps allow users to enjoy customized highlight videos of specific players or members, even from the middle of the video.

[1115] Example 1

[1116] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1117] Conventional highlight video generation systems make it difficult for users to efficiently watch the highlight scenes of specific players or team members. Furthermore, content customization based on user settings is insufficient, leading to users often missing important or interesting scenes. Furthermore, real-time highlight video generation and viewing was not possible, limiting the viewing experience.

[1118] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1119] In this invention, the server includes means for acquiring game and live video data, means for accepting information on specific players and members set by the user, means for using a generative AI model to analyze the acquired video data and detect important scenes and characteristic events, means for filtering scenes containing specific players and members from the detected scenes based on the user's settings, means for using an editing tool to edit the filtered scenes to generate a highlight video, and means for notifying and providing the generated highlight video to the user's terminal. This enables users to efficiently watch the performances of specific players and members, and the generation and viewing of highlight videos is realized in real time, significantly improving the viewing experience.

[1120] "Means for acquiring data on games and live video" refers to a combination of hardware and software for capturing game and live video in real time and transmitting the data to a server.

[1121] The "means for accepting information on specific players or members set by the user" refers to an interface and protocol for allowing the user to input information on the players or members selected by the user and transmitting that information to the server.

[1122] "Means of using a generative AI model to analyze acquired video data and detect significant scenes and distinctive events" refers to a function that uses AI technology to analyze video data and identify specific events, such as goal scenes and performance scenes.

[1123] "Means for filtering scenes that include specific players or members from the detected scenes based on user settings" refers to software processing for selecting and filtering scenes related to players or members from the analyzed data.

[1124] The "means of using an editing tool to edit the filtered scenes to generate a highlight video" refers to a video editing function for combining the selected scenes and adding transitions, background music, etc. to generate a highlight video.

[1125] The "means of notifying and providing the generated highlight video to the user's device" is a function of sending a viewing link for the generated highlight video to the user's device by means of a push notification or the like.

[1126] This invention relates to a system that allows users to set specific players or members and automatically generates highlight videos of matches and live footage based on that information. The system includes a user's device, a server, and video analysis technology using a generative AI model.

[1127] Accepting user settings

[1128] First, the user launches the application and logs in. The login process involves the user entering their email address and password and pressing the "Login" button. The user then selects the specific player or member they want to support on the settings screen. For example, if a user supports Lionel Messi, they can enter his name in the search bar and select him from the list.

[1129] The device sends this selected information to a server, which stores the information as the user's customized settings in a database, for example, using a relational database management system such as MySQL.

[1130] Start watching content

[1131] Next, the user selects the sports or live video they want to watch, for example, the "2023 Champions League Final." The device then sends this selection and configuration information to the server.

[1132] The server interacts with streaming service APIs (e.g., YouTube API, Twitch API) to set up the start of viewing. The server also obtains video data in real time and creates a viewing session.

[1133] AI model analysis

[1134] The server inputs the acquired video data into a generative AI model (e.g., OpenAI's GPT-4). The video is first broken down into frames, and the generative AI model analyzes the content of each frame and tags important scenes and distinctive events. For example, goals and important plays in a soccer game, or the start and end of a performance in a live event.

[1135] The analysis results are temporarily stored in a database and include timestamp and tag information.

[1136] Extracting your favorite scenes

[1137] The server filters the database based on specific players and team members set by the user. For example, it extracts scenes of Lionel Messi's shots. Filtering is done using SQL queries.

[1138] The filtered scenes are then sorted chronologically and subjected to a video review process, where automated checking algorithms are used to verify video continuity and content.

[1139] Highlight video generation

[1140] The server then edits the filtered scenes using an editing tool (e.g., FFmpeg), adding transition effects (fade in / out), adding background music, overlaying text, etc. The edited video is then rendered and converted to the optimal format (e.g., MP4, 1080p).

[1141] The generated video is stored in the server's storage system.

[1142] Highlights viewing available

[1143] Finally, the server sends a link to the generated highlight video to the user's device via a notification service (e.g., Firebase Cloud Messaging). The notification includes a URL for starting the video.

[1144] Users can view the video by clicking the URL provided and checking the notification on their device. The device will then stream or download the video from the server and play it. Buffering is performed during this process to prevent interruptions to viewing.

[1145] Examples and prompts

[1146] For example, when a user watches a soccer match, a highlight video including key scenes of the player selected, Lionel Messi, will be automatically generated. Similarly, when a user starts watching live footage, a highlight video focusing on the performance of their favorite member will be provided.

[1147] An example of a prompt might be "Generate highlights of Lionel Messi from the following soccer game."

[1148] This will allow users to efficiently watch the performances of specific players or team members without missing any important scenes. The viewing experience will be significantly improved, which is expected to lead to an increase in content viewing and an expansion of the service user base.

[1149] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1150] Step 1: Accepting User Settings

[1151] The user launches the application and logs in. The user selects the specific player or member they want to support on the settings screen. The input is an email address and password, and the output is a message indicating whether the login was successful or failed. The user then selects a specific player or member, and the device sends this information to the server. The input is a list of players or members, and the output is information about the selected player or member.

[1152] Step 2: Save your user settings information

[1153] The server receives the user's selection information sent from the device and stores it in a database. The input is the player and member information selected by the user, and the output is the customized settings stored in the database. Specifically, the server analyzes the received data and issues SQL queries to a relational database such as MySQL.

[1154] Step 3: Select content to watch

[1155] The user selects the sports broadcast or live video they want to watch within the application. The input is information about the content they want to watch, and the output is the identification information of the selected content. The device sends this information to the server and associates it with the settings information. The server calls the streaming service API to set up viewing. The input is the user's settings information and content identification information, and the output is the streaming URL and session ID.

[1156] Step 4: Acquiring real-time video data

[1157] The server prepares to acquire video data in real time from the streaming service. The input is the streaming URL and session ID, and the output is the acquired video data. Specifically, the server starts the real-time stream using the streaming service API.

[1158] Step 5: Analysis by AI model

[1159] The server inputs the acquired video data into a generative AI model. The input is real-time video data, and the output is analyzed scene information (tagged data). The generative AI model (e.g., GPT-4) analyzes the video frame by frame and tags important scenes and events. Specifically, the video data is divided into frames and each frame is input into the model.

[1160] Step 6: Save the analysis results

[1161] The server temporarily stores the analyzed scene information in a database. The input is tagged scene information, and the output is the analysis results stored in the database. The server issues SQL queries and stores data including timestamps and tag information.

[1162] Step 7: Filtering your favorite scenes

[1163] The server filters the saved scene information based on specific players and members set by the user. The input is the user's setting information and analysis results, and the output is the filtered scene information. Specifically, the server extracts relevant scenes from the database using SQL queries.

[1164] Step 8: Edit your highlight video

[1165] The server edits the filtered scenes using an editing tool (e.g., FFmpeg). The input is the extracted scene information, and the output is an edited highlight video. Specific operations include adding transition effects, background music, and text overlays, and rendering the video.

[1166] Step 9: Save your highlight video

[1167] The server renders the edited highlight video, converts it to the optimal format, and saves it. The input is the edited video file, and the output is the saved video file. Specifically, the server sets the video encoding settings and saves it in the server's storage system.

[1168] Step 10: Notification of highlight videos and provision of viewing

[1169] The server sends a link to the generated highlight video to the user's device via a notification service (e.g., Firebase Cloud Messaging). The input is the URL of the saved video file, and the output is a notification message. The user receives the notification and clicks the provided URL to play the video. The device streams or downloads the video from the server and plays it. The input is the video URL from the server, and the output is the video that is played.

[1170] This allows users to efficiently enjoy the performances of specific players and team members without missing anything. It also enables highlight videos to be generated and viewed in real time, improving the viewing experience.

[1171] (Application example 1)

[1172] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1173] In modern sports and music bars, customers often want to watch highlight videos of specific athletes or artists in real time, but there is a lack of an efficient system to make this possible. There is also a lack of a way to easily project the highlight videos that customers individually want. This has led to a demand for improved customer satisfaction and increased customer attraction.

[1174] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1175] In this invention, the server includes means for acquiring game and live video data, means for accepting information on specific players or members set by the user, means for analyzing the acquired video data to detect important scenes and characteristic events, means for filtering scenes including specific players or members from the detected scenes based on the user's settings, means for editing the filtered scenes to generate a highlight video, means for providing the generated highlight video to the user, and connection means for projecting the video from the user's terminal onto a display device in the store, thereby enabling users to easily view highlight videos of specific players or artists in real time in the store.

[1176] "Game and live footage" refers to video data recorded from live performances such as sports games and music concerts.

[1177] "User Customization Settings" refers to the ability for users to select specific players or artists and individually customize content based on those settings.

[1178] A "highlight video" is a shortened version of a video that is edited to include particularly important and distinctive scenes from a game or live footage.

[1179] "Specific player or member" refers to a specific individual or member of a group that the user supports.

[1180] "Server" refers to a computer system for storing and processing data.

[1181] "User's device" refers to the electronic device used by the User, such as a smartphone, tablet, or PC.

[1182] "In-store display devices" refers to displays and projectors installed in physical stores such as sports bars and music bars.

[1183] "Connection means" refers to the technology for connecting the user's terminal with the display device in the store and sending and receiving data.

[1184] "Means for analyzing video data" refers to the function of analyzing acquired video data using technologies such as AI models to detect important scenes and characteristic events.

[1185] "Filtering means" refers to the function of selecting only scenes that include specific players or members from video data.

[1186] "Editing means" refers to the function of combining filtered scenes, adding transition effects and background music, and completing the highlight video.

[1187] "Means of providing" refers to the technology used to deliver the generated highlight video to users.

[1188] This invention relates to a system that automatically generates highlight videos of matches and live footage based on user-specified players and team members. The system includes a user terminal, a server, and video analysis technology using a generative AI model.

[1189] Accepting user settings

[1190] Users launch the smartphone application and log in. They select the specific player or artist they want to support on the settings screen and enter their information. This selection information is sent from the device to the server, which then stores it in a database as the user's customized settings.

[1191] Start watching content

[1192] The user selects the game or live video they want to watch. The device sends this selection information to the server, which then works with the streaming service to prepare to obtain the video data in real time, taking the user's settings into consideration.

[1193] AI model analysis

[1194] The server inputs the acquired video data into a generative AI model and automatically analyzes important scenes and distinctive events. For example, it identifies goal scenes in a game, or performance scenes and MC talk in a live performance. The results of this analysis are temporarily stored in a database.

[1195] Extracting your favorite scenes

[1196] The server filters out relevant scenes from the analysis results based on the information of specific players or members set by the user. This allows for the extraction of only scenes featuring a specific player's shot or an artist's performance. The extracted scenes are organized in chronological order and can be reviewed.

[1197] Highlight video generation

[1198] The server then edits the filtered scenes to create a smooth highlight video, adding transition effects, background music, and clip editing including text overlays. The edited video is then rendered, converted into the optimal viewing format, and stored in a database.

[1199] Highlights viewing available

[1200] The server notifies the user's device of a link to the generated highlight video. The user receives the notification and clicks the link to play the highlight video. The device streams or downloads the video from the server and plays it. A connection means is also provided for the user's device to project the video onto a display device in the store. This allows users to watch specific videos on a large screen in physical stores such as sports bars and music bars.

[1201] Specific examples

[1202] Consider the case where a customer at a sports bar requests a highlight video of a specific soccer player. When the customer requests a Lionel Messi goal scene through a smartphone app, the server retrieves the video data, inputs it into a generative AI model for analysis, and then extracts only Messi's goal scenes from the analysis results, edits them, and generates a highlight video. A link to this video is sent to the customer's device, and the video is projected onto a large screen display in the bar. The customer can then use the app to enjoy Messi's goal scene on the big screen.

[1203] Prompt Sentence Examples

[1204] "Generate a highlight video that includes a specific player's goal."

[1205] This allows users to efficiently watch only the scenes they particularly want to support without missing any of their favorite athletes or artists.

[1206] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1207] Step 1:

[1208] The user launches the smartphone app and logs in. The user selects a specific player or artist on the settings screen. This selection information is sent from the device to the server.

[1209] Input: Selected information about players and artists that users enter into the smartphone app

[1210] Data processing: Send the selected information to the server as an HTTP request

[1211] Output: Player and artist selection information is saved in the server database.

[1212] Specific operation: The user opens the "Settings" screen of the smartphone app and selects the player or artist they want to support. The selected information is sent to the server by pressing the send button.

[1213] Step 2:

[1214] Users select the game or live video they want to watch and send the selection information from their device to the server, which then works in conjunction with the streaming service to obtain the video data in real time.

[1215] Input: Information about the game or live video you want to watch that you enter into the smartphone app

[1216] Data processing: Send desired viewing information to the server as an HTTP request

[1217] Output: The server connects to the streaming service and acquires video data in real time.

[1218] Specific operation: The user opens the "Watch" screen on the smartphone app and selects the game or live event they want to watch. The selection information is sent to the server by pressing the send button.

[1219] Step 3:

[1220] The server inputs the acquired video data into the generative AI model, which analyzes important scenes and distinctive events. For example, it identifies goal scenes in a soccer match, or performance scenes and MC talk in a live concert.

[1221] Input: Match and live video data acquired by the server

[1222] Data processing: Analyzing video data using generative AI models

[1223] Output: A list of important scenes and notable events

[1224] Specific operation: Run a generative AI model (e.g., Google Cloud Video Intelligence API) to identify important scenes in the video data and generate a scene list.

[1225] Step 4:

[1226] The server filters relevant scenes from the analysis results based on the user's settings, extracting only scenes that include specific players or artists.

[1227] Input: A list of important scenes and events output from the generative AI model, and user settings.

[1228] Data processing: Filtering the scene according to user settings

[1229] Output: A list of scenes that contain a specific player or artist

[1230] Specific operation: The server references the user's settings information and filters scenes that feature athletes or artists based on the analysis results.

[1231] Step 5:

[1232] The server then edits the filtered scenes to generate a highlight video, adding transition effects, background music, and text overlays.

[1233] Input: A filtered list of scenes

[1234] Data processing: Generate highlight videos using video editing software (e.g., FFmpeg)

[1235] Output: Highlight video file

[1236] How it works: The server uses video editing software to stitch the scenes together in order, adding transition effects and background music as needed.

[1237] Step 6:

[1238] The server provides the generated highlight video to the user's terminal, which receives the notification and plays the video. Furthermore, a connection means is provided for projecting the video from the user's terminal onto a display device within the store.

[1239] Input: Generated highlight video file

[1240] Data processing: Notifications and video streaming settings on user devices

[1241] Output: A link that can be played on the user's device

[1242] Specific operation: The server saves the highlight video in a database and sends a link to the user's device. The user clicks the notification to play the video and, if necessary, project the image on a display device in the store.

[1243] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1244] The present invention relates to a system for generating highlight videos from game or live footage based on user customization settings, and further adjusting the content of the videos by recognizing the user's emotions. The system includes a user terminal, a server, and an emotion engine.

[1245] Program processing and specific operations

[1246] Accepting user settings

[1247] 1. The user launches the application and logs in. The user selects the specific player or member they want to support and authorizes the use of the emotion engine.

[1248] 2. The device sends the user's selection information to the server and saves the setting information and permission to use the emotion engine on the server.

[1249] Start watching content

[1250] 1. The user selects the sports broadcast or live footage they want to watch.

[1251] 2. The terminal sends content selection information and setting information to the server.

[1252] Running AI models and emotion engines

[1253] 1. The server connects to the streaming service and prepares to acquire video data. The streaming data is input into the generative AI model, which begins analyzing the video data and detecting key scenes.

[1254] 2. The server starts an emotion engine that analyzes the user's facial expressions and voice in real time. The emotion engine collects emotional data using the user's camera and microphone.

[1255] Integration of favorite scenes and emotional data

[1256] 1. The server detects important scenes from the acquired video data and filters the scenes based on specific players or members set by the user.

[1257] 2. The server analyzes the user's emotional data collected in real time to detect their current emotional state (e.g., joy, excitement, sadness, etc.).

[1258] Highlight video generation and adjustment

[1259] 1. The server edits the filtered scenes, adds transitions, background music, and text overlays to generate a smooth highlight video.

[1260] 2. The server adjusts the content and order of the highlight video based on the user's emotional data. For example, if the user is excited, it will prioritize displaying more dynamic scenes.

[1261] Highlight video provided

[1262] 1. The server converts the generated highlight video into the optimal viewing format, generates a link, and notifies the user's device of the link to the highlight video stored in the database.

[1263] 2. The device receives the highlight video link sent from the server and displays a notification to the user. When the user clicks the notification, the highlight video begins playing.

[1264] Specific examples

[1265] For example, a user can watch a soccer match from the middle of the match and a highlight video will be automatically generated, including key moments (e.g., goals and assists) of Lionel Messi that the user has selected. While watching, the emotion engine analyzes the user's facial expressions, and if it detects excitement, the highlight video will be edited to include more dynamic play scenes and cheers.

[1266] Users can also watch live music concerts from the middle of the show and be provided with a highlight video of their favorite band's performance. If the emotion engine detects the user's smile, it will add more moving visual effects to that performance scene.

[1267] As described above, this system enhances the viewing experience by providing a customized highlight video that corresponds to the user's individual emotional state.

[1268] The processing flow will be explained below.

[1269] Step 1:

[1270] The user launches the application and logs in. They select the specific player or member they want to support, such as Lionel Messi or Alice, and authorize the use of the emotion engine.

[1271] Step 2:

[1272] The terminal receives the user's selection information and permission to use the emotion engine, and transmits the information to the server.

[1273] Step 3:

[1274] The server stores the received configuration information in a user database, establishing the user's customized settings.

[1275] Step 4:

[1276] The user selects the sports broadcast or live video they want to watch, and the device sends the selected information to the server.

[1277] Step 5:

[1278] The server receives the instruction to start viewing and prepares to obtain video data in real time in cooperation with the streaming service.

[1279] Step 6:

[1280] The server inputs the acquired video data into the generative AI model and begins video analysis, which then detects important scenes and characteristic events.

[1281] Step 7:

[1282] The server stores the analysis results in temporary storage and uses the data for subsequent analysis and editing.

[1283] Step 8:

[1284] The server collects the user's facial expressions and voice in real time through the user's camera and microphone, and inputs them into the emotion engine, which analyzes the emotion data and determines the user's current emotional state.

[1285] Step 9:

[1286] The server filters the scenes detected by the AI ​​model based on specific player and team information provided by the user, such as Lionel Messi's goal or Alice's performance.

[1287] Step 10:

[1288] The server organizes the video clips in chronological order and feeds them into editing software, which includes adding transitions, background music, and text overlays.

[1289] Step 11:

[1290] The server adjusts the highlight video composition based on the user's emotional data collected by the emotion engine, for example, prioritizing more action-packed scenes if the user is excited.

[1291] Step 12:

[1292] The server renders the final highlight video and converts it into the appropriate format, ensuring smooth playback on the user's device.

[1293] Step 13:

[1294] The server stores the link of the generated highlight video in a database and notifies the link information to the user's terminal.

[1295] Step 14:

[1296] The device notifies the user of the link to the highlight video received from the server. When the user clicks on the notification, the device starts streaming or downloading the highlight video and plays it.

[1297] These steps allow users to enjoy customized highlight videos that include scenes of specific players or team members. The introduction of the emotion engine further enhances the individual viewing experience, providing dynamic video content tailored to emotions.

[1298] Example 2

[1299] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1300] In modern games and live video viewing, it is difficult to quickly find specific scenes that users are interested in from vast amounts of video data. Furthermore, to improve the user's viewing experience, it is necessary not only to provide video content but also to customize it to match the user's real-time emotional state. However, conventional systems have had difficulty in recognizing emotions in real time and dynamically adjusting content based on those emotions.

[1301] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1302] In this invention, the server includes means for acquiring game and live video data, means for accepting information on specific players and members set by the user, means for analyzing the acquired video data to detect important scenes and characteristic events, means for filtering scenes including specific players and members from the detected scenes based on the user's settings, means for analyzing the collected user emotional data to detect the user's current emotional state, means for dynamically adjusting the generated highlight video based on the user's emotional state, means for generating a highlight video by editing the filtered scenes, and means for providing the generated highlight video to the user. This makes it possible to provide a customized highlight video that corresponds to the individual emotional state of the user.

[1303] "Means for acquiring game or live video data" refers to a function including communication means and interfaces for acquiring game or live event video data in real time or from storage.

[1304] "Means for accepting information on specific players and members set by the user" refers to an input means by which the user specifies the players they want to support or members they are interested in on the system, and a function for acquiring and saving that setting information.

[1305] "Means for analyzing acquired video data to detect important scenes and characteristic events" refers to a function that uses video analysis technology to automatically recognize characteristic actions and important scenes (such as goal scenes and performance scenes) during a match or event.

[1306] "Means for filtering scenes that include specific players or members from detected scenes based on user settings" is a filtering function for selecting only scenes related to specific players or members set by the user.

[1307] "Means for detecting the user's current emotional state by analyzing collected user emotional data" refers to a function for identifying the user's emotional state (e.g., joy, excitement, sadness, etc.) by analyzing data collected from the user's facial expressions and voice.

[1308] The "means for dynamically adjusting the generated highlight video based on the user's emotional state" is an editing function for changing the content and order of the highlight video in real time according to the analyzed emotional data of the user.

[1309] The "means for editing filtered scenes to generate a highlight video" is a function for editing multiple selected scenes into a single video, and then adding visual effects and music to create a visually consistent highlight video.

[1310] The "means for providing the generated highlight video to the user" is a function including a communication means and an interface for transmitting the highlight video to the user terminal so that the user can view it.

[1311] The present invention relates to a system for generating highlight videos from game or live footage based on user customization settings, and further adjusting the content of the videos by recognizing the user's emotions. The system includes a user terminal, a server, and an emotion engine.

[1312] Program processing and specific operations

[1313] Accepting user settings

[1314] Users launch a dedicated application on their smartphone or PC and log in to their account. Through the application, users can select the specific player or member they want to support and authorize the use of the emotion engine. The user's device sends this setting information to the server, which stores it in a database. For example, it is possible to set up the system so that users can "select their favorite basketball player and obtain highlight footage of that player."

[1315] Start watching content

[1316] When a user selects a game or live video they want to watch, the user's device sends the selection information and user settings information to the server. The server then works with the streaming service to prepare to acquire the video data. For example, a scenario could be that a user selects an artist's live performance and watches it in high definition.

[1317] Running AI models and emotion engines

[1318] The server inputs the streaming data into a generative AI model, which begins analyzing the video data and detecting key scenes. This AI model uses, for example, OpenAI's video analysis model. The server also activates an emotion engine, which collects data in real time from the user's camera and microphone. The emotion engine analyzes the user's facial expressions and voice to detect their current emotional state. For example, "if the user shows excitement during a game, that state is recorded."

[1319] Integration of favorite scenes and emotional data

[1320] The server detects important scenes from the analyzed video data and filters them based on specific players or team members set by the user. At the same time, it analyzes the collected user emotional data to detect the user's current emotional state. For example, if the user is "excited about a goal," the scene will be highlighted by filtering.

[1321] Highlight video generation and adjustment

[1322] The server then edits the filtered scenes, adding transitions, background music, and text overlays to create a smooth highlight video. This process uses video editing software such as Adobe Premiere Pro. The server also adjusts the content and order of the video based on the user's emotional data. For example, if the user is sad, it might add encouraging scenes.

[1323] Highlight video provided

[1324] The server encodes the generated highlight video into the optimal viewing format and uploads it to cloud storage. The user's device is then notified of the link to the generated video, which then receives the link and displays a notification to the user. When the user clicks on the notification, playback of the highlight video begins. For example, a scenario could be that a user plays a highlight video of their favorite idol during their morning commute.

[1325] Specific examples

[1326] For example, if a user starts watching a soccer match partway through, a highlight video will be automatically generated, including a goal scored by a specific player. If the emotion engine analyzes the user's facial expressions while watching and detects an excited state, the highlight video will be edited to include more dynamic playing scenes and cheers. Similarly, if a user starts watching a music concert partway through, a highlight video will be provided that includes performance scenes of their favorite band members. If the emotion engine detects the user's smile, moving visual effects will be added to the performance scenes.

[1327] Examples of prompt statements

[1328] "Please explain how you can automatically detect goals scored by specific players in real time while a user is watching a soccer game, and dynamically generate a highlight video based on the user's emotional state."

[1329] The system enhances the viewing experience by providing a customized highlight video that responds to the user's individual emotional state.

[1330] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1331] Step 1:

[1332] Users launch a dedicated application on their smartphone or computer and log in to their account. The information they enter is authentication information such as their user ID and password. The output is the authentication result indicating whether the user successfully logged in.

[1333] Step 2:

[1334] Within the application, users can select the specific player or member they want to support and authorize the use of the emotion engine. The information input is the information about the player or member they are supporting, and whether or not the emotion engine can be used. The output is the setting information and the confirmation result of permission to use the emotion engine. The device sends this setting information to the server, which stores it in a database. Specifically, the device's app screen displays a list of players and a checkbox to allow the emotion engine.

[1335] Step 3:

[1336] The user selects the game or live video they want to watch. The input information is the ID or URL of the game or live video they want to watch. The output is information about the selected content. The device sends this information and user setting information to the server.

[1337] Step 4:

[1338] The server calls the streaming service API and prepares to obtain game or live video data. The input information is the ID and URL of the game or live video to be viewed. The output is streaming data. Here, the server receives the video data in real time via the API.

[1339] Step 5:

[1340] The server inputs the acquired streaming data into the generative AI model and begins analyzing the video data and detecting key scenes. The input information is the streaming data. The output is a list of detected key scenes. The AI ​​model used could be, for example, OpenAI's video analysis model. Specifically, the server identifies goal scenes and performance scenes in real time.

[1341] Step 6:

[1342] The server runs the emotion engine to collect data in real time from the user's camera and microphone. The input information is the user's facial expressions and voice. The output is the user's emotional state (e.g., joy, excitement, sadness, etc.). Specifically, the engine captures a picture of the user's face with a webcam and analyzes their voice with a microphone.

[1343] Step 7:

[1344] The server filters out scenes related to specific players or members set by the user from among the important scenes detected by the generative AI model. The input information is a list of detected important scenes and user settings. The output is a list of filtered scenes. Specifically, the scenes in the list are organized based on the user's settings.

[1345] Step 8:

[1346] The server then analyzes and filters the video data, adjusts the content and order of the highlight video based on the user's emotional data, and begins editing. The input information is the list of filtered scenes and the user's emotional state. The output is an edited highlight video. Specifically, the server adds transitions, background music, and text overlays using video editing tools such as Adobe Premiere Pro API.

[1347] Step 9:

[1348] The server encodes the generated highlight video into an optimal viewing format and uploads it to cloud storage. The input information is the edited highlight video. The output is a link for viewing the video. Specifically, the encoded video is saved in cloud storage and a link is generated.

[1349] Step 10:

[1350] The device displays a link to view the highlight video to the user as a new arrival notification. The input information is the link to view the highlight video. The output is a notification that is displayed to the user. Specifically, the notification is displayed on the user's smartphone or computer, and when the user clicks the link, the video begins playing.

[1351] The above are the specific processing steps and operations of the program of this system.

[1352] (Application example 2)

[1353] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1354] Conventional highlight video generation systems were capable of generating highlight videos based on user customization settings, but it was difficult to reflect the user's emotions in real time. As a result, the viewing experience could not adequately respond to individual emotions, and visual satisfaction and excitement could not be maximized. The present invention aims to solve this problem and provide a more personalized and moving viewing experience by appropriately adjusting the content and order of highlight videos according to the user's emotions.

[1355] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1356] In this invention, the server includes means for acquiring game and live video data, means for accepting information on specific players and members set by the user, means for analyzing the acquired video data to detect important scenes and characteristic events, means for filtering scenes including specific players and members from the detected scenes based on the user's settings, means for editing the filtered scenes to generate a highlight video, means for providing the generated highlight video to the user, means for analyzing the user's face and voice in real time to detect their emotional state, and means for adjusting the content and order of the highlight video based on the detected emotional state, thereby making it possible to generate a customized highlight video that corresponds to the individual emotional state of the user.

[1357] "Means for acquiring game or live video data" refers to devices or software that include interfaces or protocols for acquiring frame data and audio data of games or live video in streaming or file format.

[1358] "Means for accepting information on specific players or members set by the user" refers to an interface or application that allows the user to select the specific players or members they want to support and register that information on a server or terminal.

[1359] "Means for analyzing acquired video data to detect important scenes and characteristic events" refers to hardware or software that analyzes video data using image analysis technology and machine learning algorithms to automatically detect important moments in matches and highlights from live broadcasts.

[1360] "Means for filtering scenes that include specific players or members from scenes detected based on user settings" refers to algorithms or logic for extracting scenes that feature specific players or members set by the user, and for identifying and filtering those scenes.

[1361] The "means for editing the filtered scenes to generate a highlight video" refers to editing software or an engine for editing the filtered scenes in chronological order, adding transitions and background music, overlaying text, and so on, to generate the final highlight video.

[1362] The "means for providing the generated highlight video to the user" refers to an interface or a server for distributing the generated highlight video in a streaming format or for providing a download link.

[1363] "Means for analyzing a user's face and voice in real time to detect their emotional state" refers to an emotion recognition engine or software that captures a user's facial expressions and voice tone in real time and analyzes that data to determine the user's emotional state (for example, joy, excitement, sadness, etc.).

[1364] The "means for adjusting the content and sequence of the highlight video based on the detected emotional state" refers to an algorithm or editing software for changing the selection and arrangement of scenes in the highlight video and adjusting visual effects and background music in accordance with the user's emotional state.

[1365] The present invention provides a system for generating a highlight video from a game or live video based on a user's customized settings, and the system adjusts the content of the video by recognizing the user's emotions. Detailed embodiments of the present invention will be described below.

[1366] The system's main components include a user device, a server, and an emotion engine. The server acquires game and live video data, accepts user-specified player and team member information, and analyzes the acquired video data. It then detects important scenes and distinctive events, filters them based on the user's settings, and generates a highlight video. The server then provides the generated highlight video to the user, adjusting the content and order of the video based on the user's emotional state.

[1367] The server uses a streaming interface or file acquisition protocol to acquire video data. User devices transmit information about specific players or team members using input methods implemented in the interface or application. The server then analyzes the video data using an AI model to detect key moments. This AI model utilizes image recognition technology and machine learning algorithms.

[1368] The emotion engine collects the user's facial and voice data in real time and uses emotion recognition APIs (such as the Azure Emotion API) to determine their emotional state. This emotional data is used to adjust the content and order of the highlight video. For example, if the user is excited, more dynamic scenes can be prioritized.

[1369] The filtered scenes are then edited using editing software such as MoviePy to add transitions, background music, text overlays, etc. The final highlight video is then distributed in streaming format or provided as a download link.

[1370] For example, a user can watch a soccer match and filter out key moments (e.g., goals and assists) of a specific player. While watching, the emotion engine analyzes the user's facial expressions and, if it detects excitement, edits the highlight video to include more dynamic play scenes and cheers.

[1371] Examples of prompts for generative AI models include:

[1372] "Generate dynamic highlight videos for excited users, including specific player goals and assists."

[1373] Examples include:

[1374] As mentioned above, by providing a customized highlight video that reflects the user's emotions, a more personal and moving viewing experience is possible.

[1375] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1376] Step 1: Accepting User Settings

[1377] The user launches the application and logs in. The device selects the specific player or member the user wants to support and accepts permission to use the emotion engine as input. This input information is sent to the server, which then stores it in a database.

[1378] Step 2: Start watching content

[1379] Users select the sports broadcast or live video they want to watch from the application interface. The device receives the selection information as input and sends it along with the user's settings information to the server. Based on this information, the server works with the streaming service to prepare for the acquisition of the video data.

[1380] Step 3: Running the AI ​​model and emotion engine

[1381] The server receives video data from the streaming service and passes it as input to the generative AI model. The generative AI model analyzes the video data to detect important scenes and distinctive events. At the same time, it activates an emotion engine, collecting the user's face and voice as input in real time. The emotion engine then analyzes the user's emotional state from this data and returns the results to the server as output.

[1382] Step 4: Integrating the favorite scenes and emotion data

[1383] The server filters the key scene data obtained from the generative AI model based on specific players and team members selected by the user. This filtered scene data is then combined with the analyzed emotional data to select appropriate scenes based on the user's emotional state at that time.

[1384] Step 5: Generate and adjust the highlight video

[1385] The server then passes the filtered scenes to editing software (such as MoviePy) to add transitions, background music, and text overlays. This editing process produces a smooth highlight video. The server then adjusts the content and sequence of scenes in the highlight video based on the user's emotional data. For example, if the user is excited, more dynamic scenes will be added and edited.

[1386] Step 6: Submit a highlight video

[1387] The server converts the generated highlight video into the optimal viewing format and generates a viewing link. The link stored in the database is sent to the user's device as a notification. The device receives the notification and displays it to the user. When the user clicks the notification, the highlight video is played.

[1388] The above are the specific processing steps for implementing the present invention.

[1389] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1390] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1391] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1392] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1393] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1394] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1395] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1396] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1397] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1398] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1399] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1400] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1401] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1402] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1403] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1404] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1405] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1406] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1407] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1408] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1409] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1410] The following is further disclosed regarding the above embodiment.

[1411] (Claim 1)

[1412] A system for generating highlight videos from games or live footage based on user customization settings,

[1413] A means of obtaining data on matches and live footage,

[1414] A means to receive information on specific players and members set by the user,

[1415] A means of analyzing the acquired video data to detect important scenes and characteristic events,

[1416] A means of filtering scenes containing specific players or members from the detected scenes based on user settings;

[1417] a means for editing the filtered scenes to generate a highlight video;

[1418] A means for providing the generated highlight video to a user;

[1419] A system including:

[1420] (Claim 2)

[1421] 2. The system according to claim 1, further comprising means for automatically generating highlights of a game or live video from the time the user starts watching until that time.

[1422] (Claim 3)

[1423] 10. The system according to claim 1, further comprising means for streaming the generated highlight video in real time.

[1424] "Example 1"

[1425] (Claim 1)

[1426] A system for generating highlight videos from games or live footage based on user customization settings,

[1427] A means of obtaining data on matches and live footage,

[1428] A means to receive information on specific players and members set by the user,

[1429] A means of analyzing the captured video data and using generative AI models to detect important scenes and characteristic events;

[1430] A means of filtering scenes containing specific players or members from the detected scenes based on user settings;

[1431] a means for using an editing tool to edit the filtered scenes to generate a highlight video;

[1432] a means for notifying and providing the generated highlight video to a user's terminal;

[1433] A system including:

[1434] (Claim 2)

[1435] 2. The system according to claim 1, further comprising means for automatically generating highlights of a game or live video from the time the user starts watching until that time.

[1436] (Claim 3)

[1437] 10. The system according to claim 1, further comprising means for streaming the generated highlight video in real time.

[1438] "Application Example 1"

[1439] (Claim 1)

[1440] A system for generating highlight videos from games or live footage based on user customization settings,

[1441] A means of obtaining data on matches and live footage,

[1442] A means to receive information on specific players and members set by the user,

[1443] A means of analyzing the acquired video data to detect important scenes and characteristic events,

[1444] A means of filtering scenes containing specific players or members from the detected scenes based on user settings;

[1445] a means for editing the filtered scenes to generate a highlight video;

[1446] A means for providing the generated highlight video to a user;

[1447] a connection means for projecting an image from a user's terminal onto a display device in the store;

[1448] A system including:

[1449] (Claim 2)

[1450] 2. The system according to claim 1, further comprising means for automatically generating highlights of a game or live video from the time the user starts watching until that time.

[1451] (Claim 3)

[1452] 10. The system according to claim 1, further comprising means for streaming the generated highlight video in real time.

[1453] "Example 2: Combining Emotion Engines"

[1454] (Claim 1)

[1455] A system for generating highlight videos from games or live footage based on user customization settings,

[1456] A means of obtaining data on matches and live footage,

[1457] A means to receive information on specific players and members set by the user,

[1458] A means of analyzing the acquired video data to detect important scenes and characteristic events,

[1459] A means of filtering scenes containing specific players or members from the detected scenes based on user settings;

[1460] means for analyzing the collected user emotion data to detect the user's current emotional state;

[1461] means for dynamically adjusting the generated highlight video based on the emotional state of the user;

[1462] a means for editing the filtered scenes to generate a highlight video;

[1463] A means for providing the generated highlight video to a user;

[1464] A system including:

[1465] (Claim 2)

[1466] 2. The system according to claim 1, further comprising means for automatically generating highlights of a game or live video from the time the user starts watching until that time.

[1467] (Claim 3)

[1468] 10. The system according to claim 1, further comprising means for streaming the generated highlight video in real time.

[1469] "Application example 2 when combining emotion engines"

[1470] (Claim 1)

[1471] A system for generating highlight videos from games or live footage based on user customization settings,

[1472] A means of obtaining data on matches and live footage,

[1473] A means to receive information on specific players and members set by the user,

[1474] A means of analyzing the acquired video data to detect important scenes and characteristic events,

[1475] A means of filtering scenes containing specific players or members from the detected scenes based on user settings;

[1476] a means for editing the filtered scenes to generate a highlight video;

[1477] A means for providing the generated highlight video to a user;

[1478] A means of analyzing the user's face and voice in real time to detect their emotional state;

[1479] means for adjusting the content and order of the highlight video based on the detected emotional state;

[1480] A system including:

[1481] (Claim 2)

[1482] 2. The system according to claim 1, further comprising means for automatically generating highlights of a game or live video from the time the user starts watching until that time.

[1483] (Claim 3)

[1484] 10. The system according to claim 1, further comprising means for streaming the generated highlight video in real time. [Explanation of symbols]

[1485] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. A system for generating highlight videos from games or live footage based on user customization settings, A means of obtaining data on matches and live footage, A means to receive information on specific players and members set by the user, A means of analyzing the acquired video data to detect important scenes and characteristic events, A means of filtering scenes containing specific players or members from the detected scenes based on user settings; a means for editing the filtered scenes to generate a highlight video; A means for providing the generated highlight video to a user; A system including:

2. The system according to claim 1, further comprising means for automatically generating highlights of a game or live video from the time the user starts viewing until that time.

3. The system according to claim 1 , further comprising means for streaming the generated highlight video in real time.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A