system

A system captures and analyzes game footage in real-time to generate personalized highlights, addressing the inefficiencies in creating gaming content by incorporating user emotions and providing immersive viewing experiences.

JP2026068354APending Publication Date: 2026-04-22SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-10
Publication Date
2026-04-22

Smart Images

  • Figure 2026068354000001_ABST
    Figure 2026068354000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] An input processing means for capturing game footage in real time, An analytical model means for analyzing the aforementioned game footage and evaluating the state of excitement, An editing and shaping means for identifying and editing moments of high excitement, Output means for exporting the edited highlights in an output format, A system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance as a response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In recent years, the number of content producers who provide game play videos as streaming or recordings has been increasing. However, there is a problem that it takes a great deal of time and effort to select exciting moments and important scenes that viewers demand from a vast amount of video data and create highlight videos. For this reason, there is a need for a mechanism that automatically selects exciting moments of a game in real time and efficiently and effectively generates highlights.

Means for Solving the Problems

[0005] This invention provides a system that captures game footage in real time and uses an analysis model to evaluate the level of excitement, thereby identifying moments of high excitement. By providing means for editing and shaping the identified moments and exporting them in an output format, the system enables the rapid generation of compelling highlights for viewers. This system allows content creators to reduce editing time and efficiently provide high-quality highlights.

[0006] "Game footage" refers to digital video data that visually represents the progress of an electronic game.

[0007] "Real-time" refers to a state characterized by a temporal feature where processing or responses occur almost immediately after data is generated.

[0008] "Capture" refers to the process of capturing data such as audio and video and recording it in digital format.

[0009] An "analytical model" refers to a mathematical or machine learning method used to analyze data and detect specific patterns or trends.

[0010] "Excitement" refers to a psychological or behavioral state in which the interest or emotions of users or viewers are strongly aroused under specific circumstances.

[0011] "Editing and shaping" is the process of selecting, cutting, rearranging, and processing imported data to create a final form that aligns with a specific purpose.

[0012] "Output format" refers to a specific structure or format for storing or displaying data.

[0013] "Export" refers to the process of saving or sending data that has been processed or generated internally in a format that can be used externally. [Brief explanation of the drawing]

[0014] [Figure 1]It is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] It is a conceptual diagram showing an example of the main functions of a data processing device and a smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which a plurality of emotions are mapped. [Figure 10] It shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when an emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when an emotion engine is combined.

MODE FOR CARRYING OUT THE INVENTION

[0015] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0016] First, the terms used in the following description will be explained.

[0017] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0018] In the following embodiments, the numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0019] In the following embodiments, the numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, etc.

[0020] In the following embodiments, the numbered communication I / F (Interface) is an interface including a communication processor and an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark), etc.

[0021] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0022] [First Embodiment]

[0023] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0024] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0025] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0026] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0027] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0028] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0029] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0030] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0031] As shown in Figure 2, in the data processing device 12, specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0032] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0033] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0034] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0035] Embodiments of this invention include a system for capturing and analyzing game footage in real time and automatically generating highlights. A server acquires game footage transmitted from a user's terminal in real time and analyzes the footage frame by frame using an internal analysis model. The analysis model detects moments that are likely to indicate an "excited state" based on changes in sound and video. This excitement state is characterized by significant changes in sound, visually impactful movements, and the occurrence of scores or specific events.

[0036] The moments of excitement identified by the analysis model are sent to an editing and shaping module on the server. This module rearranges the moments with high excitement ratings along a timeline, adds effects as needed, and shapes them into highlights with a narrative and entertainment value.

[0037] The generated highlights are exported in the output format specified by the server and uploaded to an online platform or saved to the user's device, based on the user's instructions. The user can also review, view, or further edit the generated highlights via their device.

[0038] As a concrete example, consider playing an online shooting game. When a user performs excellent maneuvers and achieves a high score, the server detects the sudden changes in sound and visual movements at that moment and identifies it as an excited state. Subsequently, an editing and shaping module combines this moment with other important moments to create a highlight and outputs it in a format that can be quickly shared according to the user's intentions. This entire process allows users to easily generate engaging content and deliver it to their audience without spending a lot of time. This system is expected to be in high demand from content creators because it improves content quality while saving time and effort.

[0039] The following describes the processing flow.

[0040] Step 1:

[0041] The server receives game footage from the user's terminal in real time. The video data is taken into the server in stream format and stored in a buffer to mark the starting point for necessary analysis.

[0042] Step 2:

[0043] The server divides the received video into frames and prepares them, along with audio data, for transmission to the analysis model. This preparation includes formatting and synchronizing the time axis.

[0044] Step 3:

[0045] The server's analysis model analyzes the visual impact and changes in sound within the game footage. The model detects sudden changes and specific events, and assigns an excitement level score to each frame.

[0046] Step 4:

[0047] Based on the analysis results, the server identifies moments that were highly rated as exciting and selects them as highlight candidates. The selected scenes are then passed to the editing and shaping module.

[0048] Step 5:

[0049] The server's editing and shaping module rearranges the selected moments along a timeline and adds transitions and effects between the footage as needed. This creates a highlight reel with continuity and a visually flowing narrative.

[0050] Step 6:

[0051] The server exports the final highlights to the specified output format and prepares them to be sent to the user's terminal or uploaded to an online platform.

[0052] Step 7:

[0053] Users can review the generated highlights through their device and make additional edits as needed. After editing, users can either publish the content on the platform or save it locally.

[0054] (Example 1)

[0055] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0056] In modern information processing, automatically extracting specific states of excitement from vast amounts of data and organizing and editing them in a meaningful way is crucial. However, conventional methods have made real-time analysis and effective editing difficult, hindering efficient content creation. Therefore, there is a need to establish a system that can quickly and effectively analyze information and deliver moments of excitement in an engaging format.

[0057] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0058] In this invention, the server includes data input means for acquiring information in real time, evaluation model means for analyzing the information and detecting the state of excitement, and time series formation means for organizing moments of high excitement and adding effects. This makes it possible to automatically extract moments of excitement from a vast amount of data, instantly format them, and provide them in an attractive format.

[0059] "Data input means" refers to a method or device for acquiring information in real time.

[0060] An "evaluation model means" is an algorithm or model used to detect the state of excitement based on acquired information.

[0061] A "time-series formation means" is a method or apparatus for organizing moments of high excitement and shaping information by adding visual or auditory effects.

[0062] "Format conversion means" refers to a method or apparatus for converting organized and formatted information into a specified format and providing it.

[0063] An "excited state" is a moment of heightened emotion, identified based on changes in sound or vision.

[0064] An "analysis method" is a technical means of evaluating the state of excitement based on the acoustic and visual information contained in the information.

[0065] "Rhythm and narrative" are elements that add a consistent tempo and story to edited information.

[0066] This invention features a system for acquiring, analyzing, editing, and providing information in real time. A server acquires information generated by a user via a terminal in real time using data input means. This information primarily includes game video and audio data. The server's evaluation model means uses a generation AI model to analyze the acquired information frame by frame and detect states of excitement. This analysis employs spectral analysis of acoustic information and motion detection algorithms for visual information. The analyzed data is organized by a time-series formation means, and visual and audio effects are added as needed. This generates engaging highlights that include moments based on states of excitement. Finally, a format conversion means outputs the generated highlights in a specified format, allowing the user to view or download them.

[0067] As a concrete example, when a user achieves a significant result in an online game, the server automatically detects and processes that moment, providing it as a viewable highlight. For instance, a prompt message such as "Generate highlights focusing on the moment the user achieved a high score" can be input to the AI ​​generation model. This system enables real-time data processing and efficient content generation, significantly streamlining users' content creation activities.

[0068] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0069] Step 1:

[0070] Users play games and other content using their devices, generating video and audio data in real time. The devices stream this data and send it to the server. The input is the video and audio being played, and the output is the data stream to the server.

[0071] Step 2:

[0072] The server uses a data input means to receive real-time video and audio data transmitted from the terminal. The received data is then passed directly to the evaluation model means. The input is a real-time data stream from the terminal, and the output is prepared data used for analysis.

[0073] Step 3:

[0074] The server uses an evaluation model to analyze the received video and audio data. Specifically, a generative AI model detects changes in the audio spectrum and motion within the video to identify moments indicating excitement. The input is the received data, and the output is the timestamp and metadata of the identified excitement state.

[0075] Step 4:

[0076] The server uses a time-series generation mechanism to organize identified moments of excitement on a timeline and generates highlights by adding visual and audio effects. Effects such as slow motion and text overlays are added as needed. The input is the timestamps and metadata of the excitement states, and the output is the edited highlight segments.

[0077] Step 5:

[0078] The server exports the edited highlights in a predetermined file format using a format conversion mechanism. The output format is determined by the user and may be, for example, MP4 or GIF. The input is the edited highlight segment, and the output is a format-converted file.

[0079] Step 6:

[0080] Users can view, watch, or further edit the highlights generated via the terminal. The terminal receives files from the server and provides the user with an interface for viewing or editing. The input is a formatted file, and the output, if the user makes further edits, is a further processed highlight.

[0081] (Application Example 1)

[0082] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0083] Efficiently recording and editing exciting moments during gameplay presents a challenge: manually extracting specific moments from vast amounts of video data is time-consuming and labor-intensive. Furthermore, instantly sharing the generated footage across various platforms and connecting with viewers is not easy with traditional methods. Therefore, there is a need for a system that efficiently generates and distributes game video highlights in real time.

[0084] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0085] In this invention, the server includes data input means for acquiring game footage in real time, data analysis means for analyzing the game footage and evaluating emotional states, visualization and editing means for identifying and editing moments of high emotional states, interface means for displaying and operating on an information terminal, and information sharing means for sharing over a communication network. This enables users to efficiently and quickly generate gameplay highlights and share them across various platforms.

[0086] "Data input means" is a general term for devices and methods used to acquire game footage in real time.

[0087] "Data analysis means" refers to devices or methods used to analyze acquired video data and evaluate specific emotional states.

[0088] A "visualization editing means" is a device or method for editing identified important moments and outputting them in a visually organized format.

[0089] "Data output means" refers to a device or method for exporting visualized and edited visual information in a specified format.

[0090] An "interface means" is a device or method that displays visual information on an information terminal and enables the user to operate it.

[0091] "Information sharing means" refers to devices or methods for sharing visual information generated through a communication network with other platforms or users.

[0092] The system for implementing this invention acquires, analyzes, and edits game footage in real time to automatically visualize moments that excite the user. The system mainly consists of a server, user terminals, and a communication network connecting them. Each component of this system is realized using the following specific hardware and software.

[0093] The server uses a high-performance processing unit to capture game footage in real time as a data input method. This processing uses video capture and data streaming software, specifically OpenCV, to capture frame-by-frame video data. For real-time analysis, a data analysis method is used to determine emotional states using audio and video data in order to accurately evaluate the immersion of the game. PyDub is used to analyze audio information and OpenCV is used to analyze video information, extracting moments when excitement is maximized.

[0094] As a visualization and editing method, FFmpeg is used to edit the extracted moments of high excitement into video, applying processing to give them a specific tempo and narrative structure. The edited visual information is transferred to the user's terminal via a data output means and processed so that it can be displayed and operated using the terminal's interface. This interface provides a graphical operating environment for the user to edit and modify the video.

[0095] Furthermore, the information sharing mechanism allows users to directly distribute the generated highlights via communication networks on online platforms and social media, enabling them to share them with other users. Through this system, users can easily share not only surprising moments but also enjoyable experiences with many people.

[0096] As a concrete example, if a user is playing an action-packed shooting game and a decisive moment of victory occurs, the system will reliably capture that moment, detect the audio volume and visual stimuli, and edit it into a highlight. This process flow occurs in real time, and the generated highlight can be shared immediately after gameplay.

[0097] An example of a prompt for a generative AI model is: "I want to create a highlight reel of the most exciting moments from the final battle scene of an online game. What actions would maximize the excitement?"

[0098] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0099] Step 1:

[0100] Server-based video data capture

[0101] The server captures game video data in real time via a capture device during the user's gameplay and saves it frame by frame. The input is the user's game video, and the output is a set of video frames suitable for analysis. For video streaming, OpenCV is used to process each frame.

[0102] Step 2:

[0103] Analysis of audio and video data

[0104] The server uses an analysis model to analyze the captured video and audio data and evaluate specific moments that indicate states of excitement. This process uses frame-by-frame audio and video data as input and generates metadata as output that indicates moments of high excitement. Audio data is analyzed using PyDub to identify significant changes in volume and frequency, and dynamic changes in the video are analyzed using OpenCV.

[0105] Step 3:

[0106] Identification and editing of moments of excitement

[0107] Based on the analysis results, the server rearranges the video of the identified exciting moments and edits it to create a sense of tempo and narrative. The input is the metadata and video frames obtained in step 2, and the output is a highlight video. FFmpeg is used for video editing, and effects and transitions are added.

[0108] Step 4:

[0109] Export edited highlights

[0110] The server outputs the highlights generated by the visualization and editing system in the specified format. The input is the edited highlight video, and the output is a file in a format that users can view or share. It is converted to a format usable on various devices.

[0111] Step 5:

[0112] Display and operation on the user terminal

[0113] Users view highlights on their devices and perform further editing and customization as needed. The input is an exported highlight video file, and the output is the final, shareable video. The user interface allows for intuitive editing through graphical controls.

[0114] Step 6:

[0115] Sharing highlights

[0116] Users share their edited highlights via a communication network through their device to online platforms and social media. The input is the final edited video file, and the output is shared content viewable on various platforms. This step allows users to widely showcase the exciting moments they have created using a generative AI model.

[0117] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0118] This embodiment of the invention is a system that combines an analysis model that captures game footage in real time and evaluates the state of excitement with an emotion engine that recognizes the user's emotions. The server first receives game footage from the user's terminal and prepares it for analysis frame by frame. The analysis model individually evaluates and scores the state of excitement based on changes in movement and sound within the game.

[0119] A unique element of this system is the use of an emotion engine. The server analyzes the user's facial expressions and voice tone in real time and generates emotion data. This data is integrated to more accurately assess the user's excitement level in the game. Furthermore, the emotion data allows the editing module to generate more personalized highlights. For example, if the system detects the user's surprised facial expression or excited voice tone, this is reflected in the highlight, and editing is performed to visually emphasize the emotion of that moment.

[0120] The generated highlights are exported in the specified format by the server's output module. Users can quickly review the generated highlights through their terminal and perform further editing and optimization if necessary. Finally, users can share these highlights on an online platform or save them locally.

[0121] As a concrete example, consider a scenario in a multiplayer game where users cooperate with their team to overcome a difficult challenge. The emotion engine captures the user's joy and excitement, and based on that, highlights in the video are automatically edited. This allows viewers not only to see moments from the game, but also to feel the user's emotional reactions and share in that authentic experience. This system aims to enhance the immersion of game streams and improve the visual and emotional appeal of the content.

[0122] The following describes the processing flow.

[0123] Step 1:

[0124] The server receives game video and audio from the user's terminal in real time and captures them as a stream. The data is stored in a buffer and kept ready for processing at any time.

[0125] Step 2:

[0126] The server divides the video data into frames and extracts the audio data individually. This data is then passed to an analysis model, which is prepared to evaluate the state of excitement.

[0127] Step 3:

[0128] The server's analysis model evaluates the excitement level based on changes in in-game events and audio. An excitement score is assigned to each frame, and moments with particularly high scores are selected for the next process.

[0129] Step 4:

[0130] The server activates an emotion engine to detect the user's facial expressions and voice tone in real time. This allows the system to evaluate the user's emotions and reflect them in the game's excitement level assessment.

[0131] Step 5:

[0132] By combining analysis results and emotional data, the server's editing and shaping module identifies moments of high excitement and generates highlights. Video editing is then performed with a specific tempo and a narrative structure based on the user's emotions.

[0133] Step 6:

[0134] The server exports the generated highlights in the specified output format and uploads them to an online platform or sends them to the user's device, according to the user's instructions.

[0135] Step 7:

[0136] Users can review the generated highlights through their devices and make additional edits or final adjustments as needed. Highlights can be published instantly, enhancing the user's visual and emotional experience.

[0137] (Example 2)

[0138] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0139] Traditional game content highlight generation systems often failed to accurately reflect user emotions because they evaluated excitement levels purely based on in-game movement and sound changes. As a result, the appeal and immersion of the content conveyed to viewers were limited. Furthermore, generating personalized highlights that responded to individual user emotional responses was difficult.

[0140] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0141] In this invention, the server includes data receiving means for acquiring game content in real time, data analysis means for analyzing the game content and evaluating the user's excitement level, and emotion recognition means for generating user emotion data. This makes it possible to integrate the user's emotions and excitement level to edit the content and provide viewers with more emotionally rich and personalized highlights.

[0142] "Game content" refers to interactive digital media that runs on electronic devices and is manipulated by users.

[0143] "Data receiving means" refers to a device or method for acquiring external signals or information in an appropriate format.

[0144] "Data analysis means" refers to software or hardware used to process received information and identify specific patterns or states.

[0145] "Emotion recognition means" refers to algorithms and systems that analyze a user's facial expressions, voice, etc., to estimate their emotional state.

[0146] "Information formation means" refers to a process or function that generates new information based on analyzed data and assembles it in a specific format.

[0147] "Data output means" refers to a device or method for transmitting processed information to an external source in a specified format.

[0148] "Excitement" refers to a state in which the user experiences psychological or physiological arousal or tension.

[0149] "Emotional data" refers to information collected and analyzed to represent a user's internal emotional state.

[0150] A "specific format" refers to a standardized shape or structure that information must follow when it is output.

[0151] An embodiment of this invention is a video editing system that combines real-time capture of game content with user sentiment analysis. The server acquires data using a common communication protocol to receive game content from the user's terminal. This data is organized as video frames and converted into a format that is easy to analyze using a conversion tool such as FFmpeg.

[0152] The server then utilizes machine learning libraries to analyze these frames. As a data analysis tool, it uses AI models such as TENSORFLOW® to evaluate game movements and audio information, quantifying the user's excitement level. Simultaneously, as an emotion recognition tool, it analyzes the user's facial expressions and voice tone using OpenCV and voice analysis software to generate emotion data. In this way, the system quantifies the user's psychological response.

[0153] This data is integrated and edited by information shaping tools to generate highlights. This process matches peaks of excitement with heightened emotions, creating content that reflects the user's experience. For example, it might be edited to highlight moments when a user experiences great joy after winning a game.

[0154] The server ultimately exports the generated highlights in standard formats such as MP4, making them easily accessible to users. This system also allows users to further customize the content using digital editing tools such as Adobe Premiere.

[0155] As a concrete example, consider a scenario in an online game where a user achieves a dramatic come-from-behind victory. In this case, the server identifies that moment, integrates data on excitement levels and emotion recognition, and generates a highlight that allows viewers to share in the user's inner feelings. A concrete example of a prompt to input into the generating AI model would be, "Identify the moment in this game footage when the user is most excited, and edit the scene with the strongest emotion into a highlight."

[0156] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0157] Step 1:

[0158] The server receives streaming data of game content from the user's terminal. The input is the video and audio data of the game, transmitted in real time. Upon receiving this data, the server uses a frame buffer to divide the stream into individual video frames. The output is a list of processable video frames. Specifically, the server uses the RTMP protocol to receive data efficiently.

[0159] Step 2:

[0160] The server applies data analysis tools to analyze the segmented video frames and audio data. The input is the video frames and audio data generated in step 1. An AI model using the TensorFlow library is used for the analysis, evaluating the excitement level based on the detection of movement in the game and changes in sound. The output is the excitement score for each frame. Specifically, features are extracted for each frame, and the state is quantified through a pre-trained AI model.

[0161] Step 3:

[0162] The server analyzes the user's facial expressions and voice to generate emotion data. The input consists of the user's webcam video and audio stream from the microphone. OpenCV is used to analyze facial expressions, and speech recognition software is used to evaluate the tone of the voice. The output is a dataset representing the user's emotional state. Specifically, emotion estimation is performed using facial recognition and voice pitch analysis.

[0163] Step 4:

[0164] The server integrates the analyzed excitement score and emotion data using an information-forming mechanism to edit customized highlights. The inputs are the excitement score from step 2 and the emotion data from step 3. Using this information, emotionally significant frames matching the excitement peaks are selected, and edit points are determined. The output is the final edited highlight video. Specific operations include calculating edit points and performing cut editing.

[0165] Step 5:

[0166] The server exports the edited highlights in a format based on user settings. The input is the highlight video created in step 4. An encoding tool is used to convert it to the user-specified format (e.g., MP4). The output is a video file in the final format that the user can easily play and share. Specifically, the server sets the encoding parameters and performs the encoding process with an optimal balance.

[0167] (Application Example 2)

[0168] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0169] A challenge with typical game streaming is that it's difficult for viewers to fully grasp the streamer's emotions and moments of excitement. Therefore, there's a need for a way to allow viewers to experience the streamer's emotions in real time and achieve a more immersive viewing experience.

[0170] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0171] In this invention, the server includes acquisition means for acquiring game information in real time, evaluation model means for analyzing the game information and evaluating the state of excitement, and emotion analysis means for generating and integrating emotion information based on the evaluation data. As a result, viewers can experience the excitement and emotional changes of the commentator in real time, enabling a more emotional and immersive viewing experience.

[0172] The "acquisition means" refers to a mechanism that receives game information in real time and prepares it for processing.

[0173] The "evaluation model means" is a mechanism that uses an algorithm to analyze acquired game information and quantify and evaluate the user's state of excitement.

[0174] An "emotion analysis tool" is a mechanism that recognizes emotions based on the user's facial expressions and voice data, and generates that data.

[0175] "Editing and shaping means" refers to a mechanism that identifies moments of high excitement based on data from evaluation model means and emotion analysis means, and edits them in a way that is easily understood by the viewer.

[0176] "Output method" refers to a mechanism for exporting edited highlights in a specified output format.

[0177] A "visualization means" is a mechanism for visually displaying generated emotional information in real time and conveying those emotions to the viewer.

[0178] A "conversion mechanism" is a device that converts edited highlights into a format that can be shared on online platforms.

[0179] The system for realizing this invention consists of a server and a user's client terminal. First, the user's terminal acquires game information in real time and transmits it to the server. On the server, the acquisition means is responsible for receiving the game information. The game information is analyzed by the evaluation model means, and the user's excitement level is numerically evaluated. This evaluation uses algorithms based on audio and video information.

[0180] Next, the emotion analysis means analyzes the user's facial expressions and vocal characteristics to generate emotion information. This data is integrated with evaluation data, and the editing and shaping means identifies moments of high excitement. These moments are edited by the visualization means to visually represent emotions in real time. The completed highlights can be exported in a specified format through the output means. The conversion means also formats the exported highlights so that they can be shared on an online platform.

[0181] Users can experience game footage that visualizes the streamer's emotions through devices such as VR-compatible headsets. This allows viewers to realistically feel the streamer's experience. For example, consider a scene in a game where the user overcomes a difficult challenge and feels both surprise and joy simultaneously. Their facial expressions and voice at that moment are collected as emotional data and edited in a way that conveys this to the viewer.

[0182] An example of a prompt using a generative AI model might be: "I want to create a VR experience that visualizes the emotions of a game streamer in real time and conveys them to the viewers. Please give me ideas on how to visualize exciting moments using Unity and VR technology."

[0183] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0184] Step 1:

[0185] The terminal acquires the user's game information in real time and sends it to the server. This input data includes game video frames and audio data. The terminal compresses this data and processes it to send it to the server with low latency.

[0186] Step 2:

[0187] The server analyzes the game information received using the acquisition means. In this process, game video and audio data are taken in and passed to the evaluation model means. This model means uses a generative AI model to quantify the user's excitement level for each frame and outputs it as a score.

[0188] Step 3:

[0189] The server's emotion analysis system captures the user's facial expression data and voice characteristics. Specifically, this step analyzes the facial image data and voice tone sent from the terminal to generate emotion data. The output emotion data is then integrated with evaluation data as numerical parameters.

[0190] Step 4:

[0191] The server's editing and shaping mechanism integrates evaluation data and emotional data to identify moments of high excitement. These identified moments are then extracted as video tracks and compiled for editing. Based on the integrated data, cuts are made for visualization.

[0192] Step 5:

[0193] The visualization means visually expresses the user's emotions based on the cut footage provided by the editing and shaping means. In this step, a specific tempo and narrative are added, and effects are added to enhance the visual impact. The final output is video data provided to the viewer in real time.

[0194] Step 6:

[0195] The output method exports the final highlights in the specified format. This process converts the completed video data into a shareable format such as MP4 or WebM and exports it to a file for easy access by the user.

[0196] Step 7:

[0197] The conversion tool converts the exported highlights into an appropriate format so that they can be shared on online platforms. Finally, it provides users with links or files that allow them to view, share, and save these highlights.

[0198] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0199] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0200] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0201] [Second Embodiment]

[0202] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0203] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0204] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0205] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0206] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0207] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0208] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0209] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0210] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0211] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0212] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0213] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0214] Embodiments of this invention include a system for capturing and analyzing game footage in real time and automatically generating highlights. A server acquires game footage transmitted from a user's terminal in real time and analyzes the footage frame by frame using an internal analysis model. The analysis model detects moments that are likely to indicate an "excited state" based on changes in sound and video. This excitement state is characterized by significant changes in sound, visually impactful movements, and the occurrence of scores or specific events.

[0215] The moments of excitement identified by the analysis model are sent to an editing and shaping module on the server. This module rearranges the moments with high excitement ratings along a timeline, adds effects as needed, and shapes them into highlights with a narrative and entertainment value.

[0216] The generated highlights are exported in the output format specified by the server and uploaded to an online platform or saved to the user's device, based on the user's instructions. The user can also review, view, or further edit the generated highlights via their device.

[0217] As a concrete example, consider playing an online shooting game. When a user performs excellent maneuvers and achieves a high score, the server detects the sudden changes in sound and visual movements at that moment and identifies it as an excited state. Subsequently, an editing and shaping module combines this moment with other important moments to create a highlight and outputs it in a format that can be quickly shared according to the user's intentions. This entire process allows users to easily generate engaging content and deliver it to their audience without spending a lot of time. This system is expected to be in high demand from content creators because it improves content quality while saving time and effort.

[0218] The following describes the processing flow.

[0219] Step 1:

[0220] The server receives game footage from the user's terminal in real time. The video data is taken into the server in stream format and stored in a buffer to mark the starting point for necessary analysis.

[0221] Step 2:

[0222] The server divides the received video into frames and prepares them, along with audio data, for transmission to the analysis model. This preparation includes formatting and synchronizing the time axis.

[0223] Step 3:

[0224] The server's analysis model analyzes the visual impact and changes in sound within the game footage. The model detects sudden changes and specific events, and assigns an excitement level score to each frame.

[0225] Step 4:

[0226] Based on the analysis results, the server identifies moments that were highly rated as exciting and selects them as highlight candidates. The selected scenes are then passed to the editing and shaping module.

[0227] Step 5:

[0228] The server's editing and shaping module rearranges the selected moments along a timeline and adds transitions and effects between the footage as needed. This creates a highlight reel with continuity and a visually flowing narrative.

[0229] Step 6:

[0230] The server exports the final highlights to the specified output format and prepares them to be sent to the user's terminal or uploaded to an online platform.

[0231] Step 7:

[0232] Users can review the generated highlights through their device and make additional edits as needed. After editing, users can either publish the content on the platform or save it locally.

[0233] (Example 1)

[0234] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0235] In modern information processing, automatically extracting specific states of excitement from vast amounts of data and organizing and editing them in a meaningful way is crucial. However, conventional methods have made real-time analysis and effective editing difficult, hindering efficient content creation. Therefore, there is a need to establish a system that can quickly and effectively analyze information and deliver moments of excitement in an engaging format.

[0236] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0237] In this invention, the server includes data input means for acquiring information in real time, evaluation model means for analyzing the information and detecting the state of excitement, and time series formation means for organizing moments of high excitement and adding effects. This makes it possible to automatically extract moments of excitement from a vast amount of data, instantly format them, and provide them in an attractive format.

[0238] "Data input means" refers to a method or device for acquiring information in real time.

[0239] An "evaluation model means" is an algorithm or model used to detect the state of excitement based on acquired information.

[0240] A "time-series formation means" is a method or apparatus for organizing moments of high excitement and shaping information by adding visual or auditory effects.

[0241] "Format conversion means" refers to a method or apparatus for converting organized and formatted information into a specified format and providing it.

[0242] An "excited state" is a moment of heightened emotion, identified based on changes in sound or vision.

[0243] An "analysis method" is a technical means of evaluating the state of excitement based on the acoustic and visual information contained in the information.

[0244] "Rhythm and narrative" are elements that add a consistent tempo and story to edited information.

[0245] This invention features a system for acquiring, analyzing, editing, and providing information in real time. A server acquires information generated by a user via a terminal in real time using data input means. This information primarily includes game video and audio data. The server's evaluation model means uses a generation AI model to analyze the acquired information frame by frame and detect states of excitement. This analysis employs spectral analysis of acoustic information and motion detection algorithms for visual information. The analyzed data is organized by a time-series formation means, and visual and audio effects are added as needed. This generates engaging highlights that include moments based on states of excitement. Finally, a format conversion means outputs the generated highlights in a specified format, allowing the user to view or download them.

[0246] As a concrete example, when a user achieves a significant result in an online game, the server automatically detects and processes that moment, providing it as a viewable highlight. For instance, a prompt message such as "Generate highlights focusing on the moment the user achieved a high score" can be input to the AI ​​generation model. This system enables real-time data processing and efficient content generation, significantly streamlining users' content creation activities.

[0247] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0248] Step 1:

[0249] Users play games and other content using their devices, generating video and audio data in real time. The devices stream this data and send it to the server. The input is the video and audio being played, and the output is the data stream to the server.

[0250] Step 2:

[0251] The server uses a data input means to receive real-time video and audio data transmitted from the terminal. The received data is then passed directly to the evaluation model means. The input is a real-time data stream from the terminal, and the output is prepared data used for analysis.

[0252] Step 3:

[0253] The server uses an evaluation model to analyze the received video and audio data. Specifically, a generative AI model detects changes in the audio spectrum and motion within the video to identify moments indicating excitement. The input is the received data, and the output is the timestamp and metadata of the identified excitement state.

[0254] Step 4:

[0255] The server uses a time-series generation mechanism to organize identified moments of excitement on a timeline and generates highlights by adding visual and audio effects. Effects such as slow motion and text overlays are added as needed. The input is the timestamps and metadata of the excitement states, and the output is the edited highlight segments.

[0256] Step 5:

[0257] The server exports the edited highlights in a predetermined file format using a format conversion mechanism. The output format is determined by the user and may be, for example, MP4 or GIF. The input is the edited highlight segment, and the output is a format-converted file.

[0258] Step 6:

[0259] Users can view, watch, or further edit the highlights generated via the terminal. The terminal receives files from the server and provides the user with an interface for viewing or editing. The input is a formatted file, and the output, if the user makes further edits, is a further processed highlight.

[0260] (Application Example 1)

[0261] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0262] Efficiently recording and editing exciting moments during gameplay presents a challenge: manually extracting specific moments from vast amounts of video data is time-consuming and labor-intensive. Furthermore, instantly sharing the generated footage across various platforms and connecting with viewers is not easy with traditional methods. Therefore, there is a need for a system that efficiently generates and distributes game video highlights in real time.

[0263] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0264] In this invention, the server includes data input means for acquiring game footage in real time, data analysis means for analyzing the game footage and evaluating emotional states, visualization and editing means for identifying and editing moments of high emotional states, interface means for displaying and operating on an information terminal, and information sharing means for sharing over a communication network. This enables users to efficiently and quickly generate gameplay highlights and share them across various platforms.

[0265] "Data input means" is a general term for devices and methods used to acquire game footage in real time.

[0266] "Data analysis means" refers to devices or methods used to analyze acquired video data and evaluate specific emotional states.

[0267] A "visualization editing means" is a device or method for editing identified important moments and outputting them in a visually organized format.

[0268] "Data output means" refers to a device or method for exporting visualized and edited visual information in a specified format.

[0269] An "interface means" is a device or method that displays visual information on an information terminal and enables the user to operate it.

[0270] "Information sharing means" refers to devices or methods for sharing visual information generated through a communication network with other platforms or users.

[0271] The system for implementing this invention acquires, analyzes, and edits game footage in real time to automatically visualize moments that excite the user. The system mainly consists of a server, user terminals, and a communication network connecting them. Each component of this system is realized using the following specific hardware and software.

[0272] The server uses a high-performance processing unit to capture game footage in real time as a data input method. This processing uses video capture and data streaming software, specifically OpenCV, to capture frame-by-frame video data. For real-time analysis, a data analysis method is used to determine emotional states using audio and video data in order to accurately evaluate the immersion of the game. PyDub is used to analyze audio information and OpenCV is used to analyze video information, extracting moments when excitement is maximized.

[0273] As a visualization and editing method, FFmpeg is used to edit the extracted moments of high excitement into video, applying processing to give them a specific tempo and narrative structure. The edited visual information is transferred to the user's terminal via a data output means and processed so that it can be displayed and operated using the terminal's interface. This interface provides a graphical operating environment for the user to edit and modify the video.

[0274] Furthermore, the information sharing mechanism allows users to directly distribute the generated highlights via communication networks on online platforms and social media, enabling them to share them with other users. Through this system, users can easily share not only surprising moments but also enjoyable experiences with many people.

[0275] As a concrete example, if a user is playing an action-packed shooting game and a decisive moment of victory occurs, the system will reliably capture that moment, detect the audio volume and visual stimuli, and edit it into a highlight. This process flow occurs in real time, and the generated highlight can be shared immediately after gameplay.

[0276] An example of a prompt for a generative AI model is: "I want to create a highlight reel of the most exciting moments from the final battle scene of an online game. What actions would maximize the excitement?"

[0277] The flow of the specific process in Application Example 1 will be described using FIG. 12.

[0278] Step 1:

[0279] Capture of video data by the server

[0280] During the user's game play, the server captures the game video data in real time via a capture device and saves it frame by frame. The input is the user's game video, and the output is a set of video frames suitable for analysis. For video streaming utilization, each frame is processed using OpenCV.

[0281] Step 2:

[0282] Analysis of audio and video data

[0283] The server uses an analysis model to analyze the captured video and audio data and evaluate specific moments indicating an excited state. In this process, frame-by-frame audio and video data are used as input, and metadata indicating moments with a high excited state is generated as output. The audio data is analyzed using PyDub to identify significant changes in volume and frequency, and the dynamic changes in the video are analyzed using OpenCV.

[0284] Step 3:

[0285] [[ID=�3]] Identification and editing of moments of excited state

[0286] Based on the analysis results, the server rearranges the videos of the identified excited moments and performs editing with tempo and storytelling. The input is the metadata and video frames obtained in Step 2, and the output is a highlight video. FFmpeg is used for video editing to add effects and transitions.

[0287] Step 4:

[0288] Export edited highlights

[0289] The server outputs the highlights generated by the visualization and editing system in the specified format. The input is the edited highlight video, and the output is a file in a format that users can view or share. It is converted to a format usable on various devices.

[0290] Step 5:

[0291] Display and operation on the user terminal

[0292] Users view highlights on their devices and perform further editing and customization as needed. The input is an exported highlight video file, and the output is the final, shareable video. The user interface allows for intuitive editing through graphical controls.

[0293] Step 6:

[0294] Sharing highlights

[0295] Users share their edited highlights via a communication network through their device to online platforms and social media. The input is the final edited video file, and the output is shared content viewable on various platforms. This step allows users to widely showcase the exciting moments they have created using a generative AI model.

[0296] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0297] This embodiment of the invention is a system that combines an analysis model that captures game footage in real time and evaluates the state of excitement with an emotion engine that recognizes the user's emotions. The server first receives game footage from the user's terminal and prepares it for analysis frame by frame. The analysis model individually evaluates and scores the state of excitement based on changes in movement and sound within the game.

[0298] A unique element of this system is the use of an emotion engine. The server analyzes the user's facial expressions and voice tone in real time and generates emotion data. This data is integrated to more accurately assess the user's excitement level in the game. Furthermore, the emotion data allows the editing module to generate more personalized highlights. For example, if the system detects the user's surprised facial expression or excited voice tone, this is reflected in the highlight, and editing is performed to visually emphasize the emotion of that moment.

[0299] The generated highlights are exported in the specified format by the server's output module. Users can quickly review the generated highlights through their terminal and perform further editing and optimization if necessary. Finally, users can share these highlights on an online platform or save them locally.

[0300] As a concrete example, consider a scenario in a multiplayer game where users cooperate with their team to overcome a difficult challenge. The emotion engine captures the user's joy and excitement, and based on that, highlights in the video are automatically edited. This allows viewers not only to see moments from the game, but also to feel the user's emotional reactions and share in that authentic experience. This system aims to enhance the immersion of game streams and improve the visual and emotional appeal of the content.

[0301] The following describes the processing flow.

[0302] Step 1:

[0303] The server receives game video and audio in real time from the user's terminal and captures them as a stream. The data is stored in a buffer and kept in a state where it can be processed at any time.

[0304] Step 2:

[0305] The server divides the video data into frames and also extracts the audio data separately. These data are passed to an analysis model to prepare for evaluating the excitement state.

[0306] Step 3:

[0307] The server's analysis model evaluates the excitement state based on changes in in-game events and audio changes. An excitement score is assigned to each frame, and the moments with particularly high scores are selected for the next process.

[0308] Step 4:

[0309] The server activates the emotion engine and detects the user's facial expressions and voice tones in real time. Thereby, the user's emotion is evaluated and reflected in the evaluation of the game's excitement state.

[0310] Step 5:

[0311] The server combines the analysis results and emotion data, and the editing and formation module of the server identifies the moments with a high excitement state and generates highlights. Video editing with a story nature based on a specific tempo and the user's emotion is performed.

[0312] Step 6:

[0313] The server exports the generated highlights in a specified output format and uploads them to an online platform according to the user's instructions or sends them to the user's terminal.

[0314] Step 7:

[0315] Users can review the generated highlights through their devices and make additional edits or final adjustments as needed. Highlights can be published instantly, enhancing the user's visual and emotional experience.

[0316] (Example 2)

[0317] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0318] Traditional game content highlight generation systems often failed to accurately reflect user emotions because they evaluated excitement levels purely based on in-game movement and sound changes. As a result, the appeal and immersion of the content conveyed to viewers were limited. Furthermore, generating personalized highlights that responded to individual user emotional responses was difficult.

[0319] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0320] In this invention, the server includes data receiving means for acquiring game content in real time, data analysis means for analyzing the game content and evaluating the user's excitement level, and emotion recognition means for generating user emotion data. This makes it possible to integrate the user's emotions and excitement level to edit the content and provide viewers with more emotionally rich and personalized highlights.

[0321] "Game content" refers to interactive digital media that runs on electronic devices and is manipulated by users.

[0322] "Data receiving means" refers to a device or method for acquiring external signals or information in an appropriate format.

[0323] "Data analysis means" refers to software or hardware used to process received information and identify specific patterns or states.

[0324] "Emotion recognition means" refers to algorithms and systems that analyze a user's facial expressions, voice, etc., to estimate their emotional state.

[0325] "Information formation means" refers to a process or function that generates new information based on analyzed data and assembles it in a specific format.

[0326] "Data output means" refers to a device or method for transmitting processed information to an external source in a specified format.

[0327] "Excitement" refers to a state in which the user experiences psychological or physiological arousal or tension.

[0328] "Emotional data" refers to information collected and analyzed to represent a user's internal emotional state.

[0329] A "specific format" refers to a standardized shape or structure that information must follow when it is output.

[0330] An embodiment of this invention is a video editing system that combines real-time capture of game content with user sentiment analysis. The server acquires data using a common communication protocol to receive game content from the user's terminal. This data is organized as video frames and converted into a format that is easy to analyze using a conversion tool such as FFmpeg.

[0331] The server then leverages machine learning libraries to analyze these frames. As a data analysis tool, it uses AI models like TensorFlow to evaluate game movements and audio information, quantifying the user's excitement level. Simultaneously, as an emotion recognition tool, it analyzes the user's facial expressions and voice tone using OpenCV and voice analysis software to generate emotion data. This allows the system to quantify the user's psychological response.

[0332] This data is integrated and edited by information shaping tools to generate highlights. This process matches peaks of excitement with heightened emotions, creating content that reflects the user's experience. For example, it might be edited to highlight moments when a user experiences great joy after winning a game.

[0333] The server ultimately exports the generated highlights in standard formats such as MP4, making them easily accessible to users. This system also allows users to further customize the content using digital editing tools such as Adobe Premiere.

[0334] As a concrete example, consider a scenario in an online game where a user achieves a dramatic come-from-behind victory. In this case, the server identifies that moment, integrates data on excitement levels and emotion recognition, and generates a highlight that allows viewers to share in the user's inner feelings. A concrete example of a prompt to input into the generating AI model would be, "Identify the moment in this game footage when the user is most excited, and edit the scene with the strongest emotion into a highlight."

[0335] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0336] Step 1:

[0337] The server receives streaming data of game content from the user's terminal. The input is the video and audio data of the game, transmitted in real time. Upon receiving this data, the server uses a frame buffer to divide the stream into individual video frames. The output is a list of processable video frames. Specifically, the server uses the RTMP protocol to receive data efficiently.

[0338] Step 2:

[0339] The server applies data analysis tools to analyze the segmented video frames and audio data. The input is the video frames and audio data generated in step 1. An AI model using the TensorFlow library is used for the analysis, evaluating the excitement level based on the detection of movement in the game and changes in sound. The output is the excitement score for each frame. Specifically, features are extracted for each frame, and the state is quantified through a pre-trained AI model.

[0340] Step 3:

[0341] The server analyzes the user's facial expressions and voice to generate emotion data. The input consists of the user's webcam video and audio stream from the microphone. OpenCV is used to analyze facial expressions, and speech recognition software is used to evaluate the tone of the voice. The output is a dataset representing the user's emotional state. Specifically, emotion estimation is performed using facial recognition and voice pitch analysis.

[0342] Step 4:

[0343] The server integrates the analyzed excitement score and emotion data using an information-forming mechanism to edit customized highlights. The inputs are the excitement score from step 2 and the emotion data from step 3. Using this information, emotionally significant frames matching the excitement peaks are selected, and edit points are determined. The output is the final edited highlight video. Specific operations include calculating edit points and performing cut editing.

[0344] Step 5:

[0345] The server exports the edited highlights in a format based on user settings. The input is the highlight video created in step 4. An encoding tool is used to convert it to the user-specified format (e.g., MP4). The output is a video file in the final format that the user can easily play and share. Specifically, the server sets the encoding parameters and performs the encoding process with an optimal balance.

[0346] (Application Example 2)

[0347] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0348] A challenge with typical game streaming is that it's difficult for viewers to fully grasp the streamer's emotions and moments of excitement. Therefore, there's a need for a way to allow viewers to experience the streamer's emotions in real time and achieve a more immersive viewing experience.

[0349] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0350] In this invention, the server includes acquisition means for acquiring game information in real time, evaluation model means for analyzing the game information and evaluating the state of excitement, and emotion analysis means for generating and integrating emotion information based on the evaluation data. As a result, viewers can experience the excitement and emotional changes of the commentator in real time, enabling a more emotional and immersive viewing experience.

[0351] The "acquisition means" refers to a mechanism that receives game information in real time and prepares it for processing.

[0352] The "evaluation model means" is a mechanism that uses an algorithm to analyze acquired game information and quantify and evaluate the user's state of excitement.

[0353] An "emotion analysis tool" is a mechanism that recognizes emotions based on the user's facial expressions and voice data, and generates that data.

[0354] "Editing and shaping means" refers to a mechanism that identifies moments of high excitement based on data from evaluation model means and emotion analysis means, and edits them in a way that is easily understood by the viewer.

[0355] "Output method" refers to a mechanism for exporting edited highlights in a specified output format.

[0356] A "visualization means" is a mechanism for visually displaying generated emotional information in real time and conveying those emotions to the viewer.

[0357] A "conversion mechanism" is a device that converts edited highlights into a format that can be shared on online platforms.

[0358] The system for realizing this invention consists of a server and a user's client terminal. First, the user's terminal acquires game information in real time and transmits it to the server. On the server, the acquisition means is responsible for receiving the game information. The game information is analyzed by the evaluation model means, and the user's excitement level is numerically evaluated. This evaluation uses algorithms based on audio and video information.

[0359] Next, the emotion analysis means analyzes the user's facial expressions and vocal characteristics to generate emotion information. This data is integrated with evaluation data, and the editing and shaping means identifies moments of high excitement. These moments are edited by the visualization means to visually represent emotions in real time. The completed highlights can be exported in a specified format through the output means. The conversion means also formats the exported highlights so that they can be shared on an online platform.

[0360] Users can experience game footage that visualizes the streamer's emotions through devices such as VR-compatible headsets. This allows viewers to realistically feel the streamer's experience. For example, consider a scene in a game where the user overcomes a difficult challenge and feels both surprise and joy simultaneously. Their facial expressions and voice at that moment are collected as emotional data and edited in a way that conveys this to the viewer.

[0361] An example of a prompt using a generative AI model might be: "I want to create a VR experience that visualizes the emotions of a game streamer in real time and conveys them to the viewers. Please give me ideas on how to visualize exciting moments using Unity and VR technology."

[0362] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0363] Step 1:

[0364] The terminal acquires the user's game information in real time and sends it to the server. This input data includes game video frames and audio data. The terminal compresses this data and processes it to send it to the server with low latency.

[0365] Step 2:

[0366] The server analyzes the game information received using the acquisition means. In this process, game video and audio data are taken in and passed to the evaluation model means. This model means uses a generative AI model to quantify the user's excitement level for each frame and outputs it as a score.

[0367] Step 3:

[0368] The server's emotion analysis system captures the user's facial expression data and voice characteristics. Specifically, this step analyzes the facial image data and voice tone sent from the terminal to generate emotion data. The output emotion data is then integrated with evaluation data as numerical parameters.

[0369] Step 4:

[0370] The server's editing and shaping mechanism integrates evaluation data and emotional data to identify moments of high excitement. These identified moments are then extracted as video tracks and compiled for editing. Based on the integrated data, cuts are made for visualization.

[0371] Step 5:

[0372] The visualization means visually expresses the user's emotions based on the cut footage provided by the editing and shaping means. In this step, a specific tempo and narrative are added, and effects are added to enhance the visual impact. The final output is video data provided to the viewer in real time.

[0373] Step 6:

[0374] The output method exports the final highlights in the specified format. This process converts the completed video data into a shareable format such as MP4 or WebM and exports it to a file for easy access by the user.

[0375] Step 7:

[0376] The conversion tool converts the exported highlights into an appropriate format so that they can be shared on online platforms. Finally, it provides users with links or files that allow them to view, share, and save these highlights.

[0377] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0378] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0379] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0380] [Third Embodiment]

[0381] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0382] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0383] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0384] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0385] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0386] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0387] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0388] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0389] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0390] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0391] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0392] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0393] Embodiments of this invention include a system for capturing and analyzing game footage in real time and automatically generating highlights. A server acquires game footage transmitted from a user's terminal in real time and analyzes the footage frame by frame using an internal analysis model. The analysis model detects moments that are likely to indicate an "excited state" based on changes in sound and video. This excitement state is characterized by significant changes in sound, visually impactful movements, and the occurrence of scores or specific events.

[0394] The moments of excitement identified by the analysis model are sent to an editing and shaping module on the server. This module rearranges the moments with high excitement ratings along a timeline, adds effects as needed, and shapes them into highlights with a narrative and entertainment value.

[0395] The generated highlights are exported in the output format specified by the server and uploaded to an online platform or saved to the user's device, based on the user's instructions. The user can also review, view, or further edit the generated highlights via their device.

[0396] As a concrete example, consider playing an online shooting game. When a user performs excellent maneuvers and achieves a high score, the server detects the sudden changes in sound and visual movements at that moment and identifies it as an excited state. Subsequently, an editing and shaping module combines this moment with other important moments to create a highlight and outputs it in a format that can be quickly shared according to the user's intentions. This entire process allows users to easily generate engaging content and deliver it to their audience without spending a lot of time. This system is expected to be in high demand from content creators because it improves content quality while saving time and effort.

[0397] The following describes the processing flow.

[0398] Step 1:

[0399] The server receives game footage from the user's terminal in real time. The video data is taken into the server in stream format and stored in a buffer to mark the starting point for necessary analysis.

[0400] Step 2:

[0401] The server divides the received video into frames and prepares them, along with audio data, for transmission to the analysis model. This preparation includes formatting and synchronizing the time axis.

[0402] Step 3:

[0403] The server's analysis model analyzes the visual impact and changes in sound within the game footage. The model detects sudden changes and specific events, and assigns an excitement level score to each frame.

[0404] Step 4:

[0405] Based on the analysis results, the server identifies moments that were highly rated as exciting and selects them as highlight candidates. The selected scenes are then passed to the editing and shaping module.

[0406] Step 5:

[0407] The server's editing and shaping module rearranges the selected moments along a timeline and adds transitions and effects between the footage as needed. This creates a highlight reel with continuity and a visually flowing narrative.

[0408] Step 6:

[0409] The server exports the final highlights to the specified output format and prepares them to be sent to the user's terminal or uploaded to an online platform.

[0410] Step 7:

[0411] Users can review the generated highlights through their device and make additional edits as needed. After editing, users can either publish the content on the platform or save it locally.

[0412] (Example 1)

[0413] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0414] In modern information processing, automatically extracting specific states of excitement from vast amounts of data and organizing and editing them in a meaningful way is crucial. However, conventional methods have made real-time analysis and effective editing difficult, hindering efficient content creation. Therefore, there is a need to establish a system that can quickly and effectively analyze information and deliver moments of excitement in an engaging format.

[0415] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0416] In this invention, the server includes data input means for acquiring information in real time, evaluation model means for analyzing the information and detecting the state of excitement, and time series formation means for organizing moments of high excitement and adding effects. This makes it possible to automatically extract moments of excitement from a vast amount of data, instantly format them, and provide them in an attractive format.

[0417] "Data input means" refers to a method or device for acquiring information in real time.

[0418] An "evaluation model means" is an algorithm or model used to detect the state of excitement based on acquired information.

[0419] A "time-series formation means" is a method or apparatus for organizing moments of high excitement and shaping information by adding visual or auditory effects.

[0420] "Format conversion means" refers to a method or apparatus for converting organized and formatted information into a specified format and providing it.

[0421] An "excited state" is a moment of heightened emotion, identified based on changes in sound or vision.

[0422] An "analysis method" is a technical means of evaluating the state of excitement based on the acoustic and visual information contained in the information.

[0423] "Rhythm and narrative" are elements that add a consistent tempo and story to edited information.

[0424] This invention features a system for acquiring, analyzing, editing, and providing information in real time. A server acquires information generated by a user via a terminal in real time using data input means. This information primarily includes game video and audio data. The server's evaluation model means uses a generation AI model to analyze the acquired information frame by frame and detect states of excitement. This analysis employs spectral analysis of acoustic information and motion detection algorithms for visual information. The analyzed data is organized by a time-series formation means, and visual and audio effects are added as needed. This generates engaging highlights that include moments based on states of excitement. Finally, a format conversion means outputs the generated highlights in a specified format, allowing the user to view or download them.

[0425] As a concrete example, when a user achieves a significant result in an online game, the server automatically detects and processes that moment, providing it as a viewable highlight. For instance, a prompt message such as "Generate highlights focusing on the moment the user achieved a high score" can be input to the AI ​​generation model. This system enables real-time data processing and efficient content generation, significantly streamlining users' content creation activities.

[0426] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0427] Step 1:

[0428] Users play games and other content using their devices, generating video and audio data in real time. The devices stream this data and send it to the server. The input is the video and audio being played, and the output is the data stream to the server.

[0429] Step 2:

[0430] The server uses a data input means to receive real-time video and audio data transmitted from the terminal. The received data is then passed directly to the evaluation model means. The input is a real-time data stream from the terminal, and the output is prepared data used for analysis.

[0431] Step 3:

[0432] The server uses an evaluation model to analyze the received video and audio data. Specifically, a generative AI model detects changes in the audio spectrum and motion within the video to identify moments indicating excitement. The input is the received data, and the output is the timestamp and metadata of the identified excitement state.

[0433] Step 4:

[0434] The server uses a time-series generation mechanism to organize identified moments of excitement on a timeline and generates highlights by adding visual and audio effects. Effects such as slow motion and text overlays are added as needed. The input is the timestamps and metadata of the excitement states, and the output is the edited highlight segments.

[0435] Step 5:

[0436] The server exports the edited highlights in a predetermined file format using a format conversion mechanism. The output format is determined by the user and may be, for example, MP4 or GIF. The input is the edited highlight segment, and the output is a format-converted file.

[0437] Step 6:

[0438] Users can view, watch, or further edit the highlights generated via the terminal. The terminal receives files from the server and provides the user with an interface for viewing or editing. The input is a formatted file, and the output, if the user makes further edits, is a further processed highlight.

[0439] (Application Example 1)

[0440] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0441] Efficiently recording and editing exciting moments during gameplay presents a challenge: manually extracting specific moments from vast amounts of video data is time-consuming and labor-intensive. Furthermore, instantly sharing the generated footage across various platforms and connecting with viewers is not easy with traditional methods. Therefore, there is a need for a system that efficiently generates and distributes game video highlights in real time.

[0442] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0443] In this invention, the server includes data input means for acquiring game footage in real time, data analysis means for analyzing the game footage and evaluating emotional states, visualization and editing means for identifying and editing moments of high emotional states, interface means for displaying and operating on an information terminal, and information sharing means for sharing over a communication network. This enables users to efficiently and quickly generate gameplay highlights and share them across various platforms.

[0444] "Data input means" is a general term for devices and methods used to acquire game footage in real time.

[0445] "Data analysis means" refers to devices or methods used to analyze acquired video data and evaluate specific emotional states.

[0446] A "visualization editing means" is a device or method for editing identified important moments and outputting them in a visually organized format.

[0447] "Data output means" refers to a device or method for exporting visualized and edited visual information in a specified format.

[0448] An "interface means" is a device or method that displays visual information on an information terminal and enables the user to operate it.

[0449] "Information sharing means" refers to devices or methods for sharing visual information generated through a communication network with other platforms or users.

[0450] The system for implementing this invention acquires, analyzes, and edits game footage in real time to automatically visualize moments that excite the user. The system mainly consists of a server, user terminals, and a communication network connecting them. Each component of this system is realized using the following specific hardware and software.

[0451] The server uses a high-performance processing unit to capture game footage in real time as a data input method. This processing uses video capture and data streaming software, specifically OpenCV, to capture frame-by-frame video data. For real-time analysis, a data analysis method is used to determine emotional states using audio and video data in order to accurately evaluate the immersion of the game. PyDub is used to analyze audio information and OpenCV is used to analyze video information, extracting moments when excitement is maximized.

[0452] As a visualization and editing method, FFmpeg is used to edit the extracted moments of high excitement into video, applying processing to give them a specific tempo and narrative structure. The edited visual information is transferred to the user's terminal via a data output means and processed so that it can be displayed and operated using the terminal's interface. This interface provides a graphical operating environment for the user to edit and modify the video.

[0453] Furthermore, the information sharing mechanism allows users to directly distribute the generated highlights via communication networks on online platforms and social media, enabling them to share them with other users. Through this system, users can easily share not only surprising moments but also enjoyable experiences with many people.

[0454] As a concrete example, if a user is playing an action-packed shooting game and a decisive moment of victory occurs, the system will reliably capture that moment, detect the audio volume and visual stimuli, and edit it into a highlight. This process flow occurs in real time, and the generated highlight can be shared immediately after gameplay.

[0455] An example of a prompt for a generative AI model is: "I want to create a highlight reel of the most exciting moments from the final battle scene of an online game. What actions would maximize the excitement?"

[0456] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0457] Step 1:

[0458] Server-based video data capture

[0459] The server captures game video data in real time via a capture device during the user's gameplay and saves it frame by frame. The input is the user's game video, and the output is a set of video frames suitable for analysis. For video streaming, OpenCV is used to process each frame.

[0460] Step 2:

[0461] Analysis of audio and video data

[0462] The server uses an analysis model to analyze the captured video and audio data and evaluate specific moments that indicate states of excitement. This process uses frame-by-frame audio and video data as input and generates metadata as output that indicates moments of high excitement. Audio data is analyzed using PyDub to identify significant changes in volume and frequency, and dynamic changes in the video are analyzed using OpenCV.

[0463] Step 3:

[0464] Identification and editing of moments of excitement

[0465] Based on the analysis results, the server rearranges the video of the identified exciting moments and edits it to create a sense of tempo and narrative. The input is the metadata and video frames obtained in step 2, and the output is a highlight video. FFmpeg is used for video editing, and effects and transitions are added.

[0466] Step 4:

[0467] Export edited highlights

[0468] The server outputs the highlights generated by the visualization and editing system in the specified format. The input is the edited highlight video, and the output is a file in a format that users can view or share. It is converted to a format usable on various devices.

[0469] Step 5:

[0470] Display and operation on the user terminal

[0471] Users view highlights on their devices and perform further editing and customization as needed. The input is an exported highlight video file, and the output is the final, shareable video. The user interface allows for intuitive editing through graphical controls.

[0472] Step 6:

[0473] Sharing highlights

[0474] Users share their edited highlights via a communication network through their device to online platforms and social media. The input is the final edited video file, and the output is shared content viewable on various platforms. This step allows users to widely showcase the exciting moments they have created using a generative AI model.

[0475] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0476] This embodiment of the invention is a system that combines an analysis model that captures game footage in real time and evaluates the state of excitement with an emotion engine that recognizes the user's emotions. The server first receives game footage from the user's terminal and prepares it for analysis frame by frame. The analysis model individually evaluates and scores the state of excitement based on changes in movement and sound within the game.

[0477] A unique element of this system is the use of an emotion engine. The server analyzes the user's facial expressions and voice tone in real time and generates emotion data. This data is integrated to more accurately assess the user's excitement level in the game. Furthermore, the emotion data allows the editing module to generate more personalized highlights. For example, if the system detects the user's surprised facial expression or excited voice tone, this is reflected in the highlight, and editing is performed to visually emphasize the emotion of that moment.

[0478] The generated highlights are exported in the specified format by the server's output module. Users can quickly review the generated highlights through their terminal and perform further editing and optimization if necessary. Finally, users can share these highlights on an online platform or save them locally.

[0479] As a concrete example, consider a scenario in a multiplayer game where users cooperate with their team to overcome a difficult challenge. The emotion engine captures the user's joy and excitement, and based on that, highlights in the video are automatically edited. This allows viewers not only to see moments from the game, but also to feel the user's emotional reactions and share in that authentic experience. This system aims to enhance the immersion of game streams and improve the visual and emotional appeal of the content.

[0480] The following describes the processing flow.

[0481] Step 1:

[0482] The server receives game video and audio from the user's terminal in real time and captures them as a stream. The data is stored in a buffer and kept ready for processing at any time.

[0483] Step 2:

[0484] The server divides the video data into frames and extracts the audio data individually. This data is then passed to an analysis model, which is prepared to evaluate the state of excitement.

[0485] Step 3:

[0486] The server's analysis model evaluates the excitement level based on changes in in-game events and audio. An excitement score is assigned to each frame, and moments with particularly high scores are selected for the next process.

[0487] Step 4:

[0488] The server activates an emotion engine to detect the user's facial expressions and voice tone in real time. This allows the system to evaluate the user's emotions and reflect them in the game's excitement level assessment.

[0489] Step 5:

[0490] By combining analysis results and emotional data, the server's editing and shaping module identifies moments of high excitement and generates highlights. Video editing is then performed with a specific tempo and a narrative structure based on the user's emotions.

[0491] Step 6:

[0492] The server exports the generated highlights in the specified output format and uploads them to an online platform or sends them to the user's device, according to the user's instructions.

[0493] Step 7:

[0494] Users can review the generated highlights through their devices and make additional edits or final adjustments as needed. Highlights can be published instantly, enhancing the user's visual and emotional experience.

[0495] (Example 2)

[0496] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0497] Traditional game content highlight generation systems often failed to accurately reflect user emotions because they evaluated excitement levels purely based on in-game movement and sound changes. As a result, the appeal and immersion of the content conveyed to viewers were limited. Furthermore, generating personalized highlights that responded to individual user emotional responses was difficult.

[0498] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0499] In this invention, the server includes data receiving means for acquiring game content in real time, data analysis means for analyzing the game content and evaluating the user's excitement level, and emotion recognition means for generating user emotion data. This makes it possible to integrate the user's emotions and excitement level to edit the content and provide viewers with more emotionally rich and personalized highlights.

[0500] "Game content" refers to interactive digital media that runs on electronic devices and is manipulated by users.

[0501] "Data receiving means" refers to a device or method for acquiring external signals or information in an appropriate format.

[0502] "Data analysis means" refers to software or hardware used to process received information and identify specific patterns or states.

[0503] "Emotion recognition means" refers to algorithms and systems that analyze a user's facial expressions, voice, etc., to estimate their emotional state.

[0504] "Information formation means" refers to a process or function that generates new information based on analyzed data and assembles it in a specific format.

[0505] "Data output means" refers to a device or method for transmitting processed information to an external source in a specified format.

[0506] "Excitement" refers to a state in which the user experiences psychological or physiological arousal or tension.

[0507] "Emotional data" refers to information collected and analyzed to represent a user's internal emotional state.

[0508] A "specific format" refers to a standardized shape or structure that information must follow when it is output.

[0509] An embodiment of this invention is a video editing system that combines real-time capture of game content with user sentiment analysis. The server acquires data using a common communication protocol to receive game content from the user's terminal. This data is organized as video frames and converted into a format that is easy to analyze using a conversion tool such as FFmpeg.

[0510] The server then leverages machine learning libraries to analyze these frames. As a data analysis tool, it uses AI models like TensorFlow to evaluate game movements and audio information, quantifying the user's excitement level. Simultaneously, as an emotion recognition tool, it analyzes the user's facial expressions and voice tone using OpenCV and voice analysis software to generate emotion data. This allows the system to quantify the user's psychological response.

[0511] This data is integrated and edited by information shaping tools to generate highlights. This process matches peaks of excitement with heightened emotions, creating content that reflects the user's experience. For example, it might be edited to highlight moments when a user experiences great joy after winning a game.

[0512] The server ultimately exports the generated highlights in standard formats such as MP4, making them easily accessible to users. This system also allows users to further customize the content using digital editing tools such as Adobe Premiere.

[0513] As a concrete example, consider a scenario in an online game where a user achieves a dramatic come-from-behind victory. In this case, the server identifies that moment, integrates data on excitement levels and emotion recognition, and generates a highlight that allows viewers to share in the user's inner feelings. A concrete example of a prompt to input into the generating AI model would be, "Identify the moment in this game footage when the user is most excited, and edit the scene with the strongest emotion into a highlight."

[0514] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0515] Step 1:

[0516] The server receives streaming data of game content from the user's terminal. The input is the video and audio data of the game, transmitted in real time. Upon receiving this data, the server uses a frame buffer to divide the stream into individual video frames. The output is a list of processable video frames. Specifically, the server uses the RTMP protocol to receive data efficiently.

[0517] Step 2:

[0518] The server applies data analysis tools to analyze the segmented video frames and audio data. The input is the video frames and audio data generated in step 1. An AI model using the TensorFlow library is used for the analysis, evaluating the excitement level based on the detection of movement in the game and changes in sound. The output is the excitement score for each frame. Specifically, features are extracted for each frame, and the state is quantified through a pre-trained AI model.

[0519] Step 3:

[0520] The server analyzes the user's facial expressions and voice to generate emotion data. The input consists of the user's webcam video and audio stream from the microphone. OpenCV is used to analyze facial expressions, and speech recognition software is used to evaluate the tone of the voice. The output is a dataset representing the user's emotional state. Specifically, emotion estimation is performed using facial recognition and voice pitch analysis.

[0521] Step 4:

[0522] The server integrates the analyzed excitement score and emotion data using an information-forming mechanism to edit customized highlights. The inputs are the excitement score from step 2 and the emotion data from step 3. Using this information, emotionally significant frames matching the excitement peaks are selected, and edit points are determined. The output is the final edited highlight video. Specific operations include calculating edit points and performing cut editing.

[0523] Step 5:

[0524] The server exports the edited highlights in a format based on user settings. The input is the highlight video created in step 4. An encoding tool is used to convert it to the user-specified format (e.g., MP4). The output is a video file in the final format that the user can easily play and share. Specifically, the server sets the encoding parameters and performs the encoding process with an optimal balance.

[0525] (Application Example 2)

[0526] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0527] A challenge with typical game streaming is that it's difficult for viewers to fully grasp the streamer's emotions and moments of excitement. Therefore, there's a need for a way to allow viewers to experience the streamer's emotions in real time and achieve a more immersive viewing experience.

[0528] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0529] In this invention, the server includes acquisition means for acquiring game information in real time, evaluation model means for analyzing the game information and evaluating the state of excitement, and emotion analysis means for generating and integrating emotion information based on the evaluation data. As a result, viewers can experience the excitement and emotional changes of the commentator in real time, enabling a more emotional and immersive viewing experience.

[0530] The "acquisition means" refers to a mechanism that receives game information in real time and prepares it for processing.

[0531] The "evaluation model means" is a mechanism that uses an algorithm to analyze acquired game information and quantify and evaluate the user's state of excitement.

[0532] An "emotion analysis tool" is a mechanism that recognizes emotions based on the user's facial expressions and voice data, and generates that data.

[0533] "Editing and shaping means" refers to a mechanism that identifies moments of high excitement based on data from evaluation model means and emotion analysis means, and edits them in a way that is easily understood by the viewer.

[0534] "Output method" refers to a mechanism for exporting edited highlights in a specified output format.

[0535] A "visualization means" is a mechanism for visually displaying generated emotional information in real time and conveying those emotions to the viewer.

[0536] A "conversion mechanism" is a device that converts edited highlights into a format that can be shared on online platforms.

[0537] The system for realizing this invention consists of a server and a user's client terminal. First, the user's terminal acquires game information in real time and transmits it to the server. On the server, the acquisition means is responsible for receiving the game information. The game information is analyzed by the evaluation model means, and the user's excitement level is numerically evaluated. This evaluation uses algorithms based on audio and video information.

[0538] Next, the emotion analysis means analyzes the user's facial expressions and vocal characteristics to generate emotion information. This data is integrated with evaluation data, and the editing and shaping means identifies moments of high excitement. These moments are edited by the visualization means to visually represent emotions in real time. The completed highlights can be exported in a specified format through the output means. The conversion means also formats the exported highlights so that they can be shared on an online platform.

[0539] Users can experience game footage that visualizes the streamer's emotions through devices such as VR-compatible headsets. This allows viewers to realistically feel the streamer's experience. For example, consider a scene in a game where the user overcomes a difficult challenge and feels both surprise and joy simultaneously. Their facial expressions and voice at that moment are collected as emotional data and edited in a way that conveys this to the viewer.

[0540] An example of a prompt using a generative AI model might be: "I want to create a VR experience that visualizes the emotions of a game streamer in real time and conveys them to the viewers. Please give me ideas on how to visualize exciting moments using Unity and VR technology."

[0541] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0542] Step 1:

[0543] The terminal acquires the user's game information in real time and sends it to the server. This input data includes game video frames and audio data. The terminal compresses this data and processes it to send it to the server with low latency.

[0544] Step 2:

[0545] The server analyzes the game information received using the acquisition means. In this process, game video and audio data are taken in and passed to the evaluation model means. This model means uses a generative AI model to quantify the user's excitement level for each frame and outputs it as a score.

[0546] Step 3:

[0547] The server's emotion analysis system captures the user's facial expression data and voice characteristics. Specifically, this step analyzes the facial image data and voice tone sent from the terminal to generate emotion data. The output emotion data is then integrated with evaluation data as numerical parameters.

[0548] Step 4:

[0549] The server's editing and shaping mechanism integrates evaluation data and emotional data to identify moments of high excitement. These identified moments are then extracted as video tracks and compiled for editing. Based on the integrated data, cuts are made for visualization.

[0550] Step 5:

[0551] The visualization means visually expresses the user's emotions based on the cut footage provided by the editing and shaping means. In this step, a specific tempo and narrative are added, and effects are added to enhance the visual impact. The final output is video data provided to the viewer in real time.

[0552] Step 6:

[0553] The output method exports the final highlights in the specified format. This process converts the completed video data into a shareable format such as MP4 or WebM and exports it to a file for easy access by the user.

[0554] Step 7:

[0555] The conversion tool converts the exported highlights into an appropriate format so that they can be shared on online platforms. Finally, it provides users with links or files that allow them to view, share, and save these highlights.

[0556] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0557] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0558] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0559] [Fourth Embodiment]

[0560] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0561] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0562] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0563] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0564] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0565] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0566] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0567] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0568] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0569] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0570] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0571] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0572] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0573] Embodiments of this invention include a system for capturing and analyzing game footage in real time and automatically generating highlights. A server acquires game footage transmitted from a user's terminal in real time and analyzes the footage frame by frame using an internal analysis model. The analysis model detects moments that are likely to indicate an "excited state" based on changes in sound and video. This excitement state is characterized by significant changes in sound, visually impactful movements, and the occurrence of scores or specific events.

[0574] The moments of excitement identified by the analysis model are sent to an editing and shaping module on the server. This module rearranges the moments with high excitement ratings along a timeline, adds effects as needed, and shapes them into highlights with a narrative and entertainment value.

[0575] The generated highlights are exported in the output format specified by the server and uploaded to an online platform or saved to the user's device, based on the user's instructions. The user can also review, view, or further edit the generated highlights via their device.

[0576] As a concrete example, consider playing an online shooting game. When a user performs excellent maneuvers and achieves a high score, the server detects the sudden changes in sound and visual movements at that moment and identifies it as an excited state. Subsequently, an editing and shaping module combines this moment with other important moments to create a highlight and outputs it in a format that can be quickly shared according to the user's intentions. This entire process allows users to easily generate engaging content and deliver it to their audience without spending a lot of time. This system is expected to be in high demand from content creators because it improves content quality while saving time and effort.

[0577] The following describes the processing flow.

[0578] Step 1:

[0579] The server receives game footage from the user's terminal in real time. The video data is taken into the server in stream format and stored in a buffer to mark the starting point for necessary analysis.

[0580] Step 2:

[0581] The server divides the received video into frames and prepares them, along with audio data, for transmission to the analysis model. This preparation includes formatting and synchronizing the time axis.

[0582] Step 3:

[0583] The server's analysis model analyzes the visual impact and changes in sound within the game footage. The model detects sudden changes and specific events, and assigns an excitement level score to each frame.

[0584] Step 4:

[0585] Based on the analysis results, the server identifies moments that were highly rated as exciting and selects them as highlight candidates. The selected scenes are then passed to the editing and shaping module.

[0586] Step 5:

[0587] The server's editing and shaping module rearranges the selected moments along a timeline and adds transitions and effects between the footage as needed. This creates a highlight reel with continuity and a visually flowing narrative.

[0588] Step 6:

[0589] The server exports the final highlights to the specified output format and prepares them to be sent to the user's terminal or uploaded to an online platform.

[0590] Step 7:

[0591] Users can review the generated highlights through their device and make additional edits as needed. After editing, users can either publish the content on the platform or save it locally.

[0592] (Example 1)

[0593] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0594] In modern information processing, automatically extracting specific states of excitement from vast amounts of data and organizing and editing them in a meaningful way is crucial. However, conventional methods have made real-time analysis and effective editing difficult, hindering efficient content creation. Therefore, there is a need to establish a system that can quickly and effectively analyze information and deliver moments of excitement in an engaging format.

[0595] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0596] In this invention, the server includes data input means for acquiring information in real time, evaluation model means for analyzing the information and detecting the state of excitement, and time series formation means for organizing moments of high excitement and adding effects. This makes it possible to automatically extract moments of excitement from a vast amount of data, instantly format them, and provide them in an attractive format.

[0597] "Data input means" refers to a method or device for acquiring information in real time.

[0598] An "evaluation model means" is an algorithm or model used to detect the state of excitement based on acquired information.

[0599] A "time-series formation means" is a method or apparatus for organizing moments of high excitement and shaping information by adding visual or auditory effects.

[0600] "Format conversion means" refers to a method or apparatus for converting organized and formatted information into a specified format and providing it.

[0601] An "excited state" is a moment of heightened emotion, identified based on changes in sound or vision.

[0602] An "analysis method" is a technical means of evaluating the state of excitement based on the acoustic and visual information contained in the information.

[0603] "Rhythm and narrative" are elements that add a consistent tempo and story to edited information.

[0604] This invention features a system for acquiring, analyzing, editing, and providing information in real time. A server acquires information generated by a user via a terminal in real time using data input means. This information primarily includes game video and audio data. The server's evaluation model means uses a generation AI model to analyze the acquired information frame by frame and detect states of excitement. This analysis employs spectral analysis of acoustic information and motion detection algorithms for visual information. The analyzed data is organized by a time-series formation means, and visual and audio effects are added as needed. This generates engaging highlights that include moments based on states of excitement. Finally, a format conversion means outputs the generated highlights in a specified format, allowing the user to view or download them.

[0605] As a concrete example, when a user achieves a significant result in an online game, the server automatically detects and processes that moment, providing it as a viewable highlight. For instance, a prompt message such as "Generate highlights focusing on the moment the user achieved a high score" can be input to the AI ​​generation model. This system enables real-time data processing and efficient content generation, significantly streamlining users' content creation activities.

[0606] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0607] Step 1:

[0608] Users play games and other content using their devices, generating video and audio data in real time. The devices stream this data and send it to the server. The input is the video and audio being played, and the output is the data stream to the server.

[0609] Step 2:

[0610] The server uses a data input means to receive real-time video and audio data transmitted from the terminal. The received data is then passed directly to the evaluation model means. The input is a real-time data stream from the terminal, and the output is prepared data used for analysis.

[0611] Step 3:

[0612] The server uses an evaluation model to analyze the received video and audio data. Specifically, a generative AI model detects changes in the audio spectrum and motion within the video to identify moments indicating excitement. The input is the received data, and the output is the timestamp and metadata of the identified excitement state.

[0613] Step 4:

[0614] The server uses a time-series generation mechanism to organize identified moments of excitement on a timeline and generates highlights by adding visual and audio effects. Effects such as slow motion and text overlays are added as needed. The input is the timestamps and metadata of the excitement states, and the output is the edited highlight segments.

[0615] Step 5:

[0616] The server exports the edited highlights in a predetermined file format using a format conversion mechanism. The output format is determined by the user and may be, for example, MP4 or GIF. The input is the edited highlight segment, and the output is a format-converted file.

[0617] Step 6:

[0618] Users can view, watch, or further edit the highlights generated via the terminal. The terminal receives files from the server and provides the user with an interface for viewing or editing. The input is a formatted file, and the output, if the user makes further edits, is a further processed highlight.

[0619] (Application Example 1)

[0620] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0621] Efficiently recording and editing exciting moments during gameplay presents a challenge: manually extracting specific moments from vast amounts of video data is time-consuming and labor-intensive. Furthermore, instantly sharing the generated footage across various platforms and connecting with viewers is not easy with traditional methods. Therefore, there is a need for a system that efficiently generates and distributes game video highlights in real time.

[0622] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0623] In this invention, the server includes data input means for acquiring game footage in real time, data analysis means for analyzing the game footage and evaluating emotional states, visualization and editing means for identifying and editing moments of high emotional states, interface means for displaying and operating on an information terminal, and information sharing means for sharing over a communication network. This enables users to efficiently and quickly generate gameplay highlights and share them across various platforms.

[0624] "Data input means" is a general term for devices and methods used to acquire game footage in real time.

[0625] "Data analysis means" refers to devices or methods used to analyze acquired video data and evaluate specific emotional states.

[0626] A "visualization editing means" is a device or method for editing identified important moments and outputting them in a visually organized format.

[0627] "Data output means" refers to a device or method for exporting visualized and edited visual information in a specified format.

[0628] An "interface means" is a device or method that displays visual information on an information terminal and enables the user to operate it.

[0629] "Information sharing means" refers to devices or methods for sharing visual information generated through a communication network with other platforms or users.

[0630] The system for implementing this invention acquires, analyzes, and edits game footage in real time to automatically visualize moments that excite the user. The system mainly consists of a server, user terminals, and a communication network connecting them. Each component of this system is realized using the following specific hardware and software.

[0631] The server uses a high-performance processing unit to capture game footage in real time as a data input method. This processing uses video capture and data streaming software, specifically OpenCV, to capture frame-by-frame video data. For real-time analysis, a data analysis method is used to determine emotional states using audio and video data in order to accurately evaluate the immersion of the game. PyDub is used to analyze audio information and OpenCV is used to analyze video information, extracting moments when excitement is maximized.

[0632] As a visualization and editing method, FFmpeg is used to edit the extracted moments of high excitement into video, applying processing to give them a specific tempo and narrative structure. The edited visual information is transferred to the user's terminal via a data output means and processed so that it can be displayed and operated using the terminal's interface. This interface provides a graphical operating environment for the user to edit and modify the video.

[0633] Furthermore, the information sharing mechanism allows users to directly distribute the generated highlights via communication networks on online platforms and social media, enabling them to share them with other users. Through this system, users can easily share not only surprising moments but also enjoyable experiences with many people.

[0634] As a concrete example, if a user is playing an action-packed shooting game and a decisive moment of victory occurs, the system will reliably capture that moment, detect the audio volume and visual stimuli, and edit it into a highlight. This process flow occurs in real time, and the generated highlight can be shared immediately after gameplay.

[0635] An example of a prompt for a generative AI model is: "I want to create a highlight reel of the most exciting moments from the final battle scene of an online game. What actions would maximize the excitement?"

[0636] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0637] Step 1:

[0638] Server-based video data capture

[0639] The server captures game video data in real time via a capture device during the user's gameplay and saves it frame by frame. The input is the user's game video, and the output is a set of video frames suitable for analysis. For video streaming, OpenCV is used to process each frame.

[0640] Step 2:

[0641] Analysis of audio and video data

[0642] The server uses an analysis model to analyze the captured video and audio data and evaluate specific moments that indicate states of excitement. This process uses frame-by-frame audio and video data as input and generates metadata as output that indicates moments of high excitement. Audio data is analyzed using PyDub to identify significant changes in volume and frequency, and dynamic changes in the video are analyzed using OpenCV.

[0643] Step 3:

[0644] Identification and editing of moments of excitement

[0645] Based on the analysis results, the server rearranges the video of the identified exciting moments and edits it to create a sense of tempo and narrative. The input is the metadata and video frames obtained in step 2, and the output is a highlight video. FFmpeg is used for video editing, and effects and transitions are added.

[0646] Step 4:

[0647] Export edited highlights

[0648] The server outputs the highlights generated by the visualization and editing system in the specified format. The input is the edited highlight video, and the output is a file in a format that users can view or share. It is converted to a format usable on various devices.

[0649] Step 5:

[0650] Display and operation on the user terminal

[0651] Users view highlights on their devices and perform further editing and customization as needed. The input is an exported highlight video file, and the output is the final, shareable video. The user interface allows for intuitive editing through graphical controls.

[0652] Step 6:

[0653] Sharing highlights

[0654] Users share their edited highlights via a communication network through their device to online platforms and social media. The input is the final edited video file, and the output is shared content viewable on various platforms. This step allows users to widely showcase the exciting moments they have created using a generative AI model.

[0655] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0656] This embodiment of the invention is a system that combines an analysis model that captures game footage in real time and evaluates the state of excitement with an emotion engine that recognizes the user's emotions. The server first receives game footage from the user's terminal and prepares it for analysis frame by frame. The analysis model individually evaluates and scores the state of excitement based on changes in movement and sound within the game.

[0657] A unique element of this system is the use of an emotion engine. The server analyzes the user's facial expressions and voice tone in real time and generates emotion data. This data is integrated to more accurately assess the user's excitement level in the game. Furthermore, the emotion data allows the editing module to generate more personalized highlights. For example, if the system detects the user's surprised facial expression or excited voice tone, this is reflected in the highlight, and editing is performed to visually emphasize the emotion of that moment.

[0658] The generated highlights are exported in the specified format by the server's output module. Users can quickly review the generated highlights through their terminal and perform further editing and optimization if necessary. Finally, users can share these highlights on an online platform or save them locally.

[0659] As a concrete example, consider a scenario in a multiplayer game where users cooperate with their team to overcome a difficult challenge. The emotion engine captures the user's joy and excitement, and based on that, highlights in the video are automatically edited. This allows viewers not only to see moments from the game, but also to feel the user's emotional reactions and share in that authentic experience. This system aims to enhance the immersion of game streams and improve the visual and emotional appeal of the content.

[0660] The following describes the processing flow.

[0661] Step 1:

[0662] The server receives game video and audio from the user's terminal in real time and captures them as a stream. The data is stored in a buffer and kept ready for processing at any time.

[0663] Step 2:

[0664] The server divides the video data into frames and extracts the audio data individually. This data is then passed to an analysis model, which is prepared to evaluate the state of excitement.

[0665] Step 3:

[0666] The server's analysis model evaluates the excitement level based on changes in in-game events and audio. An excitement score is assigned to each frame, and moments with particularly high scores are selected for the next process.

[0667] Step 4:

[0668] The server activates an emotion engine to detect the user's facial expressions and voice tone in real time. This allows the system to evaluate the user's emotions and reflect them in the game's excitement level assessment.

[0669] Step 5:

[0670] By combining analysis results and emotional data, the server's editing and shaping module identifies moments of high excitement and generates highlights. Video editing is then performed with a specific tempo and a narrative structure based on the user's emotions.

[0671] Step 6:

[0672] The server exports the generated highlights in the specified output format and uploads them to an online platform or sends them to the user's device, according to the user's instructions.

[0673] Step 7:

[0674] Users can review the generated highlights through their devices and make additional edits or final adjustments as needed. Highlights can be published instantly, enhancing the user's visual and emotional experience.

[0675] (Example 2)

[0676] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0677] Traditional game content highlight generation systems often failed to accurately reflect user emotions because they evaluated excitement levels purely based on in-game movement and sound changes. As a result, the appeal and immersion of the content conveyed to viewers were limited. Furthermore, generating personalized highlights that responded to individual user emotional responses was difficult.

[0678] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0679] In this invention, the server includes data receiving means for acquiring game content in real time, data analysis means for analyzing the game content and evaluating the user's excitement level, and emotion recognition means for generating user emotion data. This makes it possible to integrate the user's emotions and excitement level to edit the content and provide viewers with more emotionally rich and personalized highlights.

[0680] "Game content" refers to interactive digital media that runs on electronic devices and is manipulated by users.

[0681] "Data receiving means" refers to a device or method for acquiring external signals or information in an appropriate format.

[0682] "Data analysis means" refers to software or hardware used to process received information and identify specific patterns or states.

[0683] "Emotion recognition means" refers to algorithms and systems that analyze a user's facial expressions, voice, etc., to estimate their emotional state.

[0684] "Information formation means" refers to a process or function that generates new information based on analyzed data and assembles it in a specific format.

[0685] "Data output means" refers to a device or method for transmitting processed information to an external source in a specified format.

[0686] "Excitement" refers to a state in which the user experiences psychological or physiological arousal or tension.

[0687] "Emotional data" refers to information collected and analyzed to represent a user's internal emotional state.

[0688] A "specific format" refers to a standardized shape or structure that information must follow when it is output.

[0689] An embodiment of this invention is a video editing system that combines real-time capture of game content with user sentiment analysis. The server acquires data using a common communication protocol to receive game content from the user's terminal. This data is organized as video frames and converted into a format that is easy to analyze using a conversion tool such as FFmpeg.

[0690] The server then leverages machine learning libraries to analyze these frames. As a data analysis tool, it uses AI models like TensorFlow to evaluate game movements and audio information, quantifying the user's excitement level. Simultaneously, as an emotion recognition tool, it analyzes the user's facial expressions and voice tone using OpenCV and voice analysis software to generate emotion data. This allows the system to quantify the user's psychological response.

[0691] This data is integrated and edited by information shaping tools to generate highlights. This process matches peaks of excitement with heightened emotions, creating content that reflects the user's experience. For example, it might be edited to highlight moments when a user experiences great joy after winning a game.

[0692] The server ultimately exports the generated highlights in standard formats such as MP4, making them easily accessible to users. This system also allows users to further customize the content using digital editing tools such as Adobe Premiere.

[0693] As a concrete example, consider a scenario in an online game where a user achieves a dramatic come-from-behind victory. In this case, the server identifies that moment, integrates data on excitement levels and emotion recognition, and generates a highlight that allows viewers to share in the user's inner feelings. A concrete example of a prompt to input into the generating AI model would be, "Identify the moment in this game footage when the user is most excited, and edit the scene with the strongest emotion into a highlight."

[0694] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0695] Step 1:

[0696] The server receives streaming data of game content from the user's terminal. The input is the video and audio data of the game, transmitted in real time. Upon receiving this data, the server uses a frame buffer to divide the stream into individual video frames. The output is a list of processable video frames. Specifically, the server uses the RTMP protocol to receive data efficiently.

[0697] Step 2:

[0698] The server applies data analysis tools to analyze the segmented video frames and audio data. The input is the video frames and audio data generated in step 1. An AI model using the TensorFlow library is used for the analysis, evaluating the excitement level based on the detection of movement in the game and changes in sound. The output is the excitement score for each frame. Specifically, features are extracted for each frame, and the state is quantified through a pre-trained AI model.

[0699] Step 3:

[0700] The server analyzes the user's facial expressions and voice to generate emotion data. The input consists of the user's webcam video and audio stream from the microphone. OpenCV is used to analyze facial expressions, and speech recognition software is used to evaluate the tone of the voice. The output is a dataset representing the user's emotional state. Specifically, emotion estimation is performed using facial recognition and voice pitch analysis.

[0701] Step 4:

[0702] The server integrates the analyzed excitement score and emotion data using an information-forming mechanism to edit customized highlights. The inputs are the excitement score from step 2 and the emotion data from step 3. Using this information, emotionally significant frames matching the excitement peaks are selected, and edit points are determined. The output is the final edited highlight video. Specific operations include calculating edit points and performing cut editing.

[0703] Step 5:

[0704] The server exports the edited highlights in a format based on user settings. The input is the highlight video created in step 4. An encoding tool is used to convert it to the user-specified format (e.g., MP4). The output is a video file in the final format that the user can easily play and share. Specifically, the server sets the encoding parameters and performs the encoding process with an optimal balance.

[0705] (Application Example 2)

[0706] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0707] A challenge with typical game streaming is that it's difficult for viewers to fully grasp the streamer's emotions and moments of excitement. Therefore, there's a need for a way to allow viewers to experience the streamer's emotions in real time and achieve a more immersive viewing experience.

[0708] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0709] In this invention, the server includes acquisition means for acquiring game information in real time, evaluation model means for analyzing the game information and evaluating the state of excitement, and emotion analysis means for generating and integrating emotion information based on the evaluation data. As a result, viewers can experience the excitement and emotional changes of the commentator in real time, enabling a more emotional and immersive viewing experience.

[0710] The "acquisition means" refers to a mechanism that receives game information in real time and prepares it for processing.

[0711] The "evaluation model means" is a mechanism that uses an algorithm to analyze acquired game information and quantify and evaluate the user's state of excitement.

[0712] An "emotion analysis tool" is a mechanism that recognizes emotions based on the user's facial expressions and voice data, and generates that data.

[0713] "Editing and shaping means" refers to a mechanism that identifies moments of high excitement based on data from evaluation model means and emotion analysis means, and edits them in a way that is easily understood by the viewer.

[0714] "Output method" refers to a mechanism for exporting edited highlights in a specified output format.

[0715] A "visualization means" is a mechanism for visually displaying generated emotional information in real time and conveying those emotions to the viewer.

[0716] A "conversion mechanism" is a device that converts edited highlights into a format that can be shared on online platforms.

[0717] The system for realizing this invention consists of a server and a user's client terminal. First, the user's terminal acquires game information in real time and transmits it to the server. On the server, the acquisition means is responsible for receiving the game information. The game information is analyzed by the evaluation model means, and the user's excitement level is numerically evaluated. This evaluation uses algorithms based on audio and video information.

[0718] Next, the emotion analysis means analyzes the user's facial expressions and vocal characteristics to generate emotion information. This data is integrated with evaluation data, and the editing and shaping means identifies moments of high excitement. These moments are edited by the visualization means to visually represent emotions in real time. The completed highlights can be exported in a specified format through the output means. The conversion means also formats the exported highlights so that they can be shared on an online platform.

[0719] Users can experience game footage that visualizes the streamer's emotions through devices such as VR-compatible headsets. This allows viewers to realistically feel the streamer's experience. For example, consider a scene in a game where the user overcomes a difficult challenge and feels both surprise and joy simultaneously. Their facial expressions and voice at that moment are collected as emotional data and edited in a way that conveys this to the viewer.

[0720] An example of a prompt using a generative AI model might be: "I want to create a VR experience that visualizes the emotions of a game streamer in real time and conveys them to the viewers. Please give me ideas on how to visualize exciting moments using Unity and VR technology."

[0721] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0722] Step 1:

[0723] The terminal acquires the user's game information in real time and sends it to the server. This input data includes game video frames and audio data. The terminal compresses this data and processes it to send it to the server with low latency.

[0724] Step 2:

[0725] The server analyzes the game information received using the acquisition means. In this process, game video and audio data are taken in and passed to the evaluation model means. This model means uses a generative AI model to quantify the user's excitement level for each frame and outputs it as a score.

[0726] Step 3:

[0727] The server's emotion analysis system captures the user's facial expression data and voice characteristics. Specifically, this step analyzes the facial image data and voice tone sent from the terminal to generate emotion data. The output emotion data is then integrated with evaluation data as numerical parameters.

[0728] Step 4:

[0729] The server's editing and shaping mechanism integrates evaluation data and emotional data to identify moments of high excitement. These identified moments are then extracted as video tracks and compiled for editing. Based on the integrated data, cuts are made for visualization.

[0730] Step 5:

[0731] The visualization means visually expresses the user's emotions based on the cut footage provided by the editing and shaping means. In this step, a specific tempo and narrative are added, and effects are added to enhance the visual impact. The final output is video data provided to the viewer in real time.

[0732] Step 6:

[0733] The output method exports the final highlights in the specified format. This process converts the completed video data into a shareable format such as MP4 or WebM and exports it to a file for easy access by the user.

[0734] Step 7:

[0735] The conversion tool converts the exported highlights into an appropriate format so that they can be shared on online platforms. Finally, it provides users with links or files that allow them to view, share, and save these highlights.

[0736] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0737] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0738] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0739] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0740] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0741] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0742] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0743] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0744] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0745] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0746] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0747] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0748] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0749] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0750] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0751] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0752] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0753] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0754] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0755] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0756] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted as being incorporated by reference.

[0757] The following is further disclosed regarding the embodiments described above.

[0758] (Claim 1)

[0759] An input processing means for capturing game footage in real time,

[0760] An analytical model means for analyzing the aforementioned game footage and evaluating the state of excitement,

[0761] An editing and shaping means for identifying and editing moments of high excitement,

[0762] Output means for exporting the edited highlights in an output format,

[0763] A system that includes this.

[0764] (Claim 2)

[0765] The system according to claim 1, characterized in that the analysis model means applies an analysis algorithm for evaluating the state of excitement based on audio data and video data.

[0766] (Claim 3)

[0767] The system according to claim 1, characterized in that the editing and shaping means edits the identified moment to give it a specific tempo and storytelling quality.

[0768] "Example 1"

[0769] (Claim 1)

[0770] A data input method for acquiring information in real time,

[0771] An evaluation model means for analyzing the aforementioned information and detecting the state of excitement,

[0772] A time-series formation method for organizing and adding effects to moments of high excitement,

[0773] A format conversion means for providing the time-series information in a predetermined format,

[0774] A system that includes this.

[0775] (Claim 2)

[0776] The system according to claim 1, characterized in that the evaluation model means applies an analytical method for evaluating the state of excitement based on acoustic information and visual information.

[0777] (Claim 3)

[0778] The system according to claim 1, characterized in that the time series forming means shapes a specific moment with a specific rhythm and narrative.

[0779] "Application Example 1"

[0780] (Claim 1)

[0781] A data input method for acquiring game footage in real time,

[0782] A data analysis means for analyzing the aforementioned game footage and evaluating the emotional state,

[0783] A visualization and editing method for identifying and editing moments of high emotional state,

[0784] A data output means for exporting the visualized and edited visual information in an output format,

[0785] Interface means for making the aforementioned visual information displayable and operable on an information terminal,

[0786] Information sharing means for sharing the aforementioned visual information over a communication network,

[0787] A system that includes this.

[0788] (Claim 2)

[0789] The system according to claim 1, characterized in that the data analysis means applies an analysis algorithm for evaluating emotional states based on acoustic information and visual information.

[0790] (Claim 3)

[0791] The system according to claim 1, characterized in that the visualization editing means uses an information processing device to edit the identified moment with a specific tempo and narrative structure and output it.

[0792] "Example 2 of combining an emotion engine"

[0793] (Claim 1)

[0794] A data receiving means for acquiring game content in real time,

[0795] A data analysis means for analyzing the aforementioned game content and evaluating the state of excitement,

[0796] A means for recognizing emotions to generate user emotion data,

[0797] Information formation means for integrating and editing emotional data and excitement state,

[0798] A data output means for outputting the information-formed content in a specific format,

[0799] A system that includes this.

[0800] (Claim 2)

[0801] The system according to claim 1, characterized in that the data analysis means applies an analysis method for evaluating the state of excitement based on audio information and visual information.

[0802] (Claim 3)

[0803] The system according to claim 1, characterized in that the information forming means edits the integrated emotional data and the state of excitement to give it a specific narrative and rhythm.

[0804] "Application example 2 when combining with an emotional engine"

[0805] (Claim 1)

[0806] A means of acquiring game information in real time,

[0807] An evaluation model means for analyzing the aforementioned game information and evaluating the state of excitement,

[0808] A means of emotion analysis that generates and integrates emotional information based on evaluation data,

[0809] An editing and shaping means for identifying and editing moments of high excitement,

[0810] Output means for exporting the edited highlights in output format,

[0811] A visualization means for realizing complex visual displays,

[0812] A conversion method for converting to a format shareable on an online platform,

[0813] A system that includes this.

[0814] (Claim 2)

[0815] The system according to claim 1, characterized in that the evaluation model means and the emotion analysis means apply algorithms for evaluating the state of excitement and the state of emotion based on audio information and video information.

[0816] (Claim 3)

[0817] The system according to claim 1, characterized in that the editing and forming means edits identified moments to give them a specific rhythm and narrative quality, and uses visualization means to visually display emotional information in real time. [Explanation of Symbols]

[0818] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. An input processing means for capturing game footage in real time, An analytical model means for analyzing the aforementioned game footage and evaluating the state of excitement, An editing and shaping means for identifying and editing moments of high excitement, Output means for exporting the edited highlights in an output format, A system that includes this.

2. The system according to claim 1, characterized in that the analysis model means applies an analysis algorithm for evaluating the state of excitement based on audio data and video data.

3. The system according to claim 1, characterized in that the editing and shaping means edits the identified moment to give it a specific tempo and storytelling quality.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A