system
The system addresses the lack of personalized content in conventional broadcasts by using a generative model to analyze video data and provide interactive features for easy scene extraction and sharing, improving user engagement and experience.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-10
- Publication Date
- 2026-04-22
AI Technical Summary
Conventional television broadcasts and content distribution systems fail to provide information specialized for specific characters or artists, leading to decreased user engagement and require users to manually edit and share scenes of interest, which is time-consuming.
A system that analyzes video data using a generative model to generate metadata in real-time, delivering tailored content with interactive features allowing users to easily extract and share specific scenes.
Enhances user engagement by providing content tailored to individual interests, making the viewing experience more fulfilling and efficient.
Smart Images

Figure 2026068431000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In conventional television broadcasts and general content distribution, it is difficult for users to obtain information specialized for a specific character or artist, and there is a problem that user engagement decreases because the content includes content that viewers are not interested in. In addition, there is also a problem that it takes time for users to edit and share scenes they are interested in by themselves.
Means for Solving the Problems
[0005] This invention provides a system for delivering content tailored to user preferences. This system includes means for analyzing video data using a generative model, generating metadata related to a user-selected subject in real time, and delivering it to the user's terminal. Based on the delivered data, it generates and displays live commentary content in real time, and further provides interactive functions that allow users to easily extract and share specific scenes. In this way, users can enjoy content tailored to their interests, enhancing their fan engagement experience.
[0006] A "user" is an individual or group that utilizes a particular service or system.
[0007] A "generative model" is an algorithm or program used to learn the characteristics of data and generate new information.
[0008] "Video data" refers to digital data containing visual information in a format that can be played back as video.
[0009] "Analysis" is the process of meticulously analyzing data and information to reveal its structure and characteristics.
[0010] "Metadata" refers to additional information about data, such as its content, format, or how it was created.
[0011] A "user terminal" is a device that a user directly operates and uses to access services and data.
[0012] "Distribution" refers to the process of sending data or information to a specific location or device.
[0013] "In real time" means that processing or operations are performed instantaneously and reflected almost simultaneously with actual time.
[0014] "Live content" refers to information and media for transmitting information about events and situations to viewers in real time.
[0015] "Interactive function" refers to a function that allows users to directly interact with a system or application.
[0016] "Cutting out" refers to the process of selecting and separating specific parts from multiple contents.
[0017] "Sharing" means providing information and data to other users or devices for joint use.
Brief Description of Drawings
[0018] [Figure 1] It is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] It is a conceptual diagram showing an example of the main functions of a data processing device and a smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which multiple emotions are mapped. [Figure 10]Shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when an emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when an emotion engine is combined.
Mode for Carrying Out the Invention
[0019] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described according to the accompanying drawings.
[0020] First, the terms used in the following description will be described.
[0021] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of a plurality of arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of a plurality of types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0022] In the following embodiments, the numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0023] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.
[0024] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0025] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0026] [First Embodiment]
[0027] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0028] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0029] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0030] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0031] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0032] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0033] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0034] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0035] As shown in Figure 2, in the data processing device 12, specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0036] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0037] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0038] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0039] To implement this invention, the system employs a configuration in which a server, a terminal, and a user work in cooperation. First, the server receives video data and analyzes it using a generative model. The purpose of the analysis is to extract elements associated with a specific character, artist, or other subject selected by the user. This identifies interesting scenes and related information in the video and creates metadata.
[0040] The user's device receives the metadata delivered from the server. Based on this metadata, the device generates customized live commentary content for the user. This commentary is implemented, for example, through real-time updated text overlays or synthesized speech explanations.
[0041] Users can take advantage of interactive features while watching this live content provided through their devices. Specifically, users are given the means to cut out specific scenes from the video and share those scenes on other social media platforms. This allows users to easily share content related to their favorite things and enjoy it with their friends.
[0042] As a concrete example, suppose a user is watching a music program. The server receives the program's video, analyzes the scenes featuring the user's favorite artist, and generates metadata for those scenes. The device then uses this metadata to generate live commentary content about the artist's performance and provides it to the user in real time. The user can then select specific parts of the performance while watching this commentary and share those memorable moments on social media.
[0043] In this way, the system can provide content tailored to the information that users are interested in, making the experience of supporting one's favorite idols even more fulfilling.
[0044] The following describes the processing flow.
[0045] Step 1:
[0046] The server receives video data in real time from television and streaming services. The received video data is then sent directly to the analysis process.
[0047] Step 2:
[0048] The server analyzes the video data using a generative model. This analysis utilizes facial recognition and feature analysis to identify the user's favorite characters or artists.
[0049] Step 3:
[0050] Based on the analysis results, the server generates metadata related to the identified target. This metadata includes a timestamp, the target's name, and related episode information.
[0051] Step 4:
[0052] The server delivers the generated metadata to the user's device. The delivery is performed with low latency to avoid affecting the user's viewing experience.
[0053] Step 5:
[0054] The device generates live commentary content in real time based on the received metadata. Specifically, it provides text information related to the target scene and narration generated by speech synthesis.
[0055] Step 6:
[0056] Users can select scenes of interest while watching the live commentary content provided on their device. A clipping option is available at this time.
[0057] Step 7:
[0058] Users can easily share the clipped scenes on social media. The device provides a dedicated UI for this purpose, facilitating sharing.
[0059] In this way, the entire system works together to provide a content experience tailored to the user's preferences.
[0060] (Example 1)
[0061] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0062] Traditional systems struggled to efficiently extract information related to specific user selections and provide customized content in real time, resulting in users expending considerable effort to find the necessary information amidst overwhelming amounts of data. Solving this problem is crucial.
[0063] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0064] In this invention, the server includes means for using a generative model that analyzes visual information specifically for an object selected by the user, means for generating supplementary information related to the object based on the analysis, and means for distributing the generated supplementary information to the user's device. This makes it possible to effectively and quickly provide specific information of interest to the user and improve usability.
[0065] The term "user" refers to a person who uses a system to receive and manipulate visual information.
[0066] "Visual information" is a broad term referring to information formats that include video data such as moving images and still images.
[0067] A "generative model" refers to an algorithm or machine learning technique used to analyze elements related to a specific subject.
[0068] "Incidental information" refers to additional data or metadata related to the subject, generated based on the analyzed visual information.
[0069] "User equipment" refers to devices such as computers and mobile devices that users use to interface with the system.
[0070] "Real-time supplementary content" refers to explanatory and annotated information that is generated in real time and provided to the user.
[0071] "Interactive features" refer to means of operation that allow users to interact with the system and utilize specific information or functions.
[0072] This invention provides an information processing system in which a server, a terminal, and a user work in cooperation. Specific embodiments are described below.
[0073] Server Functions
[0074] The server receives video data and is equipped with a generative AI model to analyze it. This model is used to recognize specific subjects (characters or artists selected by the user) contained in the video data and extract related elements. Based on the extracted information, it generates supplementary information related to the subject. As a concrete example, one could input a prompt message into the generative AI model such as, "Analyze the scenes featuring a specific artist contained in this video data and generate related metadata."
[0075] Device functions
[0076] The terminal receives supplementary information delivered from the server and generates timely, user-oriented support content based on that information. This involves using software to display subtitles and supplementary information in real time. It is also possible to provide commentary on the video using speech synthesis technology.
[0077] User functions
[0078] Users can view supplementary content provided through their devices while utilizing interactive features. For example, they can easily extract specific video scenes and share them on social media platforms. This feature allows users to quickly and easily share information they are interested in with others.
[0079] The entire system aims to efficiently process information related to specific targets selected by the user and deliver customized content in real time.
[0080] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0081] Step 1:
[0082] The server receives visual information specified by the user. This visual information is used as input to send prompt messages to a generating AI model, which analyzes elements related to a specific object within the video data. Specifically, the received video is divided into frames, which are then input into the machine learning model for analysis. The output generates supplementary information based on the extracted elements and data.
[0083] Step 2:
[0084] The server structures the generated supplementary information and distributes it to the user's device. The specific actions performed in this process involve packaging the supplementary information in JSON or XML format and sending it to the terminal via the network. The input is the supplementary information resulting from the analysis, and the output is a data packet receivable by the terminal.
[0085] Step 3:
[0086] The terminal receives supplementary information sent from the server. Using the received data as input, it generates immediate supplementary content (e.g., text overlays, audio descriptions). Specifically, it analyzes the received information and generates a script to display it to the user at the appropriate time. As output, it provides visual or audio content that the user can view.
[0087] Step 4:
[0088] The user uses interactive features while viewing supplementary content provided through the device. As input, the user provides the device with instructions to select a specific scene. Based on these instructions, the device extracts the selected scene and converts it into a format that the user can share on other platforms. The output is a shareable media file.
[0089] Through the process described above, the system aims to effectively provide users with information of interest and further improve the user experience.
[0090] (Application Example 1)
[0091] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0092] Traditional content distribution methods had limited means of effectively providing relevant information when users focused on specific topics during viewing. As a result, the viewing experience was limited due to insufficient information tailored to users' interests, and the convenience of sharing specific scenes was also reduced.
[0093] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0094] In this invention, the server includes means for using a generative model that analyzes video information specifically for a target selected by the user, means for generating information related to the target based on the analysis, and means for distributing the generated information to the user's device. This makes it possible to provide the user with information related to a specific target in real time and to improve the viewing experience through an interactive function based on that information.
[0095] A "user" is an individual who views content provided by the system and utilizes its interactive features.
[0096] "Target" refers to characters, artists, or other entities that users are interested in and want to obtain specific information about.
[0097] "Visual information" refers to media containing visual or audio data, and includes real-time or recorded content.
[0098] A "generative model" is an algorithm or artificial intelligence used to analyze video information and extract information related to a specific subject.
[0099] "Related information" refers to additional information such as data and explanations generated from the analyzed video information in a way that is relevant to the subject.
[0100] "User device" refers to a device used by a user to view content and receive explanations, and includes smartphones, smart glasses, and other similar devices.
[0101] "Distribution" refers to the act of sending generated and related information to a user's device via a network.
[0102] "Explanatory content" refers to additional information, including text presentations and audio output, that is provided in real time.
[0103] The "dialogue function" refers to a feature that allows users to extract specific scenes from content and share them on other platforms.
[0104] To implement this invention, the server first receives video information and extracts information related to the subject using a generative model. Specifically, it focuses on a particular character or artist within the video information and uses a generative AI model to analyze interesting scenes and related data. At this time, the server uses a cloud server equipped with a high-performance GPU to achieve efficient video analysis.
[0105] The information obtained through analysis is stored on a server as metadata representing the generated state, and this metadata is distributed to the user's device via the network. The user's device is a smartphone or smart glasses, which generates and displays explanatory content in real time based on the received metadata. The explanatory content includes text presentation and audio output, providing information through the user's sight and hearing.
[0106] Based on this explanatory content, users can easily extract specific scenes on their devices. These extracted scenes can then be shared with other platforms through a dialogue function. This allows users to share their interests and enjoy a richer viewing experience.
[0107] For example, if a user is watching a live music stream, the server analyzes the live video in real time and detects scenes in which the featured artist appears. The generated metadata is sent to the user's device, providing a commentary on the artist's performance in that scene. Based on this commentary, the user can capture memorable moments and share them on social media, following a prompt such as, "Detect the artist's solo section and generate related information."
[0108] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0109] Step 1:
[0110] The server receives video information transmitted from the user's device. At this stage, the input is streaming video data, and the output is video data converted into a format that can be processed in the next analysis step. Specifically, data preprocessing is performed, such as noise reduction and extraction of necessary parts of the video.
[0111] Step 2:
[0112] The server inputs the pre-processed video information into a generative AI model for analysis. This input is the video data transformed in step 1. The generative AI model detects specific characters or artists within the video and analyzes interesting scenes. The output generates metadata related to the subject. This metadata includes timestamps and features of specific scenes.
[0113] Step 3:
[0114] The server distributes the generated metadata to the user's device. The input is the metadata obtained in step 2. This is transmitted to the user's device via the network, resulting in output that enables the delivery of rich content. Specifically, a low-latency and highly efficient transmission protocol is used to maintain real-time performance.
[0115] Step 4:
[0116] The user's device generates and displays explanatory content in real time using the received metadata. The input at this stage is metadata received from the server. The device provides the user with relevant information through text and speech synthesis. The output is a customized explanation that is conveyed through the user's sight and hearing.
[0117] Step 5:
[0118] The user selects a specific scene from a video they are watching on their device and extracts that scene. The input is the video scene the user is interested in, and the output is a clip of the selected scene. The interface responds to the user's actions and provides feedback until the selection is complete.
[0119] Step 6:
[0120] Users share the clipped scenes to other platforms through a dialogue function. The input is the video clip obtained in step 5. The output is the shared video clip, which is then posted to social media and other video platforms. Specifically, the clip is encoded and converted to the appropriate format.
[0121] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0122] In this embodiment of the invention, the system is designed to function in cooperation with three parties: a server, a terminal, and a user. First, the server receives video data in real time and analyzes it using a generative model. This analysis process identifies specific objects within the video and generates associated metadata.
[0123] Next, this system incorporates an emotion engine to identify the user's emotions. When a user interacts with the system through a terminal, the emotion engine analyzes the user's tone of voice, facial expressions, and input data. This analysis is used to estimate the user's emotions.
[0124] Metadata generated on the server is delivered to the user's terminal in real time, and the terminal generates commentary content based on this data and the output of the emotion engine. This commentary content is optimized to provide information that takes into account the user's current emotional state. For example, if the user is expressing joy, content that emphasizes positive commentary and musical elements will be displayed.
[0125] Users can view live streaming content displayed on their devices, clip specific scenes of interest, and share them on social media. Furthermore, the system suggests content recommendations that take into account the user's emotions. This enables a flexible content experience tailored to the user's interests and feelings.
[0126] As a concrete example, consider a scenario where a user is watching live video. The server identifies scenes featuring the artist in real time and generates relevant metadata. Furthermore, when the user reacts using their device, the emotion engine analyzes the user's emotions based on that reaction. If the user shows an excited reaction, the device adjusts the live commentary content by enhancing the audio and changing the camera angle to match that excitement, enriching the user's viewing experience.
[0127] In this way, the system can provide an interactive content experience optimized for both the user's choices and their emotions.
[0128] The following describes the processing flow.
[0129] Step 1:
[0130] The server receives video data in real time. After receiving the data, it starts analyzing the video using a generative model to identify the object specified by the user. After the analysis, it generates metadata related to the object.
[0131] Step 2:
[0132] The server delivers the identified metadata to the user's terminal. Simultaneously, the analysis results are transferred with low latency so that they can be used on the terminal.
[0133] Step 3:
[0134] The device receives the delivered metadata and generates live commentary content to display to the user. This content is customized according to the user's selection.
[0135] Step 4:
[0136] The device utilizes an emotion engine. It analyzes data (voice and facial expressions) collected through the user's microphone and camera to estimate the user's current emotions.
[0137] Step 5:
[0138] The device dynamically adjusts the live commentary content based on the results of the emotion engine. For example, if the user is surprised, it will emphasize more fun content and recommendations that match that emotion.
[0139] Step 6:
[0140] Users can enjoy customized live commentary content provided on their devices in real time. Furthermore, they can easily cut out scenes of interest and share them on social media through interactive operations.
[0141] Step 7:
[0142] The server collects feedback data for future content delivery based on user sentiment and interaction data. This data is used for analysis to further improve the user experience.
[0143] Through this series of steps, the system delivers an immersive content experience tailored to the user's emotions and preferences.
[0144] (Example 2)
[0145] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0146] Modern viewing platforms fail to adequately provide content tailored to the individual user's emotions and state of mind, making it difficult to deliver an experience that aligns with user interests and feelings. Generating relevant information in real time from massive amounts of video data and dynamically adjusting content using that information is particularly challenging. Furthermore, features for easily extracting and sharing scenes of interest are limited, highlighting the need for methods to enhance user engagement.
[0147] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0148] In this invention, the server includes means for using a generative model that analyzes visual information specifically for an object selected by the user, means for generating information related to the object based on the analysis, and means for distributing the generated information to the user terminal. This makes it possible to provide a customized, interactive viewing experience for each user in real time.
[0149] "Visual information" refers to image and video data obtained from users, and the process involves analyzing this data to extract meaningful information.
[0150] A "generative model" is an algorithm or system that learns specific patterns or features from input data and generates new information or predictions based on them.
[0151] "Information" refers to data obtained as a result of analysis and processing, and is a concept that includes metadata and analysis results.
[0152] A "user terminal" refers to a device used by a user for direct operation, and includes a group of devices such as smartphones, tablets, and personal computers.
[0153] An "interactive viewing experience" means dynamically changing content in real time in response to user input and reactions, providing an experience optimized for each individual user.
[0154] "Dynamic adjustment" refers to the process by which the system automatically changes the elements and characteristics of content in response to the user's emotions and circumstances.
[0155] "Operational functions" refer to the interfaces and tools provided to users to select or share specific scenes.
[0156] The system of this invention functions through the cooperation of three parties: a server, a terminal, and a user. The server receives visual information in real time from an external source. Then, utilizing a generative AI model, it analyzes this visual information to generate information tailored to the user's interests and choices. This generative AI model employs advanced pattern recognition algorithms, enabling it to quickly and accurately identify specific objects within the visual information and generate related metadata.
[0157] Information generated on the server is delivered to the user's device in real time. Along with the received information, the device utilizes audio and visual data acquired from its built-in microphone and camera to analyze the user's emotional state. The emotion engine analyzes this data to estimate the user's emotions. This allows the device to dynamically adjust visual and audio information to provide the user with the most suitable interactive viewing experience.
[0158] For example, consider a scenario where a user is watching a live music performance. The server identifies the artist's appearance in real time and generates relevant information. The device detects the user's smile and dynamically adjusts the audio and video effects based on that positive emotion. This functionality allows the user to enjoy a more immersive content experience.
[0159] An example of a prompt message would be: "React sensitively to the artist's appearance in the live video. Describe in three lines why the user would be excited, and then generate commentary content that matches that reaction." By providing this prompt message to the server, the generation AI model will perform the optimal content generation.
[0160] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0161] Step 1:
[0162] The server receives visual information from an external source in real time. The received video data is preprocessed for use in a generating AI model. This preprocessing includes resolution adjustment and noise filtering. The input is raw video data, and the output is data in a format suitable for analysis.
[0163] Step 2:
[0164] The server inputs pre-processed visual information into a generating AI model to identify specific objects. This identification process generates metadata, including symmetrical appearance timing and position information, such as the artist's entrance scene. The input is pre-processed visual information, and the output is metadata based on the identified information.
[0165] Step 3:
[0166] The server sends the generated metadata to the user's terminal. Compression techniques may be used during this process to improve data transfer efficiency. The input is the generated metadata, and the output is the metadata sent to the user's terminal.
[0167] Step 4:
[0168] The device receives metadata sent from the server and acquires audio and visual data from its built-in microphone and camera to analyze the user's emotional state. The emotion engine analyzes this data and estimates the user's emotions. The input is the audio and visual data acquired by the user's device, and the output is the estimated result regarding the user's emotional state.
[0169] Step 5:
[0170] The device generates an interactive viewing experience based on received metadata and sentiment analysis results. It dynamically adjusts the visual and audio presentation according to the user's emotional state, providing optimal content. The input is metadata and sentiment analysis results, while the output is the optimized content provided to the user.
[0171] Step 6:
[0172] Users watch optimized content and select scenes that interest them. They can then clip these selected scenes and share them with other users on social media. The input is the scenes the user selects as of viewing, and the output is the content snippet that is shared.
[0173] (Application Example 2)
[0174] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0175] When users watch video content, general content is provided uniformly to all users, making it difficult to provide an optimal viewing experience tailored to individual emotions and interests. Furthermore, there is no means to dynamically adjust content in real time in response to changes in user emotions, which can lead to decreased viewing satisfaction. Therefore, there is a need for technology that enables interactive and optimized content delivery in response to user emotions.
[0176] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0177] In this invention, the server includes means for using a generative model that analyzes video data specifically for a user-selected subject, means for generating metadata related to the subject based on the analysis, and means for distributing the generated metadata to the user terminal. This enables the generation and real-time distribution of optimal live commentary content that is tailored to the user's emotions.
[0178] A "generative model" is a mathematical or algorithmic structure used to analyze input data and extract specific patterns or features.
[0179] Metadata is additional information related to data, representing attributes and conditions concerning a specific subject or content.
[0180] A "user terminal" is an electronic device used by a user for direct operation, enabling them to view and manipulate content.
[0181] "Live commentary content" refers to dynamic viewing content that is provided in real time and generated in response to user interaction.
[0182] An "emotion engine" is an algorithm or system for analyzing and inferring a user's emotions, and for analyzing data based on the user's responses.
[0183] "Dynamic optimization" is the process of adjusting content based on real-time data and conditions to maintain an optimized state for the user.
[0184] "Interactive features" are functions that allow users to interact with and respond to the system in a two-way manner, improving participation and responsiveness.
[0185] The system implementing this invention mainly consists of three components: a server, a terminal, and a user. The server receives video data, analyzes a specific object in real time using a generation AI model, and generates metadata based on the results. This generated metadata is then transmitted from the server to the user's terminal.
[0186] The device generates live commentary content based on received metadata and displays it to the user in real time. Furthermore, the device has a built-in emotion engine that analyzes the user's facial expressions and tone of voice to estimate the user's emotions. Based on this emotion analysis, the live commentary content is dynamically optimized according to the user's current emotions.
[0187] For example, if the device detects that a user is excited while watching live video, it can enhance the audio and video of the live content accordingly, thereby enriching the viewing experience.
[0188] As a concrete example, consider a scenario where a user is watching a live concert of their favorite artist on their smartphone. The server recognizes the artist's appearance from the live video, generates relevant metadata, and automatically sends it to the user's device. At this time, the smartphone's camera and microphone capture the user's facial expressions and voice, which are then analyzed by an emotion engine. As a result, if the user is smiling, live commentary content that emphasizes upbeat music and positive comments will be displayed.
[0189] An example of a prompt statement might be, "How to optimize content when the user is happy." This prompt statement serves as a concrete example of how the generative model contributes to optimizing content in response to the user's emotions.
[0190] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0191] Step 1:
[0192] The server receives video data and begins analysis using a generative AI model. The server receives streaming video data as input, recognizes the object, and generates metadata related to the object as a result. The generative AI model detects specific patterns and features in the video and constructs metadata based on these.
[0193] Step 2:
[0194] Metadata generated from the server is delivered to the terminal. This metadata contains information about a specific video scene, and the terminal receives this metadata. The delivered metadata is then used for real-time content generation.
[0195] Step 3:
[0196] The device uses an emotion engine to analyze the user's emotions. It acquires facial expressions and voice tone as input data from the user's facial recognition camera and microphone, and analyzes this data using the emotion engine. This analysis process estimates the user's emotions (e.g., joy, excitement).
[0197] Step 4:
[0198] The device generates commentary content based on received metadata and sentiment analysis results. The real-time optimized commentary content dynamically adjusts video and audio, taking into account metadata information and user sentiment. For example, if the system determines that the user is happy, positive commentary and music will be emphasized in the commentary content.
[0199] Step 5:
[0200] Users can watch live commentary content generated on their devices, select and clip scenes of interest, and share them on social media. The device automatically edits the selected scenes based on user interaction and outputs them in a sharing format.
[0201] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0202] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0203] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0204] [Second Embodiment]
[0205] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0206] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0207] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0208] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0209] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0210] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0211] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0212] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0213] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0214] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0215] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0216] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0217] To implement this invention, the system employs a configuration in which a server, a terminal, and a user work in cooperation. First, the server receives video data and analyzes it using a generative model. The purpose of the analysis is to extract elements associated with a specific character, artist, or other subject selected by the user. This identifies interesting scenes and related information in the video and creates metadata.
[0218] The user's device receives the metadata delivered from the server. Based on this metadata, the device generates customized live commentary content for the user. This commentary is implemented, for example, through real-time updated text overlays or synthesized speech explanations.
[0219] Users can take advantage of interactive features while watching this live content provided through their devices. Specifically, users are given the means to cut out specific scenes from the video and share those scenes on other social media platforms. This allows users to easily share content related to their favorite things and enjoy it with their friends.
[0220] As a concrete example, suppose a user is watching a music program. The server receives the program's video, analyzes the scenes featuring the user's favorite artist, and generates metadata for those scenes. The device then uses this metadata to generate live commentary content about the artist's performance and provides it to the user in real time. The user can then select specific parts of the performance while watching this commentary and share those memorable moments on social media.
[0221] In this way, the system can provide content tailored to the information that users are interested in, making the experience of supporting one's favorite idols even more fulfilling.
[0222] The following describes the processing flow.
[0223] Step 1:
[0224] The server receives video data in real time from television and streaming services. The received video data is then sent directly to the analysis process.
[0225] Step 2:
[0226] The server analyzes the video data using a generative model. This analysis utilizes facial recognition and feature analysis to identify the user's favorite characters or artists.
[0227] Step 3:
[0228] Based on the analysis results, the server generates metadata related to the identified target. This metadata includes a timestamp, the target's name, and related episode information.
[0229] Step 4:
[0230] The server delivers the generated metadata to the user's device. The delivery is performed with low latency to avoid affecting the user's viewing experience.
[0231] Step 5:
[0232] The device generates live commentary content in real time based on the received metadata. Specifically, it provides text information related to the target scene and narration generated by speech synthesis.
[0233] Step 6:
[0234] Users can select scenes of interest while watching the live commentary content provided on their device. A clipping option is available at this time.
[0235] Step 7:
[0236] Users can easily share the clipped scenes on social media. The device provides a dedicated UI for this purpose, facilitating sharing.
[0237] In this way, the entire system works together to provide a content experience tailored to the user's preferences.
[0238] (Example 1)
[0239] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0240] Traditional systems struggled to efficiently extract information related to specific user selections and provide customized content in real time, resulting in users expending considerable effort to find the necessary information amidst overwhelming amounts of data. Solving this problem is crucial.
[0241] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0242] In this invention, the server includes means for using a generative model that analyzes visual information specifically for an object selected by the user, means for generating supplementary information related to the object based on the analysis, and means for distributing the generated supplementary information to the user's device. This makes it possible to effectively and quickly provide specific information of interest to the user and improve usability.
[0243] The term "user" refers to a person who uses a system to receive and manipulate visual information.
[0244] "Visual information" is a broad term referring to information formats that include video data such as moving images and still images.
[0245] A "generative model" refers to an algorithm or machine learning technique used to analyze elements related to a specific subject.
[0246] "Incidental information" refers to additional data or metadata related to the subject, generated based on the analyzed visual information.
[0247] "User equipment" refers to devices such as computers and mobile devices that users use to interface with the system.
[0248] "Real-time supplementary content" refers to explanatory and annotated information that is generated in real time and provided to the user.
[0249] "Interactive features" refer to means of operation that allow users to interact with the system and utilize specific information or functions.
[0250] This invention provides an information processing system in which a server, a terminal, and a user work in cooperation. Specific embodiments are described below.
[0251] Server Functions
[0252] The server receives video data and is equipped with a generative AI model to analyze it. This model is used to recognize specific subjects (characters or artists selected by the user) contained in the video data and extract related elements. Based on the extracted information, it generates supplementary information related to the subject. As a concrete example, one could input a prompt message into the generative AI model such as, "Analyze the scenes featuring a specific artist contained in this video data and generate related metadata."
[0253] Device functions
[0254] The terminal receives supplementary information delivered from the server and generates timely, user-oriented support content based on that information. This involves using software to display subtitles and supplementary information in real time. It is also possible to provide commentary on the video using speech synthesis technology.
[0255] User functions
[0256] Users can view supplementary content provided through their devices while utilizing interactive features. For example, they can easily extract specific video scenes and share them on social media platforms. This feature allows users to quickly and easily share information they are interested in with others.
[0257] The entire system aims to efficiently process information related to specific targets selected by the user and deliver customized content in real time.
[0258] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0259] Step 1:
[0260] The server receives visual information specified by the user. This visual information is used as input to send prompt messages to a generating AI model, which analyzes elements related to a specific object within the video data. Specifically, the received video is divided into frames, which are then input into the machine learning model for analysis. The output generates supplementary information based on the extracted elements and data.
[0261] Step 2:
[0262] The server structures the generated supplementary information and distributes it to the user's device. The specific actions performed in this process involve packaging the supplementary information in JSON or XML format and sending it to the terminal via the network. The input is the supplementary information resulting from the analysis, and the output is a data packet receivable by the terminal.
[0263] Step 3:
[0264] The terminal receives supplementary information sent from the server. Using the received data as input, it generates immediate supplementary content (e.g., text overlays, audio descriptions). Specifically, it analyzes the received information and generates a script to display it to the user at the appropriate time. As output, it provides visual or audio content that the user can view.
[0265] Step 4:
[0266] The user uses interactive features while viewing supplementary content provided through the device. As input, the user provides the device with instructions to select a specific scene. Based on these instructions, the device extracts the selected scene and converts it into a format that the user can share on other platforms. The output is a shareable media file.
[0267] Through the process described above, the system aims to effectively provide users with information of interest and further improve the user experience.
[0268] (Application Example 1)
[0269] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0270] Traditional content distribution methods had limited means of effectively providing relevant information when users focused on specific topics during viewing. As a result, the viewing experience was limited due to insufficient information tailored to users' interests, and the convenience of sharing specific scenes was also reduced.
[0271] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0272] In this invention, the server includes means for using a generative model that analyzes video information specifically for a user-selected object, means for generating information related to the object based on the analysis, and means for distributing the generated information to the user's device. This makes it possible to provide the user with information related to a specific object in real time and to improve the viewing experience through an interactive function based on that information.
[0273] A "user" is an individual who views content provided by the system and utilizes its interactive features.
[0274] "Target" refers to characters, artists, or other entities that users are interested in and want to obtain specific information about.
[0275] "Visual information" refers to media containing visual or audio data, and includes real-time or recorded content.
[0276] A "generative model" is an algorithm or artificial intelligence used to analyze video information and extract information related to a specific subject.
[0277] "Related information" refers to additional information such as data and explanations generated from the analyzed video information in a way that is relevant to the subject.
[0278] "User device" refers to a device used by a user to view content and receive explanations, and includes smartphones, smart glasses, and other similar devices.
[0279] "Distribution" refers to the act of sending generated and related information to a user's device via a network.
[0280] "Explanatory content" refers to additional information, including text presentations and audio output, that is provided in real time.
[0281] The "dialogue function" refers to the function that allows users to extract specific scenes of content and share them with other platforms.
[0282] To implement this invention, first, the server receives video information and extracts information related to the object using a generation model. Specifically, it focuses on specific characters or artists in the video information and utilizes the generative AI model to analyze interesting scenes and related data. At this time, the server uses a cloud server equipped with a high-performance GPU to achieve efficient video analysis.
[0283] The information obtained through the analysis is stored in the server as metadata in a generated state, and this metadata is distributed to the user's device via the network. The user device can be a smartphone or smart glasses, and it generates and displays explanatory content in real time based on the received metadata. The explanatory content includes text presentation and audio output, providing information through the user's vision and hearing.
[0284] Based on this explanatory content, the user can easily extract a specific scene on their own device. The extracted scene can be shared with other platforms through the dialogue function. This allows users to share their interests and enjoy a rich viewing experience.
[0285] As a specific example, when a user is watching a certain music live stream, the server analyzes the live video in real time and detects the scene where the target artist appears. The generated metadata is sent to the user's device, and an explanation regarding the artist's performance in that scene is provided. Based on this explanation, the user can capture an impressive moment and share it on SNS, etc., following an example prompt such as "Detect the solo part of the artist and generate related information."
[0286] The flow of the specific process in Application Example 1 will be described using FIG. 12.
[0287] Step 1:
[0288] The server receives video information transmitted from the user's device. At this stage, the input is streaming video data, and the output is video data converted into a format that can be processed in the next analysis step. Specifically, data preprocessing is performed, such as noise reduction and extraction of necessary parts of the video.
[0289] Step 2:
[0290] The server inputs the pre-processed video information into a generative AI model for analysis. This input is the video data transformed in step 1. The generative AI model detects specific characters or artists within the video and analyzes interesting scenes. The output generates metadata related to the subject. This metadata includes timestamps and features of specific scenes.
[0291] Step 3:
[0292] The server distributes the generated metadata to the user's device. The input is the metadata obtained in step 2. This is transmitted to the user's device via the network, resulting in output that enables the delivery of rich content. Specifically, a low-latency and highly efficient transmission protocol is used to maintain real-time performance.
[0293] Step 4:
[0294] The user's device generates and displays explanatory content in real time using the received metadata. The input at this stage is metadata received from the server. The device provides the user with relevant information through text and speech synthesis. The output is a customized explanation that is conveyed through the user's sight and hearing.
[0295] Step 5:
[0296] The user selects a specific scene from a video they are watching on their device and extracts that scene. The input is the video scene the user is interested in, and the output is a clip of the selected scene. The interface responds to the user's actions and provides feedback until the selection is complete.
[0297] Step 6:
[0298] Users share the clipped scenes to other platforms through a dialogue function. The input is the video clip obtained in step 5. The output is the shared video clip, which is then posted to social media and other video platforms. Specifically, the clip is encoded and converted to the appropriate format.
[0299] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0300] In this embodiment of the invention, the system is designed to function in cooperation with three parties: a server, a terminal, and a user. First, the server receives video data in real time and analyzes it using a generative model. This analysis process identifies specific objects within the video and generates associated metadata.
[0301] Next, this system incorporates an emotion engine to identify the user's emotions. When a user interacts with the system through a terminal, the emotion engine analyzes the user's tone of voice, facial expressions, and input data. This analysis is used to estimate the user's emotions.
[0302] The metadata generated by the server is delivered to the user terminal in real time, and the terminal generates live content based on this data and the output of the emotion engine. The live content here takes into account the user's current emotional state and provides optimized information. For example, when the user shows joy, content with positive explanations and emphasized music elements is displayed.
[0303] While browsing the live content displayed on the terminal, the user can clip specific scenes of interest and share them on SNS. Furthermore, the system also proposes recommended content considering the user's emotions. This enables a flexible content experience according to the user's interests and emotions.
[0304] As a specific example, consider the situation of watching a live video. The server identifies in real time the scene where the artist appears and generates relevant metadata. Furthermore, when the user shows a reaction using the terminal, the emotion engine analyzes the user's emotion based on this. When the user shows an excited reaction, the terminal enhances the live content according to the excitement by emphasizing the sound or changing the camera angle, enriching the user's viewing experience.
[0305] In this way, the system can provide an interactive content experience optimized for both the user's selection target and emotions.
[0306] The following explains the processing flow.
[0307] Step 1:
[0308] The server receives video data in real time. After reception, it starts analyzing the video using a generation model to identify the target specified by the user. After analysis, it generates metadata related to the target.
[0309] Step 2:
[0310] The server delivers the identified metadata to the user's terminal. Simultaneously, the analysis results are transferred with low latency so that they can be used on the terminal.
[0311] Step 3:
[0312] The device receives the delivered metadata and generates live commentary content to display to the user. This content is customized according to the user's selection.
[0313] Step 4:
[0314] The device utilizes an emotion engine. It analyzes data (voice and facial expressions) collected through the user's microphone and camera to estimate the user's current emotions.
[0315] Step 5:
[0316] The device dynamically adjusts the live commentary content based on the results of the emotion engine. For example, if the user is surprised, it will emphasize more fun content and recommendations that match that emotion.
[0317] Step 6:
[0318] Users can enjoy customized live commentary content provided on their devices in real time. Furthermore, they can easily cut out scenes of interest and share them on social media through interactive operations.
[0319] Step 7:
[0320] The server collects feedback data for future content delivery based on user sentiment and interaction data. This data is used for analysis to further improve the user experience.
[0321] Through this series of steps, the system delivers an immersive content experience tailored to the user's emotions and preferences.
[0322] (Example 2)
[0323] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0324] Modern viewing platforms fail to adequately provide content tailored to the individual user's emotions and state of mind, making it difficult to deliver an experience that aligns with user interests and feelings. Generating relevant information in real time from massive amounts of video data and dynamically adjusting content using that information is particularly challenging. Furthermore, features for easily extracting and sharing scenes of interest are limited, highlighting the need for methods to enhance user engagement.
[0325] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0326] In this invention, the server includes means for using a generative model that analyzes visual information specifically for an object selected by the user, means for generating information related to the object based on the analysis, and means for distributing the generated information to the user terminal. This makes it possible to provide a customized, interactive viewing experience for each user in real time.
[0327] "Visual information" refers to image and video data obtained from users, and the process involves analyzing this data to extract meaningful information.
[0328] A "generative model" is an algorithm or system that learns specific patterns or features from input data and generates new information or predictions based on them.
[0329] "Information" refers to data obtained as a result of analysis and processing, and is a concept that includes metadata and analysis results.
[0330] A "user terminal" refers to a device used by a user for direct operation, and includes a group of devices such as smartphones, tablets, and personal computers.
[0331] An "interactive viewing experience" is one in which content dynamically changes in real time in response to user input and reactions, providing an experience optimized for each individual user.
[0332] "Dynamic adjustment" refers to the process by which the system automatically changes the elements and characteristics of content in response to the user's emotions and circumstances.
[0333] "Operational functions" refer to the interfaces and tools provided to users to select or share specific scenes.
[0334] The system of this invention functions through the cooperation of three parties: a server, a terminal, and a user. The server receives visual information in real time from an external source. Then, utilizing a generative AI model, it analyzes this visual information to generate information tailored to the user's interests and choices. This generative AI model employs advanced pattern recognition algorithms, enabling it to quickly and accurately identify specific objects within the visual information and generate related metadata.
[0335] Information generated on the server is delivered to the user's device in real time. Along with the received information, the device utilizes audio and visual data acquired from its built-in microphone and camera to analyze the user's emotional state. The emotion engine analyzes this data to estimate the user's emotions. This allows the device to dynamically adjust visual and audio information to provide the user with the most suitable interactive viewing experience.
[0336] For example, consider a scenario where a user is watching a live music performance. The server identifies the artist's appearance in real time and generates relevant information. The device detects the user's smile and dynamically adjusts the audio and video effects based on that positive emotion. This functionality allows the user to enjoy a more immersive content experience.
[0337] An example of a prompt message would be: "React sensitively to the artist's appearance in the live video. Describe in three lines why the user would be excited, and then generate commentary content that matches that reaction." By providing this prompt message to the server, the generation AI model will perform the optimal content generation.
[0338] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0339] Step 1:
[0340] The server receives visual information from an external source in real time. The received video data is preprocessed for use in a generating AI model. This preprocessing includes resolution adjustment and noise filtering. The input is raw video data, and the output is data in a format suitable for analysis.
[0341] Step 2:
[0342] The server inputs pre-processed visual information into a generating AI model to identify specific objects. This identification process generates metadata, including symmetrical appearance timing and position information, such as the artist's entrance scene. The input is pre-processed visual information, and the output is metadata based on the identified information.
[0343] Step 3:
[0344] The server sends the generated metadata to the user's terminal. Compression techniques may be used during this process to improve data transfer efficiency. The input is the generated metadata, and the output is the metadata sent to the user's terminal.
[0345] Step 4:
[0346] The device receives metadata sent from the server and acquires audio and visual data from its built-in microphone and camera to analyze the user's emotional state. The emotion engine analyzes this data and estimates the user's emotions. The input is the audio and visual data acquired by the user's device, and the output is the estimated result regarding the user's emotional state.
[0347] Step 5:
[0348] The device generates an interactive viewing experience based on received metadata and sentiment analysis results. It dynamically adjusts the visual and audio presentation according to the user's emotional state, providing optimal content. The input is metadata and sentiment analysis results, while the output is the optimized content provided to the user.
[0349] Step 6:
[0350] Users watch optimized content and select scenes that interest them. They can then clip these selected scenes and share them with other users on social media. The input is the scenes the user selects as of viewing, and the output is the content snippet that is shared.
[0351] (Application Example 2)
[0352] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0353] When users watch video content, general content is provided uniformly to all users, making it difficult to provide an optimal viewing experience tailored to individual emotions and interests. Furthermore, there is no means to dynamically adjust content in real time in response to changes in user emotions, which can lead to decreased viewing satisfaction. Therefore, there is a need for technology that enables interactive and optimized content delivery in response to user emotions.
[0354] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0355] In this invention, the server includes means for using a generative model that analyzes video data specifically for a user-selected subject, means for generating metadata related to the subject based on the analysis, and means for distributing the generated metadata to the user terminal. This enables the generation and real-time distribution of optimal live commentary content that is tailored to the user's emotions.
[0356] A "generative model" is a mathematical or algorithmic structure used to analyze input data and extract specific patterns or features.
[0357] Metadata is additional information related to data, representing attributes and conditions concerning a specific subject or content.
[0358] A "user terminal" is an electronic device used by a user for direct operation, enabling them to view and manipulate content.
[0359] "Live commentary content" refers to dynamic viewing content that is provided in real time and generated in response to user interaction.
[0360] An "emotion engine" is an algorithm or system for analyzing and inferring a user's emotions, and for analyzing data based on the user's responses.
[0361] "Dynamic optimization" is the process of adjusting content based on real-time data and conditions to maintain an optimized state for the user.
[0362] "Interactive features" are functions that allow users to interact with and respond to the system in a two-way manner, improving participation and responsiveness.
[0363] The system implementing this invention mainly consists of three components: a server, a terminal, and a user. The server receives video data, analyzes a specific object in real time using a generation AI model, and generates metadata based on the results. This generated metadata is then transmitted from the server to the user's terminal.
[0364] The device generates live commentary content based on received metadata and displays it to the user in real time. Furthermore, the device has a built-in emotion engine that analyzes the user's facial expressions and tone of voice to estimate the user's emotions. Based on this emotion analysis, the live commentary content is dynamically optimized according to the user's current emotions.
[0365] For example, if the device detects that a user is excited while watching live video, it can enhance the audio and video of the live content accordingly, thereby enriching the viewing experience.
[0366] As a concrete example, consider a scenario where a user is watching a live concert of their favorite artist on their smartphone. The server recognizes the artist's appearance from the live video, generates relevant metadata, and automatically sends it to the user's device. At this time, the smartphone's camera and microphone capture the user's facial expressions and voice, which are then analyzed by an emotion engine. As a result, if the user is smiling, live commentary content that emphasizes upbeat music and positive comments will be displayed.
[0367] An example of a prompt statement might be, "How to optimize content when the user is happy." This prompt statement serves as a concrete example of how the generative model contributes to optimizing content in response to the user's emotions.
[0368] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0369] Step 1:
[0370] The server receives video data and begins analysis using a generative AI model. The server receives streaming video data as input, recognizes the object, and generates metadata related to the object as a result. The generative AI model detects specific patterns and features in the video and constructs metadata based on these.
[0371] Step 2:
[0372] Metadata generated from the server is delivered to the terminal. This metadata contains information about a specific video scene, and the terminal receives this metadata. The delivered metadata is then used for real-time content generation.
[0373] Step 3:
[0374] The device uses an emotion engine to analyze the user's emotions. It acquires facial expressions and voice tone as input data from the user's facial recognition camera and microphone, and analyzes this data using the emotion engine. This analysis process estimates the user's emotions (e.g., joy, excitement).
[0375] Step 4:
[0376] The device generates commentary content based on received metadata and sentiment analysis results. The real-time optimized commentary content dynamically adjusts video and audio, taking into account metadata information and user sentiment. For example, if the system determines that the user is happy, positive commentary and music will be emphasized in the commentary content.
[0377] Step 5:
[0378] Users can watch live commentary content generated on their devices, select and clip scenes of interest, and share them on social media. The device automatically edits the selected scenes based on user interaction and outputs them in a sharing format.
[0379] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0380] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0381] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0382] [Third Embodiment]
[0383] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0384] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0385] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0386] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0387] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0388] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0389] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0390] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0391] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0392] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0393] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0394] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0395] To implement this invention, the system employs a configuration in which a server, a terminal, and a user work in cooperation. First, the server receives video data and analyzes it using a generative model. The purpose of the analysis is to extract elements associated with a specific character, artist, or other subject selected by the user. This identifies interesting scenes and related information in the video and creates metadata.
[0396] The user's device receives the metadata delivered from the server. Based on this metadata, the device generates customized live commentary content for the user. This commentary is implemented, for example, through real-time updated text overlays or synthesized speech explanations.
[0397] Users can take advantage of interactive features while watching this live content provided through their devices. Specifically, users are given the means to cut out specific scenes from the video and share those scenes on other social media platforms. This allows users to easily share content related to their favorite things and enjoy it with their friends.
[0398] As a concrete example, suppose a user is watching a music program. The server receives the program's video, analyzes the scenes featuring the user's favorite artist, and generates metadata for those scenes. The device then uses this metadata to generate live commentary content about the artist's performance and provides it to the user in real time. The user can then select specific parts of the performance while watching this commentary and share those memorable moments on social media.
[0399] In this way, the system can provide content tailored to the information that users are interested in, making the experience of supporting one's favorite idols even more fulfilling.
[0400] The following describes the processing flow.
[0401] Step 1:
[0402] The server receives video data in real time from television and streaming services. The received video data is then sent directly to the analysis process.
[0403] Step 2:
[0404] The server analyzes the video data using a generative model. This analysis utilizes facial recognition and feature analysis to identify the user's favorite characters or artists.
[0405] Step 3:
[0406] Based on the analysis results, the server generates metadata related to the identified target. This metadata includes a timestamp, the target's name, and related episode information.
[0407] Step 4:
[0408] The server delivers the generated metadata to the user's device. The delivery is performed with low latency to avoid affecting the user's viewing experience.
[0409] Step 5:
[0410] The device generates live commentary content in real time based on the received metadata. Specifically, it provides text information related to the target scene and narration generated by speech synthesis.
[0411] Step 6:
[0412] Users can select scenes of interest while watching the live commentary content provided on their device. A clipping option is available at this time.
[0413] Step 7:
[0414] Users can easily share the clipped scenes on social media. The device provides a dedicated UI for this purpose, facilitating sharing.
[0415] In this way, the entire system works together to provide a content experience tailored to the user's preferences.
[0416] (Example 1)
[0417] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0418] Traditional systems struggled to efficiently extract information related to specific user selections and provide customized content in real time, resulting in users expending considerable effort to find the necessary information amidst overwhelming amounts of data. Solving this problem is crucial.
[0419] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0420] In this invention, the server includes means for using a generative model that analyzes visual information specifically for an object selected by the user, means for generating supplementary information related to the object based on the analysis, and means for distributing the generated supplementary information to the user's device. This makes it possible to effectively and quickly provide specific information of interest to the user and improve usability.
[0421] The term "user" refers to a person who uses a system to receive and manipulate visual information.
[0422] "Visual information" is a broad term referring to information formats that include video data such as moving images and still images.
[0423] A "generative model" refers to an algorithm or machine learning technique used to analyze elements related to a specific subject.
[0424] "Incidental information" refers to additional data or metadata related to the subject, generated based on the analyzed visual information.
[0425] "User equipment" refers to devices such as computers and mobile devices that users use to interface with the system.
[0426] "Real-time supplementary content" refers to explanatory and annotated information that is generated in real time and provided to the user.
[0427] "Interactive features" refer to means of operation that allow users to interact with the system and utilize specific information or functions.
[0428] This invention provides an information processing system in which a server, a terminal, and a user work in cooperation. Specific embodiments are described below.
[0429] Server Functions
[0430] The server receives video data and is equipped with a generative AI model to analyze it. This model is used to recognize specific subjects (characters or artists selected by the user) contained in the video data and extract related elements. Based on the extracted information, it generates supplementary information related to the subject. As a concrete example, one could input a prompt message into the generative AI model such as, "Analyze the scenes featuring a specific artist contained in this video data and generate related metadata."
[0431] Device functions
[0432] The terminal receives supplementary information delivered from the server and generates timely, user-oriented support content based on that information. This involves using software to display subtitles and supplementary information in real time. It is also possible to provide commentary on the video using speech synthesis technology.
[0433] User functions
[0434] Users can view supplementary content provided through their devices while utilizing interactive features. For example, they can easily extract specific video scenes and share them on social media platforms. This feature allows users to quickly and easily share information they are interested in with others.
[0435] The entire system aims to efficiently process information related to specific targets selected by the user and deliver customized content in real time.
[0436] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0437] Step 1:
[0438] The server receives visual information specified by the user. This visual information is used as input to send prompt messages to a generating AI model, which analyzes elements related to a specific object within the video data. Specifically, the received video is divided into frames, which are then input into the machine learning model for analysis. The output generates supplementary information based on the extracted elements and data.
[0439] Step 2:
[0440] The server structures the generated supplementary information and distributes it to the user's device. The specific actions performed in this process involve packaging the supplementary information in JSON or XML format and sending it to the terminal via the network. The input is the supplementary information resulting from the analysis, and the output is a data packet receivable by the terminal.
[0441] Step 3:
[0442] The terminal receives supplementary information sent from the server. Using the received data as input, it generates immediate supplementary content (e.g., text overlays, audio descriptions). Specifically, it analyzes the received information and generates a script to display it to the user at the appropriate time. As output, it provides visual or audio content that the user can view.
[0443] Step 4:
[0444] The user uses interactive features while viewing supplementary content provided through the device. As input, the user provides the device with instructions to select a specific scene. Based on these instructions, the device extracts the selected scene and converts it into a format that the user can share on other platforms. The output is a shareable media file.
[0445] Through the process described above, the system aims to effectively provide users with information of interest and further improve the user experience.
[0446] (Application Example 1)
[0447] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0448] Traditional content distribution methods had limited means of effectively providing relevant information when users focused on specific topics during viewing. As a result, the viewing experience was limited due to insufficient information tailored to users' interests, and the convenience of sharing specific scenes was also reduced.
[0449] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0450] In this invention, the server includes means for using a generative model that analyzes video information specifically for a user-selected object, means for generating information related to the object based on the analysis, and means for distributing the generated information to the user's device. This makes it possible to provide the user with information related to a specific object in real time and to improve the viewing experience through an interactive function based on that information.
[0451] A "user" is an individual who views content provided by the system and utilizes its interactive features.
[0452] "Target" refers to characters, artists, or other entities that users are interested in and want to obtain specific information about.
[0453] "Visual information" refers to media containing visual or audio data, and includes real-time or recorded content.
[0454] A "generative model" is an algorithm or artificial intelligence used to analyze video information and extract information related to a specific subject.
[0455] "Related information" refers to additional information such as data and explanations generated from the analyzed video information in a way that is relevant to the subject.
[0456] "User device" refers to a device used by a user to view content and receive explanations, and includes smartphones, smart glasses, and other similar devices.
[0457] "Distribution" refers to the act of sending generated and related information to a user's device via a network.
[0458] "Explanatory content" refers to additional information, including text presentations and audio output, that is provided in real time.
[0459] The "dialogue function" refers to a feature that allows users to extract specific scenes from content and share them on other platforms.
[0460] To implement this invention, the server first receives video information and extracts information related to the subject using a generative model. Specifically, it focuses on a particular character or artist within the video information and uses a generative AI model to analyze interesting scenes and related data. At this time, the server uses a cloud server equipped with a high-performance GPU to achieve efficient video analysis.
[0461] The information obtained through analysis is stored on a server as metadata representing the generated state, and this metadata is distributed to the user's device via the network. The user's device is a smartphone or smart glasses, which generates and displays explanatory content in real time based on the received metadata. The explanatory content includes text presentation and audio output, providing information through the user's sight and hearing.
[0462] Based on this explanatory content, users can easily extract specific scenes on their devices. These extracted scenes can then be shared with other platforms through a dialogue function. This allows users to share their interests and enjoy a richer viewing experience.
[0463] For example, if a user is watching a live music stream, the server analyzes the live video in real time and detects scenes in which the featured artist appears. The generated metadata is sent to the user's device, providing a commentary on the artist's performance in that scene. Based on this commentary, the user can capture memorable moments and share them on social media, following a prompt such as, "Detect the artist's solo section and generate related information."
[0464] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0465] Step 1:
[0466] The server receives video information transmitted from the user's device. At this stage, the input is streaming video data, and the output is video data converted into a format that can be processed in the next analysis step. Specifically, data preprocessing is performed, such as noise reduction and extraction of necessary parts of the video.
[0467] Step 2:
[0468] The server inputs the pre-processed video information into a generative AI model for analysis. This input is the video data transformed in step 1. The generative AI model detects specific characters or artists within the video and analyzes interesting scenes. The output generates metadata related to the subject. This metadata includes timestamps and features of specific scenes.
[0469] Step 3:
[0470] The server distributes the generated metadata to the user's device. The input is the metadata obtained in step 2. This is transmitted to the user's device via the network, resulting in output that enables the delivery of rich content. Specifically, a low-latency and highly efficient transmission protocol is used to maintain real-time performance.
[0471] Step 4:
[0472] The user's device generates and displays explanatory content in real time using the received metadata. The input at this stage is metadata received from the server. The device provides the user with relevant information through text and speech synthesis. The output is a customized explanation that is conveyed through the user's sight and hearing.
[0473] Step 5:
[0474] The user selects a specific scene from a video they are watching on their device and extracts that scene. The input is the video scene the user is interested in, and the output is a clip of the selected scene. The interface responds to the user's actions and provides feedback until the selection is complete.
[0475] Step 6:
[0476] Users share the clipped scenes to other platforms through a dialogue function. The input is the video clip obtained in step 5. The output is the shared video clip, which is then posted to social media and other video platforms. Specifically, the clip is encoded and converted to the appropriate format.
[0477] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0478] In this embodiment of the invention, the system is designed to function in cooperation with three parties: a server, a terminal, and a user. First, the server receives video data in real time and analyzes it using a generative model. This analysis process identifies specific objects within the video and generates associated metadata.
[0479] Next, this system incorporates an emotion engine to identify the user's emotions. When a user interacts with the system through a terminal, the emotion engine analyzes the user's tone of voice, facial expressions, and input data. This analysis is used to estimate the user's emotions.
[0480] Metadata generated on the server is delivered to the user's terminal in real time, and the terminal generates commentary content based on this data and the output of the emotion engine. This commentary content is optimized to provide information that takes into account the user's current emotional state. For example, if the user is expressing joy, content that emphasizes positive commentary and musical elements will be displayed.
[0481] Users can view live streaming content displayed on their devices, clip specific scenes of interest, and share them on social media. Furthermore, the system suggests content recommendations that take into account the user's emotions. This enables a flexible content experience tailored to the user's interests and feelings.
[0482] As a concrete example, consider a scenario where a user is watching live video. The server identifies scenes featuring the artist in real time and generates relevant metadata. Furthermore, when the user reacts using their device, the emotion engine analyzes the user's emotions based on that reaction. If the user shows an excited reaction, the device adjusts the live commentary content by enhancing the audio and changing the camera angle to match that excitement, enriching the user's viewing experience.
[0483] In this way, the system can provide an interactive content experience optimized for both the user's choices and their emotions.
[0484] The following describes the processing flow.
[0485] Step 1:
[0486] The server receives video data in real time. After receiving the data, it starts analyzing the video using a generative model to identify the object specified by the user. After the analysis, it generates metadata related to the object.
[0487] Step 2:
[0488] The server delivers the identified metadata to the user's terminal. Simultaneously, the analysis results are transferred with low latency so that they can be used on the terminal.
[0489] Step 3:
[0490] The device receives the delivered metadata and generates live commentary content to display to the user. This content is customized according to the user's selection.
[0491] Step 4:
[0492] The device utilizes an emotion engine. It analyzes data (voice and facial expressions) collected through the user's microphone and camera to estimate the user's current emotions.
[0493] Step 5:
[0494] The device dynamically adjusts the live commentary content based on the results of the emotion engine. For example, if the user is surprised, it will emphasize more fun content and recommendations that match that emotion.
[0495] Step 6:
[0496] Users can enjoy customized live commentary content provided on their devices in real time. Furthermore, they can easily cut out scenes of interest and share them on social media through interactive operations.
[0497] Step 7:
[0498] The server collects feedback data for future content delivery based on user sentiment and interaction data. This data is used for analysis to further improve the user experience.
[0499] Through this series of steps, the system delivers an immersive content experience tailored to the user's emotions and preferences.
[0500] (Example 2)
[0501] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0502] Modern viewing platforms fail to adequately provide content tailored to the individual user's emotions and state of mind, making it difficult to deliver an experience that aligns with user interests and feelings. Generating relevant information in real time from massive amounts of video data and dynamically adjusting content using that information is particularly challenging. Furthermore, features for easily extracting and sharing scenes of interest are limited, highlighting the need for methods to enhance user engagement.
[0503] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0504] In this invention, the server includes means for using a generative model that analyzes visual information specifically for an object selected by the user, means for generating information related to the object based on the analysis, and means for distributing the generated information to the user terminal. This makes it possible to provide a customized, interactive viewing experience for each user in real time.
[0505] "Visual information" refers to image and video data obtained from users, and the process involves analyzing this data to extract meaningful information.
[0506] A "generative model" is an algorithm or system that learns specific patterns or features from input data and generates new information or predictions based on them.
[0507] "Information" refers to data obtained as a result of analysis and processing, and is a concept that includes metadata and analysis results.
[0508] A "user terminal" refers to a device used by a user for direct operation, and includes a group of devices such as smartphones, tablets, and personal computers.
[0509] An "interactive viewing experience" means dynamically changing content in real time in response to user input and reactions, providing an experience optimized for each individual user.
[0510] "Dynamic adjustment" refers to the process by which the system automatically changes the elements and characteristics of content in response to the user's emotions and circumstances.
[0511] "Operational functions" refer to the interfaces and tools provided to users to select or share specific scenes.
[0512] The system of this invention functions through the cooperation of three parties: a server, a terminal, and a user. The server receives visual information in real time from an external source. Then, utilizing a generative AI model, it analyzes this visual information to generate information tailored to the user's interests and choices. This generative AI model employs advanced pattern recognition algorithms, enabling it to quickly and accurately identify specific objects within the visual information and generate related metadata.
[0513] Information generated on the server is delivered to the user's device in real time. Along with the received information, the device utilizes audio and visual data acquired from its built-in microphone and camera to analyze the user's emotional state. The emotion engine analyzes this data to estimate the user's emotions. This allows the device to dynamically adjust visual and audio information to provide the user with the most suitable interactive viewing experience.
[0514] For example, consider a scenario where a user is watching a live music performance. The server identifies the artist's appearance in real time and generates relevant information. The device detects the user's smile and dynamically adjusts the audio and video effects based on that positive emotion. This functionality allows the user to enjoy a more immersive content experience.
[0515] An example of a prompt message would be: "React sensitively to the artist's appearance in the live video. Describe in three lines why the user would be excited, and then generate commentary content that matches that reaction." By providing this prompt message to the server, the generation AI model will perform the optimal content generation.
[0516] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0517] Step 1:
[0518] The server receives visual information from an external source in real time. The received video data is preprocessed for use in a generating AI model. This preprocessing includes resolution adjustment and noise filtering. The input is raw video data, and the output is data in a format suitable for analysis.
[0519] Step 2:
[0520] The server inputs pre-processed visual information into a generating AI model to identify specific objects. This identification process generates metadata, including symmetrical appearance timing and position information, such as the artist's entrance scene. The input is pre-processed visual information, and the output is metadata based on the identified information.
[0521] Step 3:
[0522] The server sends the generated metadata to the user's terminal. Compression techniques may be used during this process to improve data transfer efficiency. The input is the generated metadata, and the output is the metadata sent to the user's terminal.
[0523] Step 4:
[0524] The device receives metadata sent from the server and acquires audio and visual data from its built-in microphone and camera to analyze the user's emotional state. The emotion engine analyzes this data and estimates the user's emotions. The input is the audio and visual data acquired by the user's device, and the output is the estimated result regarding the user's emotional state.
[0525] Step 5:
[0526] The device generates an interactive viewing experience based on received metadata and sentiment analysis results. It dynamically adjusts the visual and audio presentation according to the user's emotional state, providing optimal content. The input is metadata and sentiment analysis results, while the output is the optimized content provided to the user.
[0527] Step 6:
[0528] Users watch optimized content and select scenes that interest them. They can then clip these selected scenes and share them with other users on social media. The input is the scenes the user selects as of viewing, and the output is the content snippet that is shared.
[0529] (Application Example 2)
[0530] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0531] When users watch video content, general content is provided uniformly to all users, making it difficult to provide an optimal viewing experience tailored to individual emotions and interests. Furthermore, there is no means to dynamically adjust content in real time in response to changes in user emotions, which can lead to decreased viewing satisfaction. Therefore, there is a need for technology that enables interactive and optimized content delivery in response to user emotions.
[0532] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0533] In this invention, the server includes means for using a generative model that analyzes video data specifically for a user-selected subject, means for generating metadata related to the subject based on the analysis, and means for distributing the generated metadata to the user terminal. This enables the generation and real-time distribution of optimal live commentary content that is tailored to the user's emotions.
[0534] A "generative model" is a mathematical or algorithmic structure used to analyze input data and extract specific patterns or features.
[0535] Metadata is additional information related to data, representing attributes and conditions concerning a specific subject or content.
[0536] A "user terminal" is an electronic device used by a user for direct operation, enabling them to view and manipulate content.
[0537] "Live commentary content" refers to dynamic viewing content that is provided in real time and generated in response to user interaction.
[0538] An "emotion engine" is an algorithm or system for analyzing and inferring a user's emotions, and for analyzing data based on the user's responses.
[0539] "Dynamic optimization" is the process of adjusting content based on real-time data and conditions to maintain an optimized state for the user.
[0540] "Interactive features" are functions that allow users to interact with and respond to the system in a two-way manner, improving participation and responsiveness.
[0541] The system implementing this invention mainly consists of three components: a server, a terminal, and a user. The server receives video data, analyzes a specific object in real time using a generation AI model, and generates metadata based on the results. This generated metadata is then transmitted from the server to the user's terminal.
[0542] The device generates live commentary content based on received metadata and displays it to the user in real time. Furthermore, the device has a built-in emotion engine that analyzes the user's facial expressions and tone of voice to estimate the user's emotions. Based on this emotion analysis, the live commentary content is dynamically optimized according to the user's current emotions.
[0543] For example, if the device detects that a user is excited while watching live video, it can enhance the audio and video of the live content accordingly, thereby enriching the viewing experience.
[0544] As a concrete example, consider a scenario where a user is watching a live concert of their favorite artist on their smartphone. The server recognizes the artist's appearance from the live video, generates relevant metadata, and automatically sends it to the user's device. At this time, the smartphone's camera and microphone capture the user's facial expressions and voice, which are then analyzed by an emotion engine. As a result, if the user is smiling, live commentary content that emphasizes upbeat music and positive comments will be displayed.
[0545] An example of a prompt statement might be, "How to optimize content when the user is happy." This prompt statement serves as a concrete example of how the generative model contributes to optimizing content in response to the user's emotions.
[0546] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0547] Step 1:
[0548] The server receives video data and begins analysis using a generative AI model. The server receives streaming video data as input, recognizes the object, and generates metadata related to the object as a result. The generative AI model detects specific patterns and features in the video and constructs metadata based on these.
[0549] Step 2:
[0550] Metadata generated from the server is delivered to the terminal. This metadata contains information about a specific video scene, and the terminal receives this metadata. The delivered metadata is then used for real-time content generation.
[0551] Step 3:
[0552] The device uses an emotion engine to analyze the user's emotions. It acquires facial expressions and voice tone as input data from the user's facial recognition camera and microphone, and analyzes this data using the emotion engine. This analysis process estimates the user's emotions (e.g., joy, excitement).
[0553] Step 4:
[0554] The device generates commentary content based on received metadata and sentiment analysis results. The real-time optimized commentary content dynamically adjusts video and audio, taking into account metadata information and user sentiment. For example, if the system determines that the user is happy, positive commentary and music will be emphasized in the commentary content.
[0555] Step 5:
[0556] Users can watch live commentary content generated on their devices, select and clip scenes of interest, and share them on social media. The device automatically edits the selected scenes based on user interaction and outputs them in a sharing format.
[0557] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0558] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0559] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[0560] [Fourth Embodiment]
[0561] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0562] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0563] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0564] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0565] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0566] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0567] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0568] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0569] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0570] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0571] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0572] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0573] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0574] To implement this invention, the system employs a configuration in which a server, a terminal, and a user work in cooperation. First, the server receives video data and analyzes it using a generative model. The purpose of the analysis is to extract elements associated with a specific character, artist, or other subject selected by the user. This identifies interesting scenes and related information in the video and creates metadata.
[0575] The user's device receives the metadata delivered from the server. Based on this metadata, the device generates customized live commentary content for the user. This commentary is implemented, for example, through real-time updated text overlays or synthesized speech explanations.
[0576] Users can take advantage of interactive features while watching this live content provided through their devices. Specifically, users are given the means to cut out specific scenes from the video and share those scenes on other social media platforms. This allows users to easily share content related to their favorite things and enjoy it with their friends.
[0577] As a concrete example, suppose a user is watching a music program. The server receives the program's video, analyzes the scenes featuring the user's favorite artist, and generates metadata for those scenes. The device then uses this metadata to generate live commentary content about the artist's performance and provides it to the user in real time. The user can then select specific parts of the performance while watching this commentary and share those memorable moments on social media.
[0578] In this way, the system can provide content tailored to the information that users are interested in, making the experience of supporting one's favorite idols even more fulfilling.
[0579] The following describes the processing flow.
[0580] Step 1:
[0581] The server receives video data in real time from television and streaming services. The received video data is then sent directly to the analysis process.
[0582] Step 2:
[0583] The server analyzes the video data using a generative model. This analysis utilizes facial recognition and feature analysis to identify the user's favorite characters or artists.
[0584] Step 3:
[0585] Based on the analysis results, the server generates metadata related to the identified target. This metadata includes a timestamp, the target's name, and related episode information.
[0586] Step 4:
[0587] The server delivers the generated metadata to the user's device. The delivery is performed with low latency to avoid affecting the user's viewing experience.
[0588] Step 5:
[0589] The device generates live commentary content in real time based on the received metadata. Specifically, it provides text information related to the target scene and narration generated by speech synthesis.
[0590] Step 6:
[0591] Users can select scenes of interest while watching the live commentary content provided on their device. A clipping option is available at this time.
[0592] Step 7:
[0593] Users can easily share the clipped scenes on social media. The device provides a dedicated UI for this purpose, facilitating sharing.
[0594] In this way, the entire system works together to provide a content experience tailored to the user's preferences.
[0595] (Example 1)
[0596] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0597] Traditional systems struggled to efficiently extract information related to specific user selections and provide customized content in real time, resulting in users expending considerable effort to find the necessary information amidst overwhelming amounts of data. Solving this problem is crucial.
[0598] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0599] In this invention, the server includes means for using a generative model that analyzes visual information specifically for an object selected by the user, means for generating supplementary information related to the object based on the analysis, and means for distributing the generated supplementary information to the user's device. This makes it possible to effectively and quickly provide specific information of interest to the user and improve usability.
[0600] The term "user" refers to a person who uses a system to receive and manipulate visual information.
[0601] "Visual information" is a broad term referring to information formats that include video data such as moving images and still images.
[0602] A "generative model" refers to an algorithm or machine learning technique used to analyze elements related to a specific subject.
[0603] "Incidental information" refers to additional data or metadata related to the subject, generated based on the analyzed visual information.
[0604] "User equipment" refers to devices such as computers and mobile devices that users use to interface with the system.
[0605] "Real-time supplementary content" refers to explanatory and annotated information that is generated in real time and provided to the user.
[0606] "Interactive features" refer to means of operation that allow users to interact with the system and utilize specific information or functions.
[0607] This invention provides an information processing system in which a server, a terminal, and a user work in cooperation. Specific embodiments are described below.
[0608] Server Functions
[0609] The server receives video data and is equipped with a generative AI model to analyze it. This model is used to recognize specific subjects (characters or artists selected by the user) contained in the video data and extract related elements. Based on the extracted information, it generates supplementary information related to the subject. As a concrete example, one could input a prompt message into the generative AI model such as, "Analyze the scenes featuring a specific artist contained in this video data and generate related metadata."
[0610] Device functions
[0611] The terminal receives supplementary information delivered from the server and generates timely, user-oriented support content based on that information. This involves using software to display subtitles and supplementary information in real time. It is also possible to provide commentary on the video using speech synthesis technology.
[0612] User functions
[0613] Users can view supplementary content provided through their devices while utilizing interactive features. For example, they can easily extract specific video scenes and share them on social media platforms. This feature allows users to quickly and easily share information they are interested in with others.
[0614] The entire system aims to efficiently process information related to specific targets selected by the user and deliver customized content in real time.
[0615] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0616] Step 1:
[0617] The server receives visual information specified by the user. This visual information is used as input to send prompt messages to a generating AI model, which analyzes elements related to a specific object within the video data. Specifically, the received video is divided into frames, which are then input into the machine learning model for analysis. The output generates supplementary information based on the extracted elements and data.
[0618] Step 2:
[0619] The server structures the generated supplementary information and distributes it to the user's device. The specific actions performed in this process involve packaging the supplementary information in JSON or XML format and sending it to the terminal via the network. The input is the supplementary information resulting from the analysis, and the output is a data packet receivable by the terminal.
[0620] Step 3:
[0621] The terminal receives supplementary information sent from the server. Using the received data as input, it generates immediate supplementary content (e.g., text overlays, audio descriptions). Specifically, it analyzes the received information and generates a script to display it to the user at the appropriate time. As output, it provides visual or audio content that the user can view.
[0622] Step 4:
[0623] The user uses interactive features while viewing supplementary content provided through the device. As input, the user provides the device with instructions to select a specific scene. Based on these instructions, the device extracts the selected scene and converts it into a format that the user can share on other platforms. The output is a shareable media file.
[0624] Through the process described above, the system aims to effectively provide users with information of interest and further improve the user experience.
[0625] (Application Example 1)
[0626] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0627] Traditional content distribution methods had limited means of effectively providing relevant information when users focused on specific topics during viewing. As a result, the viewing experience was limited due to insufficient information tailored to users' interests, and the convenience of sharing specific scenes was also reduced.
[0628] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0629] In this invention, the server includes means for using a generative model that analyzes video information specifically for a user-selected object, means for generating information related to the object based on the analysis, and means for distributing the generated information to the user's device. This makes it possible to provide the user with information related to a specific object in real time and to improve the viewing experience through an interactive function based on that information.
[0630] A "user" is an individual who views content provided by the system and utilizes its interactive features.
[0631] "Target" refers to characters, artists, or other entities that users are interested in and want to obtain specific information about.
[0632] "Visual information" refers to media containing visual or audio data, and includes real-time or recorded content.
[0633] A "generative model" is an algorithm or artificial intelligence used to analyze video information and extract information related to a specific subject.
[0634] "Related information" refers to additional information such as data and explanations generated from the analyzed video information in a way that is relevant to the subject.
[0635] "User device" refers to a device used by a user to view content and receive explanations, and includes smartphones, smart glasses, and other similar devices.
[0636] "Distribution" refers to the act of sending generated and related information to a user's device via a network.
[0637] "Explanatory content" refers to additional information, including text presentations and audio output, that is provided in real time.
[0638] The "dialogue function" refers to a feature that allows users to extract specific scenes from content and share them on other platforms.
[0639] To implement this invention, the server first receives video information and extracts information related to the subject using a generative model. Specifically, it focuses on a particular character or artist within the video information and uses a generative AI model to analyze interesting scenes and related data. At this time, the server uses a cloud server equipped with a high-performance GPU to achieve efficient video analysis.
[0640] The information obtained through analysis is stored on a server as metadata representing the generated state, and this metadata is distributed to the user's device via the network. The user's device is a smartphone or smart glasses, which generates and displays explanatory content in real time based on the received metadata. The explanatory content includes text presentation and audio output, providing information through the user's sight and hearing.
[0641] Based on this explanatory content, users can easily extract specific scenes on their devices. These extracted scenes can then be shared with other platforms through a dialogue function. This allows users to share their interests and enjoy a richer viewing experience.
[0642] For example, if a user is watching a live music stream, the server analyzes the live video in real time and detects scenes in which the featured artist appears. The generated metadata is sent to the user's device, providing a commentary on the artist's performance in that scene. Based on this commentary, the user can capture memorable moments and share them on social media, following a prompt such as, "Detect the artist's solo section and generate related information."
[0643] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0644] Step 1:
[0645] The server receives video information transmitted from the user's device. At this stage, the input is streaming video data, and the output is video data converted into a format that can be processed in the next analysis step. Specifically, data preprocessing is performed, such as noise reduction and extraction of necessary parts of the video.
[0646] Step 2:
[0647] The server inputs the pre-processed video information into a generative AI model for analysis. This input is the video data transformed in step 1. The generative AI model detects specific characters or artists within the video and analyzes interesting scenes. The output generates metadata related to the subject. This metadata includes timestamps and features of specific scenes.
[0648] Step 3:
[0649] The server distributes the generated metadata to the user's device. The input is the metadata obtained in step 2. This is transmitted to the user's device via the network, resulting in output that enables the delivery of rich content. Specifically, a low-latency and highly efficient transmission protocol is used to maintain real-time performance.
[0650] Step 4:
[0651] The user's device generates and displays explanatory content in real time using the received metadata. The input at this stage is metadata received from the server. The device provides the user with relevant information through text and speech synthesis. The output is a customized explanation that is conveyed through the user's sight and hearing.
[0652] Step 5:
[0653] The user selects a specific scene from a video they are watching on their device and extracts that scene. The input is the video scene the user is interested in, and the output is a clip of the selected scene. The interface responds to the user's actions and provides feedback until the selection is complete.
[0654] Step 6:
[0655] Users share the clipped scenes to other platforms through a dialogue function. The input is the video clip obtained in step 5. The output is the shared video clip, which is then posted to social media and other video platforms. Specifically, the clip is encoded and converted to the appropriate format.
[0656] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0657] In this embodiment of the invention, the system is designed to function in cooperation with three parties: a server, a terminal, and a user. First, the server receives video data in real time and analyzes it using a generative model. This analysis process identifies specific objects within the video and generates associated metadata.
[0658] Next, this system incorporates an emotion engine to identify the user's emotions. When a user interacts with the system through a terminal, the emotion engine analyzes the user's tone of voice, facial expressions, and input data. This analysis is used to estimate the user's emotions.
[0659] Metadata generated on the server is delivered to the user's terminal in real time, and the terminal generates commentary content based on this data and the output of the emotion engine. This commentary content is optimized to provide information that takes into account the user's current emotional state. For example, if the user is expressing joy, content that emphasizes positive commentary and musical elements will be displayed.
[0660] Users can view live streaming content displayed on their devices, clip specific scenes of interest, and share them on social media. Furthermore, the system suggests content recommendations that take into account the user's emotions. This enables a flexible content experience tailored to the user's interests and feelings.
[0661] As a concrete example, consider a scenario where a user is watching live video. The server identifies scenes featuring the artist in real time and generates relevant metadata. Furthermore, when the user reacts using their device, the emotion engine analyzes the user's emotions based on that reaction. If the user shows an excited reaction, the device adjusts the live commentary content by enhancing the audio and changing the camera angle to match that excitement, enriching the user's viewing experience.
[0662] In this way, the system can provide an interactive content experience optimized for both the user's choices and their emotions.
[0663] The following describes the processing flow.
[0664] Step 1:
[0665] The server receives video data in real time. After receiving the data, it starts analyzing the video using a generative model to identify the object specified by the user. After the analysis, it generates metadata related to the object.
[0666] Step 2:
[0667] The server delivers the identified metadata to the user's terminal. Simultaneously, the analysis results are transferred with low latency so that they can be used on the terminal.
[0668] Step 3:
[0669] The device receives the delivered metadata and generates live commentary content to display to the user. This content is customized according to the user's selection.
[0670] Step 4:
[0671] The device utilizes an emotion engine. It analyzes data (voice and facial expressions) collected through the user's microphone and camera to estimate the user's current emotions.
[0672] Step 5:
[0673] The device dynamically adjusts the live commentary content based on the results of the emotion engine. For example, if the user is surprised, it will emphasize more fun content and recommendations that match that emotion.
[0674] Step 6:
[0675] Users can enjoy customized live commentary content provided on their devices in real time. Furthermore, they can easily cut out scenes of interest and share them on social media through interactive operations.
[0676] Step 7:
[0677] The server collects feedback data for future content delivery based on user sentiment and interaction data. This data is used for analysis to further improve the user experience.
[0678] Through this series of steps, the system delivers an immersive content experience tailored to the user's emotions and preferences.
[0679] (Example 2)
[0680] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0681] Modern viewing platforms fail to adequately provide content tailored to the individual user's emotions and state of mind, making it difficult to deliver an experience that aligns with user interests and feelings. Generating relevant information in real time from massive amounts of video data and dynamically adjusting content using that information is particularly challenging. Furthermore, features for easily extracting and sharing scenes of interest are limited, highlighting the need for methods to enhance user engagement.
[0682] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0683] In this invention, the server includes means for using a generative model that analyzes visual information specifically for an object selected by the user, means for generating information related to the object based on the analysis, and means for distributing the generated information to the user terminal. This makes it possible to provide a customized, interactive viewing experience for each user in real time.
[0684] "Visual information" refers to image and video data obtained from users, and the process involves analyzing this data to extract meaningful information.
[0685] A "generative model" is an algorithm or system that learns specific patterns or features from input data and generates new information or predictions based on them.
[0686] "Information" refers to data obtained as a result of analysis and processing, and is a concept that includes metadata and analysis results.
[0687] A "user terminal" refers to a device used by a user for direct operation, and includes a group of devices such as smartphones, tablets, and personal computers.
[0688] An "interactive viewing experience" means dynamically changing content in real time in response to user input and reactions, providing an experience optimized for each individual user.
[0689] "Dynamic adjustment" refers to the process by which the system automatically changes the elements and characteristics of content in response to the user's emotions and circumstances.
[0690] "Operational functions" refer to the interfaces and tools provided to users to select or share specific scenes.
[0691] The system of this invention functions through the cooperation of three parties: a server, a terminal, and a user. The server receives visual information in real time from an external source. Then, utilizing a generative AI model, it analyzes this visual information to generate information tailored to the user's interests and choices. This generative AI model employs advanced pattern recognition algorithms, enabling it to quickly and accurately identify specific objects within the visual information and generate related metadata.
[0692] Information generated on the server is delivered to the user's device in real time. Along with the received information, the device utilizes audio and visual data acquired from its built-in microphone and camera to analyze the user's emotional state. The emotion engine analyzes this data to estimate the user's emotions. This allows the device to dynamically adjust visual and audio information to provide the user with the most suitable interactive viewing experience.
[0693] For example, consider a scenario where a user is watching a live music performance. The server identifies the artist's appearance in real time and generates relevant information. The device detects the user's smile and dynamically adjusts the audio and video effects based on that positive emotion. This functionality allows the user to enjoy a more immersive content experience.
[0694] An example of a prompt message would be: "React sensitively to the artist's appearance in the live video. Describe in three lines why the user would be excited, and then generate commentary content that matches that reaction." By providing this prompt message to the server, the generation AI model will perform the optimal content generation.
[0695] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0696] Step 1:
[0697] The server receives visual information from an external source in real time. The received video data is preprocessed for use in a generating AI model. This preprocessing includes resolution adjustment and noise filtering. The input is raw video data, and the output is data in a format suitable for analysis.
[0698] Step 2:
[0699] The server inputs pre-processed visual information into a generating AI model to identify specific objects. This identification process generates metadata, including symmetrical appearance timing and position information, such as the artist's entrance scene. The input is pre-processed visual information, and the output is metadata based on the identified information.
[0700] Step 3:
[0701] The server sends the generated metadata to the user's terminal. Compression techniques may be used during this process to improve data transfer efficiency. The input is the generated metadata, and the output is the metadata sent to the user's terminal.
[0702] Step 4:
[0703] The device receives metadata sent from the server and acquires audio and visual data from its built-in microphone and camera to analyze the user's emotional state. The emotion engine analyzes this data and estimates the user's emotions. The input is the audio and visual data acquired by the user's device, and the output is the estimated result regarding the user's emotional state.
[0704] Step 5:
[0705] The device generates an interactive viewing experience based on received metadata and sentiment analysis results. It dynamically adjusts the visual and audio presentation according to the user's emotional state, providing optimal content. The input is metadata and sentiment analysis results, while the output is the optimized content provided to the user.
[0706] Step 6:
[0707] Users watch optimized content and select scenes that interest them. They can then clip these selected scenes and share them with other users on social media. The input is the scenes the user selects as of viewing, and the output is the content snippet that is shared.
[0708] (Application Example 2)
[0709] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0710] When users watch video content, general content is provided uniformly to all users, making it difficult to provide an optimal viewing experience tailored to individual emotions and interests. Furthermore, there is no means to dynamically adjust content in real time in response to changes in user emotions, which can lead to decreased viewing satisfaction. Therefore, there is a need for technology that enables interactive and optimized content delivery in response to user emotions.
[0711] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0712] In this invention, the server includes means for using a generative model that analyzes video data specifically for a user-selected subject, means for generating metadata related to the subject based on the analysis, and means for distributing the generated metadata to the user terminal. This enables the generation and real-time distribution of optimal live commentary content that is tailored to the user's emotions.
[0713] A "generative model" is a mathematical or algorithmic structure used to analyze input data and extract specific patterns or features.
[0714] Metadata is additional information related to data, representing attributes and conditions concerning a specific subject or content.
[0715] A "user terminal" is an electronic device used by a user for direct operation, enabling them to view and manipulate content.
[0716] "Live commentary content" refers to dynamic viewing content that is provided in real time and generated in response to user interaction.
[0717] An "emotion engine" is an algorithm or system for analyzing and inferring a user's emotions, and for analyzing data based on the user's responses.
[0718] "Dynamic optimization" is the process of adjusting content based on real-time data and conditions to maintain an optimized state for the user.
[0719] "Interactive features" are functions that allow users to interact with and respond to the system in a two-way manner, improving participation and responsiveness.
[0720] The system implementing this invention mainly consists of three components: a server, a terminal, and a user. The server receives video data, analyzes a specific object in real time using a generation AI model, and generates metadata based on the results. This generated metadata is then transmitted from the server to the user's terminal.
[0721] The device generates live commentary content based on received metadata and displays it to the user in real time. Furthermore, the device has a built-in emotion engine that analyzes the user's facial expressions and tone of voice to estimate the user's emotions. Based on this emotion analysis, the live commentary content is dynamically optimized according to the user's current emotions.
[0722] For example, if the device detects that a user is excited while watching live video, it can enhance the audio and video of the live content accordingly, thereby enriching the viewing experience.
[0723] As a concrete example, consider a scenario where a user is watching a live concert of their favorite artist on their smartphone. The server recognizes the artist's appearance from the live video, generates relevant metadata, and automatically sends it to the user's device. At this time, the smartphone's camera and microphone capture the user's facial expressions and voice, which are then analyzed by an emotion engine. As a result, if the user is smiling, live commentary content that emphasizes upbeat music and positive comments will be displayed.
[0724] An example of a prompt statement might be, "How to optimize content when the user is happy." This prompt statement serves as a concrete example of how the generative model contributes to optimizing content in response to the user's emotions.
[0725] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0726] Step 1:
[0727] The server receives video data and begins analysis using a generative AI model. The server receives streaming video data as input, recognizes the object, and generates metadata related to the object as a result. The generative AI model detects specific patterns and features in the video and constructs metadata based on these.
[0728] Step 2:
[0729] Metadata generated from the server is delivered to the terminal. This metadata contains information about a specific video scene, and the terminal receives this metadata. The delivered metadata is then used for real-time content generation.
[0730] Step 3:
[0731] The device uses an emotion engine to analyze the user's emotions. It acquires facial expressions and voice tone as input data from the user's facial recognition camera and microphone, and analyzes this data using the emotion engine. This analysis process estimates the user's emotions (e.g., joy, excitement).
[0732] Step 4:
[0733] The device generates live commentary content based on received metadata and sentiment analysis results. The real-time optimized live commentary content dynamically adjusts video and audio, taking into account metadata information and user sentiment. For example, if the system determines that the user is happy, positive commentary and music will be emphasized in the live commentary content.
[0734] Step 5:
[0735] Users can watch live commentary content generated on their devices, select and clip scenes of interest, and share them on social media. The device automatically edits the selected scenes based on user interaction and outputs them in a sharing format.
[0736] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0737] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0738] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[0739] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0740] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0741] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0742] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0743] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0744] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0745] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0746] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[0747] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[0748] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0749] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0750] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0751] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0752] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0753] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0754] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0755] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0756] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted as being incorporated by reference.
[0757] The following is further disclosed regarding the embodiments described above.
[0758] (Claim 1)
[0759] A method using a generative model that analyzes video data specifically for the target selected by the user,
[0760] A means for generating metadata related to the subject based on the aforementioned analysis,
[0761] A means of delivering the generated metadata to the user's terminal,
[0762] A means for generating and displaying live commentary content in real time based on distributed metadata,
[0763] A means of providing an interactive function for users to crop and share selected scenes,
[0764] A system that includes this.
[0765] (Claim 2)
[0766] The system according to claim 1, comprising image recognition technology for identifying an object from analyzed video data.
[0767] (Claim 3)
[0768] The system according to claim 1, further comprising means for displaying text and synthesizing speech based on generated live commentary content.
[0769] "Example 1"
[0770] (Claim 1)
[0771] A method that uses a generative model to analyze visual information specifically for the object selected by the user,
[0772] A means for generating supplementary information related to the subject based on the aforementioned analysis,
[0773] A means for distributing the generated supplementary information to the user's device,
[0774] A means for generating and displaying timely supplementary content based on the distributed supplementary information,
[0775] A means of providing an interactive function for users to crop and share selected scenes,
[0776] A system that includes this.
[0777] (Claim 2)
[0778] The system according to claim 1, comprising image recognition technology for identifying an object from analyzed visual information.
[0779] (Claim 3)
[0780] The system according to claim 1, further comprising means for outputting and displaying the generated auxiliary content and for speech synthesis.
[0781] "Application Example 1"
[0782] (Claim 1)
[0783] A method using a generative model that analyzes video information specifically for the target selected by the user,
[0784] A means for generating information related to the subject based on the aforementioned analysis,
[0785] A means for distributing the generated information to the user's device,
[0786] A means of generating and displaying explanatory content in real time based on the distributed information,
[0787] A means of providing a dialogue function for users to crop and share selected scenes,
[0788] The added explanatory content includes means for displaying text and outputting audio in specific scenes,
[0789] A system that includes this.
[0790] (Claim 2)
[0791] The system according to claim 1, comprising image analysis technology for identifying an object from analyzed video information.
[0792] (Claim 3)
[0793] The system according to claim 1, further comprising means for displaying text and outputting audio based on generated explanatory content.
[0794] "Example 2 of combining an emotion engine"
[0795] (Claim 1)
[0796] A method that uses a generative model to analyze visual information specifically for the object selected by the user,
[0797] A means for generating information related to the subject based on the aforementioned analysis,
[0798] A means of distributing the generated information to the user's terminal,
[0799] A means for generating and displaying a real-time interactive viewing experience based on the distributed information,
[0800] A means of acquiring audio and visual data to analyze the user's emotional state and dynamically adjusting content based on the analysis results,
[0801] A means of providing operational functions for cropping and sharing scenes selected by the user,
[0802] A system that includes this.
[0803] (Claim 2)
[0804] The system according to claim 1, comprising pattern recognition technology for identifying an object from analyzed visual information.
[0805] (Claim 3)
[0806] The system according to claim 1, comprising means for displaying linguistic information and generating speech based on the generated interactive experience.
[0807] "Application example 2 when combining with an emotional engine"
[0808] (Claim 1)
[0809] A method using a generative model that analyzes video data specifically for the target selected by the user,
[0810] A means for generating metadata related to the subject based on the aforementioned analysis,
[0811] A means of delivering the generated metadata to the user's terminal,
[0812] A means for generating and displaying live commentary content in real time based on distributed metadata,
[0813] A means of using an emotion engine to analyze user emotions,
[0814] A means of dynamically optimizing live commentary content based on analyzed user sentiment,
[0815] A means of providing an interactive function for users to crop and share selected scenes,
[0816] A system that includes this.
[0817] (Claim 2)
[0818] The system according to claim 1, comprising image recognition technology for identifying an object from analyzed video data.
[0819] (Claim 3)
[0820] The system according to claim 1, further comprising means for displaying text and synthesizing speech based on generated live commentary content. [Explanation of Symbols]
[0821] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. A method using a generative model that analyzes video data specifically for the target selected by the user, A means for generating metadata related to the subject based on the aforementioned analysis, A means of delivering the generated metadata to the user's terminal, A means for generating and displaying live commentary content in real time based on distributed metadata, A means of providing an interactive function for users to crop and share selected scenes, A system that includes this.
2. The system according to claim 1, comprising image recognition technology for identifying an object from analyzed video data.
3. The system according to claim 1, further comprising means for displaying text and synthesizing speech based on the generated live commentary content.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A