System

The system addresses the challenge of saving video content information by detecting user interactions, using AI to identify and save objects or information, and integrating with user services, providing a seamless experience.

JP2026021000APending Publication Date: 2026-02-10SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024122682
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-29
Publication Date
2026-02-10

AI Technical Summary

Technical Problem

Conventional methods for identifying and saving information about objects, places, or songs while watching videos require users to pause the video or manually search, leading to interruptions and loss of information.

Method used

A system that detects user interactions with video content, captures frames and audio data, uses generative AI to identify objects or information, saves them in user services, and notifies the user of successful saving, allowing seamless integration with other services.

Benefits of technology

Enables users to easily save and access information of interest without interrupting their viewing experience, enhancing user convenience and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026021000000001_ABST
    Figure 2026021000000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system comprising: means for detecting that a user performs an operation related to a specific object or information while viewing a moving image; means for acquiring a specific frame of the moving image; means for identifying the object or information using a generative AI model; means for storing the identified object or information in a database of another service used by the user; and means for notifying the user of storage completion of the information.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] In today's world, an increasing number of users are watching videos on smartphones and tablets. However, it is difficult to easily identify and save information about objects, places, songs, and other items that interest them while watching. Conventional methods require users to pause the video or manually search for specific information using a separate application, resulting in problems such as interruptions to the viewing experience and loss of information due to forgetting. The present invention aims to enable users to easily save objects or information that interest them while watching a video and use them later without interrupting their viewing. [Means for solving the problem]

[0005] The present invention is a system that includes means for detecting when a user performs an operation related to a specific object or information while watching a video, means for acquiring specific frames or audio data from the video, and video metadata, means for identifying the object or information using a generative AI model based on the acquired data, means for saving the identified object or information in a database of another service used by the user, and means for notifying the user that the information has been saved. This allows users to easily save information about objects, places, songs, and other items of interest without interrupting their video viewing, and seamlessly use it later.

[0006] "Watching a video" refers to a state in which a user is playing video content using a device.

[0007] "Means for detecting when a user performs an action related to a specific object or information" refers to a technical device that identifies actions such as taps and clicks performed by a user while a video is being played.

[0008] "Means for obtaining specific frames of video, audio data, and video metadata" refers to a technical device that automatically collects still image data and audio data from video at the moment a user performs an operation, as well as the video's title and playback time.

[0009] "Means for identifying objects or information using generative AI models" refers to technical devices that use generative AI (such as generative adversarial networks) to identify objects (e.g., costumes, places, songs, etc.) or information from collected data.

[0010] "Means for storing in the databases of other services used by the user" refers to technical devices that store identified objects and information in the databases of external services that the user regularly uses (e.g. shopping sites, music streaming services, etc.).

[0011] "Means for notifying the user that the information has been saved" refers to a technical device that has a notification function to inform the user that the saving operation has been successful. [Brief explanation of the drawings]

[0012] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14]FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0013] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0014] First, the terms used in the following description will be explained.

[0015] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0016] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0017] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0018] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0019] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0020] [First embodiment]

[0021] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0022] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0023] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0024] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0025] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0026] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0027] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0028] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0029] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0030] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0031] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0032] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0033] The present invention relates to a system that identifies objects or information that a user finds interesting while watching a video and saves them in the user's favorites list. This system detects the user's actions, analyzes the video data being watched, and provides a function to save the identified information in other services used by the user. Specific embodiments of the system are described below.

[0034] Capturing user actions

[0035] Users: While watching a video, if they come across an outfit, place, song, etc. that catches their eye, they tap the screen to show their interest. Taps can be single, double, or triple taps, and each type indicates interest in a different category (e.g., object, place, song, etc.).

[0036] Device: Detects the user's tap and determines whether it is a single tap or a double tap. Depending on the detected tap, it collects information such as the video frame at the moment of the tap, audio data, current playback time, and video title.

[0037] Sending tap information

[0038] Device: The collected tap information (frame image, audio data, playback time, video title) is packaged. This package contains all the data necessary for subsequent information analysis.

[0039] Terminal: Sends packaged information to the server.

[0040] Identifying and locating information

[0041] Server: Receives the package information and begins analyzing it. The server is equipped with powerful generative AI models that use techniques such as image and voice recognition to identify specific outfits, locations, songs, etc.

[0042] Server: Based on the analysis results, the server searches for resources on the Internet related to the identified object or information. For example, for an identified outfit, the server retrieves the product page of an online shopping site, and for an identified song, the server retrieves track information from a music streaming service.

[0043] Save to Favorites list

[0044] Server: The acquired information is used to call the API of other services the user uses (e.g., e-commerce sites, music services) and added to the favorites list. For example, the identified outfit is added to the user's favorite item list on a shopping site.

[0045] Order completion notification

[0046] Server: Confirms that the information has been saved successfully and notifies the device of the result.

[0047] On your device: The user will be notified that their information has been saved, allowing them to easily access the information they are interested in later without interrupting their video viewing.

[0048] Specific examples

[0049] Scenario 1: User identifies and purchases an outfit from a video

[0050] 1. User: While watching a video, find an outfit you like and single-tap the screen.

[0051] 2. Device: Detects a single tap and sends the frame image and video title at that time to the server.

[0052] 3. Server: Analyzes the frame images and identifies the outfits using a generative AI model.

[0053] 4. Server: Search the online shopping site and retrieve the product page for the relevant costume.

[0054] 5. Server: Calls the API and registers the identified outfit in the user's favorites list.

[0055] 6. Device: Send a notification to the user saying "The outfit has been added to your favorites list."

[0056] 7. User: After watching, you can check and purchase the costumes from your favorite list.

[0057] Scenario 2: User identifies a song in a video and adds it to a playlist

[0058] 1. User: When a song you like comes on while watching a video, double tap the screen.

[0059] 2. Device: Detects the double-tap operation and sends the audio data and playback time at that time to the server.

[0060] 3. Server: Analyzes the audio data and identifies the song using a generative AI model.

[0061] 4. Server: Searches music streaming services and retrieves track information for identified songs.

[0062] 5. Server: Calls the API and adds the identified song to the user's playlist.

[0063] 6. Device: Send a notification to the user saying "Song has been added to your playlist."

[0064] 7. User: After listening, you can check the song from the playlist and play it.

[0065] As described above, the present invention allows users to easily identify objects or information of interest while watching a video, and to seamlessly use the objects or information later.

[0066] The processing flow will be explained below.

[0067] Step 1:

[0068] User: While watching a video, if they come across an outfit, location, song, etc. that catches their eye, they tap the screen. This tapping action shows the user's interest.

[0069] Step 2:

[0070] Device: Detects user taps in real time and distinguishes between different types of taps (single tap, double tap, triple tap, etc.). For example, a single tap indicates an outfit, a double tap indicates a song, and a triple tap indicates a location.

[0071] Step 3:

[0072] On the device: Obtain the video frame, audio data, playback time, and metadata such as the video title at the moment the touch occurred.

[0073] Step 4:

[0074] Terminal: Packages information such as acquired frame images, audio data, playback time, and video title.

[0075] Step 5:

[0076] Terminal: Sends packaged information to the server.

[0077] Step 6:

[0078] Server: Analyzes the received package information (frame image, audio data, playback time, video title).

[0079] Step 7:

[0080] Server: Using a generative AI model based on the analysis results, it identifies costumes and locations from frame images and identifies songs from audio data.

[0081] Step 8:

[0082] Server: Searches internet resources for the identified object or information and retrieves related product pages or music track information.

[0083] Step 9:

[0084] Server: Calls the API to store the acquired information in the database of other services used by the user.

[0085] Step 10:

[0086] Server: Confirms that the information has been successfully stored on the other service and generates a notification message indicating the result.

[0087] Step 11:

[0088] Server: Sends the generated notification message to the terminal.

[0089] Step 12:

[0090] On the device: Display a notification to the user that your information has been saved.

[0091] Step 13:

[0092] User: After receiving the notification, the user can access their account for the relevant service and check and reuse the information added to their favorites list, playlist, etc.

[0093] Example 1

[0094] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0095] Conventional video viewing systems have the drawback of making it difficult for users to easily identify objects or information that interest them while viewing and use them later. Furthermore, identifying the target object or information often requires a lot of manual effort and complex operations. This results in a poor user experience and reduced convenience. Furthermore, the lack of an efficient way to link the acquired information with other services makes seamless information use difficult.

[0096] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0097] In this invention, the server includes means for detecting when a user performs an operation on a specific object or information while watching a video, means for acquiring specific frames of the video, audio data, and video metadata, means for identifying the object or information using a generative AI model based on the acquired data, means for saving the identified object or information in a database of another service used by the user, means for notifying the user that the information has been saved, means for identifying the user's tap operation and determining the type of operation, means for packaging and transmitting the information collected in response to the tap operation to the server, means for the server to search for resources on the Internet based on the analysis results, and means for saving the acquired information to the user's favorites list using the API of another service. This allows the user to easily identify objects or information of interest without interrupting video viewing and seamlessly use them later.

[0098] "Video viewing" refers to a user continuously playing video content using a digital device.

[0099] "Operations related to specific objects or information" refers to a user performing specific actions or instructions regarding specific objects or data while watching a video.

[0100] "Tap operation" refers to input actions such as a single tap, double tap, or triple tap that a user performs by touching the screen of a digital device with their finger.

[0101] A "specific frame of a video" refers to a still image taken at a specific moment from a continuously played video.

[0102] "Audio data" refers to data that records audio information contained in video in digital format.

[0103] "Video metadata" refers to supplemental information contained in a video file, such as the title, playing time, and creator.

[0104] "Packaging" refers to the process of organizing and consolidating collected data into a certain format.

[0105] "Server" refers to a remote computer system that processes and stores data over a network.

[0106] A "generative AI model" refers to a program model that uses artificial intelligence technology to analyze and recognize data.

[0107] "Identification" refers to the act of identifying objects or information using technologies such as generative AI models.

[0108] A "database" refers to a system for storing large amounts of data and efficiently managing, searching, and updating it.

[0109] "API" refers to a standardized interface used by applications to interact with each other.

[0110] "Internet resources" refers to data and services accessible online.

[0111] A "frame image" refers to a still image that captures a specific moment in a video.

[0112] "Notification" refers to the act of a system informing a user of specific information.

[0113] The present invention relates to a system that identifies objects or information that a user finds interesting while watching a video and saves them in the user's favorites list. This system detects the user's actions, analyzes the video data being watched, and provides a function to save the identified information in other services used by the user. Specific embodiments of the system are described below.

[0114] Capturing user actions

[0115] Users: While watching a video, if they come across an outfit, place, song, etc. that catches their eye, they tap the screen to show their interest. Taps can be single, double, or triple taps, and each type indicates interest in a different category (e.g., object, place, song, etc.).

[0116] Device:

[0117] To detect user tap operations, an event listener is installed in the user interface, allowing you to detect user tap operations in real time.

[0118] To determine the type of tap, the event listener records the time interval and number of taps.

[0119] It captures a frame of the video at the moment of the tap and collects the following data:

[0120] Frame image

[0121] Audio data

[0122] Current playback time

[0123] Video title

[0124] Sending tap information

[0125] Device:

[0126] The collected information (frame images, audio data, playback time, video title) is compiled into a package.

[0127] An example of sending packaged information to a server as an HTTP request:

[0128] "Send a JSON package containing binary data of the frame image, the current playback time, audio data, and the video title."

[0129] Identifying and locating information

[0130] server:

[0131] The server receives the package information and begins analyzing it. The server is equipped with a generative AI model that uses image and voice recognition to identify specific objects and information.

[0132] Based on the analysis results, resources on the Internet are searched for. For example, the product page of an online shopping site is retrieved for the identified costume, and track information of a music streaming service is retrieved for the identified song.

[0133] Save to Favorites list

[0134] server:

[0135] The acquired information is used to call the API of other services the user uses (e.g., e-commerce sites, music services) and register it in a favorites list. For example, the identified outfit is added to the user's favorites list on a shopping site.

[0136] Order completion notification

[0137] server:

[0138] The system confirms that the information has been saved successfully and notifies the device of the result.

[0139] Device:

[0140] Users will be notified that their information has been saved, allowing them to easily access the information they are interested in later without having to interrupt their video viewing.

[0141] Specific examples

[0142] Scenario 1: User identifies and purchases an outfit from a video

[0143] 1. User: While watching a video, find an outfit you like and single-tap the screen.

[0144] 2. Device: Detects a single tap and sends the frame image and video title at that time to the server.

[0145] 3. Server: Analyzes the frame images and identifies the outfits using a generative AI model.

[0146] 4. Server: Search the online shopping site and retrieve the product page for the relevant costume.

[0147] 5. Server: Calls the API and registers the identified outfit in the user's favorites list.

[0148] 6. Device: Send a notification to the user saying "The outfit has been added to your favorites list."

[0149] 7. User: After watching, you can check and purchase the costumes from your favorite list.

[0150] Scenario 2: User identifies a song in a video and adds it to a playlist

[0151] 1. User: When a song you like comes on while watching a video, double tap the screen.

[0152] 2. Device: Detects the double-tap operation and sends the audio data and playback time at that time to the server.

[0153] 3. Server: Analyzes the audio data and identifies the song using a generative AI model.

[0154] 4. Server: Searches music streaming services and retrieves track information for identified songs.

[0155] 5. Server: Calls the API and adds the identified song to the user's playlist.

[0156] 6. Device: Send a notification to the user saying "Song has been added to your playlist."

[0157] 7. User: After listening, you can check the song from the playlist and play it.

[0158] summary

[0159] As described above, the present invention allows a user to easily identify objects or information that interest them while watching a video, and to seamlessly use the objects or information later.

[0160] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0161] Step 1:

[0162] Capturing user actions

[0163] Users: While watching a video, they find an outfit, location, song, or other item that catches their eye and tap the screen to show their interest. Different types of taps, such as single taps, double taps, and triple taps, indicate interest in different categories.

[0164] Device:

[0165] Input: User taps.

[0166] What it does: An event listener detects a tap action. For example, if a user single-taps the screen, the event listener records this action.

[0167] Processing: Identify the type of tap and collect the video frame at the moment of the tap, audio data, playback time, and video title.

[0168] Output: Type of tap and various collected data (frame image, audio data, playback time, video title).

[0169] Step 2:

[0170] Sending tap information

[0171] Device:

[0172] Input: Various collected data (frame images, audio data, playback time, video title).

[0173] What it does: Packages information and constructs an HTTP request, such as binary data for frame images and playback times, into JSON format.

[0174] Processing: Send the packaged information to the server.

[0175] Output: Packaged information sent to the server.

[0176] Step 3:

[0177] Identifying and locating information

[0178] server:

[0179] Input: Packaged information sent from the terminal.

[0180] How it works: It receives package information and begins analyzing it. It uses a generative AI model to perform image and audio recognition. For example, it inputs frame images into the AI ​​model to identify specific outfits, locations, songs, etc.

[0181] Processing: Search for resources on the Internet based on the analysis results, for example, searching for product pages on an online shopping site.

[0182] Output: Links and metadata about the identified objects and information.

[0183] Step 4:

[0184] Save to Favorites list

[0185] server:

[0186] Input: Links and metadata about the identified object or information.

[0187] What it does: Calls the API of another service and saves the information it retrieves to a favorites list. For example, it uses a shopping site API to add product information to a user's favorites list.

[0188] Processing: Use the API to store the information.

[0189] Output: The result of the save operation (success or failure).

[0190] Step 5:

[0191] Order completion notification

[0192] server:

[0193] Input: The result of the save operation (success or failure).

[0194] Behavior: If the save is successful, send a notification to the device.

[0195] Processing: Constructs a notification message and sends it to the device as an HTTP response.

[0196] Output: Notification messages sent to the terminal.

[0197] Device:

[0198] Input: Notification message from the server.

[0199] Behavior: Displays a "Your information has been saved" notification to the user, for example as a pop-up message.

[0200] Processing: Parse the received message and display a notification to the user.

[0201] Output: The notification message that is displayed to the user.

[0202] The above are the processing steps of the program of this system, and a specific explanation of the processing flow.

[0203] (Application example 1)

[0204] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0205] In today's video viewing environment, it takes a lot of effort for users to identify objects or information that interest them while watching and use them later. Specifically, users must manually search for and save information, which often interrupts the viewing experience. In addition, technology to efficiently identify specific information and store it in the appropriate category is not yet fully developed. This makes it difficult for users to easily retrieve and access costumes, music, locations, and other items that interest them while watching a video.

[0206] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0207] In this invention, the server includes means for detecting when a user performs an operation related to a specific object or information, means for acquiring specific frames of a video, audio data, and video metadata, means for identifying the object or information using a generative AI model based on the acquired data, means for saving the identified object or information in a database of another service used by the user, means for notifying the user that the information has been saved, means for detecting user operations in cooperation with a smartphone, smart glasses, or head-mounted display, means for identifying different categories based on the type of tap operation, and means for searching for resources on the Internet based on the identified information. This allows users to easily identify objects or information that interest them while watching a video and seamlessly save and use them.

[0208] "User operation" refers to an operation performed by a user on a specific object or information while watching a video.

[0209] A "specific frame" refers to the frame of the video at the moment the user performs an operation.

[0210] "Audio data" refers to audio signal data captured during video playback.

[0211] "Video metadata" refers to additional information such as the video title, playback time, and creator information.

[0212] A "generative AI model" refers to an artificial intelligence model that uses machine learning to identify specific patterns or features from data.

[0213] "Object identification" refers to identifying specific objects or information within a video based on specific frames or audio data.

[0214] "Database" refers to a collection of information for storing and managing identified objects or information.

[0215] "Saving completion notification" refers to a notification that informs the user that the data has been successfully saved.

[0216] "Smart devices" refers to devices with internet connectivity and advanced computing capabilities, such as smartphones, smart glasses, and head-mounted displays.

[0217] "Tap operation" refers to an operation in which a user touches the screen of a smart device with their finger, and includes single taps, double taps, triple taps, etc.

[0218] "Internet resource searching" refers to the process of searching for and retrieving specific information on the Internet.

[0219] The present invention is a system that identifies objects or information that a user finds interesting while watching a video and saves them in the user's favorites list. This system detects the user's actions, analyzes the video data being watched, and provides a function to save the identified information in other services used by the user. Specific embodiments of the system are described below.

[0220] First, a user watches a video using a smart device (e.g., smartphone, smart glasses, head-mounted display, etc.). If the user finds an object of interest (e.g., an outfit, a location, or a song) while watching, the user taps the screen. This tapping action can be of several types, such as single tap, double tap, or triple tap, each of which indicates interest in a different category (e.g., outfit, location, or song).

[0221] Capturing user actions

[0222] The device detects the user's tap and determines whether it is a single tap, double tap, or triple tap, while simultaneously collecting the video frame, audio data, playback time, and video metadata at the moment of the tap.

[0223] Sending tap information

[0224] The device packages the collected tap information (frame images, audio data, playback time, video metadata) and sends this package to the server.

[0225] Identifying and locating information

[0226] The server receives the package information and begins analyzing it. It is equipped with a powerful generative AI model (e.g., based on OpenAI's GPT-4) and uses image and voice recognition to identify specific objects and information. For example, for a specific outfit, it analyzes the image and searches for product information on an online shopping site.

[0227] Save to Favorites list

[0228] The server registers the acquired information in a favorites list through the API of other services used by the user (e.g., e-commerce site, music service). For example, the identified outfit is added to the user's favorites list on the shopping site, and the identified song is added to the playlist on the music service.

[0229] Save completion notification

[0230] The server confirms that the information has been successfully saved and notifies the device of the result. The device then displays a notification to the user saying "Information saved." This allows the user to easily access the information of interest later without interrupting their video viewing.

[0231] Specific examples

[0232] Scenario 1: User identifies and purchases an outfit from a video

[0233] User: While watching a video, find an outfit you like and single-tap the screen.

[0234] Device: Detects a single tap and sends the frame image and video metadata at that time to the server.

[0235] Server: Analyzes the frame image and identifies the outfit using a generative AI model. For example, use the following prompt:

[0236] Analyze the frame image that the user taps once and identify the outfit contained in this image. Then, search online shopping sites (e.g., Amazon, Rakuten) to retrieve the corresponding product page.

[0237] Server: Calls the API of the online shopping site and registers the identified outfit in the user's favorites list.

[0238] Device: Sends a notification to the user saying "Outfit has been added to your favorites list."

[0239] Users: After watching, they can check and purchase the costumes from their favorite list.

[0240] Scenario 2: User identifies a song in a video and adds it to a playlist

[0241] User: When you are watching a video and a song you like is playing, double tap the screen.

[0242] Device: Detects the double-tap operation and sends the audio data and playback time at that time to the server.

[0243] Server: Analyzes the audio data and identifies the song using a generative AI model, for example using the following prompt:

[0244] Analyze the audio data at the time the user double-tap, identify the song contained in this audio, and then search music streaming services (e.g., Spotify, Apple Music) to retrieve the corresponding track information.

[0245] Server: Calls the music streaming service's API and adds the identified song to the user's playlist.

[0246] On your device: Notify the user that a song has been added to your playlist.

[0247] User: After listening, you can check the song from the playlist and play it.

[0248] This allows users to easily identify objects or information that interest them while watching a video, and then seamlessly save and use them.

[0249] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0250] Step 1:

[0251] When a user finds an object of interest (e.g., an outfit, a location, or a song) while watching a video, they tap the screen. This tap action can be a single tap, double tap, or triple tap to indicate the category of interest.

[0252] Input: User tap action (type of tap)

[0253] Output: Detecting touch operations and determining their types

[0254] Step 2:

[0255] The device detects the user's tap operation and determines whether it is a single tap, double tap, or triple tap. At the same time, it collects the video frame image, audio data, playback time, and video metadata at the moment of the tap.

[0256] Input: Tap operation detection result

[0257] Output: collected frame images, audio data, playback time, video metadata

[0258] Step 3:

[0259] The device packages the collected tap information (frame images, audio data, playback time, and video metadata), which includes processing the data to ensure that all information is included.

[0260] Input: Frame images, audio data, playback time, video metadata

[0261] Output: Packaged tap information

[0262] Step 4:

[0263] The device transmits the packaged tap information to the server, which involves encoding the data and using a communication protocol.

[0264] Input: Packaged tap information

[0265] Output: Sending information to the server

[0266] Step 5:

[0267] The server receives the package information and begins analysis. It uses a generative AI model to perform image and audio recognition. Specifically, it identifies specific costumes from frame images and specific songs from audio data.

[0268] Input: Packaged tap information

[0269] Output: Identified objects and information (outfits, songs, etc.)

[0270] Step 6:

[0271] The server searches resources on the Internet based on the identified object or information, for example, online shopping sites for clothing or music streaming services for songs.

[0272] Input: Identified objects or information

[0273] Output: Search results (product pages, track information, etc.)

[0274] Step 7:

[0275] The server uses the APIs of other services (e.g., e-commerce sites, music services) that the user uses to add the acquired information to a favorites list or playlist.

[0276] Input: Search results (product page, track information, etc.)

[0277] Output: Favorites list and playlist registration results

[0278] Step 8:

[0279] The server confirms that the information has been saved successfully and notifies the terminal of the result.

[0280] Input: Favorites list and playlist registration results

[0281] Output: Save completion notification

[0282] Step 9:

[0283] The device will display a notification to the user that the information has been saved, allowing the user to easily access the information later without interrupting their video viewing.

[0284] Input:Save completion notification

[0285] Output: Display a notification to the user

[0286] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0287] The present invention relates to a system that identifies objects or information that a user finds interesting while watching a video, saves them in the user's favorites list, and combines an emotion engine to recognize the user's emotions and provide a more personalized experience. This system detects user operations, analyzes video data and emotion data being watched, and provides functions to save the identified information in other services used by the user. Specific embodiments of the system are described below.

[0288] Capturing user actions

[0289] Users: While watching a video, if they come across an outfit, place, song, etc. that catches their eye, they tap the screen to show their interest. Taps can be single, double, or triple taps, and each type indicates interest in a different category (e.g., object, place, song, etc.).

[0290] Device: Detects user taps in real time and distinguishes between different types of taps (single tap, double tap, triple tap, etc.). For example, a single tap indicates an outfit, a double tap indicates a song, and a triple tap indicates a location.

[0291] Capturing user emotions

[0292] Device: The device takes a picture of the user's face with a camera and analyzes their facial expressions. It also uses voice recognition technology to analyze the tone and pitch of their voice and recognize the user's current emotional state (happiness, surprise, sadness, etc.).

[0293] Sending tap information and emotion data

[0294] Terminal: Tap information (frame image, audio data, playback time, video title) and emotion data are packaged and sent to the server.

[0295] Identifying and locating information

[0296] Server: Analyzes the received package information and uses a generative AI model to identify costumes and locations from frame images and songs from audio data.

[0297] Server: Searches internet resources for the identified object or information and retrieves related product pages or music track information.

[0298] Save to Favorites list

[0299] Server: Calls an API to save the acquired information in the database of other services used by the user. For example, the identified outfit is added to the user's favorite item list on a shopping site.

[0300] Order completion notification

[0301] Server: Confirms that the information has been successfully stored on the other service and generates a notification message indicating the result.

[0302] Creating Emotion-Based Notifications

[0303] Server: Customize the content of the notification message based on the user's emotional state. For example, if the user is surprised, send a message like "Wow! Do you like this outfit?"

[0304] On Device: Show users an emotionally sensitive "Information Saved" notification, which allows for a more personalized experience.

[0305] Specific examples

[0306] Scenario 1: User identifies and purchases an outfit from a video

[0307] 1. User: While watching a video, find an outfit you like and single-tap the screen.

[0308] 2. Device: Detects a single tap and also captures and analyzes the user's facial expression with a camera.

[0309] 3. Terminal: Sends the current frame image, video title, and user emotion data to the server.

[0310] 4. Server: Analyzes the frame images and identifies the outfits using a generative AI model.

[0311] 5. Server: Search the online shopping site and retrieve the product page of the relevant costume.

[0312] 6. Server: Calls the API and registers the identified outfit in the user's favorites list.

[0313] 7. Server: Customize notification messages such as "Outfit has been added to your favorite list" based on emotion data.

[0314] 8. Device: Show notifications to users and convey messages based on their emotions.

[0315] 9. User: After watching, you can check and purchase the costumes from your favorite list.

[0316] Scenario 2: User identifies a song in a video and adds it to a playlist

[0317] 1. User: When a song you like comes on while watching a video, double tap the screen.

[0318] 2. Device: Detects double-tap gestures and also analyzes the user's voice to recognize emotions.

[0319] 3. Terminal: Sends the current audio data, playback time, and user emotion data to the server.

[0320] 4. Server: Analyzes the audio data and identifies the song using a generative AI model.

[0321] 5. Server: Searches music streaming services and retrieves track information for identified songs.

[0322] 6. Server: Calls the API and adds the identified song to the user's playlist.

[0323] 7. Server: Customize notification messages such as "A song has been added to your playlist" based on sentiment data.

[0324] 8. Device: Show notifications to users and convey messages based on their emotions.

[0325] 9. User: After listening, you can check the song from the playlist and play it.

[0326] As described above, the present invention provides a more personalized experience by allowing users to easily identify objects or information that interest them while watching a video, and by recognizing the user's emotions.

[0327] The processing flow will be explained below.

[0328] Step 1:

[0329] User: While watching a video, if they come across an outfit, location, song, etc. that catches their eye, they tap the screen. This tapping action shows the user's interest.

[0330] Step 2:

[0331] Device: Detects user taps in real time and distinguishes between different types of taps (single tap, double tap, triple tap, etc.). For example, a single tap indicates an outfit, a double tap indicates a song, and a triple tap indicates a location.

[0332] Step 3:

[0333] On the device: Obtain the video frame, audio data, playback time, and metadata such as the video title at the moment the touch occurred.

[0334] Step 4:

[0335] Device: The device captures a picture of the user's face with a camera and uses a facial expression analysis engine to recognize emotions in real time, and also uses voice recognition to identify emotions from the tone and pitch of the user's voice.

[0336] Step 5:

[0337] Device: Packages the acquired tap information (frame image, audio data, playback time, video title) and emotion data.

[0338] Step 6:

[0339] Terminal: Sends packaged information to the server.

[0340] Step 7:

[0341] Server: Analyzes the received package information, specifically using a generative AI model to identify costumes and locations from frame images and songs from audio data.

[0342] Step 8:

[0343] Server: Searches internet resources for the identified object or information and retrieves related product pages or music track information.

[0344] Step 9:

[0345] Server: Calls the API to store the acquired information in the database of other services used by the user.

[0346] Step 10:

[0347] Server: Ensures that the information is successfully stored in other services and customizes the content of notification messages based on emotion data.

[0348] Step 11:

[0349] Server: Sends the generated notification message to the terminal.

[0350] Step 12:

[0351] Device: Display a notification to the user stating "Information has been saved" based on the results of the sentiment analysis.

[0352] Step 13:

[0353] Users: After receiving the notification, they can access their account for the relevant service and check and reuse the information added to their favorites list or playlist.

[0354] Example 1: A user specifically purchases an outfit from a video

[0355] Step 1:

[0356] User: While watching a video, find an outfit you like and single-tap the screen.

[0357] Step 2:

[0358] Device: Detects a single tap and obtains the frame image and video title at that time.

[0359] Step 3:

[0360] Device: Takes a picture of the user's face and uses an expression analysis engine to recognize the user's emotions.

[0361] Step 4:

[0362] Terminal: The acquired frame images, video title, and emotion data are packaged and sent to the server.

[0363] Step 5:

[0364] Server: Analyzes frame images and identifies outfits using a generative AI model.

[0365] Step 6:

[0366] Server: Searches the online shopping site and retrieves the product page for the relevant costume.

[0367] Step 7:

[0368] Server: Calls the API to register the identified outfit in the user's favorites list.

[0369] Step 8:

[0370] Server: Generate a notification message such as "The outfit that surprised you has been added to your favorites list" based on the emotion data.

[0371] Step 9:

[0372] Server: Sends the generated notification message to the terminal.

[0373] Step 10:

[0374] Device: Display a notification to the user and convey a message based on their emotion.

[0375] Step 11:

[0376] User: You can later check and purchase the outfit from your favorites list.

[0377] Example 2: User identifies a song in a video and adds it to a playlist

[0378] Step 1:

[0379] User: When a song you like comes on while watching a video, double tap the screen.

[0380] Step 2:

[0381] Device: Detects a double-tap operation and obtains the audio data and playback time at that time.

[0382] Step 3:

[0383] On the device: Records the user's voice and uses a voice analysis engine to recognize the user's emotions.

[0384] Step 4:

[0385] Terminal: The acquired audio data, playback time, and emotion data are packaged and sent to the server.

[0386] Step 5:

[0387] Server: Analyzes the audio data and identifies the song using a generative AI model.

[0388] Step 6:

[0389] Server: Searches music streaming services and retrieves track information for identified songs.

[0390] Step 7:

[0391] Server: Calls the API to add the identified song to the user's playlist.

[0392] Step 8:

[0393] Server: Generate a notification message based on the emotion data, such as "A song you enjoyed has been added to your playlist."

[0394] Step 9:

[0395] Server: Sends the generated notification message to the terminal.

[0396] Step 10:

[0397] Device: Display a notification to the user and convey a message based on their emotion.

[0398] Step 11:

[0399] User: Later, you can check and play songs from the playlist.

[0400] As described above, the present invention provides a more personalized experience by allowing users to easily identify objects or information that interest them while watching a video, and by recognizing the user's emotions.

[0401] Example 2

[0402] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0403] In conventional video viewing systems, when a user shows interest in an object or piece of information in a video, it is time-consuming to identify and save that information. Furthermore, personalization that takes user emotions into account is lacking, leaving a need for improved user experience. Furthermore, the system lacks the ability to distinguish the type of user operation (e.g., type of tap), making it difficult to quickly identify the specific information the user is looking for.

[0404] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0405] In this invention, the server includes means for detecting when a user performs an operation related to a specific object or information while watching a video, means for analyzing the user's face and voice to recognize emotions, means for acquiring specific frames of the video, audio data, and video metadata, means for identifying objects and information using a generative artificial intelligence model based on the acquired data, means for saving the identified objects and information in a database of another service used by the user, means for notifying the user that the information has been saved, and means for customizing a notification message based on the user's emotions. This enables users to easily identify objects and information of interest while watching a video and receive personalized notifications based on their emotions.

[0406] "User action capture" is the process of detecting when a user performs an action (such as a tap) on a specific object or piece of information while watching a video.

[0407] "Tap action" refers to actions such as single tapping, double tapping, and triple tapping by a user on the screen, and each type of action indicates interest in a different category.

[0408] "Emotion recognition" is a technology that analyzes a user's facial expressions and vocal tone to identify emotional states such as joy, surprise, and sadness in real time.

[0409] "Frame image" means a still image taken at a particular point in time in a video, and is used to identify an object or piece of information in which a user has shown interest.

[0410] "Audio Data" refers to recordings of a user's voice and background sounds in a video, and is data that is analyzed to identify specific songs or information.

[0411] "Metadata" refers to additional information that identifies a video, such as the video's title and duration, and is used to search and identify data.

[0412] A "generative artificial intelligence model" is an AI algorithm used to automatically identify objects and information based on frame images, audio data, etc.

[0413] "Databases of other services" refers to data storage for storing user-specified objects and information in external databases, such as shopping sites or music streaming services.

[0414] A "notification message" is a message that notifies a user that specific information has been saved, and includes content that is customized based on the user's emotions.

[0415] "Packaging" refers to the process of combining video frame images, audio data, metadata, and user emotion data into a single data packet, a formatting process that allows for efficient transmission to a server.

[0416] This invention relates to a system that identifies objects or information that a user finds interesting while watching a video, saves them in the user's favorites list, and combines an emotion engine to recognize the user's emotions and provide a more personalized experience. This system detects the user's actions, analyzes video data and emotion data being watched, and provides the function of saving the identified information to other services used by the user.

[0417] Hardware and software used

[0418] The system is implemented using the following hardware and software.

[0419] Hardware: Devices such as smartphones, tablets, and PCs, servers connected to the internet, and cameras (in-cameras or webcams)

[0420] Software: Facial recognition software (e.g., OpenCV), speech recognition software (e.g., Google Cloud Speech-to-Text), generative AI models (e.g., TensorFlow, PyTorch)

[0421] Data processing and calculation methods

[0422] 1. User interaction detection: When a user finds an object or piece of information that interests them while watching a video, they respond by tapping the screen. Taps can be single taps, double taps, triple taps, etc., and each type indicates interest in a different category (e.g., object, place, song, etc.). The device detects these actions in real time and distinguishes the type of action.

[0423] 2. Emotion Recognition: The device captures a picture of the user's face with a camera and analyzes their facial expressions. It also uses voice recognition technology to analyze the tone and pitch of the voice to recognize the user's current emotional state (happiness, surprise, sadness, etc.).

[0424] 3. Packaging and sending data: The device packages the tap information (frame image, audio data, playback time, video title) and emotion data and sends them to the server.

[0425] 4. Information Identification and Retrieval: The server analyzes the received package information and uses a generative AI model to identify costumes and locations from frame images and songs from audio data, using prompts such as:

[0426] Example prompt: "What outfit is in this picture?"

[0427] Sample prompt: "Which song does this audio data match?"

[0428] 5. Save to Favorites List: The server searches online shopping sites and music streaming services for the identified objects and information, retrieves related product pages and music track information, and saves the retrieved information by calling APIs to save it in the databases of other services used by the user.

[0429] 6. Notification and Personalization: The server verifies that the information has been successfully stored in other services and generates a notification message based on the result. Furthermore, the content of the notification message is customized based on the user's emotional state and displayed on the device.

[0430] Specific examples

[0431] Scenario 1: User identifies and purchases an outfit from a video

[0432] 1. User: While watching a video, find an outfit you like and single-tap the screen.

[0433] 2. Device: Detects a single tap and also captures and analyzes the user's facial expression with a camera.

[0434] 3. Terminal: Sends the current frame image, video title, and user emotion data to the server.

[0435] 4. Server: Analyzes the frame image and identifies the outfit using a generative AI model. Example: "What outfit is in this image?"

[0436] 5. Server: Search the online shopping site and retrieve the product page of the relevant costume.

[0437] 6. Server: Calls the API and registers the identified outfit in the user's favorites list.

[0438] 7. Server: Customize notification messages such as "Outfit has been added to your favorite list" based on emotion data.

[0439] 8. Device: Show notifications to users and convey messages based on their emotions.

[0440] 9. User: After watching, you can check and purchase the costumes from your favorite list.

[0441] Scenario 2: User identifies a song in a video and adds it to a playlist

[0442] 1. User: When a song you like comes on while watching a video, double tap the screen.

[0443] 2. Device: Detects double-tap gestures and also analyzes the user's voice to recognize emotions.

[0444] 3. Terminal: Sends the current audio data, playback time, and user emotion data to the server.

[0445] 4. Server: Analyzes the audio data and uses a generative AI model to identify the song. Example: "Which song does this audio data match?"

[0446] 5. Server: Searches music streaming services and retrieves track information for identified songs.

[0447] 6. Server: Calls the API and adds the identified song to the user's playlist.

[0448] 7. Server: Customize notification messages such as "A song has been added to your playlist" based on sentiment data.

[0449] 8. Device: Show notifications to users and convey messages based on their emotions.

[0450] 9. User: After listening, you can check the song from the playlist and play it.

[0451] The present invention allows users to easily identify objects or information of interest while watching videos, and further provides a personalized experience based on emotion recognition.

[0452] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0453] Step 1:

[0454] Capturing user actions

[0455] User: When they find an object that catches their eye while watching a video, they tap the screen. For example, they tap once when an outfit that catches their eye is displayed.

[0456] Terminal: Detects the user's tap operations in real time. As input, it acquires the tap operation and its timing (playback time). Specifically, it determines the type of tap (single tap, double tap, triple tap, etc.) and identifies the category (e.g., single tap is an outfit). As output, it generates the type of tap operation and the corresponding category information.

[0457] Step 2:

[0458] Capturing user emotions

[0459] Terminal: The device takes a picture of the user's face with a camera and analyzes their facial expressions. It also uses voice recognition technology to analyze the tone and pitch of their voice. It receives facial images and voice data as input. Specifically, it processes the user's facial expressions in real time using face recognition software (e.g., OpenCV) and analyzes the tone of their voice using voice recognition software (e.g., Google Cloud Speech-to-Text). It outputs the user's emotional state (e.g., joy, surprise, sadness) as text data.

[0460] Step 3:

[0461] Sending tap information and emotion data

[0462] Terminal: Packages tap information and emotion data. As input, it acquires tap information (category, playback time), emotion data, frame image (screenshot at the time of tap), audio data, video title, etc. As a concrete example, it packages this data in JSON format and generates an HTTP POST request to send it to the server. As output, it sends the packaged data to the server.

[0463] Step 4:

[0464] Identifying and locating information

[0465] Server: Analyzes the received package information. As input, it receives packaged data. Specifically, it analyzes the received JSON data and uses a generative AI model to identify costumes and locations from frame images and identify songs from audio data. An example prompt is "What costume is in this image?". As output, it generates the identified objects and information as data.

[0466] Step 5:

[0467] Save to Favorites list

[0468] Server: Saves the identified information in the database of the user's other services. Receives the identified objects and information as input. Specific operations include calling the API of an online shopping site or music streaming service to add the identified outfits and songs to the user's list. Receives the status of the save completion as output.

[0469] Step 6:

[0470] Order completion notification

[0471] Server: Confirms that the information has been successfully saved to another service. Receives the save completion status as input. For example, generates a success message and creates a notification message such as "Information has been saved." Prepares the generated notification message as output.

[0472] Step 7:

[0473] Creating Emotion-Based Notifications

[0474] Server: Customizes the content of the notification message based on the recognized emotional state of the user. Receives the user's emotional data and order completion status as input. Specific behavior is to create a personalized message such as "Ah, I'm surprised! Do you like this outfit?" depending on the emotional state (e.g., surprise). Generates a customized notification message as output.

[0475] Terminal: Displays a customized notification message to the user. As input, it receives the notification message received from the server. The specific behavior is to display the message to the user using a push notification or an alert dialog. As output, it performs the action required for the user to receive the notification.

[0476] (Application example 2)

[0477] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0478] Conventional video viewing systems have had problems such as making it difficult for users to easily save objects or information that interest them while viewing, and lacking means for analyzing users' emotions to provide a more personalized experience. The present invention aims to solve these problems and provide a system for improving the video viewing experience.

[0479] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0480] In this invention, the server includes means for detecting when a user performs an operation on a specific object or information while watching a video, means for acquiring specific frames or audio data of the video, and video metadata, means for identifying the object or information using a generative AI model based on the acquired data, means for storing the identified object or information in a database of another service used by the user, means for notifying the user that the information has been saved, means for acquiring and analyzing emotion data, and means for customizing a notification message based on the emotion data. This not only enables the user to easily save objects or information that interest them, but also enables the analysis of emotions to provide personalized notification messages.

[0481] "Means for detecting when a user performs an operation on a specific object or information while watching a video" refers to technology that recognizes when a user performs an action on a specific object or information, such as by tapping the screen, while watching a video.

[0482] "Means for obtaining specific frames and audio data from a video, and video metadata" refers to technology for extracting frames and audio from a video that a user is interested in, as well as metadata related to that video.

[0483] "Means for identifying objects and information using a generative AI model based on acquired data" refers to a technique for applying a generative AI model using extracted data to identify objects and information of interest.

[0484] "Means for storing identified objects or information in the databases of other services used by the user" refers to technology for recording identified objects or information in the databases of other online services used by the user.

[0485] The "means for notifying the user that the information has been saved" is a technique for notifying the user that the specified object or information has been saved successfully.

[0486] "Means for acquiring and analyzing emotional data" refers to technology for acquiring and analyzing emotions from the user's facial expressions, voice, etc.

[0487] "Means for customizing notification messages based on emotional data" refers to a technology for generating notification messages tailored to individual users based on analyzed emotional data of the users.

[0488] The present invention relates to a system that identifies objects or information that a user finds interesting while watching a video, saves them in the user's favorites list, and recognizes the user's emotions to provide a personalized experience. A system for implementing the present invention will be described in detail below.

[0489] System Overview

[0490] The system includes the following main features:

[0491] 1. Capturing user actions

[0492] 2. Capturing user emotions

[0493] 3. Data submission and analysis

[0494] 4. Identifying and storing objects and information

[0495] 5. Generating notifications based on emotions

[0496] Capturing user actions

[0497] While watching a video, if a user comes across an object or piece of information that interests them (such as an outfit, song, or location), they can tap the screen to show their interest. Different types of taps, such as a single tap, double tap, or triple tap, indicate interest in different categories. This feature is detected by the smartphone's touchscreen.

[0498] Capturing user emotions

[0499] The device (smartphone) captures the user's facial expression with its front camera and analyzes it using the OpenCV library. It also uses the voice recognition function to analyze the voice tone and pitch with the Google Cloud Speech-to-Text API to recognize the user's current emotional state. This allows it to obtain emotional data such as joy, surprise, and sadness.

[0500] Data transmission and analysis

[0501] The user's tap information (frame image, audio data, playback time, video title) and emotion data are sent from the device to the server. The server receives the data using an API server that uses Flask and analyzes the packaged data.

[0502] Identifying and storing objects and information

[0503] The server uses PyTorch to run a generative AI model, identifying objects (such as costumes) from the received frame images and songs from the audio data. The identified objects and information are stored in databases for other services used by the user.

[0504] Generate notifications based on emotions

[0505] The server generates a customized notification message based on the emotion data. For example, if the user is surprised, it generates a message like "Wow! Do you like this outfit?". The above process allows the user to have a more personalized experience.

[0506] Specific examples

[0507] Example 1: Identifying an outfit while watching a video

[0508] The user is interested in a particular outfit and single-taps the screen. This operation information and the user's emotional data (e.g., a happy expression) are sent from the device to the server. The server analyzes the frame image using a generative AI model and adds the identified outfit to the user's favorites list from the online shopping site. The notification includes the message, "Great choice!"

[0509] Example prompt for a generative AI model:

[0510] Identify specific outfits from frame images in your video. These images contain the fashion items shown in the frame. Identify outfits that viewers particularly like and provide links to online shopping sites.

[0511] As described above, the present invention provides a specific method for analyzing a user's interests and emotions and providing a personalized experience.

[0512] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0513] Step 1:

[0514] Detecting user taps

[0515] Input: A user single-tap, double-tap, or triple-tap the screen while watching a video.

[0516] How it works: The device uses the touchscreen sensor to detect taps and distinguish between single, double, and triple taps.

[0517] Output: Generates data about the type of tap.

[0518] Step 2:

[0519] User Emotion Capture

[0520] Input: Video and audio data of the user's face at the moment the touch action is performed.

[0521] How it works: The device's front camera captures a picture of the user's face and analyzes their facial expressions using OpenCV. It also analyzes audio data acquired from the microphone using the Google Cloud Speech-to-Text API to detect tone and pitch of the voice and identify emotions.

[0522] Output: Parsed emotion data (happiness, surprise, sadness, etc.).

[0523] Step 3:

[0524] Sending and packaging data

[0525] Input: Tap type data, emotion data, frame image, audio data, playback time, video title.

[0526] How it works: The device packages this data and sends it to an API server using Flask, securely transmitting it using network protocols.

[0527] Output: A packaged dataset.

[0528] Step 4:

[0529] Data reception and analysis by the server

[0530] Input: Packaged dataset.

[0531] How it works: The server receives data via the Flask API, inputs frame images into a PyTorch-based generative AI model to identify objects (e.g., costumes), and analyzes audio data to identify songs.

[0532] Output: Data on identified objects and songs.

[0533] Step 5:

[0534] Storing objects and information

[0535] Input: Identified object and song data.

[0536] How it works: The server calls the database API of another service to add these data to the user's favorites list.

[0537] Output: A status indicating the information has been saved.

[0538] Step 6:

[0539] Customize and generate notifications based on emotions

[0540] Input: Saved status and emotion data.

[0541] How it works: The server customizes the notification message based on the emotion data and generates a message like "Your information has been saved." For example, if the user's emotion is joy, it creates a message like "Great choice!"

[0542] Output: A customized notification message.

[0543] Step 7:

[0544] Display notifications on your device

[0545] Input: Customized notification message.

[0546] What it does: The device displays a notification to the user, conveying the stored information along with an emotionally appropriate message.

[0547] Output: The user receives a notification.

[0548] Example prompt sentence:

[0549] "Identify specific outfits from framed images in the video. These images contain the fashion items shown in the frame. Identify outfits that viewers particularly liked and provide links to online shopping sites."

[0550] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0551] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0552] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0553] [Second embodiment]

[0554] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0555] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0556] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0557] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0558] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0559] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0560] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0561] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0562] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0563] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0564] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0565] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0566] The present invention relates to a system that identifies objects or information that a user finds interesting while watching a video and saves them in the user's favorites list. This system detects the user's actions, analyzes the video data being watched, and provides a function to save the identified information in other services used by the user. Specific embodiments of the system are described below.

[0567] Capturing user actions

[0568] Users: While watching a video, if they come across an outfit, place, song, etc. that catches their eye, they tap the screen to show their interest. Taps can be single, double, or triple taps, and each type indicates interest in a different category (e.g., object, place, song, etc.).

[0569] Device: Detects the user's tap and determines whether it is a single tap or a double tap. Depending on the detected tap, it collects information such as the video frame at the moment of the tap, audio data, current playback time, and video title.

[0570] Sending tap information

[0571] Device: The collected tap information (frame image, audio data, playback time, video title) is packaged. This package contains all the data necessary for subsequent information analysis.

[0572] Terminal: Sends packaged information to the server.

[0573] Identifying and locating information

[0574] Server: Receives the package information and begins analyzing it. The server is equipped with powerful generative AI models that use techniques such as image and voice recognition to identify specific outfits, locations, songs, etc.

[0575] Server: Based on the analysis results, the server searches for resources on the Internet related to the identified object or information. For example, for an identified outfit, the server retrieves the product page of an online shopping site, and for an identified song, the server retrieves track information from a music streaming service.

[0576] Save to Favorites list

[0577] Server: The acquired information is used to call the API of other services the user uses (e.g., e-commerce sites, music services) and added to the favorites list. For example, the identified outfit is added to the user's favorite item list on a shopping site.

[0578] Order completion notification

[0579] Server: Confirms that the information has been saved successfully and notifies the device of the result.

[0580] On your device: The user will be notified that their information has been saved, allowing them to easily access the information they are interested in later without interrupting their video viewing.

[0581] Specific examples

[0582] Scenario 1: User identifies and purchases an outfit from a video

[0583] 1. User: While watching a video, find an outfit you like and single-tap the screen.

[0584] 2. Device: Detects a single tap and sends the frame image and video title at that time to the server.

[0585] 3. Server: Analyzes the frame images and identifies the outfits using a generative AI model.

[0586] 4. Server: Search the online shopping site and retrieve the product page for the relevant costume.

[0587] 5. Server: Calls the API and registers the identified outfit in the user's favorites list.

[0588] 6. Device: Send a notification to the user saying "The outfit has been added to your favorites list."

[0589] 7. User: After watching, you can check and purchase the costumes from your favorite list.

[0590] Scenario 2: User identifies a song in a video and adds it to a playlist

[0591] 1. User: When a song you like comes on while watching a video, double tap the screen.

[0592] 2. Device: Detects the double-tap operation and sends the audio data and playback time at that time to the server.

[0593] 3. Server: Analyzes the audio data and identifies the song using a generative AI model.

[0594] 4. Server: Searches music streaming services and retrieves track information for identified songs.

[0595] 5. Server: Calls the API and adds the identified song to the user's playlist.

[0596] 6. Device: Send a notification to the user saying "Song has been added to your playlist."

[0597] 7. User: After listening, you can check the song from the playlist and play it.

[0598] As described above, the present invention allows users to easily identify objects or information of interest while watching a video, and to seamlessly use the objects or information later.

[0599] The processing flow will be explained below.

[0600] Step 1:

[0601] User: While watching a video, if they come across an outfit, location, song, etc. that catches their eye, they tap the screen. This tapping action shows the user's interest.

[0602] Step 2:

[0603] Device: Detects user taps in real time and distinguishes between different types of taps (single tap, double tap, triple tap, etc.). For example, a single tap indicates an outfit, a double tap indicates a song, and a triple tap indicates a location.

[0604] Step 3:

[0605] On the device: Obtain the video frame, audio data, playback time, and metadata such as the video title at the moment the touch occurred.

[0606] Step 4:

[0607] Terminal: Packages information such as acquired frame images, audio data, playback time, and video title.

[0608] Step 5:

[0609] Terminal: Sends packaged information to the server.

[0610] Step 6:

[0611] Server: Analyzes the received package information (frame image, audio data, playback time, video title).

[0612] Step 7:

[0613] Server: Using a generative AI model based on the analysis results, it identifies costumes and locations from frame images and identifies songs from audio data.

[0614] Step 8:

[0615] Server: Searches internet resources for the identified object or information and retrieves related product pages or music track information.

[0616] Step 9:

[0617] Server: Calls the API to store the acquired information in the database of other services used by the user.

[0618] Step 10:

[0619] Server: Confirms that the information has been successfully stored on the other service and generates a notification message indicating the result.

[0620] Step 11:

[0621] Server: Sends the generated notification message to the terminal.

[0622] Step 12:

[0623] On the device: Display a notification to the user that your information has been saved.

[0624] Step 13:

[0625] User: After receiving the notification, the user can access their account for the relevant service and check and reuse the information added to their favorites list, playlist, etc.

[0626] Example 1

[0627] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0628] Conventional video viewing systems have the drawback of making it difficult for users to easily identify objects or information that interest them while viewing and use them later. Furthermore, identifying the target object or information often requires a lot of manual effort and complex operations. This results in a poor user experience and reduced convenience. Furthermore, the lack of an efficient way to link the acquired information with other services makes seamless information use difficult.

[0629] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0630] In this invention, the server includes means for detecting when a user performs an operation on a specific object or information while watching a video, means for acquiring specific frames of the video, audio data, and video metadata, means for identifying the object or information using a generative AI model based on the acquired data, means for saving the identified object or information in a database of another service used by the user, means for notifying the user that the information has been saved, means for identifying the user's tap operation and determining the type of operation, means for packaging and transmitting the information collected in response to the tap operation to the server, means for the server to search for resources on the Internet based on the analysis results, and means for saving the acquired information to the user's favorites list using the API of another service. This allows the user to easily identify objects or information of interest without interrupting video viewing and seamlessly use them later.

[0631] "Video viewing" refers to a user continuously playing video content using a digital device.

[0632] "Operations related to specific objects or information" refers to a user performing specific actions or instructions regarding specific objects or data while watching a video.

[0633] "Tap operation" refers to input actions such as a single tap, double tap, or triple tap that a user performs by touching the screen of a digital device with their finger.

[0634] A "specific frame of a video" refers to a still image taken at a specific moment from a continuously played video.

[0635] "Audio data" refers to data that records audio information contained in video in digital format.

[0636] "Video metadata" refers to supplemental information contained in a video file, such as the title, playing time, and creator.

[0637] "Packaging" refers to the process of organizing and consolidating collected data into a certain format.

[0638] "Server" refers to a remote computer system that processes and stores data over a network.

[0639] A "generative AI model" refers to a program model that uses artificial intelligence technology to analyze and recognize data.

[0640] "Identification" refers to the act of identifying objects or information using technologies such as generative AI models.

[0641] A "database" refers to a system for storing large amounts of data and efficiently managing, searching, and updating it.

[0642] "API" refers to a standardized interface used by applications to interact with each other.

[0643] "Internet resources" refers to data and services accessible online.

[0644] A "frame image" refers to a still image that captures a specific moment in a video.

[0645] "Notification" refers to the act of a system informing a user of specific information.

[0646] The present invention relates to a system that identifies objects or information that a user finds interesting while watching a video and saves them in the user's favorites list. This system detects the user's actions, analyzes the video data being watched, and provides a function to save the identified information in other services used by the user. Specific embodiments of the system are described below.

[0647] Capturing user actions

[0648] Users: While watching a video, if they come across an outfit, place, song, etc. that catches their eye, they tap the screen to show their interest. Taps can be single, double, or triple taps, and each type indicates interest in a different category (e.g., object, place, song, etc.).

[0649] Device:

[0650] To detect user tap operations, an event listener is installed in the user interface, allowing you to detect user tap operations in real time.

[0651] To determine the type of tap, the event listener records the time interval and number of taps.

[0652] It captures a frame of the video at the moment of the tap and collects the following data:

[0653] Frame image

[0654] Audio data

[0655] Current playback time

[0656] Video title

[0657] Sending tap information

[0658] Device:

[0659] The collected information (frame images, audio data, playback time, video title) is compiled into a package.

[0660] An example of sending packaged information to a server as an HTTP request:

[0661] "Send a JSON package containing binary data of the frame image, the current playback time, audio data, and the video title."

[0662] Identifying and locating information

[0663] server:

[0664] The server receives the package information and begins analyzing it. The server is equipped with a generative AI model that uses image and voice recognition to identify specific objects and information.

[0665] Based on the analysis results, resources on the Internet are searched for. For example, the product page of an online shopping site is retrieved for the identified costume, and track information of a music streaming service is retrieved for the identified song.

[0666] Save to Favorites list

[0667] server:

[0668] The acquired information is used to call the API of other services the user uses (e.g., e-commerce sites, music services) and register it in a favorites list. For example, the identified outfit is added to the user's favorites list on a shopping site.

[0669] Order completion notification

[0670] server:

[0671] The system confirms that the information has been saved successfully and notifies the device of the result.

[0672] Device:

[0673] Users will be notified that their information has been saved, allowing them to easily access the information they are interested in later without having to interrupt their video viewing.

[0674] Specific examples

[0675] Scenario 1: User identifies and purchases an outfit from a video

[0676] 1. User: While watching a video, find an outfit you like and single-tap the screen.

[0677] 2. Device: Detects a single tap and sends the frame image and video title at that time to the server.

[0678] 3. Server: Analyzes the frame images and identifies the outfits using a generative AI model.

[0679] 4. Server: Search the online shopping site and retrieve the product page for the relevant costume.

[0680] 5. Server: Calls the API and registers the identified outfit in the user's favorites list.

[0681] 6. Device: Send a notification to the user saying "The outfit has been added to your favorites list."

[0682] 7. User: After watching, you can check and purchase the costumes from your favorite list.

[0683] Scenario 2: User identifies a song in a video and adds it to a playlist

[0684] 1. User: When a song you like comes on while watching a video, double tap the screen.

[0685] 2. Device: Detects the double-tap operation and sends the audio data and playback time at that time to the server.

[0686] 3. Server: Analyzes the audio data and identifies the song using a generative AI model.

[0687] 4. Server: Searches music streaming services and retrieves track information for identified songs.

[0688] 5. Server: Calls the API and adds the identified song to the user's playlist.

[0689] 6. Device: Send a notification to the user saying "Song has been added to your playlist."

[0690] 7. User: After listening, you can check the song from the playlist and play it.

[0691] summary

[0692] As described above, the present invention allows a user to easily identify objects or information that interest them while watching a video, and to seamlessly use the objects or information later.

[0693] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0694] Step 1:

[0695] Capturing user actions

[0696] Users: While watching a video, they find an outfit, location, song, or other item that catches their eye and tap the screen to show their interest. Different types of taps, such as single taps, double taps, and triple taps, indicate interest in different categories.

[0697] Device:

[0698] Input: User taps.

[0699] What it does: An event listener detects a tap action. For example, if a user single-taps the screen, the event listener records this action.

[0700] Processing: Identify the type of tap and collect the video frame at the moment of the tap, audio data, playback time, and video title.

[0701] Output: Type of tap and various collected data (frame image, audio data, playback time, video title).

[0702] Step 2:

[0703] Sending tap information

[0704] Device:

[0705] Input: Various collected data (frame images, audio data, playback time, video title).

[0706] What it does: Packages information and constructs an HTTP request, such as binary data for frame images and playback times, into JSON format.

[0707] Processing: Send the packaged information to the server.

[0708] Output: Packaged information sent to the server.

[0709] Step 3:

[0710] Identifying and locating information

[0711] server:

[0712] Input: Packaged information sent from the terminal.

[0713] How it works: It receives package information and begins analyzing it. It uses a generative AI model to perform image and audio recognition. For example, it inputs frame images into the AI ​​model to identify specific outfits, locations, songs, etc.

[0714] Processing: Search for resources on the Internet based on the analysis results, for example, searching for product pages on an online shopping site.

[0715] Output: Links and metadata about the identified objects and information.

[0716] Step 4:

[0717] Save to Favorites list

[0718] server:

[0719] Input: Links and metadata about the identified object or information.

[0720] What it does: Calls the API of another service and saves the retrieved information to a favorites list. For example, it uses a shopping site API to add product information to a user's favorites list.

[0721] Processing: Use the API to store the information.

[0722] Output: The result of the save operation (success or failure).

[0723] Step 5:

[0724] Order completion notification

[0725] server:

[0726] Input: The result of the save operation (success or failure).

[0727] Behavior: If the save is successful, send a notification to the device.

[0728] Processing: Constructs a notification message and sends it to the device as an HTTP response.

[0729] Output: Notification messages sent to the terminal.

[0730] Device:

[0731] Input: Notification message from the server.

[0732] Behavior: Displays a "Your information has been saved" notification to the user, for example as a pop-up message.

[0733] Processing: Parse the received message and display a notification to the user.

[0734] Output: The notification message that is displayed to the user.

[0735] The above are the processing steps of the program of this system, and a specific explanation of the processing flow.

[0736] (Application example 1)

[0737] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0738] In today's video viewing environment, it takes a lot of effort for users to identify objects or information that interest them while watching and use them later. Specifically, users must manually search for and save information, which often interrupts the viewing experience. In addition, technology to efficiently identify specific information and store it in the appropriate category is not yet fully developed. This makes it difficult for users to easily retrieve and access costumes, music, locations, and other items that interest them while watching a video.

[0739] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0740] In this invention, the server includes means for detecting when a user performs an operation related to a specific object or information, means for acquiring specific frames of a video, audio data, and video metadata, means for identifying the object or information using a generative AI model based on the acquired data, means for saving the identified object or information in a database of another service used by the user, means for notifying the user that the information has been saved, means for detecting user operations in cooperation with a smartphone, smart glasses, or head-mounted display, means for identifying different categories based on the type of tap operation, and means for searching for resources on the Internet based on the identified information. This allows users to easily identify objects or information that interest them while watching a video and seamlessly save and use them.

[0741] "User operation" refers to an operation performed by a user on a specific object or information while watching a video.

[0742] A "specific frame" refers to the frame of the video at the moment the user performs an operation.

[0743] "Audio data" refers to audio signal data captured during video playback.

[0744] "Video metadata" refers to additional information such as the video title, playback time, and creator information.

[0745] A "generative AI model" refers to an artificial intelligence model that uses machine learning to identify specific patterns or features from data.

[0746] "Object identification" refers to identifying specific objects or information within a video based on specific frames or audio data.

[0747] "Database" refers to a collection of information for storing and managing identified objects or information.

[0748] "Saving completion notification" refers to a notification that informs the user that the data has been successfully saved.

[0749] "Smart devices" refers to devices with internet connectivity and advanced computing capabilities, such as smartphones, smart glasses, and head-mounted displays.

[0750] "Tap operation" refers to an operation in which a user touches the screen of a smart device with their finger, and includes single taps, double taps, triple taps, etc.

[0751] "Internet resource searching" refers to the process of searching for and retrieving specific information on the Internet.

[0752] The present invention is a system that identifies objects or information that a user finds interesting while watching a video and saves them in the user's favorites list. This system detects the user's actions, analyzes the video data being watched, and provides a function to save the identified information in other services used by the user. Specific embodiments of the system are described below.

[0753] First, a user watches a video using a smart device (e.g., smartphone, smart glasses, head-mounted display, etc.). If the user finds an object of interest (e.g., an outfit, a location, or a song) while watching, the user taps the screen. This tapping action can be of several types, such as single tap, double tap, or triple tap, each of which indicates interest in a different category (e.g., outfit, location, or song).

[0754] Capturing user actions

[0755] The device detects the user's tap and determines whether it is a single tap, double tap, or triple tap, while simultaneously collecting the video frame, audio data, playback time, and video metadata at the moment of the tap.

[0756] Sending tap information

[0757] The device packages the collected tap information (frame images, audio data, playback time, video metadata) and sends this package to the server.

[0758] Identifying and locating information

[0759] The server receives the package information and begins analyzing it. It is equipped with a powerful generative AI model (e.g., based on OpenAI's GPT-4) and uses image and voice recognition to identify specific objects and information. For example, for a specific outfit, it analyzes the image and searches for product information on an online shopping site.

[0760] Save to Favorites list

[0761] The server registers the acquired information in a favorites list through the API of other services used by the user (e.g., e-commerce site, music service). For example, the identified outfit is added to the user's favorites list on the shopping site, and the identified song is added to the playlist on the music service.

[0762] Save completion notification

[0763] The server confirms that the information has been successfully saved and notifies the device of the result. The device then displays a notification to the user saying "Information saved." This allows the user to easily access the information of interest later without interrupting their video viewing.

[0764] Specific examples

[0765] Scenario 1: User identifies and purchases an outfit from a video

[0766] User: While watching a video, find an outfit you like and single-tap the screen.

[0767] Device: Detects a single tap and sends the frame image and video metadata at that time to the server.

[0768] Server: Analyzes the frame image and identifies the outfit using a generative AI model. For example, use the following prompt:

[0769] Analyze the frame image that the user taps once and identify the outfit contained in this image. Then, search online shopping sites (e.g., Amazon, Rakuten) to retrieve the corresponding product page.

[0770] Server: Calls the API of the online shopping site and registers the identified outfit in the user's favorites list.

[0771] Device: Sends a notification to the user saying "Outfit has been added to your favorites list."

[0772] Users: After watching, they can check and purchase the costumes from their favorite list.

[0773] Scenario 2: User identifies a song in a video and adds it to a playlist

[0774] User: When you are watching a video and a song you like is playing, double tap the screen.

[0775] Device: Detects the double-tap operation and sends the audio data and playback time at that time to the server.

[0776] Server: Analyzes the audio data and identifies the song using a generative AI model, for example using the following prompt:

[0777] Analyze the audio data at the time the user double-tap, identify the song contained in this audio, and then search music streaming services (e.g., Spotify, Apple Music) to retrieve the corresponding track information.

[0778] Server: Calls the music streaming service's API and adds the identified song to the user's playlist.

[0779] On your device: Notify the user that a song has been added to your playlist.

[0780] User: After listening, you can check the song from the playlist and play it.

[0781] This allows users to easily identify objects or information that interest them while watching a video, and then seamlessly save and use them.

[0782] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0783] Step 1:

[0784] When a user finds an object of interest (e.g., an outfit, a location, or a song) while watching a video, they tap the screen. This tap action can be a single tap, double tap, or triple tap to indicate the category of interest.

[0785] Input: User tap action (type of tap)

[0786] Output: Detecting touch operations and determining their types

[0787] Step 2:

[0788] The device detects the user's tap operation and determines whether it is a single tap, double tap, or triple tap. At the same time, it collects the video frame image, audio data, playback time, and video metadata at the moment of the tap.

[0789] Input: Tap operation detection result

[0790] Output: collected frame images, audio data, playback time, video metadata

[0791] Step 3:

[0792] The device packages the collected tap information (frame images, audio data, playback time, and video metadata), which includes processing the data to ensure that all information is included.

[0793] Input: Frame images, audio data, playback time, video metadata

[0794] Output: Packaged tap information

[0795] Step 4:

[0796] The device transmits the packaged tap information to the server, which involves encoding the data and using a communication protocol.

[0797] Input: Packaged tap information

[0798] Output: Sending information to the server

[0799] Step 5:

[0800] The server receives the package information and begins analysis. It uses a generative AI model to perform image and audio recognition. Specifically, it identifies specific costumes from frame images and specific songs from audio data.

[0801] Input: Packaged tap information

[0802] Output: Identified objects and information (outfits, songs, etc.)

[0803] Step 6:

[0804] The server searches resources on the Internet based on the identified object or information, for example, online shopping sites for clothing or music streaming services for songs.

[0805] Input: Identified objects or information

[0806] Output: Search results (product pages, track information, etc.)

[0807] Step 7:

[0808] The server uses the APIs of other services (e.g., e-commerce sites, music services) that the user uses to add the acquired information to a favorites list or playlist.

[0809] Input: Search results (product page, track information, etc.)

[0810] Output: Favorites list and playlist registration results

[0811] Step 8:

[0812] The server confirms that the information has been saved successfully and notifies the terminal of the result.

[0813] Input: Favorites list and playlist registration results

[0814] Output: Save completion notification

[0815] Step 9:

[0816] The device will display a notification to the user that the information has been saved, allowing the user to easily access the information later without interrupting their video viewing.

[0817] Input:Save completion notification

[0818] Output: Display a notification to the user

[0819] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0820] The present invention relates to a system that identifies objects or information that a user finds interesting while watching a video, saves them in the user's favorites list, and combines an emotion engine to recognize the user's emotions and provide a more personalized experience. This system detects user operations, analyzes video data and emotion data being watched, and provides functions to save the identified information in other services used by the user. Specific embodiments of the system are described below.

[0821] Capturing user actions

[0822] Users: While watching a video, if they come across an outfit, place, song, etc. that catches their eye, they tap the screen to show their interest. Taps can be single, double, or triple taps, and each type indicates interest in a different category (e.g., object, place, song, etc.).

[0823] Device: Detects user taps in real time and distinguishes between different types of taps (single tap, double tap, triple tap, etc.). For example, a single tap indicates an outfit, a double tap indicates a song, and a triple tap indicates a location.

[0824] Capturing user emotions

[0825] Device: The device takes a picture of the user's face with a camera and analyzes their facial expressions. It also uses voice recognition technology to analyze the tone and pitch of their voice and recognize the user's current emotional state (happiness, surprise, sadness, etc.).

[0826] Sending tap information and emotion data

[0827] Terminal: Tap information (frame image, audio data, playback time, video title) and emotion data are packaged and sent to the server.

[0828] Identifying and locating information

[0829] Server: Analyzes the received package information and uses a generative AI model to identify costumes and locations from frame images and songs from audio data.

[0830] Server: Searches internet resources for the identified object or information and retrieves related product pages or music track information.

[0831] Save to Favorites list

[0832] Server: Calls an API to save the acquired information in the database of other services used by the user. For example, the identified outfit is added to the user's favorite item list on a shopping site.

[0833] Order completion notification

[0834] Server: Confirms that the information has been successfully stored on the other service and generates a notification message indicating the result.

[0835] Creating Emotion-Based Notifications

[0836] Server: Customize the content of the notification message based on the user's emotional state. For example, if the user is surprised, send a message like "Wow! Do you like this outfit?"

[0837] On Device: Show users an emotionally sensitive "Information Saved" notification, which allows for a more personalized experience.

[0838] Specific examples

[0839] Scenario 1: User identifies and purchases an outfit from a video

[0840] 1. User: While watching a video, find an outfit you like and single-tap the screen.

[0841] 2. Device: Detects a single tap and also captures and analyzes the user's facial expression with a camera.

[0842] 3. Terminal: Sends the current frame image, video title, and user emotion data to the server.

[0843] 4. Server: Analyzes the frame images and identifies the outfits using a generative AI model.

[0844] 5. Server: Search the online shopping site and retrieve the product page of the relevant costume.

[0845] 6. Server: Calls the API and registers the identified outfit in the user's favorites list.

[0846] 7. Server: Customize notification messages such as "Outfit has been added to your favorite list" based on emotion data.

[0847] 8. Device: Show notifications to users and convey messages based on their emotions.

[0848] 9. User: After watching, you can check and purchase the costumes from your favorite list.

[0849] Scenario 2: User identifies a song in a video and adds it to a playlist

[0850] 1. User: When a song you like comes on while watching a video, double tap the screen.

[0851] 2. Device: Detects double-tap gestures and also analyzes the user's voice to recognize emotions.

[0852] 3. Terminal: Sends the current audio data, playback time, and user emotion data to the server.

[0853] 4. Server: Analyzes the audio data and identifies the song using a generative AI model.

[0854] 5. Server: Searches music streaming services and retrieves track information for identified songs.

[0855] 6. Server: Calls the API and adds the identified song to the user's playlist.

[0856] 7. Server: Customize notification messages such as "A song has been added to your playlist" based on sentiment data.

[0857] 8. Device: Show notifications to users and convey messages based on their emotions.

[0858] 9. User: After listening, you can check the song from the playlist and play it.

[0859] As described above, the present invention provides a more personalized experience by allowing users to easily identify objects or information that interest them while watching a video, and by recognizing the user's emotions.

[0860] The processing flow will be explained below.

[0861] Step 1:

[0862] User: While watching a video, if they come across an outfit, location, song, etc. that catches their eye, they tap the screen. This tapping action shows the user's interest.

[0863] Step 2:

[0864] Device: Detects user taps in real time and distinguishes between different types of taps (single tap, double tap, triple tap, etc.). For example, a single tap indicates an outfit, a double tap indicates a song, and a triple tap indicates a location.

[0865] Step 3:

[0866] On the device: Obtain the video frame, audio data, playback time, and metadata such as the video title at the moment the touch occurred.

[0867] Step 4:

[0868] Device: The device captures a picture of the user's face with a camera and uses a facial expression analysis engine to recognize emotions in real time, and also uses voice recognition to identify emotions from the tone and pitch of the user's voice.

[0869] Step 5:

[0870] Device: Packages the acquired tap information (frame image, audio data, playback time, video title) and emotion data.

[0871] Step 6:

[0872] Terminal: Sends packaged information to the server.

[0873] Step 7:

[0874] Server: Analyzes the received package information, specifically using a generative AI model to identify costumes and locations from frame images and songs from audio data.

[0875] Step 8:

[0876] Server: Searches internet resources for the identified object or information and retrieves related product pages or music track information.

[0877] Step 9:

[0878] Server: Calls the API to store the acquired information in the database of other services used by the user.

[0879] Step 10:

[0880] Server: Ensures that the information is successfully stored in other services and customizes the content of notification messages based on emotion data.

[0881] Step 11:

[0882] Server: Sends the generated notification message to the terminal.

[0883] Step 12:

[0884] Device: Display a notification to the user stating "Information has been saved" based on the results of the sentiment analysis.

[0885] Step 13:

[0886] Users: After receiving the notification, they can access their account for the relevant service and check and reuse the information added to their favorites list or playlist.

[0887] Example 1: A user specifically purchases an outfit from a video

[0888] Step 1:

[0889] User: While watching a video, find an outfit you like and single-tap the screen.

[0890] Step 2:

[0891] Device: Detects a single tap and obtains the frame image and video title at that time.

[0892] Step 3:

[0893] Device: Takes a picture of the user's face and uses an expression analysis engine to recognize the user's emotions.

[0894] Step 4:

[0895] Terminal: The acquired frame images, video title, and emotion data are packaged and sent to the server.

[0896] Step 5:

[0897] Server: Analyzes frame images and identifies outfits using a generative AI model.

[0898] Step 6:

[0899] Server: Searches the online shopping site and retrieves the product page for the relevant costume.

[0900] Step 7:

[0901] Server: Calls the API to register the identified outfit in the user's favorites list.

[0902] Step 8:

[0903] Server: Generate a notification message such as "The outfit that surprised you has been added to your favorites list" based on the emotion data.

[0904] Step 9:

[0905] Server: Sends the generated notification message to the terminal.

[0906] Step 10:

[0907] Device: Display a notification to the user and convey a message based on their emotion.

[0908] Step 11:

[0909] User: You can later check and purchase the outfit from your favorites list.

[0910] Example 2: User identifies a song in a video and adds it to a playlist

[0911] Step 1:

[0912] User: When a song you like comes on while watching a video, double tap the screen.

[0913] Step 2:

[0914] Device: Detects a double-tap operation and obtains the audio data and playback time at that time.

[0915] Step 3:

[0916] On the device: Records the user's voice and uses a voice analysis engine to recognize the user's emotions.

[0917] Step 4:

[0918] Terminal: The acquired audio data, playback time, and emotion data are packaged and sent to the server.

[0919] Step 5:

[0920] Server: Analyzes the audio data and identifies the song using a generative AI model.

[0921] Step 6:

[0922] Server: Searches music streaming services and retrieves track information for identified songs.

[0923] Step 7:

[0924] Server: Calls the API to add the identified song to the user's playlist.

[0925] Step 8:

[0926] Server: Generate a notification message based on the emotion data, such as "A song you enjoyed has been added to your playlist."

[0927] Step 9:

[0928] Server: Sends the generated notification message to the terminal.

[0929] Step 10:

[0930] Device: Display a notification to the user and convey a message based on their emotion.

[0931] Step 11:

[0932] User: Later, you can check and play songs from the playlist.

[0933] As described above, the present invention provides a more personalized experience by allowing users to easily identify objects or information that interest them while watching a video, and by recognizing the user's emotions.

[0934] Example 2

[0935] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0936] In conventional video viewing systems, when a user shows interest in an object or piece of information in a video, it is time-consuming to identify and save that information. Furthermore, personalization that takes user emotions into account is lacking, leaving a need for improved user experience. Furthermore, the system lacks the ability to distinguish the type of user operation (e.g., type of tap), making it difficult to quickly identify the specific information the user is looking for.

[0937] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0938] In this invention, the server includes means for detecting when a user performs an operation related to a specific object or information while watching a video, means for analyzing the user's face and voice to recognize emotions, means for acquiring specific frames of the video, audio data, and video metadata, means for identifying objects and information using a generative artificial intelligence model based on the acquired data, means for saving the identified objects and information in a database of another service used by the user, means for notifying the user that the information has been saved, and means for customizing a notification message based on the user's emotions. This enables users to easily identify objects and information of interest while watching a video and receive personalized notifications based on their emotions.

[0939] "User action capture" is the process of detecting when a user performs an action (such as a tap) on a specific object or piece of information while watching a video.

[0940] "Tap action" refers to actions such as single tapping, double tapping, and triple tapping by a user on the screen, and each type of action indicates interest in a different category.

[0941] "Emotion recognition" is a technology that analyzes a user's facial expressions and vocal tone to identify emotional states such as joy, surprise, and sadness in real time.

[0942] "Frame image" means a still image taken at a particular point in time in a video, and is used to identify an object or piece of information in which a user has shown interest.

[0943] "Audio Data" refers to recordings of a user's voice and background sounds in a video, and is data that is analyzed to identify specific songs or information.

[0944] "Metadata" refers to additional information that identifies a video, such as the video's title and duration, and is used to search and identify data.

[0945] A "generative artificial intelligence model" is an AI algorithm used to automatically identify objects and information based on frame images, audio data, etc.

[0946] "Databases of other services" refers to data storage for storing user-specified objects and information in external databases, such as shopping sites or music streaming services.

[0947] A "notification message" is a message that notifies a user that specific information has been saved, and includes content that is customized based on the user's emotions.

[0948] "Packaging" refers to the process of combining video frame images, audio data, metadata, and user emotion data into a single data packet, a formatting process that allows for efficient transmission to a server.

[0949] This invention relates to a system that identifies objects or information that a user finds interesting while watching a video, saves them in the user's favorites list, and combines an emotion engine to recognize the user's emotions and provide a more personalized experience. This system detects the user's actions, analyzes video data and emotion data being watched, and provides the function of saving the identified information to other services used by the user.

[0950] Hardware and software used

[0951] The system is implemented using the following hardware and software.

[0952] Hardware: Devices such as smartphones, tablets, and PCs, servers connected to the internet, and cameras (in-cameras or webcams)

[0953] Software: Facial recognition software (e.g., OpenCV), speech recognition software (e.g., Google Cloud Speech-to-Text), generative AI models (e.g., TensorFlow, PyTorch)

[0954] Data processing and calculation methods

[0955] 1. User interaction detection: When a user finds an object or piece of information that interests them while watching a video, they respond by tapping the screen. Taps can be single taps, double taps, triple taps, etc., and each type indicates interest in a different category (e.g., object, place, song, etc.). The device detects these actions in real time and distinguishes the type of action.

[0956] 2. Emotion Recognition: The device captures a picture of the user's face with a camera and analyzes their facial expressions. It also uses voice recognition technology to analyze the tone and pitch of the voice to recognize the user's current emotional state (happiness, surprise, sadness, etc.).

[0957] 3. Packaging and sending data: The device packages the tap information (frame image, audio data, playback time, video title) and emotion data and sends them to the server.

[0958] 4. Information Identification and Retrieval: The server analyzes the received package information and uses a generative AI model to identify costumes and locations from frame images and songs from audio data, using prompts such as:

[0959] Example prompt: "What outfit is in this picture?"

[0960] Sample prompt: "Which song does this audio data match?"

[0961] 5. Save to Favorites List: The server searches online shopping sites and music streaming services for the identified objects and information, retrieves related product pages and music track information, and saves the retrieved information by calling APIs to save it in the databases of other services used by the user.

[0962] 6. Notification and Personalization: The server verifies that the information has been successfully stored in other services and generates a notification message based on the result. Furthermore, the content of the notification message is customized based on the user's emotional state and displayed on the device.

[0963] Specific examples

[0964] Scenario 1: User identifies and purchases an outfit from a video

[0965] 1. User: While watching a video, find an outfit you like and single-tap the screen.

[0966] 2. Device: Detects a single tap and also captures and analyzes the user's facial expression with a camera.

[0967] 3. Terminal: Sends the current frame image, video title, and user emotion data to the server.

[0968] 4. Server: Analyzes the frame image and identifies the outfit using a generative AI model. Example: "What outfit is in this image?"

[0969] 5. Server: Search the online shopping site and retrieve the product page of the relevant costume.

[0970] 6. Server: Calls the API and registers the identified outfit in the user's favorites list.

[0971] 7. Server: Customize notification messages such as "Outfit has been added to your favorite list" based on emotion data.

[0972] 8. Device: Show notifications to users and convey messages based on their emotions.

[0973] 9. User: After watching, you can check and purchase the costumes from your favorite list.

[0974] Scenario 2: User identifies a song in a video and adds it to a playlist

[0975] 1. User: When a song you like comes on while watching a video, double tap the screen.

[0976] 2. Device: Detects double-tap gestures and also analyzes the user's voice to recognize emotions.

[0977] 3. Terminal: Sends the current audio data, playback time, and user emotion data to the server.

[0978] 4. Server: Analyzes the audio data and uses a generative AI model to identify the song. Example: "Which song does this audio data match?"

[0979] 5. Server: Searches music streaming services and retrieves track information for identified songs.

[0980] 6. Server: Calls the API and adds the identified song to the user's playlist.

[0981] 7. Server: Customize notification messages such as "A song has been added to your playlist" based on sentiment data.

[0982] 8. Device: Show notifications to users and convey messages based on their emotions.

[0983] 9. User: After listening, you can check the song from the playlist and play it.

[0984] The present invention allows users to easily identify objects or information of interest while watching videos, and further provides a personalized experience based on emotion recognition.

[0985] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0986] Step 1:

[0987] Capturing user actions

[0988] User: When they find an object that catches their eye while watching a video, they tap the screen. For example, they tap once when an outfit that catches their eye is displayed.

[0989] Terminal: Detects the user's tap operations in real time. As input, it acquires the tap operation and its timing (playback time). Specifically, it determines the type of tap (single tap, double tap, triple tap, etc.) and identifies the category (e.g., single tap is an outfit). As output, it generates the type of tap operation and the corresponding category information.

[0990] Step 2:

[0991] Capturing user emotions

[0992] Terminal: The device takes a picture of the user's face with a camera and analyzes their facial expressions. It also uses voice recognition technology to analyze the tone and pitch of their voice. It receives facial images and voice data as input. Specifically, it processes the user's facial expressions in real time using face recognition software (e.g., OpenCV) and analyzes the tone of their voice using voice recognition software (e.g., Google Cloud Speech-to-Text). It outputs the user's emotional state (e.g., joy, surprise, sadness) as text data.

[0993] Step 3:

[0994] Sending tap information and emotion data

[0995] Terminal: Packages tap information and emotion data. As input, it acquires tap information (category, playback time), emotion data, frame image (screenshot at the time of tap), audio data, video title, etc. As a concrete example, it packages this data in JSON format and generates an HTTP POST request to send it to the server. As output, it sends the packaged data to the server.

[0996] Step 4:

[0997] Identifying and locating information

[0998] Server: Analyzes the received package information. As input, it receives packaged data. Specifically, it analyzes the received JSON data and uses a generative AI model to identify costumes and locations from frame images and identify songs from audio data. An example prompt is "What costume is in this image?". As output, it generates the identified objects and information as data.

[0999] Step 5:

[1000] Save to Favorites list

[1001] Server: Saves the identified information in the database of the user's other services. Receives the identified objects and information as input. Specific operations include calling the API of an online shopping site or music streaming service to add the identified outfits and songs to the user's list. Receives the status of the save completion as output.

[1002] Step 6:

[1003] Order completion notification

[1004] Server: Confirms that the information has been successfully saved to another service. Receives the save completion status as input. For example, generates a success message and creates a notification message such as "Information has been saved." Prepares the generated notification message as output.

[1005] Step 7:

[1006] Creating Emotion-Based Notifications

[1007] Server: Customizes the content of the notification message based on the recognized emotional state of the user. Receives the user's emotional data and order completion status as input. Specific behavior is to create a personalized message such as "Ah, I'm surprised! Do you like this outfit?" depending on the emotional state (e.g., surprise). Generates a customized notification message as output.

[1008] Terminal: Displays a customized notification message to the user. As input, it receives the notification message received from the server. The specific behavior is to display the message to the user using a push notification or an alert dialog. As output, it performs the action of allowing the user to receive the notification.

[1009] (Application example 2)

[1010] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[1011] Conventional video viewing systems have had problems such as making it difficult for users to easily save objects or information that interest them while viewing, and lacking means for analyzing users' emotions to provide a more personalized experience. The present invention aims to solve these problems and provide a system for improving the video viewing experience.

[1012] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1013] In this invention, the server includes means for detecting when a user performs an operation on a specific object or information while watching a video, means for acquiring specific frames or audio data of the video, and video metadata, means for identifying the object or information using a generative AI model based on the acquired data, means for storing the identified object or information in a database of another service used by the user, means for notifying the user that the information has been saved, means for acquiring and analyzing emotion data, and means for customizing a notification message based on the emotion data. This not only enables the user to easily save objects or information that interest them, but also enables the analysis of emotions to provide personalized notification messages.

[1014] "Means for detecting when a user performs an operation on a specific object or information while watching a video" refers to technology that recognizes when a user performs an action on a specific object or information, such as by tapping the screen, while watching a video.

[1015] "Means for obtaining specific frames and audio data from a video, and video metadata" refers to technology for extracting frames and audio from a video that a user is interested in, as well as metadata related to that video.

[1016] "Means for identifying objects and information using a generative AI model based on acquired data" refers to a technique for applying a generative AI model using extracted data to identify objects and information of interest.

[1017] "Means for storing identified objects or information in the databases of other services used by the user" refers to technology for recording identified objects or information in the databases of other online services used by the user.

[1018] The "means for notifying the user that the information has been saved" is a technique for notifying the user that the specified object or information has been saved successfully.

[1019] "Means for acquiring and analyzing emotional data" refers to technology for acquiring and analyzing emotions from the user's facial expressions, voice, etc.

[1020] "Means for customizing notification messages based on emotional data" refers to a technology for generating notification messages tailored to individual users based on analyzed emotional data of the users.

[1021] The present invention relates to a system that identifies objects or information that a user finds interesting while watching a video, saves them in the user's favorites list, and recognizes the user's emotions to provide a personalized experience. A system for implementing the present invention will be described in detail below.

[1022] System Overview

[1023] The system includes the following main features:

[1024] 1. Capturing user actions

[1025] 2. Capturing user emotions

[1026] 3. Data submission and analysis

[1027] 4. Identifying and storing objects and information

[1028] 5. Generating notifications based on emotions

[1029] Capturing user actions

[1030] While watching a video, if a user comes across an object or piece of information that interests them (such as an outfit, song, or location), they can tap the screen to show their interest. Different types of taps, such as a single tap, double tap, or triple tap, indicate interest in different categories. This feature is detected by the smartphone's touchscreen.

[1031] Capturing user emotions

[1032] The device (smartphone) captures the user's facial expression with its front camera and analyzes it using the OpenCV library. It also uses the voice recognition function to analyze the voice tone and pitch with the Google Cloud Speech-to-Text API to recognize the user's current emotional state. This allows it to obtain emotional data such as joy, surprise, and sadness.

[1033] Data transmission and analysis

[1034] The user's tap information (frame image, audio data, playback time, video title) and emotion data are sent from the device to the server. The server receives the data using an API server that uses Flask and analyzes the packaged data.

[1035] Identifying and storing objects and information

[1036] The server uses PyTorch to run a generative AI model, identifying objects (such as costumes) from the received frame images and songs from the audio data. The identified objects and information are stored in databases for other services used by the user.

[1037] Generate notifications based on emotions

[1038] The server generates a customized notification message based on the emotion data. For example, if the user is surprised, it generates a message like "Wow! Do you like this outfit?". The above process allows the user to have a more personalized experience.

[1039] Specific examples

[1040] Example 1: Identifying an outfit while watching a video

[1041] The user is interested in a particular outfit and single-taps the screen. This operation information and the user's emotional data (e.g., a happy expression) are sent from the device to the server. The server analyzes the frame image using a generative AI model and adds the identified outfit to the user's favorites list from the online shopping site. The notification includes the message, "Great choice!"

[1042] Example prompt for a generative AI model:

[1043] Identify specific outfits from frame images in your video. These images contain the fashion items shown in the frame. Identify outfits that viewers particularly like and provide links to online shopping sites.

[1044] As described above, the present invention provides a specific method for analyzing a user's interests and emotions and providing a personalized experience.

[1045] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1046] Step 1:

[1047] Detecting user taps

[1048] Input: A user single-tap, double-tap, or triple-tap the screen while watching a video.

[1049] How it works: The device uses the touchscreen sensor to detect taps and distinguish between single, double, and triple taps.

[1050] Output: Generates data about the type of tap.

[1051] Step 2:

[1052] User Emotion Capture

[1053] Input: Video and audio data of the user's face at the moment the touch action is performed.

[1054] How it works: The device's front camera captures a picture of the user's face and analyzes their facial expressions using OpenCV. It also analyzes audio data acquired from the microphone using the Google Cloud Speech-to-Text API to detect tone and pitch of the voice and identify emotions.

[1055] Output: Parsed emotion data (happiness, surprise, sadness, etc.).

[1056] Step 3:

[1057] Sending and packaging data

[1058] Input: Tap type data, emotion data, frame image, audio data, playback time, video title.

[1059] How it works: The device packages this data and sends it to an API server using Flask, securely transmitting it using network protocols.

[1060] Output: A packaged dataset.

[1061] Step 4:

[1062] Data reception and analysis by the server

[1063] Input: Packaged dataset.

[1064] How it works: The server receives data via the Flask API, inputs frame images into a PyTorch-based generative AI model to identify objects (e.g., costumes), and analyzes audio data to identify songs.

[1065] Output: Data on identified objects and songs.

[1066] Step 5:

[1067] Storing objects and information

[1068] Input: Identified object and song data.

[1069] How it works: The server calls the database API of another service to add these data to the user's favorites list.

[1070] Output: A status indicating the information has been saved.

[1071] Step 6:

[1072] Customize and generate notifications based on emotions

[1073] Input: Saved status and emotion data.

[1074] How it works: The server customizes the notification message based on the emotion data and generates a message like "Your information has been saved." For example, if the user's emotion is joy, it creates a message like "Great choice!"

[1075] Output: A customized notification message.

[1076] Step 7:

[1077] Display notifications on your device

[1078] Input: Customized notification message.

[1079] What it does: The device displays a notification to the user, conveying the stored information along with an emotionally appropriate message.

[1080] Output: The user receives a notification.

[1081] Example prompt sentence:

[1082] "Identify specific outfits from framed images in the video. These images contain the fashion items shown in the frame. Identify outfits that viewers particularly liked and provide links to online shopping sites."

[1083] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1084] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1085] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[1086] [Third embodiment]

[1087] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[1088] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[1089] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1090] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[1091] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1092] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1093] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1094] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1095] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1096] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1097] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1098] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[1099] The present invention relates to a system that identifies objects or information that a user finds interesting while watching a video and saves them in the user's favorites list. This system detects the user's actions, analyzes the video data being watched, and provides a function to save the identified information in other services used by the user. Specific embodiments of the system are described below.

[1100] Capturing user actions

[1101] Users: While watching a video, if they come across an outfit, place, song, etc. that catches their eye, they tap the screen to show their interest. Taps can be single, double, or triple taps, and each type indicates interest in a different category (e.g., object, place, song, etc.).

[1102] Device: Detects the user's tap and determines whether it is a single tap or a double tap. Depending on the detected tap, it collects information such as the video frame at the moment of the tap, audio data, current playback time, and video title.

[1103] Sending tap information

[1104] Device: The collected tap information (frame image, audio data, playback time, video title) is packaged. This package contains all the data necessary for subsequent information analysis.

[1105] Terminal: Sends packaged information to the server.

[1106] Identifying and locating information

[1107] Server: Receives the package information and begins analyzing it. The server is equipped with powerful generative AI models that use techniques such as image and voice recognition to identify specific outfits, locations, songs, etc.

[1108] Server: Based on the analysis results, the server searches for resources on the Internet related to the identified object or information. For example, for an identified outfit, the server retrieves the product page of an online shopping site, and for an identified song, the server retrieves track information from a music streaming service.

[1109] Save to Favorites list

[1110] Server: The acquired information is used to call the API of other services the user uses (e.g., e-commerce sites, music services) and added to the favorites list. For example, the identified outfit is added to the user's favorite item list on a shopping site.

[1111] Order completion notification

[1112] Server: Confirms that the information has been saved successfully and notifies the device of the result.

[1113] On your device: The user will be notified that their information has been saved, allowing them to easily access the information they are interested in later without interrupting their video viewing.

[1114] Specific examples

[1115] Scenario 1: User identifies and purchases an outfit from a video

[1116] 1. User: While watching a video, find an outfit you like and single-tap the screen.

[1117] 2. Device: Detects a single tap and sends the frame image and video title at that time to the server.

[1118] 3. Server: Analyzes the frame images and identifies the outfits using a generative AI model.

[1119] 4. Server: Search the online shopping site and retrieve the product page for the relevant costume.

[1120] 5. Server: Calls the API and registers the identified outfit in the user's favorites list.

[1121] 6. Device: Send a notification to the user saying "The outfit has been added to your favorites list."

[1122] 7. User: After watching, you can check and purchase the costumes from your favorite list.

[1123] Scenario 2: User identifies a song in a video and adds it to a playlist

[1124] 1. User: When a song you like comes on while watching a video, double tap the screen.

[1125] 2. Device: Detects the double-tap operation and sends the audio data and playback time at that time to the server.

[1126] 3. Server: Analyzes the audio data and identifies the song using a generative AI model.

[1127] 4. Server: Searches music streaming services and retrieves track information for identified songs.

[1128] 5. Server: Calls the API and adds the identified song to the user's playlist.

[1129] 6. Device: Send a notification to the user saying "Song has been added to your playlist."

[1130] 7. User: After listening, you can check the song from the playlist and play it.

[1131] As described above, the present invention allows users to easily identify objects or information of interest while watching a video, and to seamlessly use the objects or information later.

[1132] The processing flow will be explained below.

[1133] Step 1:

[1134] User: While watching a video, if they come across an outfit, location, song, etc. that catches their eye, they tap the screen. This tapping action shows the user's interest.

[1135] Step 2:

[1136] Device: Detects user taps in real time and distinguishes between different types of taps (single tap, double tap, triple tap, etc.). For example, a single tap indicates an outfit, a double tap indicates a song, and a triple tap indicates a location.

[1137] Step 3:

[1138] On the device: Obtain the video frame, audio data, playback time, and metadata such as the video title at the moment the touch occurred.

[1139] Step 4:

[1140] Terminal: Packages information such as acquired frame images, audio data, playback time, and video title.

[1141] Step 5:

[1142] Terminal: Sends packaged information to the server.

[1143] Step 6:

[1144] Server: Analyzes the received package information (frame image, audio data, playback time, video title).

[1145] Step 7:

[1146] Server: Using a generative AI model based on the analysis results, it identifies costumes and locations from frame images and identifies songs from audio data.

[1147] Step 8:

[1148] Server: Searches internet resources for the identified object or information and retrieves related product pages or music track information.

[1149] Step 9:

[1150] Server: Calls the API to store the acquired information in the database of other services used by the user.

[1151] Step 10:

[1152] Server: Confirms that the information has been successfully stored on the other service and generates a notification message indicating the result.

[1153] Step 11:

[1154] Server: Sends the generated notification message to the terminal.

[1155] Step 12:

[1156] On the device: Display a notification to the user that your information has been saved.

[1157] Step 13:

[1158] User: After receiving the notification, the user can access their account for the relevant service and check and reuse the information added to their favorites list, playlist, etc.

[1159] Example 1

[1160] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1161] Conventional video viewing systems have the drawback of making it difficult for users to easily identify objects or information that interest them while viewing and use them later. Furthermore, identifying the target object or information often requires a lot of manual effort and complex operations. This results in a poor user experience and reduced convenience. Furthermore, the lack of an efficient way to link the acquired information with other services makes seamless information use difficult.

[1162] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1163] In this invention, the server includes means for detecting when a user performs an operation on a specific object or information while watching a video, means for acquiring specific frames of the video, audio data, and video metadata, means for identifying the object or information using a generative AI model based on the acquired data, means for saving the identified object or information in a database of another service used by the user, means for notifying the user that the information has been saved, means for identifying the user's tap operation and determining the type of operation, means for packaging and transmitting the information collected in response to the tap operation to the server, means for the server to search for resources on the Internet based on the analysis results, and means for saving the acquired information to the user's favorites list using the API of another service. This allows the user to easily identify objects or information of interest without interrupting video viewing and seamlessly use them later.

[1164] "Video viewing" refers to a user continuously playing video content using a digital device.

[1165] "Operations related to specific objects or information" refers to a user performing specific actions or instructions regarding specific objects or data while watching a video.

[1166] "Tap operation" refers to input actions such as a single tap, double tap, or triple tap that a user performs by touching the screen of a digital device with their finger.

[1167] A "specific frame of a video" refers to a still image taken at a specific moment from a continuously played video.

[1168] "Audio data" refers to data that records audio information contained in video in digital format.

[1169] "Video metadata" refers to supplemental information contained in a video file, such as the title, playing time, and creator.

[1170] "Packaging" refers to the process of organizing and consolidating collected data into a certain format.

[1171] "Server" refers to a remote computer system that processes and stores data over a network.

[1172] A "generative AI model" refers to a program model that uses artificial intelligence technology to analyze and recognize data.

[1173] "Identification" refers to the act of identifying objects or information using technologies such as generative AI models.

[1174] A "database" refers to a system for storing large amounts of data and efficiently managing, searching, and updating it.

[1175] "API" refers to a standardized interface used by applications to interact with each other.

[1176] "Internet resources" refers to data and services accessible online.

[1177] A "frame image" refers to a still image that captures a specific moment in a video.

[1178] "Notification" refers to the act of a system informing a user of specific information.

[1179] The present invention relates to a system that identifies objects or information that a user finds interesting while watching a video and saves them in the user's favorites list. This system detects the user's actions, analyzes the video data being watched, and provides a function to save the identified information in other services used by the user. Specific embodiments of the system are described below.

[1180] Capturing user actions

[1181] Users: While watching a video, if they come across an outfit, place, song, etc. that catches their eye, they tap the screen to show their interest. Taps can be single, double, or triple taps, and each type indicates interest in a different category (e.g., object, place, song, etc.).

[1182] Device:

[1183] To detect user tap operations, an event listener is installed in the user interface, allowing you to detect user tap operations in real time.

[1184] To determine the type of tap, the event listener records the time interval and number of taps.

[1185] It captures a frame of the video at the moment of the tap and collects the following data:

[1186] Frame image

[1187] Audio data

[1188] Current playback time

[1189] Video title

[1190] Sending tap information

[1191] Device:

[1192] The collected information (frame images, audio data, playback time, video title) is compiled into a package.

[1193] An example of sending packaged information to a server as an HTTP request:

[1194] "Send a JSON package containing binary data of the frame image, the current playback time, audio data, and the video title."

[1195] Identifying and locating information

[1196] server:

[1197] The server receives the package information and begins analyzing it. The server is equipped with a generative AI model that uses image and voice recognition to identify specific objects and information.

[1198] Based on the analysis results, resources on the Internet are searched for. For example, the product page of an online shopping site is retrieved for the identified costume, and track information of a music streaming service is retrieved for the identified song.

[1199] Save to Favorites list

[1200] server:

[1201] The acquired information is used to call the API of other services the user uses (e.g., e-commerce sites, music services) and register it in a favorites list. For example, the identified outfit is added to the user's favorites list on a shopping site.

[1202] Order completion notification

[1203] server:

[1204] The system confirms that the information has been saved successfully and notifies the device of the result.

[1205] Device:

[1206] Users will be notified that their information has been saved, allowing them to easily access the information they are interested in later without having to interrupt their video viewing.

[1207] Specific examples

[1208] Scenario 1: User identifies and purchases an outfit from a video

[1209] 1. User: While watching a video, find an outfit you like and single-tap the screen.

[1210] 2. Device: Detects a single tap and sends the frame image and video title at that time to the server.

[1211] 3. Server: Analyzes the frame images and identifies the outfits using a generative AI model.

[1212] 4. Server: Search the online shopping site and retrieve the product page for the relevant costume.

[1213] 5. Server: Calls the API and registers the identified outfit in the user's favorites list.

[1214] 6. Device: Send a notification to the user saying "The outfit has been added to your favorites list."

[1215] 7. User: After watching, you can check and purchase the costumes from your favorite list.

[1216] Scenario 2: User identifies a song in a video and adds it to a playlist

[1217] 1. User: When a song you like comes on while watching a video, double tap the screen.

[1218] 2. Device: Detects the double-tap operation and sends the audio data and playback time at that time to the server.

[1219] 3. Server: Analyzes the audio data and identifies the song using a generative AI model.

[1220] 4. Server: Searches music streaming services and retrieves track information for identified songs.

[1221] 5. Server: Calls the API and adds the identified song to the user's playlist.

[1222] 6. Device: Send a notification to the user saying "Song has been added to your playlist."

[1223] 7. User: After listening, you can check the song from the playlist and play it.

[1224] summary

[1225] As described above, the present invention allows a user to easily identify objects or information that interest them while watching a video, and to seamlessly use the objects or information later.

[1226] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1227] Step 1:

[1228] Capturing user actions

[1229] Users: While watching a video, they find an outfit, location, song, or other item that catches their eye and tap the screen to show their interest. Different types of taps, such as single taps, double taps, and triple taps, indicate interest in different categories.

[1230] Device:

[1231] Input: User taps.

[1232] What it does: An event listener detects a tap action. For example, if a user single-taps the screen, the event listener records this action.

[1233] Processing: Identify the type of tap and collect the video frame at the moment of the tap, audio data, playback time, and video title.

[1234] Output: Type of tap and various collected data (frame image, audio data, playback time, video title).

[1235] Step 2:

[1236] Sending tap information

[1237] Device:

[1238] Input: Various collected data (frame images, audio data, playback time, video title).

[1239] What it does: Packages information and constructs an HTTP request, such as binary data for frame images and playback times, into JSON format.

[1240] Processing: Send the packaged information to the server.

[1241] Output: Packaged information sent to the server.

[1242] Step 3:

[1243] Identifying and locating information

[1244] server:

[1245] Input: Packaged information sent from the terminal.

[1246] How it works: It receives package information and begins analyzing it. It uses a generative AI model to perform image and audio recognition. For example, it inputs frame images into the AI ​​model to identify specific outfits, locations, songs, etc.

[1247] Processing: Search for resources on the Internet based on the analysis results, for example, searching for product pages on an online shopping site.

[1248] Output: Links and metadata about the identified objects and information.

[1249] Step 4:

[1250] Save to Favorites list

[1251] server:

[1252] Input: Links and metadata about the identified object or information.

[1253] What it does: Calls the API of another service and saves the information it retrieves to a favorites list. For example, it uses a shopping site API to add product information to a user's favorites list.

[1254] Processing: Use the API to store the information.

[1255] Output: The result of the save operation (success or failure).

[1256] Step 5:

[1257] Order completion notification

[1258] server:

[1259] Input: The result of the save operation (success or failure).

[1260] Behavior: If the save is successful, send a notification to the device.

[1261] Processing: Constructs a notification message and sends it to the device as an HTTP response.

[1262] Output: Notification messages sent to the terminal.

[1263] Device:

[1264] Input: Notification message from the server.

[1265] Behavior: Displays a "Your information has been saved" notification to the user, for example as a pop-up message.

[1266] Processing: Parse the received message and display a notification to the user.

[1267] Output: The notification message that is displayed to the user.

[1268] The above are the processing steps of the program of this system, and a specific explanation of the processing flow.

[1269] (Application example 1)

[1270] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1271] In today's video viewing environment, it takes a lot of effort for users to identify objects or information that interest them while watching and use them later. Specifically, users must manually search for and save information, which often interrupts the viewing experience. In addition, technology to efficiently identify specific information and store it in the appropriate category is not yet fully developed. This makes it difficult for users to easily retrieve and access costumes, music, locations, and other items that interest them while watching a video.

[1272] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1273] In this invention, the server includes means for detecting when a user performs an operation related to a specific object or information, means for acquiring specific frames of a video, audio data, and video metadata, means for identifying the object or information using a generative AI model based on the acquired data, means for saving the identified object or information in a database of another service used by the user, means for notifying the user that the information has been saved, means for detecting user operations in cooperation with a smartphone, smart glasses, or head-mounted display, means for identifying different categories based on the type of tap operation, and means for searching for resources on the Internet based on the identified information. This allows users to easily identify objects or information that interest them while watching a video and seamlessly save and use them.

[1274] "User operation" refers to an operation performed by a user on a specific object or information while watching a video.

[1275] A "specific frame" refers to the frame of the video at the moment the user performs an operation.

[1276] "Audio data" refers to audio signal data captured during video playback.

[1277] "Video metadata" refers to additional information such as the video title, playback time, and creator information.

[1278] A "generative AI model" refers to an artificial intelligence model that uses machine learning to identify specific patterns or features from data.

[1279] "Object identification" refers to identifying specific objects or information within a video based on specific frames or audio data.

[1280] "Database" refers to a collection of information for storing and managing identified objects or information.

[1281] "Saving completion notification" refers to a notification that informs the user that the data has been successfully saved.

[1282] "Smart devices" refers to devices with internet connectivity and advanced computing capabilities, such as smartphones, smart glasses, and head-mounted displays.

[1283] "Tap operation" refers to an operation in which a user touches the screen of a smart device with their finger, and includes single taps, double taps, triple taps, etc.

[1284] "Internet resource searching" refers to the process of searching for and retrieving specific information on the Internet.

[1285] The present invention is a system that identifies objects or information that a user finds interesting while watching a video and saves them in the user's favorites list. This system detects the user's actions, analyzes the video data being watched, and provides a function to save the identified information in other services used by the user. Specific embodiments of the system are described below.

[1286] First, a user watches a video using a smart device (e.g., smartphone, smart glasses, head-mounted display, etc.). If the user finds an object of interest (e.g., an outfit, a location, or a song) while watching, the user taps the screen. This tapping action can be of several types, such as single tap, double tap, or triple tap, each of which indicates interest in a different category (e.g., outfit, location, or song).

[1287] Capturing user actions

[1288] The device detects the user's tap and determines whether it is a single tap, double tap, or triple tap, while simultaneously collecting the video frame, audio data, playback time, and video metadata at the moment of the tap.

[1289] Sending tap information

[1290] The device packages the collected tap information (frame images, audio data, playback time, video metadata) and sends this package to the server.

[1291] Identifying and locating information

[1292] The server receives the package information and begins analyzing it. It is equipped with a powerful generative AI model (e.g., based on OpenAI's GPT-4) and uses image and voice recognition to identify specific objects and information. For example, for a specific outfit, it analyzes the image and searches for product information on an online shopping site.

[1293] Save to Favorites list

[1294] The server registers the acquired information in a favorites list through the API of other services used by the user (e.g., e-commerce site, music service). For example, the identified outfit is added to the user's favorites list on the shopping site, and the identified song is added to the playlist on the music service.

[1295] Save completion notification

[1296] The server confirms that the information has been successfully saved and notifies the device of the result. The device then displays a notification to the user saying "Information saved." This allows the user to easily access the information of interest later without interrupting their video viewing.

[1297] Specific examples

[1298] Scenario 1: User identifies and purchases an outfit from a video

[1299] User: While watching a video, find an outfit you like and single-tap the screen.

[1300] Device: Detects a single tap and sends the frame image and video metadata at that time to the server.

[1301] Server: Analyzes the frame image and identifies the outfit using a generative AI model. For example, use the following prompt:

[1302] Analyze the frame image that the user taps once and identify the outfit contained in this image. Then, search online shopping sites (e.g., Amazon, Rakuten) to retrieve the corresponding product page.

[1303] Server: Calls the API of the online shopping site and registers the identified outfit in the user's favorites list.

[1304] Device: Sends a notification to the user saying "Outfit has been added to your favorites list."

[1305] Users: After watching, they can check and purchase the costumes from their favorite list.

[1306] Scenario 2: User identifies a song in a video and adds it to a playlist

[1307] User: When you are watching a video and a song you like is playing, double tap the screen.

[1308] Device: Detects the double-tap operation and sends the audio data and playback time at that time to the server.

[1309] Server: Analyzes the audio data and identifies the song using a generative AI model, for example using the following prompt:

[1310] Analyze the audio data at the time the user double-tap, identify the song contained in this audio, and then search music streaming services (e.g., Spotify, Apple Music) to retrieve the corresponding track information.

[1311] Server: Calls the music streaming service's API and adds the identified song to the user's playlist.

[1312] On your device: Notify the user that a song has been added to your playlist.

[1313] User: After listening, you can check the song from the playlist and play it.

[1314] This allows users to easily identify objects or information that interest them while watching a video, and then seamlessly save and use them.

[1315] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1316] Step 1:

[1317] When a user finds an object of interest (e.g., an outfit, a location, or a song) while watching a video, they tap the screen. This tap action can be a single tap, double tap, or triple tap to indicate the category of interest.

[1318] Input: User tap action (type of tap)

[1319] Output: Detecting touch operations and determining their types

[1320] Step 2:

[1321] The device detects the user's tap operation and determines whether it is a single tap, double tap, or triple tap. At the same time, it collects the video frame image, audio data, playback time, and video metadata at the moment of the tap.

[1322] Input: Tap operation detection result

[1323] Output: collected frame images, audio data, playback time, video metadata

[1324] Step 3:

[1325] The device packages the collected tap information (frame images, audio data, playback time, and video metadata), which includes processing the data to ensure that all information is included.

[1326] Input: Frame images, audio data, playback time, video metadata

[1327] Output: Packaged tap information

[1328] Step 4:

[1329] The device transmits the packaged tap information to the server, which involves encoding the data and using a communication protocol.

[1330] Input: Packaged tap information

[1331] Output: Sending information to the server

[1332] Step 5:

[1333] The server receives the package information and begins analysis. It uses a generative AI model to perform image and audio recognition. Specifically, it identifies specific costumes from frame images and specific songs from audio data.

[1334] Input: Packaged tap information

[1335] Output: Identified objects and information (outfits, songs, etc.)

[1336] Step 6:

[1337] The server searches resources on the Internet based on the identified object or information, for example, online shopping sites for clothing or music streaming services for songs.

[1338] Input: Identified objects or information

[1339] Output: Search results (product pages, track information, etc.)

[1340] Step 7:

[1341] The server uses the APIs of other services (e.g., e-commerce sites, music services) that the user uses to add the acquired information to a favorites list or playlist.

[1342] Input: Search results (product page, track information, etc.)

[1343] Output: Favorites list and playlist registration results

[1344] Step 8:

[1345] The server confirms that the information has been saved successfully and notifies the terminal of the result.

[1346] Input: Favorites list and playlist registration results

[1347] Output: Save completion notification

[1348] Step 9:

[1349] The device will display a notification to the user that the information has been saved, allowing the user to easily access the information later without interrupting their video viewing.

[1350] Input:Save completion notification

[1351] Output: Display a notification to the user

[1352] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1353] The present invention relates to a system that identifies objects or information that a user finds interesting while watching a video, saves them in the user's favorites list, and combines an emotion engine to recognize the user's emotions and provide a more personalized experience. This system detects user operations, analyzes video data and emotion data being watched, and provides functions to save the identified information in other services used by the user. Specific embodiments of the system are described below.

[1354] Capturing user actions

[1355] Users: While watching a video, if they come across an outfit, place, song, etc. that catches their eye, they tap the screen to show their interest. Taps can be single, double, or triple taps, and each type indicates interest in a different category (e.g., object, place, song, etc.).

[1356] Device: Detects user taps in real time and distinguishes between different types of taps (single tap, double tap, triple tap, etc.). For example, a single tap indicates an outfit, a double tap indicates a song, and a triple tap indicates a location.

[1357] Capturing user emotions

[1358] Device: The device takes a picture of the user's face with a camera and analyzes their facial expressions. It also uses voice recognition technology to analyze the tone and pitch of their voice and recognize the user's current emotional state (happiness, surprise, sadness, etc.).

[1359] Sending tap information and emotion data

[1360] Terminal: Tap information (frame image, audio data, playback time, video title) and emotion data are packaged and sent to the server.

[1361] Identifying and locating information

[1362] Server: Analyzes the received package information and uses a generative AI model to identify costumes and locations from frame images and songs from audio data.

[1363] Server: Searches internet resources for the identified object or information and retrieves related product pages or music track information.

[1364] Save to Favorites list

[1365] Server: Calls an API to save the acquired information in the database of other services used by the user. For example, the identified outfit is added to the user's favorite item list on a shopping site.

[1366] Order completion notification

[1367] Server: Confirms that the information has been successfully stored on the other service and generates a notification message indicating the result.

[1368] Creating Emotion-Based Notifications

[1369] Server: Customize the content of the notification message based on the user's emotional state. For example, if the user is surprised, send a message like "Wow! Do you like this outfit?"

[1370] On Device: Show users an emotionally sensitive "Information Saved" notification, which allows for a more personalized experience.

[1371] Specific examples

[1372] Scenario 1: User identifies and purchases an outfit from a video

[1373] 1. User: While watching a video, find an outfit you like and single-tap the screen.

[1374] 2. Device: Detects a single tap and also captures and analyzes the user's facial expression with a camera.

[1375] 3. Terminal: Sends the current frame image, video title, and user emotion data to the server.

[1376] 4. Server: Analyzes the frame images and identifies the outfits using a generative AI model.

[1377] 5. Server: Search the online shopping site and retrieve the product page of the relevant costume.

[1378] 6. Server: Calls the API and registers the identified outfit in the user's favorites list.

[1379] 7. Server: Customize notification messages such as "Outfit has been added to your favorite list" based on emotion data.

[1380] 8. Device: Show notifications to users and convey messages based on their emotions.

[1381] 9. User: After watching, you can check and purchase the costumes from your favorite list.

[1382] Scenario 2: User identifies a song in a video and adds it to a playlist

[1383] 1. User: When a song you like comes on while watching a video, double tap the screen.

[1384] 2. Device: Detects double-tap gestures and also analyzes the user's voice to recognize emotions.

[1385] 3. Terminal: Sends the current audio data, playback time, and user emotion data to the server.

[1386] 4. Server: Analyzes the audio data and identifies the song using a generative AI model.

[1387] 5. Server: Searches music streaming services and retrieves track information for identified songs.

[1388] 6. Server: Calls the API and adds the identified song to the user's playlist.

[1389] 7. Server: Customize notification messages such as "A song has been added to your playlist" based on sentiment data.

[1390] 8. Device: Show notifications to users and convey messages based on their emotions.

[1391] 9. User: After listening, you can check the song from the playlist and play it.

[1392] As described above, the present invention provides a more personalized experience by allowing users to easily identify objects or information that interest them while watching a video, and by recognizing the user's emotions.

[1393] The processing flow will be explained below.

[1394] Step 1:

[1395] User: While watching a video, if they come across an outfit, location, song, etc. that catches their eye, they tap the screen. This tapping action shows the user's interest.

[1396] Step 2:

[1397] Device: Detects user taps in real time and distinguishes between different types of taps (single tap, double tap, triple tap, etc.). For example, a single tap indicates an outfit, a double tap indicates a song, and a triple tap indicates a location.

[1398] Step 3:

[1399] On the device: Obtain the video frame, audio data, playback time, and metadata such as the video title at the moment the touch occurred.

[1400] Step 4:

[1401] Device: The device captures a picture of the user's face with a camera and uses a facial expression analysis engine to recognize emotions in real time, and also uses voice recognition to identify emotions from the tone and pitch of the user's voice.

[1402] Step 5:

[1403] Device: Packages the acquired tap information (frame image, audio data, playback time, video title) and emotion data.

[1404] Step 6:

[1405] Terminal: Sends packaged information to the server.

[1406] Step 7:

[1407] Server: Analyzes the received package information, specifically using a generative AI model to identify costumes and locations from frame images and songs from audio data.

[1408] Step 8:

[1409] Server: Searches internet resources for the identified object or information and retrieves related product pages or music track information.

[1410] Step 9:

[1411] Server: Calls the API to store the acquired information in the database of other services used by the user.

[1412] Step 10:

[1413] Server: Ensures that the information is successfully stored in other services and customizes the content of notification messages based on emotion data.

[1414] Step 11:

[1415] Server: Sends the generated notification message to the terminal.

[1416] Step 12:

[1417] Device: Display a notification to the user stating "Information has been saved" based on the results of the sentiment analysis.

[1418] Step 13:

[1419] Users: After receiving the notification, they can access their account for the relevant service and check and reuse the information added to their favorites list or playlist.

[1420] Example 1: A user specifically purchases an outfit from a video

[1421] Step 1:

[1422] User: While watching a video, find an outfit you like and single-tap the screen.

[1423] Step 2:

[1424] Device: Detects a single tap and obtains the frame image and video title at that time.

[1425] Step 3:

[1426] Device: Takes a picture of the user's face and uses an expression analysis engine to recognize the user's emotions.

[1427] Step 4:

[1428] Terminal: The acquired frame images, video title, and emotion data are packaged and sent to the server.

[1429] Step 5:

[1430] Server: Analyzes frame images and identifies outfits using a generative AI model.

[1431] Step 6:

[1432] Server: Searches the online shopping site and retrieves the product page for the relevant costume.

[1433] Step 7:

[1434] Server: Calls the API to register the identified outfit in the user's favorites list.

[1435] Step 8:

[1436] Server: Generate a notification message such as "The outfit that surprised you has been added to your favorites list" based on the emotion data.

[1437] Step 9:

[1438] Server: Sends the generated notification message to the terminal.

[1439] Step 10:

[1440] Device: Display a notification to the user and convey a message based on their emotion.

[1441] Step 11:

[1442] User: You can later check and purchase the outfit from your favorites list.

[1443] Example 2: User identifies a song in a video and adds it to a playlist

[1444] Step 1:

[1445] User: When a song you like comes on while watching a video, double tap the screen.

[1446] Step 2:

[1447] Device: Detects a double-tap operation and obtains the audio data and playback time at that time.

[1448] Step 3:

[1449] On the device: Records the user's voice and uses a voice analysis engine to recognize the user's emotions.

[1450] Step 4:

[1451] Terminal: The acquired audio data, playback time, and emotion data are packaged and sent to the server.

[1452] Step 5:

[1453] Server: Analyzes the audio data and identifies the song using a generative AI model.

[1454] Step 6:

[1455] Server: Searches music streaming services and retrieves track information for identified songs.

[1456] Step 7:

[1457] Server: Calls the API to add the identified song to the user's playlist.

[1458] Step 8:

[1459] Server: Generate a notification message based on the emotion data, such as "A song you enjoyed has been added to your playlist."

[1460] Step 9:

[1461] Server: Sends the generated notification message to the terminal.

[1462] Step 10:

[1463] Device: Display a notification to the user and convey a message based on their emotion.

[1464] Step 11:

[1465] User: Later, you can check and play songs from the playlist.

[1466] As described above, the present invention provides a more personalized experience by allowing users to easily identify objects or information that interest them while watching a video, and by recognizing the user's emotions.

[1467] Example 2

[1468] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1469] In conventional video viewing systems, when a user shows interest in an object or piece of information in a video, it is time-consuming to identify and save that information. Furthermore, personalization that takes user emotions into account is lacking, leaving a need for improved user experience. Furthermore, the system lacks the ability to distinguish the type of user operation (e.g., type of tap), making it difficult to quickly identify the specific information the user is looking for.

[1470] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1471] In this invention, the server includes means for detecting when a user performs an operation related to a specific object or information while watching a video, means for analyzing the user's face and voice to recognize emotions, means for acquiring specific frames of the video, audio data, and video metadata, means for identifying objects and information using a generative artificial intelligence model based on the acquired data, means for saving the identified objects and information in a database of another service used by the user, means for notifying the user that the information has been saved, and means for customizing a notification message based on the user's emotions. This enables users to easily identify objects and information of interest while watching a video and receive personalized notifications based on their emotions.

[1472] "User action capture" is the process of detecting when a user performs an action (such as a tap) on a specific object or piece of information while watching a video.

[1473] "Tap action" refers to actions such as single tapping, double tapping, and triple tapping by a user on the screen, and each type of action indicates interest in a different category.

[1474] "Emotion recognition" is a technology that analyzes a user's facial expressions and vocal tone to identify emotional states such as joy, surprise, and sadness in real time.

[1475] "Frame image" means a still image taken at a particular point in time in a video, and is used to identify an object or piece of information in which a user has shown interest.

[1476] "Audio Data" refers to recordings of a user's voice and background sounds in a video, and is data that is analyzed to identify specific songs or information.

[1477] "Metadata" refers to additional information that identifies a video, such as the video's title and duration, and is used to search and identify data.

[1478] A "generative artificial intelligence model" is an AI algorithm used to automatically identify objects and information based on frame images, audio data, etc.

[1479] "Databases of other services" refers to data storage for storing user-specified objects and information in external databases, such as shopping sites or music streaming services.

[1480] A "notification message" is a message that notifies a user that specific information has been saved, and includes content that is customized based on the user's emotions.

[1481] "Packaging" refers to the process of combining video frame images, audio data, metadata, and user emotion data into a single data packet, a formatting process that allows for efficient transmission to a server.

[1482] This invention relates to a system that identifies objects or information that a user finds interesting while watching a video, saves them in the user's favorites list, and combines an emotion engine to recognize the user's emotions and provide a more personalized experience. This system detects the user's actions, analyzes video data and emotion data being watched, and provides the function of saving the identified information to other services used by the user.

[1483] Hardware and software used

[1484] The system is implemented using the following hardware and software.

[1485] Hardware: Devices such as smartphones, tablets, and PCs, servers connected to the internet, and cameras (in-cameras or webcams)

[1486] Software: Facial recognition software (e.g., OpenCV), speech recognition software (e.g., Google Cloud Speech-to-Text), generative AI models (e.g., TensorFlow, PyTorch)

[1487] Data processing and calculation methods

[1488] 1. User interaction detection: When a user finds an object or piece of information that interests them while watching a video, they respond by tapping the screen. Taps can be single taps, double taps, triple taps, etc., and each type indicates interest in a different category (e.g., object, place, song, etc.). The device detects these actions in real time and distinguishes the type of action.

[1489] 2. Emotion Recognition: The device captures a picture of the user's face with a camera and analyzes their facial expressions. It also uses voice recognition technology to analyze the tone and pitch of the voice to recognize the user's current emotional state (happiness, surprise, sadness, etc.).

[1490] 3. Packaging and sending data: The device packages the tap information (frame image, audio data, playback time, video title) and emotion data and sends them to the server.

[1491] 4. Information Identification and Retrieval: The server analyzes the received package information and uses a generative AI model to identify costumes and locations from frame images and songs from audio data, using prompts such as:

[1492] Example prompt: "What outfit is in this picture?"

[1493] Sample prompt: "Which song does this audio data match?"

[1494] 5. Save to Favorites List: The server searches online shopping sites and music streaming services for the identified objects and information, retrieves related product pages and music track information, and saves the retrieved information by calling APIs to save it in the databases of other services used by the user.

[1495] 6. Notification and Personalization: The server verifies that the information has been successfully stored in other services and generates a notification message based on the result. Furthermore, the content of the notification message is customized based on the user's emotional state and displayed on the device.

[1496] Specific examples

[1497] Scenario 1: User identifies and purchases an outfit from a video

[1498] 1. User: While watching a video, find an outfit you like and single-tap the screen.

[1499] 2. Device: Detects a single tap and also captures and analyzes the user's facial expression with a camera.

[1500] 3. Terminal: Sends the current frame image, video title, and user emotion data to the server.

[1501] 4. Server: Analyzes the frame image and identifies the outfit using a generative AI model. Example: "What outfit is in this image?"

[1502] 5. Server: Search the online shopping site and retrieve the product page of the relevant costume.

[1503] 6. Server: Calls the API and registers the identified outfit in the user's favorites list.

[1504] 7. Server: Customize notification messages such as "Outfit has been added to your favorite list" based on emotion data.

[1505] 8. Device: Show notifications to users and convey messages based on their emotions.

[1506] 9. User: After watching, you can check and purchase the costumes from your favorite list.

[1507] Scenario 2: User identifies a song in a video and adds it to a playlist

[1508] 1. User: When a song you like comes on while watching a video, double tap the screen.

[1509] 2. Device: Detects double-tap gestures and also analyzes the user's voice to recognize emotions.

[1510] 3. Terminal: Sends the current audio data, playback time, and user emotion data to the server.

[1511] 4. Server: Analyzes the audio data and uses a generative AI model to identify the song. Example: "Which song does this audio data match?"

[1512] 5. Server: Searches music streaming services and retrieves track information for identified songs.

[1513] 6. Server: Calls the API and adds the identified song to the user's playlist.

[1514] 7. Server: Customize notification messages such as "A song has been added to your playlist" based on sentiment data.

[1515] 8. Device: Show notifications to users and convey messages based on their emotions.

[1516] 9. User: After listening, you can check the song from the playlist and play it.

[1517] The present invention allows users to easily identify objects or information of interest while watching videos, and further provides a personalized experience based on emotion recognition.

[1518] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1519] Step 1:

[1520] Capturing user actions

[1521] User: When they find an object that catches their eye while watching a video, they tap the screen. For example, they tap once when an outfit that catches their eye is displayed.

[1522] Terminal: Detects the user's tap operations in real time. As input, it acquires the tap operation and its timing (playback time). Specifically, it determines the type of tap (single tap, double tap, triple tap, etc.) and identifies the category (e.g., single tap is an outfit). As output, it generates the type of tap operation and the corresponding category information.

[1523] Step 2:

[1524] Capturing user emotions

[1525] Terminal: The device takes a picture of the user's face with a camera and analyzes their facial expressions. It also uses voice recognition technology to analyze the tone and pitch of their voice. It acquires facial images and voice data as input. Specifically, it processes the user's facial expressions in real time using face recognition software (e.g., OpenCV) and analyzes the tone of their voice using voice recognition software (e.g., Google Cloud Speech-to-Text). It outputs the user's emotional state (e.g., joy, surprise, sadness) as text data.

[1526] Step 3:

[1527] Sending tap information and emotion data

[1528] Terminal: Packages tap information and emotion data. As input, it acquires tap information (category, playback time), emotion data, frame image (screenshot at the time of tap), audio data, video title, etc. As a concrete example, it packages this data in JSON format and generates an HTTP POST request to send it to the server. As output, it sends the packaged data to the server.

[1529] Step 4:

[1530] Identifying and locating information

[1531] Server: Analyzes the received package information. As input, it receives packaged data. Specifically, it analyzes the received JSON data and uses a generative AI model to identify costumes and locations from frame images and identify songs from audio data. An example prompt is "What costume is in this image?". As output, it generates the identified objects and information as data.

[1532] Step 5:

[1533] Save to Favorites list

[1534] Server: Saves the identified information in the database of the user's other services. Receives the identified objects and information as input. Specific operations include calling the API of an online shopping site or music streaming service to add the identified outfits and songs to the user's list. Receives the status of the save completion as output.

[1535] Step 6:

[1536] Order completion notification

[1537] Server: Confirms that the information has been successfully saved to another service. Receives the save completion status as input. For example, generates a success message and creates a notification message such as "Information has been saved." Prepares the generated notification message as output.

[1538] Step 7:

[1539] Creating Emotion-Based Notifications

[1540] Server: Customizes the content of the notification message based on the recognized emotional state of the user. Receives the user's emotional data and order completion status as input. Specific behavior is to create a personalized message such as "Ah, I'm surprised! Do you like this outfit?" depending on the emotional state (e.g., surprise). Generates a customized notification message as output.

[1541] Terminal: Displays a customized notification message to the user. As input, it receives the notification message received from the server. The specific behavior is to display the message to the user using a push notification or an alert dialog. As output, it performs the action of allowing the user to receive the notification.

[1542] (Application example 2)

[1543] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1544] Conventional video viewing systems have had problems such as making it difficult for users to easily save objects or information that interest them while viewing, and lacking means for analyzing users' emotions to provide a more personalized experience. The present invention aims to solve these problems and provide a system for improving the video viewing experience.

[1545] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1546] In this invention, the server includes means for detecting when a user performs an operation on a specific object or information while watching a video, means for acquiring specific frames or audio data of the video, and video metadata, means for identifying the object or information using a generative AI model based on the acquired data, means for storing the identified object or information in a database of another service used by the user, means for notifying the user that the information has been saved, means for acquiring and analyzing emotion data, and means for customizing a notification message based on the emotion data. This not only enables the user to easily save objects or information that interest them, but also enables the analysis of emotions to provide personalized notification messages.

[1547] "Means for detecting when a user performs an operation on a specific object or information while watching a video" refers to technology that recognizes when a user performs an action on a specific object or information, such as by tapping the screen, while watching a video.

[1548] "Means for obtaining specific frames and audio data from a video, and video metadata" refers to technology for extracting frames and audio from a video that a user is interested in, as well as metadata related to that video.

[1549] "Means for identifying objects and information using a generative AI model based on acquired data" refers to a technique for applying a generative AI model using extracted data to identify objects and information of interest.

[1550] "Means for storing identified objects or information in the databases of other services used by the user" refers to technology for recording identified objects or information in the databases of other online services used by the user.

[1551] The "means for notifying the user that the information has been saved" is a technique for notifying the user that the specified object or information has been saved successfully.

[1552] "Means for acquiring and analyzing emotional data" refers to technology for acquiring and analyzing emotions from the user's facial expressions, voice, etc.

[1553] "Means for customizing notification messages based on emotional data" refers to a technology for generating notification messages tailored to individual users based on analyzed emotional data of the users.

[1554] The present invention relates to a system that identifies objects or information that a user finds interesting while watching a video, saves them in the user's favorites list, and recognizes the user's emotions to provide a personalized experience. A system for implementing the present invention will be described in detail below.

[1555] System Overview

[1556] The system includes the following main features:

[1557] 1. Capturing user actions

[1558] 2. Capturing user emotions

[1559] 3. Data submission and analysis

[1560] 4. Identifying and storing objects and information

[1561] 5. Generating notifications based on emotions

[1562] Capturing user actions

[1563] While watching a video, if a user comes across an object or piece of information that interests them (such as an outfit, song, or location), they can tap the screen to show their interest. Different types of taps, such as a single tap, double tap, or triple tap, indicate interest in different categories. This feature is detected by the smartphone's touchscreen.

[1564] Capturing user emotions

[1565] The device (smartphone) captures the user's facial expression with its front camera and analyzes it using the OpenCV library. It also uses the voice recognition function to analyze the voice tone and pitch with the Google Cloud Speech-to-Text API to recognize the user's current emotional state. This allows it to obtain emotional data such as joy, surprise, and sadness.

[1566] Data transmission and analysis

[1567] The user's tap information (frame image, audio data, playback time, video title) and emotion data are sent from the device to the server. The server receives the data using an API server that uses Flask and analyzes the packaged data.

[1568] Identifying and storing objects and information

[1569] The server uses PyTorch to run a generative AI model, identifying objects (such as costumes) from the received frame images and songs from the audio data. The identified objects and information are stored in databases for other services used by the user.

[1570] Generate notifications based on emotions

[1571] The server generates a customized notification message based on the emotion data. For example, if the user is surprised, it generates a message like "Wow! Do you like this outfit?". The above process allows the user to have a more personalized experience.

[1572] Specific examples

[1573] Example 1: Identifying an outfit while watching a video

[1574] The user is interested in a particular outfit and single-taps the screen. This operation information and the user's emotional data (e.g., a happy expression) are sent from the device to the server. The server analyzes the frame image using a generative AI model and adds the identified outfit to the user's favorites list from the online shopping site. The notification includes the message, "Great choice!"

[1575] Example prompt for a generative AI model:

[1576] Identify specific outfits from frame images in your video. These images contain the fashion items shown in the frame. Identify outfits that viewers particularly like and provide links to online shopping sites.

[1577] As described above, the present invention provides a specific method for analyzing a user's interests and emotions and providing a personalized experience.

[1578] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1579] Step 1:

[1580] Detecting user taps

[1581] Input: A user single-tap, double-tap, or triple-tap the screen while watching a video.

[1582] How it works: The device uses the touchscreen sensor to detect taps and distinguish between single, double, and triple taps.

[1583] Output: Generates data about the type of tap.

[1584] Step 2:

[1585] User Emotion Capture

[1586] Input: Video and audio data of the user's face at the moment the touch action is performed.

[1587] How it works: The device's front camera captures a picture of the user's face and analyzes their facial expressions using OpenCV. It also analyzes audio data acquired from the microphone using the Google Cloud Speech-to-Text API to detect tone and pitch of the voice and identify emotions.

[1588] Output: Parsed emotion data (happiness, surprise, sadness, etc.).

[1589] Step 3:

[1590] Sending and packaging data

[1591] Input: Tap type data, emotion data, frame image, audio data, playback time, video title.

[1592] How it works: The device packages this data and sends it to an API server using Flask, securely transmitting it using network protocols.

[1593] Output: A packaged dataset.

[1594] Step 4:

[1595] Data reception and analysis by the server

[1596] Input: Packaged dataset.

[1597] How it works: The server receives data via the Flask API, inputs frame images into a PyTorch-based generative AI model to identify objects (e.g., costumes), and analyzes audio data to identify songs.

[1598] Output: Data on identified objects and songs.

[1599] Step 5:

[1600] Storing objects and information

[1601] Input: Identified object and song data.

[1602] How it works: The server calls the database API of another service to add these data to the user's favorites list.

[1603] Output: A status indicating the information has been saved.

[1604] Step 6:

[1605] Customize and generate notifications based on emotions

[1606] Input: Saved status and emotion data.

[1607] How it works: The server customizes the notification message based on the emotion data and generates a message like "Your information has been saved." For example, if the user's emotion is joy, it creates a message like "Great choice!"

[1608] Output: A customized notification message.

[1609] Step 7:

[1610] Display notifications on your device

[1611] Input: Customized notification message.

[1612] What it does: The device displays a notification to the user, conveying the stored information along with an emotionally appropriate message.

[1613] Output: The user receives a notification.

[1614] Example prompt sentence:

[1615] "Identify specific outfits from framed images in the video. These images contain the fashion items shown in the frame. Identify outfits that viewers particularly liked and provide links to online shopping sites."

[1616] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1617] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1618] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1619] [Fourth embodiment]

[1620] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1621] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1622] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1623] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1624] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1625] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1626] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1627] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1628] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1629] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1630] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1631] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1632] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1633] The present invention relates to a system that identifies objects or information that a user finds interesting while watching a video and saves them in the user's favorites list. This system detects the user's actions, analyzes the video data being watched, and provides a function to save the identified information in other services used by the user. Specific embodiments of the system are described below.

[1634] Capturing user actions

[1635] Users: While watching a video, if they come across an outfit, place, song, etc. that catches their eye, they tap the screen to show their interest. Taps can be single, double, or triple taps, and each type indicates interest in a different category (e.g., object, place, song, etc.).

[1636] Device: Detects the user's tap and determines whether it is a single tap or a double tap. Depending on the detected tap, it collects information such as the video frame at the moment of the tap, audio data, current playback time, and video title.

[1637] Sending tap information

[1638] Device: The collected tap information (frame image, audio data, playback time, video title) is packaged. This package contains all the data necessary for subsequent information analysis.

[1639] Terminal: Sends packaged information to the server.

[1640] Identifying and locating information

[1641] Server: Receives the package information and begins analyzing it. The server is equipped with powerful generative AI models that use techniques such as image and voice recognition to identify specific outfits, locations, songs, etc.

[1642] Server: Based on the analysis results, the server searches for resources on the Internet related to the identified object or information. For example, for an identified outfit, the server retrieves the product page of an online shopping site, and for an identified song, the server retrieves track information from a music streaming service.

[1643] Save to Favorites list

[1644] Server: The acquired information is used to call the API of other services the user uses (e.g., e-commerce sites, music services) and added to the favorites list. For example, the identified outfit is added to the user's favorite item list on a shopping site.

[1645] Order completion notification

[1646] Server: Confirms that the information has been saved successfully and notifies the device of the result.

[1647] On your device: The user will be notified that their information has been saved, allowing them to easily access the information they are interested in later without interrupting their video viewing.

[1648] Specific examples

[1649] Scenario 1: User identifies and purchases an outfit from a video

[1650] 1. User: While watching a video, find an outfit you like and single-tap the screen.

[1651] 2. Device: Detects a single tap and sends the frame image and video title at that time to the server.

[1652] 3. Server: Analyzes the frame images and identifies the outfits using a generative AI model.

[1653] 4. Server: Search the online shopping site and retrieve the product page for the relevant costume.

[1654] 5. Server: Calls the API and registers the identified outfit in the user's favorites list.

[1655] 6. Device: Send a notification to the user saying "The outfit has been added to your favorites list."

[1656] 7. User: After watching, you can check and purchase the costumes from your favorite list.

[1657] Scenario 2: User identifies a song in a video and adds it to a playlist

[1658] 1. User: When a song you like comes on while watching a video, double tap the screen.

[1659] 2. Device: Detects the double-tap operation and sends the audio data and playback time at that time to the server.

[1660] 3. Server: Analyzes the audio data and identifies the song using a generative AI model.

[1661] 4. Server: Searches music streaming services and retrieves track information for identified songs.

[1662] 5. Server: Calls the API and adds the identified song to the user's playlist.

[1663] 6. Device: Send a notification to the user saying "Song has been added to your playlist."

[1664] 7. User: After listening, you can check the song from the playlist and play it.

[1665] As described above, the present invention allows users to easily identify objects or information of interest while watching a video, and to seamlessly use the objects or information later.

[1666] The processing flow will be explained below.

[1667] Step 1:

[1668] User: While watching a video, if they come across an outfit, location, song, etc. that catches their eye, they tap the screen. This tapping action shows the user's interest.

[1669] Step 2:

[1670] Device: Detects user taps in real time and distinguishes between different types of taps (single tap, double tap, triple tap, etc.). For example, a single tap indicates an outfit, a double tap indicates a song, and a triple tap indicates a location.

[1671] Step 3:

[1672] On the device: Obtain the video frame, audio data, playback time, and metadata such as the video title at the moment the touch occurred.

[1673] Step 4:

[1674] Terminal: Packages information such as acquired frame images, audio data, playback time, and video title.

[1675] Step 5:

[1676] Terminal: Sends packaged information to the server.

[1677] Step 6:

[1678] Server: Analyzes the received package information (frame image, audio data, playback time, video title).

[1679] Step 7:

[1680] Server: Using a generative AI model based on the analysis results, it identifies costumes and locations from frame images and identifies songs from audio data.

[1681] Step 8:

[1682] Server: Searches internet resources for the identified object or information and retrieves related product pages or music track information.

[1683] Step 9:

[1684] Server: Calls the API to store the acquired information in the database of other services used by the user.

[1685] Step 10:

[1686] Server: Confirms that the information has been successfully stored on the other service and generates a notification message indicating the result.

[1687] Step 11:

[1688] Server: Sends the generated notification message to the terminal.

[1689] Step 12:

[1690] On the device: Display a notification to the user that your information has been saved.

[1691] Step 13:

[1692] User: After receiving the notification, the user can access their account for the relevant service and check and reuse the information added to their favorites list, playlist, etc.

[1693] Example 1

[1694] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1695] Conventional video viewing systems have the drawback of making it difficult for users to easily identify objects or information that interest them while viewing and use them later. Furthermore, identifying the target object or information often requires a lot of manual effort and complex operations. This results in a poor user experience and reduced convenience. Furthermore, the lack of an efficient way to link the acquired information with other services makes seamless information use difficult.

[1696] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1697] In this invention, the server includes means for detecting when a user performs an operation on a specific object or information while watching a video, means for acquiring specific frames of the video, audio data, and video metadata, means for identifying the object or information using a generative AI model based on the acquired data, means for saving the identified object or information in a database of another service used by the user, means for notifying the user that the information has been saved, means for identifying the user's tap operation and determining the type of operation, means for packaging and transmitting the information collected in response to the tap operation to the server, means for the server to search for resources on the Internet based on the analysis results, and means for saving the acquired information to the user's favorites list using the API of another service. This allows the user to easily identify objects or information of interest without interrupting video viewing and seamlessly use them later.

[1698] "Video viewing" refers to a user continuously playing video content using a digital device.

[1699] "Operations related to specific objects or information" refers to a user performing specific actions or instructions regarding specific objects or data while watching a video.

[1700] "Tap operation" refers to input actions such as a single tap, double tap, or triple tap that a user performs by touching the screen of a digital device with their finger.

[1701] A "specific frame of a video" refers to a still image taken at a specific moment from a continuously played video.

[1702] "Audio data" refers to data that records audio information contained in video in digital format.

[1703] "Video metadata" refers to supplemental information contained in a video file, such as the title, playing time, and creator.

[1704] "Packaging" refers to the process of organizing and consolidating collected data into a certain format.

[1705] "Server" refers to a remote computer system that processes and stores data over a network.

[1706] A "generative AI model" refers to a program model that uses artificial intelligence technology to analyze and recognize data.

[1707] "Identification" refers to the act of identifying objects or information using technologies such as generative AI models.

[1708] A "database" refers to a system for storing large amounts of data and efficiently managing, searching, and updating it.

[1709] "API" refers to a standardized interface used by applications to interact with each other.

[1710] "Internet resources" refers to data and services accessible online.

[1711] A "frame image" refers to a still image that captures a specific moment in a video.

[1712] "Notification" refers to the act of a system informing a user of specific information.

[1713] The present invention relates to a system that identifies objects or information that a user finds interesting while watching a video and saves them in the user's favorites list. This system detects the user's actions, analyzes the video data being watched, and provides a function to save the identified information in other services used by the user. Specific embodiments of the system are described below.

[1714] Capturing user actions

[1715] Users: While watching a video, if they come across an outfit, place, song, etc. that catches their eye, they tap the screen to show their interest. Taps can be single, double, or triple taps, and each type indicates interest in a different category (e.g., object, place, song, etc.).

[1716] Device:

[1717] To detect user tap operations, an event listener is installed in the user interface, which allows you to detect user tap operations in real time.

[1718] To determine the type of tap, the event listener records the time interval and number of taps.

[1719] It captures a frame of the video at the moment of the tap and collects the following data:

[1720] Frame image

[1721] Audio data

[1722] Current playback time

[1723] Video title

[1724] Sending tap information

[1725] Device:

[1726] The collected information (frame images, audio data, playback time, video title) is compiled into a package.

[1727] An example of sending packaged information to a server as an HTTP request:

[1728] "Send a JSON package containing binary data of the frame image, the current playback time, audio data, and the video title."

[1729] Identifying and locating information

[1730] server:

[1731] The server receives the package information and begins analyzing it. The server is equipped with a generative AI model that uses image and voice recognition to identify specific objects and information.

[1732] Based on the analysis results, resources on the Internet are searched for. For example, the product page of an online shopping site is retrieved for the identified costume, and track information of a music streaming service is retrieved for the identified song.

[1733] Save to Favorites list

[1734] server:

[1735] The acquired information is used to call the API of other services the user uses (e.g., e-commerce sites, music services) and register it in a favorites list. For example, the identified outfit is added to the user's favorites list on a shopping site.

[1736] Order completion notification

[1737] server:

[1738] The system confirms that the information has been saved successfully and notifies the device of the result.

[1739] Device:

[1740] Users will be notified that their information has been saved, allowing them to easily access the information they are interested in later without having to interrupt their video viewing.

[1741] Specific examples

[1742] Scenario 1: User identifies and purchases an outfit from a video

[1743] 1. User: While watching a video, find an outfit you like and single-tap the screen.

[1744] 2. Device: Detects a single tap and sends the frame image and video title at that time to the server.

[1745] 3. Server: Analyzes the frame images and identifies the outfits using a generative AI model.

[1746] 4. Server: Search the online shopping site and retrieve the product page for the relevant costume.

[1747] 5. Server: Calls the API and registers the identified outfit in the user's favorites list.

[1748] 6. Device: Send a notification to the user saying "The outfit has been added to your favorites list."

[1749] 7. User: After watching, you can check and purchase the costumes from your favorite list.

[1750] Scenario 2: User identifies a song in a video and adds it to a playlist

[1751] 1. User: When a song you like comes on while watching a video, double tap the screen.

[1752] 2. Device: Detects the double-tap operation and sends the audio data and playback time at that time to the server.

[1753] 3. Server: Analyzes the audio data and identifies the song using a generative AI model.

[1754] 4. Server: Searches music streaming services and retrieves track information for identified songs.

[1755] 5. Server: Calls the API and adds the identified song to the user's playlist.

[1756] 6. Device: Send a notification to the user saying "Song has been added to your playlist."

[1757] 7. User: After listening, you can check the song from the playlist and play it.

[1758] summary

[1759] As described above, the present invention allows a user to easily identify objects or information that interest them while watching a video, and to seamlessly use the objects or information later.

[1760] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1761] Step 1:

[1762] Capturing user actions

[1763] Users: While watching a video, they find an outfit, location, song, or other item that catches their eye and tap the screen to show their interest. Different types of taps, such as single taps, double taps, and triple taps, indicate interest in different categories.

[1764] Device:

[1765] Input: User taps.

[1766] What it does: An event listener detects a tap action. For example, if a user single-taps the screen, the event listener records this action.

[1767] Processing: Identify the type of tap and collect the video frame at the moment of the tap, audio data, playback time, and video title.

[1768] Output: Type of tap and various collected data (frame image, audio data, playback time, video title).

[1769] Step 2:

[1770] Sending tap information

[1771] Device:

[1772] Input: Various collected data (frame images, audio data, playback time, video title).

[1773] What it does: Packages information and constructs an HTTP request, such as binary data for frame images and playback times, into JSON format.

[1774] Processing: Send the packaged information to the server.

[1775] Output: Packaged information sent to the server.

[1776] Step 3:

[1777] Identifying and locating information

[1778] server:

[1779] Input: Packaged information sent from the terminal.

[1780] How it works: It receives package information and begins analyzing it. It uses a generative AI model to perform image and audio recognition. For example, it inputs frame images into the AI ​​model to identify specific outfits, locations, songs, etc.

[1781] Processing: Search for resources on the Internet based on the analysis results, for example, searching for product pages on an online shopping site.

[1782] Output: Links and metadata about the identified objects and information.

[1783] Step 4:

[1784] Save to Favorites list

[1785] server:

[1786] Input: Links and metadata about the identified object or information.

[1787] What it does: Calls the API of another service and saves the information it retrieves to a favorites list. For example, it uses a shopping site API to add product information to a user's favorites list.

[1788] Processing: Use the API to store the information.

[1789] Output: The result of the save operation (success or failure).

[1790] Step 5:

[1791] Order completion notification

[1792] server:

[1793] Input: The result of the save operation (success or failure).

[1794] Behavior: If the save is successful, send a notification to the device.

[1795] Processing: Constructs a notification message and sends it to the device as an HTTP response.

[1796] Output: Notification messages sent to the terminal.

[1797] Device:

[1798] Input: Notification message from the server.

[1799] Behavior: Displays a "Your information has been saved" notification to the user, for example as a pop-up message.

[1800] Processing: Parse the received message and display a notification to the user.

[1801] Output: The notification message that is displayed to the user.

[1802] The above are the processing steps of the program of this system, and a specific explanation of the processing flow.

[1803] (Application example 1)

[1804] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1805] In today's video viewing environment, it takes a lot of effort for users to identify objects or information that interest them while watching and use them later. Specifically, users must manually search for and save information, which often interrupts the viewing experience. In addition, technology to efficiently identify specific information and store it in the appropriate category is not yet fully developed. This makes it difficult for users to easily retrieve and access costumes, music, locations, and other items that interest them while watching a video.

[1806] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1807] In this invention, the server includes means for detecting when a user performs an operation related to a specific object or information, means for acquiring specific frames of a video, audio data, and video metadata, means for identifying the object or information using a generative AI model based on the acquired data, means for saving the identified object or information in a database of another service used by the user, means for notifying the user that the information has been saved, means for detecting user operations in cooperation with a smartphone, smart glasses, or head-mounted display, means for identifying different categories based on the type of tap operation, and means for searching for resources on the Internet based on the identified information. This allows users to easily identify objects or information that interest them while watching a video and seamlessly save and use them.

[1808] "User operation" refers to an operation performed by a user on a specific object or information while watching a video.

[1809] A "specific frame" refers to the frame of the video at the moment the user performs an operation.

[1810] "Audio data" refers to audio signal data captured during video playback.

[1811] "Video metadata" refers to additional information such as the video title, playback time, and creator information.

[1812] A "generative AI model" refers to an artificial intelligence model that uses machine learning to identify specific patterns or features from data.

[1813] "Object identification" refers to identifying specific objects or information within a video based on specific frames or audio data.

[1814] "Database" refers to a collection of information for storing and managing identified objects or information.

[1815] "Saving completion notification" refers to a notification that informs the user that the data has been successfully saved.

[1816] "Smart devices" refers to devices with internet connectivity and advanced computing capabilities, such as smartphones, smart glasses, and head-mounted displays.

[1817] "Tap operation" refers to an operation in which a user touches the screen of a smart device with their finger, and includes single taps, double taps, triple taps, etc.

[1818] "Internet resource searching" refers to the process of searching for and retrieving specific information on the Internet.

[1819] The present invention is a system that identifies objects or information that a user finds interesting while watching a video and saves them in the user's favorites list. This system detects the user's actions, analyzes the video data being watched, and provides a function to save the identified information in other services used by the user. Specific embodiments of the system are described below.

[1820] First, a user watches a video using a smart device (e.g., smartphone, smart glasses, head-mounted display, etc.). If the user finds an object of interest (e.g., an outfit, a location, or a song) while watching, the user taps the screen. This tapping action can be of several types, such as single tap, double tap, or triple tap, each of which indicates interest in a different category (e.g., outfit, location, or song).

[1821] Capturing user actions

[1822] The device detects the user's tap and determines whether it is a single tap, double tap, or triple tap, while simultaneously collecting the video frame, audio data, playback time, and video metadata at the moment of the tap.

[1823] Sending tap information

[1824] The device packages the collected tap information (frame images, audio data, playback time, video metadata) and sends this package to the server.

[1825] Identifying and locating information

[1826] The server receives the package information and begins analyzing it. It is equipped with a powerful generative AI model (e.g., based on OpenAI's GPT-4) and uses image and voice recognition to identify specific objects and information. For example, for a specific outfit, it analyzes the image and searches for product information on an online shopping site.

[1827] Save to Favorites list

[1828] The server registers the acquired information in a favorites list through the API of other services used by the user (e.g., e-commerce site, music service). For example, the identified outfit is added to the user's favorites list on the shopping site, and the identified song is added to the playlist on the music service.

[1829] Save completion notification

[1830] The server confirms that the information has been successfully saved and notifies the device of the result. The device then displays a notification to the user saying "Information saved." This allows the user to easily access the information of interest later without interrupting their video viewing.

[1831] Specific examples

[1832] Scenario 1: User identifies and purchases an outfit from a video

[1833] User: While watching a video, find an outfit you like and single-tap the screen.

[1834] Device: Detects a single tap and sends the frame image and video metadata at that time to the server.

[1835] Server: Analyzes the frame image and identifies the outfit using a generative AI model. For example, use the following prompt:

[1836] Analyze the frame image that the user taps once and identify the outfit contained in this image. Then, search online shopping sites (e.g., Amazon, Rakuten) to retrieve the corresponding product page.

[1837] Server: Calls the API of the online shopping site and registers the identified outfit in the user's favorites list.

[1838] Device: Sends a notification to the user saying "Outfit has been added to your favorites list."

[1839] Users: After watching, they can check and purchase the costumes from their favorite list.

[1840] Scenario 2: User identifies a song in a video and adds it to a playlist

[1841] User: When you are watching a video and a song you like is playing, double tap the screen.

[1842] Device: Detects the double-tap operation and sends the audio data and playback time at that time to the server.

[1843] Server: Analyzes the audio data and identifies the song using a generative AI model, for example using the following prompt:

[1844] Analyze the audio data at the time the user double-tap, identify the song contained in this audio, and then search music streaming services (e.g., Spotify, Apple Music) to retrieve the corresponding track information.

[1845] Server: Calls the music streaming service's API and adds the identified song to the user's playlist.

[1846] On your device: Notify the user that a song has been added to your playlist.

[1847] User: After listening, you can check the song from the playlist and play it.

[1848] This allows users to easily identify objects or information that interest them while watching a video, and then seamlessly save and use them.

[1849] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1850] Step 1:

[1851] When a user finds an object of interest (e.g., an outfit, a location, or a song) while watching a video, they tap the screen. This tap action can be a single tap, double tap, or triple tap to indicate the category of interest.

[1852] Input: User tap action (type of tap)

[1853] Output: Detecting touch operations and determining their types

[1854] Step 2:

[1855] The device detects the user's tap operation and determines whether it is a single tap, double tap, or triple tap. At the same time, it collects the video frame image, audio data, playback time, and video metadata at the moment of the tap.

[1856] Input: Tap operation detection result

[1857] Output: collected frame images, audio data, playback time, video metadata

[1858] Step 3:

[1859] The device packages the collected tap information (frame images, audio data, playback time, and video metadata), which includes processing the data to ensure that all information is included.

[1860] Input: Frame images, audio data, playback time, video metadata

[1861] Output: Packaged tap information

[1862] Step 4:

[1863] The device transmits the packaged tap information to the server, which involves encoding the data and using a communication protocol.

[1864] Input: Packaged tap information

[1865] Output: Sending information to the server

[1866] Step 5:

[1867] The server receives the package information and begins analysis. It uses a generative AI model to perform image and audio recognition. Specifically, it identifies specific costumes from frame images and specific songs from audio data.

[1868] Input: Packaged tap information

[1869] Output: Identified objects and information (outfits, songs, etc.)

[1870] Step 6:

[1871] The server searches resources on the Internet based on the identified object or information, for example, online shopping sites for clothing or music streaming services for songs.

[1872] Input: Identified objects or information

[1873] Output: Search results (product pages, track information, etc.)

[1874] Step 7:

[1875] The server uses the APIs of other services (e.g., e-commerce sites, music services) that the user uses to add the acquired information to a favorites list or playlist.

[1876] Input: Search results (product page, track information, etc.)

[1877] Output: Favorites list and playlist registration results

[1878] Step 8:

[1879] The server confirms that the information has been saved successfully and notifies the terminal of the result.

[1880] Input: Favorites list and playlist registration results

[1881] Output: Save completion notification

[1882] Step 9:

[1883] The device will display a notification to the user that the information has been saved, allowing the user to easily access the information later without interrupting their video viewing.

[1884] Input:Save completion notification

[1885] Output: Display a notification to the user

[1886] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1887] The present invention relates to a system that identifies objects or information that a user finds interesting while watching a video, saves them in the user's favorites list, and combines an emotion engine to recognize the user's emotions and provide a more personalized experience. This system detects user operations, analyzes video data and emotion data being watched, and provides functions to save the identified information in other services used by the user. Specific embodiments of the system are described below.

[1888] Capturing user actions

[1889] Users: While watching a video, if they come across an outfit, place, song, etc. that catches their eye, they tap the screen to show their interest. Taps can be single, double, or triple taps, and each type indicates interest in a different category (e.g., object, place, song, etc.).

[1890] Device: Detects user taps in real time and distinguishes between different types of taps (single tap, double tap, triple tap, etc.). For example, a single tap indicates an outfit, a double tap indicates a song, and a triple tap indicates a location.

[1891] Capturing user emotions

[1892] Device: The device takes a picture of the user's face with a camera and analyzes their facial expressions. It also uses voice recognition technology to analyze the tone and pitch of their voice and recognize the user's current emotional state (happiness, surprise, sadness, etc.).

[1893] Sending tap information and emotion data

[1894] Terminal: Tap information (frame image, audio data, playback time, video title) and emotion data are packaged and sent to the server.

[1895] Identifying and locating information

[1896] Server: Analyzes the received package information and uses a generative AI model to identify costumes and locations from frame images and songs from audio data.

[1897] Server: Searches internet resources for the identified object or information and retrieves related product pages or music track information.

[1898] Save to Favorites list

[1899] Server: Calls an API to save the acquired information in the database of other services used by the user. For example, the identified outfit is added to the user's favorite item list on a shopping site.

[1900] Order completion notification

[1901] Server: Confirms that the information has been successfully stored on the other service and generates a notification message indicating the result.

[1902] Creating Emotion-Based Notifications

[1903] Server: Customize the content of the notification message based on the user's emotional state. For example, if the user is surprised, send a message like "Wow! Do you like this outfit?"

[1904] On Device: Show users an emotionally sensitive "Information Saved" notification, which allows for a more personalized experience.

[1905] Specific examples

[1906] Scenario 1: User identifies and purchases an outfit from a video

[1907] 1. User: While watching a video, find an outfit you like and single-tap the screen.

[1908] 2. Device: Detects a single tap and also captures and analyzes the user's facial expression with a camera.

[1909] 3. Terminal: Sends the current frame image, video title, and user emotion data to the server.

[1910] 4. Server: Analyzes the frame images and identifies the outfits using a generative AI model.

[1911] 5. Server: Search the online shopping site and retrieve the product page of the relevant costume.

[1912] 6. Server: Calls the API and registers the identified outfit in the user's favorites list.

[1913] 7. Server: Customize notification messages such as "Outfit has been added to your favorite list" based on emotion data.

[1914] 8. Device: Show notifications to users and convey messages based on their emotions.

[1915] 9. User: After watching, you can check and purchase the costumes from your favorite list.

[1916] Scenario 2: User identifies a song in a video and adds it to a playlist

[1917] 1. User: When a song you like comes on while watching a video, double tap the screen.

[1918] 2. Device: Detects double-tap gestures and also analyzes the user's voice to recognize emotions.

[1919] 3. Terminal: Sends the current audio data, playback time, and user emotion data to the server.

[1920] 4. Server: Analyzes the audio data and identifies the song using a generative AI model.

[1921] 5. Server: Searches music streaming services and retrieves track information for identified songs.

[1922] 6. Server: Calls the API and adds the identified song to the user's playlist.

[1923] 7. Server: Customize notification messages such as "A song has been added to your playlist" based on sentiment data.

[1924] 8. Device: Show notifications to users and convey messages based on their emotions.

[1925] 9. User: After listening, you can check the song from the playlist and play it.

[1926] As described above, the present invention provides a more personalized experience by allowing users to easily identify objects or information that interest them while watching a video, and by recognizing the user's emotions.

[1927] The processing flow will be explained below.

[1928] Step 1:

[1929] User: While watching a video, if they come across an outfit, location, song, etc. that catches their eye, they tap the screen. This tapping action shows the user's interest.

[1930] Step 2:

[1931] Device: Detects user taps in real time and distinguishes between different types of taps (single tap, double tap, triple tap, etc.). For example, a single tap indicates an outfit, a double tap indicates a song, and a triple tap indicates a location.

[1932] Step 3:

[1933] On the device: Obtain the video frame, audio data, playback time, and metadata such as the video title at the moment the touch occurred.

[1934] Step 4:

[1935] Device: The device captures a picture of the user's face with a camera and uses a facial expression analysis engine to recognize emotions in real time, and also uses voice recognition to identify emotions from the tone and pitch of the user's voice.

[1936] Step 5:

[1937] Device: Packages the acquired tap information (frame image, audio data, playback time, video title) and emotion data.

[1938] Step 6:

[1939] Terminal: Sends packaged information to the server.

[1940] Step 7:

[1941] Server: Analyzes the received package information, specifically using a generative AI model to identify costumes and locations from frame images and songs from audio data.

[1942] Step 8:

[1943] Server: Searches internet resources for the identified object or information and retrieves related product pages or music track information.

[1944] Step 9:

[1945] Server: Calls the API to store the acquired information in the database of other services used by the user.

[1946] Step 10:

[1947] Server: Ensures that the information is successfully stored in other services and customizes the content of notification messages based on emotion data.

[1948] Step 11:

[1949] Server: Sends the generated notification message to the terminal.

[1950] Step 12:

[1951] Device: Display a notification to the user stating "Information has been saved" based on the results of the sentiment analysis.

[1952] Step 13:

[1953] Users: After receiving the notification, they can access their account for the relevant service and check and reuse the information added to their favorites list or playlist.

[1954] Example 1: A user specifically purchases an outfit from a video

[1955] Step 1:

[1956] User: While watching a video, find an outfit you like and single-tap the screen.

[1957] Step 2:

[1958] Device: Detects a single tap and obtains the frame image and video title at that time.

[1959] Step 3:

[1960] Device: Takes a picture of the user's face and uses an expression analysis engine to recognize the user's emotions.

[1961] Step 4:

[1962] Terminal: The acquired frame images, video title, and emotion data are packaged and sent to the server.

[1963] Step 5:

[1964] Server: Analyzes frame images and identifies outfits using a generative AI model.

[1965] Step 6:

[1966] Server: Searches the online shopping site and retrieves the product page for the relevant costume.

[1967] Step 7:

[1968] Server: Calls the API to register the identified outfit in the user's favorites list.

[1969] Step 8:

[1970] Server: Generate a notification message such as "The outfit that surprised you has been added to your favorites list" based on the emotion data.

[1971] Step 9:

[1972] Server: Sends the generated notification message to the terminal.

[1973] Step 10:

[1974] Device: Display a notification to the user and convey a message based on their emotion.

[1975] Step 11:

[1976] User: You can later check and purchase the outfit from your favorites list.

[1977] Example 2: User identifies a song in a video and adds it to a playlist

[1978] Step 1:

[1979] User: When a song you like comes on while watching a video, double tap the screen.

[1980] Step 2:

[1981] Device: Detects a double-tap operation and obtains the audio data and playback time at that time.

[1982] Step 3:

[1983] On the device: Records the user's voice and uses a voice analysis engine to recognize the user's emotions.

[1984] Step 4:

[1985] Terminal: The acquired audio data, playback time, and emotion data are packaged and sent to the server.

[1986] Step 5:

[1987] Server: Analyzes the audio data and identifies the song using a generative AI model.

[1988] Step 6:

[1989] Server: Searches music streaming services and retrieves track information for identified songs.

[1990] Step 7:

[1991] Server: Calls the API to add the identified song to the user's playlist.

[1992] Step 8:

[1993] Server: Generate a notification message based on the emotion data, such as "A song you enjoyed has been added to your playlist."

[1994] Step 9:

[1995] Server: Sends the generated notification message to the terminal.

[1996] Step 10:

[1997] Device: Display a notification to the user and convey a message based on their emotion.

[1998] Step 11:

[1999] User: Later, you can check and play songs from the playlist.

[2000] As described above, the present invention provides a more personalized experience by allowing users to easily identify objects or information that interest them while watching a video, and by recognizing the user's emotions.

[2001] Example 2

[2002] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[2003] In conventional video viewing systems, when a user shows interest in an object or piece of information in a video, it is time-consuming to identify and save that information. Furthermore, personalization that takes user emotions into account is lacking, leaving a need for improved user experience. Furthermore, the system lacks the ability to distinguish the type of user operation (e.g., type of tap), making it difficult to quickly identify the specific information the user is looking for.

[2004] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[2005] In this invention, the server includes means for detecting when a user performs an operation related to a specific object or information while watching a video, means for analyzing the user's face and voice to recognize emotions, means for acquiring specific frames of the video, audio data, and video metadata, means for identifying objects and information using a generative artificial intelligence model based on the acquired data, means for saving the identified objects and information in a database of another service used by the user, means for notifying the user that the information has been saved, and means for customizing a notification message based on the user's emotions. This enables users to easily identify objects and information of interest while watching a video and receive personalized notifications based on their emotions.

[2006] "User action capture" is the process of detecting when a user performs an action (such as a tap) on a specific object or piece of information while watching a video.

[2007] "Tap action" refers to actions such as single tapping, double tapping, and triple tapping by a user on the screen, and each type of action indicates interest in a different category.

[2008] "Emotion recognition" is a technology that analyzes a user's facial expressions and vocal tone to identify emotional states such as joy, surprise, and sadness in real time.

[2009] "Frame image" means a still image taken at a particular point in time in a video, and is used to identify an object or piece of information in which a user has shown interest.

[2010] "Audio Data" refers to recordings of a user's voice and background sounds in a video, and is data that is analyzed to identify specific songs or information.

[2011] "Metadata" refers to additional information that identifies a video, such as the video's title and duration, and is used to search and identify data.

[2012] A "generative artificial intelligence model" is an AI algorithm used to automatically identify objects and information based on frame images, audio data, etc.

[2013] "Databases of other services" refers to data storage for storing user-specified objects and information in external databases, such as shopping sites or music streaming services.

[2014] A "notification message" is a message that notifies a user that specific information has been saved, and includes content that is customized based on the user's emotions.

[2015] "Packaging" refers to the process of combining video frame images, audio data, metadata, and user emotion data into a single data packet, a formatting process that allows for efficient transmission to a server.

[2016] This invention relates to a system that identifies objects or information that a user finds interesting while watching a video, saves them in the user's favorites list, and combines an emotion engine to recognize the user's emotions and provide a more personalized experience. This system detects the user's actions, analyzes video data and emotion data being watched, and provides the function of saving the identified information to other services used by the user.

[2017] Hardware and software used

[2018] The system is implemented using the following hardware and software.

[2019] Hardware: Devices such as smartphones, tablets, and PCs, servers connected to the internet, and cameras (in-cameras or webcams)

[2020] Software: Facial recognition software (e.g., OpenCV), speech recognition software (e.g., Google Cloud Speech-to-Text), generative AI models (e.g., TensorFlow, PyTorch)

[2021] Data processing and calculation methods

[2022] 1. User interaction detection: When a user finds an object or piece of information that interests them while watching a video, they respond by tapping the screen. Taps can be single taps, double taps, triple taps, etc., and each type indicates interest in a different category (e.g., object, place, song, etc.). The device detects these actions in real time and distinguishes the type of action.

[2023] 2. Emotion Recognition: The device captures a picture of the user's face with a camera and analyzes their facial expressions. It also uses voice recognition technology to analyze the tone and pitch of the voice to recognize the user's current emotional state (happiness, surprise, sadness, etc.).

[2024] 3. Packaging and sending data: The device packages the tap information (frame image, audio data, playback time, video title) and emotion data and sends them to the server.

[2025] 4. Information Identification and Retrieval: The server analyzes the received package information and uses a generative AI model to identify costumes and locations from frame images and songs from audio data, using prompts such as:

[2026] Example prompt: "What outfit is in this picture?"

[2027] Sample prompt: "Which song does this audio data match?"

[2028] 5. Save to Favorites List: The server searches online shopping sites and music streaming services for the identified objects and information, retrieves related product pages and music track information, and saves the retrieved information by calling APIs to save it in the databases of other services used by the user.

[2029] 6. Notification and Personalization: The server verifies that the information has been successfully stored in other services and generates a notification message based on the result. Furthermore, the content of the notification message is customized based on the user's emotional state and displayed on the device.

[2030] Specific examples

[2031] Scenario 1: User identifies and purchases an outfit from a video

[2032] 1. User: While watching a video, find an outfit you like and single-tap the screen.

[2033] 2. Device: Detects a single tap and also captures and analyzes the user's facial expression with a camera.

[2034] 3. Terminal: Sends the current frame image, video title, and user emotion data to the server.

[2035] 4. Server: Analyzes the frame image and identifies the outfit using a generative AI model. Example: "What outfit is in this image?"

[2036] 5. Server: Search the online shopping site and retrieve the product page of the relevant costume.

[2037] 6. Server: Calls the API and registers the identified outfit in the user's favorites list.

[2038] 7. Server: Customize notification messages such as "Outfit has been added to your favorite list" based on emotion data.

[2039] 8. Device: Show notifications to users and convey messages based on their emotions.

[2040] 9. User: After watching, you can check and purchase the costumes from your favorite list.

[2041] Scenario 2: User identifies a song in a video and adds it to a playlist

[2042] 1. User: When a song you like comes on while watching a video, double tap the screen.

[2043] 2. Device: Detects double-tap gestures and also analyzes the user's voice to recognize emotions.

[2044] 3. Terminal: Sends the current audio data, playback time, and user emotion data to the server.

[2045] 4. Server: Analyzes the audio data and uses a generative AI model to identify the song. Example: "Which song does this audio data match?"

[2046] 5. Server: Searches music streaming services and retrieves track information for identified songs.

[2047] 6. Server: Calls the API and adds the identified song to the user's playlist.

[2048] 7. Server: Customize notification messages such as "A song has been added to your playlist" based on sentiment data.

[2049] 8. Device: Show notifications to users and convey messages based on their emotions.

[2050] 9. User: After listening, you can check the song from the playlist and play it.

[2051] The present invention allows users to easily identify objects or information of interest while watching videos, and further provides a personalized experience based on emotion recognition.

[2052] The flow of the identification process in the second embodiment will be described with reference to FIG.

[2053] Step 1:

[2054] Capturing user actions

[2055] User: When they find an object that catches their eye while watching a video, they tap the screen. For example, they tap once when an outfit that catches their eye is displayed.

[2056] Terminal: Detects the user's tap operations in real time. As input, it acquires the tap operation and its timing (playback time). Specifically, it determines the type of tap (single tap, double tap, triple tap, etc.) and identifies the category (e.g., single tap is an outfit). As output, it generates the type of tap operation and the corresponding category information.

[2057] Step 2:

[2058] Capturing user emotions

[2059] Terminal: The device takes a picture of the user's face with a camera and analyzes their facial expressions. It also uses voice recognition technology to analyze the tone and pitch of their voice. It receives facial images and voice data as input. Specifically, it processes the user's facial expressions in real time using face recognition software (e.g., OpenCV) and analyzes the tone of their voice using voice recognition software (e.g., Google Cloud Speech-to-Text). It outputs the user's emotional state (e.g., joy, surprise, sadness) as text data.

[2060] Step 3:

[2061] Sending tap information and emotion data

[2062] Terminal: Packages tap information and emotion data. As input, it acquires tap information (category, playback time), emotion data, frame image (screenshot at the time of tap), audio data, video title, etc. As a concrete example, it packages this data in JSON format and generates an HTTP POST request to send it to the server. As output, it sends the packaged data to the server.

[2063] Step 4:

[2064] Identifying and locating information

[2065] Server: Analyzes the received package information. As input, it receives packaged data. Specifically, it analyzes the received JSON data and uses a generative AI model to identify costumes and locations from frame images and identify songs from audio data. An example prompt is "What costume is in this image?". As output, it generates the identified objects and information as data.

[2066] Step 5:

[2067] Save to Favorites list

[2068] Server: Saves the identified information in the database of the user's other services. Receives the identified objects and information as input. Specific operations include calling the API of an online shopping site or music streaming service to add the identified outfits and songs to the user's list. Receives the status of the save completion as output.

[2069] Step 6:

[2070] Order completion notification

[2071] Server: Confirms that the information has been successfully saved to another service. Receives the save completion status as input. For example, generates a success message and creates a notification message such as "Information has been saved." Prepares the generated notification message as output.

[2072] Step 7:

[2073] Creating Emotion-Based Notifications

[2074] Server: Customizes the content of the notification message based on the recognized emotional state of the user. Receives the user's emotional data and order completion status as input. Specific behavior is to create a personalized message such as "Ah, I'm surprised! Do you like this outfit?" depending on the emotional state (e.g., surprise). Generates a customized notification message as output.

[2075] Terminal: Displays a customized notification message to the user. As input, it receives the notification message received from the server. The specific behavior is to display the message to the user using a push notification or an alert dialog. As output, it performs the action of allowing the user to receive the notification.

[2076] (Application example 2)

[2077] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[2078] Conventional video viewing systems have had problems such as making it difficult for users to easily save objects or information that interest them while viewing, and lacking means for analyzing users' emotions to provide a more personalized experience. The present invention aims to solve these problems and provide a system for improving the video viewing experience.

[2079] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[2080] In this invention, the server includes means for detecting when a user performs an operation on a specific object or information while watching a video, means for acquiring specific frames or audio data of the video, and video metadata, means for identifying the object or information using a generative AI model based on the acquired data, means for storing the identified object or information in a database of another service used by the user, means for notifying the user that the information has been saved, means for acquiring and analyzing emotion data, and means for customizing a notification message based on the emotion data. This not only enables the user to easily save objects or information that interest them, but also enables the analysis of emotions to provide personalized notification messages.

[2081] "Means for detecting when a user performs an operation on a specific object or information while watching a video" refers to technology that recognizes when a user performs an action on a specific object or information, such as by tapping the screen, while watching a video.

[2082] "Means for obtaining specific frames and audio data from a video, and video metadata" refers to technology for extracting frames and audio from a video that a user is interested in, as well as metadata related to that video.

[2083] "Means for identifying objects and information using a generative AI model based on acquired data" refers to a technique for applying a generative AI model using extracted data to identify objects and information of interest.

[2084] "Means for storing identified objects or information in the databases of other services used by the user" refers to technology for recording identified objects or information in the databases of other online services used by the user.

[2085] The "means for notifying the user that the information has been saved" is a technique for notifying the user that the specified object or information has been saved successfully.

[2086] "Means for acquiring and analyzing emotional data" refers to technology for acquiring and analyzing emotions from the user's facial expressions, voice, etc.

[2087] "Means for customizing notification messages based on emotional data" refers to a technology for generating notification messages tailored to individual users based on analyzed emotional data of the users.

[2088] The present invention relates to a system that identifies objects or information that a user finds interesting while watching a video, saves them in the user's favorites list, and recognizes the user's emotions to provide a personalized experience. A system for implementing the present invention will be described in detail below.

[2089] System Overview

[2090] The system includes the following main features:

[2091] 1. Capturing user actions

[2092] 2. Capturing user emotions

[2093] 3. Data submission and analysis

[2094] 4. Identifying and storing objects and information

[2095] 5. Generating notifications based on emotions

[2096] Capturing user actions

[2097] While watching a video, if a user comes across an object or piece of information that interests them (such as an outfit, song, or location), they can tap the screen to show their interest. Different types of taps, such as a single tap, double tap, or triple tap, indicate interest in different categories. This feature is detected by the smartphone's touchscreen.

[2098] Capturing user emotions

[2099] The device (smartphone) captures the user's facial expression with its front camera and analyzes it using the OpenCV library. It also uses the voice recognition function to analyze the voice tone and pitch with the Google Cloud Speech-to-Text API to recognize the user's current emotional state. This allows it to obtain emotional data such as joy, surprise, and sadness.

[2100] Data transmission and analysis

[2101] The user's tap information (frame image, audio data, playback time, video title) and emotion data are sent from the device to the server. The server receives the data using an API server that uses Flask and analyzes the packaged data.

[2102] Identifying and storing objects and information

[2103] The server uses PyTorch to run a generative AI model, identifying objects (such as costumes) from the received frame images and songs from the audio data. The identified objects and information are stored in databases for other services used by the user.

[2104] Generate notifications based on emotions

[2105] The server generates a customized notification message based on the emotion data. For example, if the user is surprised, it generates a message like "Wow! Do you like this outfit?". The above process allows the user to have a more personalized experience.

[2106] Specific examples

[2107] Example 1: Identifying an outfit while watching a video

[2108] The user is interested in a particular outfit and single-taps the screen. This operation information and the user's emotional data (e.g., a happy expression) are sent from the device to the server. The server analyzes the frame image using a generative AI model and adds the identified outfit to the user's favorites list from the online shopping site. The notification includes the message, "Great choice!"

[2109] Example prompt for a generative AI model:

[2110] Identify specific outfits from frame images in your video. These images contain the fashion items shown in the frame. Identify outfits that viewers particularly like and provide links to online shopping sites.

[2111] As described above, the present invention provides a specific method for analyzing a user's interests and emotions and providing a personalized experience.

[2112] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[2113] Step 1:

[2114] Detecting user taps

[2115] Input: A user single-tap, double-tap, or triple-tap the screen while watching a video.

[2116] How it works: The device uses the touchscreen sensor to detect taps and distinguish between single, double, and triple taps.

[2117] Output: Generates data about the type of tap.

[2118] Step 2:

[2119] User Emotion Capture

[2120] Input: Video and audio data of the user's face at the moment the touch action is performed.

[2121] How it works: The device's front camera captures a picture of the user's face and analyzes their facial expressions using OpenCV. It also analyzes audio data acquired from the microphone using the Google Cloud Speech-to-Text API to detect tone and pitch of the voice and identify emotions.

[2122] Output: Parsed emotion data (happiness, surprise, sadness, etc.).

[2123] Step 3:

[2124] Sending and packaging data

[2125] Input: Tap type data, emotion data, frame image, audio data, playback time, video title.

[2126] How it works: The device packages this data and sends it to an API server using Flask, securely transmitting it using network protocols.

[2127] Output: A packaged dataset.

[2128] Step 4:

[2129] Data reception and analysis by the server

[2130] Input: Packaged dataset.

[2131] How it works: The server receives data via the Flask API, inputs frame images into a PyTorch-based generative AI model to identify objects (e.g., costumes), and analyzes audio data to identify songs.

[2132] Output: Data on identified objects and songs.

[2133] Step 5:

[2134] Storing objects and information

[2135] Input: Identified object and song data.

[2136] How it works: The server calls the database API of another service to add these data to the user's favorites list.

[2137] Output: A status indicating the information has been saved.

[2138] Step 6:

[2139] Customize and generate notifications based on emotions

[2140] Input: Saved status and emotion data.

[2141] How it works: The server customizes the notification message based on the emotion data and generates a message like "Your information has been saved." For example, if the user's emotion is joy, it creates a message like "Great choice!"

[2142] Output: A customized notification message.

[2143] Step 7:

[2144] Display notifications on your device

[2145] Input: Customized notification message.

[2146] What it does: The device displays a notification to the user, conveying the stored information along with an emotionally appropriate message.

[2147] Output: The user receives a notification.

[2148] Example prompt sentence:

[2149] "Identify specific outfits from framed images in the video. These images contain the fashion items shown in the frame. Identify outfits that viewers particularly liked and provide links to online shopping sites."

[2150] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[2151] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[2152] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[2153] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[2154] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[2155] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[2156] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[2157] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, motorcycles, and other devices, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[2158] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[2159] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[2160] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[2161] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[2162] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[2163] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[2164] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[2165] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[2166] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[2167] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[2168] Furthermore, the har...

Claims

1. A means for detecting when a user performs an operation on a specific object or information while watching a video; A means of obtaining specific frames of video, audio data, and video metadata; A means of identifying objects and information using a generative AI model based on the acquired data; and means for storing identified objects and information in the databases of other services used by the user; a means for notifying the user that the information has been saved; A system including:

2. The system of claim 1 further comprising means for identifying the type of tap a user performs.

3. 2. The system of claim 1, further comprising means for packaging the acquired data and transmitting the packaged data to a server.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A