system
The system addresses the need for instant product information by allowing users to tap on video subjects, using AI to extract and provide purchase links, thereby improving the shopping experience.
Patent Information
- Application Number
- JP2024181694
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-10-17
- Publication Date
- 2026-04-30
AI Technical Summary
There is a lack of a method for viewers to instantly obtain information about products they are interested in within video content and smoothly connect to the purchase process.
A system that allows users to tap on specific subjects within a video frame, using an AI model to extract information, which is sent to a server to search for similar products and provide product information and purchase links.
Enables users to intuitively purchase products of interest by providing immediate and personalized product information, enhancing the shopping experience.
Smart Images

Figure 2026071656000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance as a response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] There is a need for a new marketing method to stimulate viewers' purchasing desire through video content. However, currently, there is a lack of a method for viewers to instantly obtain information about products they are interested in and smoothly connect to the purchase. To solve this problem, a mechanism is needed to quickly link the user's interest in items within a video to product purchases.
Means for Solving the Problems
[0005] This invention provides a system that, when a user taps on a specific subject within a video frame on their device while watching video content, uses an AI model to extract information about that subject. The extracted information is sent to a server, which searches an online database for information on similar products. Furthermore, the server provides the user with product information and a purchase link, enabling the user to intuitively purchase products that interest them.
[0006] A "user" refers to an individual who uses the system to view video content and interacts with specific items out of interest.
[0007] "Video content" refers to a series of video data distributed to viewers.
[0008] A "device" refers to an electronic device that a user uses to watch video content.
[0009] "Subject" refers to a specific item or product that the user focuses on and tries to obtain information about within the video.
[0010] A "server" refers to a central computer system that receives data transmitted from devices, performs analysis, and searches for product information.
[0011] An "AI model" refers to an algorithm that uses artificial intelligence technology to perform image analysis and is responsible for recognizing subjects and extracting features.
[0012] An "online database" refers to a database system that stores product information and allows users to search and provide that information as needed.
[0013] "Similar products" refer to products that are related to a subject the user is interested in and that are available for purchase online. [Brief explanation of the drawing]
[0014] [Figure 1]It is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] It is a conceptual diagram showing an example of the main functions of a data processing device and a smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. <00神仙道0065> [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] [[ID=1神仙道8]]It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [[ID=2神仙道]] [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which multiple emotions are mapped. [Figure 10] It shows an emotion map to which multiple emotions are mapped. <000008神仙道>[[ID=3神仙道]]It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when an emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when an emotion engine is combined. [[ID=4神仙道]]
Mode for Carrying Out the Invention
[0015] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described according to the accompanying drawings.
[0016] First, the terms used in the following description will be explained.
[0017] In the following embodiments, the labeled processor (hereinafter simply referred to as "processor") may be one arithmetic unit or a combination of a plurality of arithmetic units. Also, the processor may be one type of arithmetic unit or a combination of a plurality of types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0018] In the following embodiments, the labeled RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0019] In the following embodiments, the labeled storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, and the like.
[0020] In the following embodiments, the labeled communication I / F (Interface) is an interface including a communication processor and an antenna, etc. The communication I / F controls communication between a plurality of computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark), and the like.
[0021] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0022] [First Embodiment]
[0023] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0024] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0025] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0026] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0027] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0028] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0029] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0030] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0031] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0032] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0033] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0034] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0035] This invention constructs an interactive system that directly provides product information from video content, thereby stimulating viewers' purchasing intent. Users can view video content using mobile devices, tablets, etc. If a user becomes interested in a particular item on this device, they tap the screen, and the system collects information about that item.
[0036] An AI model running on the device recognizes subjects within the video frame, analyzes tapped locations, and extracts features related to those subjects. This feature data is sent from the device to a server, where it performs a more detailed product information search. The server accesses an online database and applies an algorithm to find similar products. Information on similar products includes product names, prices, and purchase links, and the server formats this information and sends it back to the user's device.
[0037] As a concrete example, consider a scenario where a user is watching a fashion-related video and becomes interested in a particular bag. When the user taps on the bag, the device identifies the bag within the video frame and extracts its features. Based on the feature data sent to the server, the server retrieves information on similar bags from shopping websites. This information, along with a link to the bag's details page, is sent to the user, allowing them to browse and purchase the product.
[0038] Furthermore, this system records user behavior data, which is used to improve AI models and target advertisements. This system allows viewers to gain direct inspiration from videos and instantly access detailed product information, dramatically improving the shopping experience.
[0039] The following describes the processing flow.
[0040] Step 1:
[0041] Users watch videos and tap on products or items that interest them on the screen. The tap locations are then recorded by the device.
[0042] Step 2:
[0043] The device captures a video frame at the tapped location and uses an AI model to recognize the subject within that frame. Feature data of the subject is then extracted.
[0044] Step 3:
[0045] The device sends the extracted feature data to the server. The data includes information such as color, shape, and size.
[0046] Step 4:
[0047] Based on the received data, the server uses an AI algorithm to search online databases and identify information on similar products.
[0048] Step 5:
[0049] The server organizes the retrieved data on similar products and generates an information package that includes product name, price, image, and purchase link.
[0050] Step 6:
[0051] The server sends the generated information package to the user's terminal.
[0052] Step 7:
[0053] The terminal displays the received product information on the user interface and provides the user with product details and a purchase link.
[0054] Step 8:
[0055] Users can review the displayed information and, if interested, tap the provided link to access the product purchase page.
[0056] (Example 1)
[0057] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0058] A system is needed that can quickly and intuitively provide users viewing video information on communication terminals with detailed information about objects they are interested in. Conventional systems require users to manually search for information when they want to know more about items in the video they are watching, lacking convenience and speed. Furthermore, there were challenges in optimizing marketing using user selection history.
[0059] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0060] In this invention, the server includes means for the terminal to analyze an image and extract feature information of an object when the user taps an object while viewing video information; means for transmitting the feature information to a processing unit, which then searches for similar item information from an online storage area; and means for generating information about items related to the object and providing the user with purchase guidance. This allows the user to easily obtain information about related items from the video they are viewing and to proceed with the purchase process interactively.
[0061] "Video information" refers to video data that users can view using their communication devices.
[0062] An "image" is a static visual image extracted from any frame within a video.
[0063] An "object" is a specific object of interest that exists within the video information.
[0064] A "terminal" is a portable or stationary electronic device used by a user to view video information.
[0065] "Feature information" refers to data that represents the characteristics of an object, including its shape, color, and size.
[0066] A "processing device" is a computer system that receives characteristic information and searches for information on similar items.
[0067] "Online storage" refers to a database or data storage that is accessible over a network.
[0068] "Item information" refers to detailed data including the name, price, and purchase link of similar objects.
[0069] A "purchase guide" is an interface that provides users with information and links that enable them to purchase goods.
[0070] This invention constructs an interactive system that directly provides product information from video information, thereby stimulating viewers' purchasing intent. Users can view the video information on a portable or stationary electronic device. For example, suppose a user is watching a fashion-related video and becomes interested in a particular bag. The user can then tap the portion of the image showing that bag on their device.
[0071] The device, upon receiving a tap from the user, uses a generated AI model to recognize objects in the tapped image. The AI model extracts feature information from objects within the image and analyzes attributes such as shape, color, and size. This identifies objects of interest to the user, and that information is sent to the server.
[0072] The server inputs the feature information received from the terminal into the processing unit to access online storage. The processing unit applies an algorithm to search for similar items and finds item information. This item information includes the product name, price, and purchase link. The server returns the formatted information to the user's terminal, allowing the user to view the item details and make a purchase.
[0073] As an example of a prompt, you could instruct the generative AI model to "explain a system that, when a user taps on an item they are interested in within a video, extracts the features of that item and provides information on similar products."
[0074] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0075] Step 1:
[0076] The user views video information using a communication device. They tap on the part of the image that contains an object of interest. The input is the coordinates of the user's tap, and the output is image data of the tapped location. This action initiates the detection of the target object.
[0077] Step 2:
[0078] The device receives image data from the tapped location and uses a generative AI model to recognize the object. The input is the tapped image data, and the output is the object's feature information. The AI model extracts features such as the object's shape, color, and size based on the image data. In this step, the device collects specific attribute information of the detected object.
[0079] Step 3:
[0080] The terminal transmits feature information to the processing unit. This processing unit is connected to a server and performs initial data processing. The input is the feature information of an object, and the output is formatted data for database searching. This process formats the feature information into a format that can be processed by the server.
[0081] Step 4:
[0082] The server performs searches by comparing formatted feature information with a database in online storage. The input is formatted feature data, and the output is information about similar items. The server uses an algorithm to identify items that match the database contents and extracts the necessary information.
[0083] Step 5:
[0084] The server receives item information and formats it into a user-friendly format. The input is raw data of similar items, and the output is formatted information including product name, price, and purchase link. This information formatting allows users to efficiently obtain the necessary information.
[0085] Step 6:
[0086] The server sends formatted item information to the terminal. The terminal presents the received information to the user and displays details of the object tapped within the video information. The input is formatted item information, and the output is the user's visual interface. Based on this information, the user can access the online store and purchase items that interest them.
[0087] (Application Example 1)
[0088] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0089] During video viewing, there is a lack of means for viewers to immediately grasp detailed information about specific products or services and translate that information into purchasing decisions. As a result, viewers' purchasing intent cannot be immediately satisfied, and the shopping experience is limited.
[0090] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0091] In this invention, the server includes means for an information processing device to analyze information and extract information about a specific object within video information when a user taps on that object while watching the video information; means for transmitting the information about the object to the information processing device, which then searches for similar product information from online information sources; and means for generating information about products related to the object and providing the user with purchase-inducing information. This makes it possible to instantly obtain product information about an object of interest in the video being watched and to quickly engage in purchasing activities.
[0092] "User is watching video information" refers to a situation where a user is using their device to play a video in real time and review its content.
[0093] "A subject within specific video information" refers to an object or person that the user is interested in within the image displayed at a specific point in time during video playback.
[0094] An "information processing device" is a general term for hardware and software that have the function of receiving data, analyzing it, and providing appropriate information to the user.
[0095] "Extracting information about a target" refers to the process of collecting detailed data about a specific target, such as its shape, color, and characteristics.
[0096] "Online information sources" refer to databases and web services accessible via the internet, and serve as a source of product information.
[0097] "Searching for similar product information" is the act of finding data on products with similar characteristics to the extracted target information from online information sources.
[0098] "Providing purchase-inducing information" means showing users details about products they are interested in and presenting links and pricing information to facilitate the purchase process.
[0099] "Operates as an application installed on a smartphone or tablet device" means that it is a program developed to run on a mobile device.
[0100] This invention is a system that, when a user watches a video using a mobile device such as a smartphone or tablet, instantly obtains product information related to a specific object if the user becomes interested in that object, thereby promoting purchasing activity.
[0101] When a user taps on a specific object in a video they are watching, the device begins processing the information. The device is equipped with an AI model that recognizes the object in the tapped video information and extracts its features. AI modeling software such as TENSORFLOW® Lite is used for this process. The extracted feature information is sent from the device to a server via a network protocol.
[0102] Based on the received feature information, the server accesses a cloud database (e.g., Firebase) to search for similar product information from online sources. This uses REST API communication technology. When similar product information is found, the server composes the information and sends it back to the user's device. The user can then view the product name, price, and purchase link through the device's interface and proceed with the purchase.
[0103] For example, if a user sees a bag in a scene from a drama and becomes interested in it, they can use this feature to check information about the bag on the spot and easily access the purchase page.
[0104] Example of a prompt:
[0105] "How can we easily recognize specific items from videos a user is watching and retrieve information about them? How does an AI app work to extract item features and provide purchase links for similar products?"
[0106] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0107] Step 1:
[0108] While watching a video, the user taps on an object of interest on the screen. The input is the user's tap location, and the device detects this location and captures the frame of the tapped video information. The output is the video frame of the object in question.
[0109] Step 2:
[0110] The system inputs video frames captured by the device into an AI model to identify and recognize objects. The input is video frames, and the output is feature information of the recognized objects. TensorFlow Lite is used as the AI model, and its calculations extract features such as the shape, color, and size of specific objects within the video.
[0111] Step 3:
[0112] The terminal extracts feature information and sends it to the server via a REST API. The input is feature information, and the output is a log of successful communication to the server. This communication allows the server to confirm that the data has been successfully delivered.
[0113] Step 4:
[0114] Based on the feature information received by the server, it accesses a cloud database to search for similar product information. The input is feature information, and the output is a dataset of similar products. The server queries online information sources (e.g., Firebase) to collect information such as the characteristics and prices of similar products.
[0115] Step 5:
[0116] The server searches for product information and sends it back to the user's terminal. The input is product information, and the output is a notification that data transmission to the user's terminal is complete. The server formats the data and returns it for a smooth user experience.
[0117] Step 6:
[0118] The received product information is displayed on the user's device interface. Input is product information sent from the server, and output is a visual display. Users can view product details, pricing, and purchase links, and proceed directly with the purchase.
[0119] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0120] This invention constructs an interactive system that provides product information directly from video content while taking into account the viewer's emotions. Users can view the video content using mobile devices or tablet devices. These devices are equipped with an emotion engine that further enhances the user's viewing experience.
[0121] When a user taps on a product or item they are interested in within a video, an AI model running on the device recognizes the subject within that video frame and extracts its feature data. Additionally, an emotion engine analyzes the user's facial expressions and actions to recognize their emotions while watching. This resulting emotion data, along with the subject's feature data, is then sent to the server.
[0122] Based on the received feature and sentiment data, the server searches an online database to find information on similar products that best match the user's emotions. Specifically, it incorporates sentiment data into the analysis to customize product suggestions based on emotions. The server organizes the acquired product information and generates an information package that includes product names, prices, and purchase links optimized based on emotions.
[0123] For example, if a user is watching a travel-related video and makes a happy face while tapping their favorite suitcase, the emotion engine will sense that happiness. The device will identify the suitcase at the tapped location and extract its features. The server will then perform a search based on this data, retrieving information about suitcases that particularly resonate with the excitement of travel, and providing this information to the user.
[0124] This system not only allows users to get direct inspiration from videos and instantly obtain detailed product information, but also enables emotionally-driven product selection, providing a more personalized shopping experience.
[0125] The following describes the processing flow.
[0126] Step 1:
[0127] The user watches a video and taps on products or items that interest them on the screen. The device records the tap position and obtains the video frame at that point.
[0128] Step 2:
[0129] The device analyzes the acquired video frames using an AI model to identify the subject at the tapped location. This process extracts characteristic data of the subject.
[0130] Step 3:
[0131] The device uses an emotion engine to evaluate the user's current emotions. Facial recognition technology analyzes the type and intensity of emotions from the user's facial expressions and voice tone.
[0132] Step 4:
[0133] The device transmits subject characteristic data and user emotion data to the server. The transmitted data includes subject identification information and the user's emotional state.
[0134] Step 5:
[0135] The server searches online databases based on the received data to find information on similar products related to the subject. During this process, it optimizes product recommendations by taking into account the user's emotional data.
[0136] Step 6:
[0137] The server organizes product information and generates an information package that reflects sentiment-based customization. This information package includes product details, images, pricing, and a purchase link.
[0138] Step 7:
[0139] The server sends the generated information package to the user's terminal.
[0140] Step 8:
[0141] The device displays the received product information to the user. The user can check the product details and easily purchase items they are interested in through the purchase link.
[0142] (Example 2)
[0143] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0144] Traditional video viewing presents challenges such as users being unable to quickly obtain detailed information about specific objects within a video while watching, and being unable to receive personalized product recommendations based on their emotions. Furthermore, there is a lack of effective methods for providing information that takes user emotions into consideration.
[0145] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0146] In this invention, the server includes means for the device to analyze video data and obtain attribute data of an object when the user selects an object in the video data during the viewing experience; means for transmitting the attribute data of the object and the user's emotion data, and for the data processing device to retrieve related information from an external storage device; and means for generating suggested content related to the object and presenting a transaction link to the user. As a result, the user can instantly obtain information about products that have caught their interest while watching a video and receive personalized product suggestions that match their emotions.
[0147] A "user" refers to an individual who uses this system to view digital content and operate specific functions.
[0148] "Viewing experience" refers to the process by which a user views video content through a digital device.
[0149] "Video data" refers to viewable video information that is stored or distributed in digital format.
[0150] "Object" refers to an object in the video data that the user has focused their attention on and which is the subject of identification or analysis.
[0151] "Device" refers to a hardware device used by the user, specifically equipment for analyzing video and processing data.
[0152] "Attribute data" refers to information about the features and characteristics associated with an object, and is obtained through image analysis.
[0153] "Emotional data" refers to information about a user's emotional state and is collected through sentiment analysis.
[0154] A "data processing device" refers to a computer system used to process received data and retrieve or generate information.
[0155] "External storage device" refers to a database or storage system that a data processing device accesses to retrieve information.
[0156] "Proposed content" refers to information about products and services offered to the user, and is customized based on the user's choices and feelings.
[0157] A "transaction link" refers to a web link that a user uses to purchase or use a suggested product or service.
[0158] This invention is built as a system that takes into account the viewer's emotions while they are watching video content and provides product information interactively. Users watch video content using a mobile device or tablet and select items that interest them. This system has the following hardware and software configuration.
[0159] The device is equipped with a high-resolution camera and touchscreen, and runs an emotion engine and machine learning models for object recognition. For example, OpenCV is used for the emotion engine, and YOLO or TensorFlow are used for the machine learning models. This allows the device to analyze the user's facial expressions in real time, recognize objects in videos the moment they are tapped, and acquire attribute data for those objects.
[0160] When a user taps a specific object, the device sends the object's attribute data and the user's sentiment data to a server. The server processes the received data and searches for information stored in external storage devices. The server generates information about related products and services based on the user's sentiment and sends it to the device as a suggestion.
[0161] For example, if a user is enjoying watching a travel video and taps on a suitcase in the video, the emotion engine will determine the user's emotion as "happy." The device retrieves the suitcase's attribute data and sends it to the server. The server uses this data to search its database for travel suitcase information and generates suggestions, including a purchase link. The user can then purchase the suitcase through the information displayed on the screen.
[0162] A concrete example of a prompt message would be: "Explain how to provide related product information when the user taps on a product with a happy expression."
[0163] In this way, this system instantly provides personalized product information based on the objects that the user shows interest in while watching videos, realizing a personalized shopping experience that responds to the user's emotions.
[0164] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0165] Step 1:
[0166] Users view video content on mobile devices or tablets. The input is video data selected by the user. This data is played on the device, and the user's facial expressions are captured by the camera during viewing. The output provides a video viewing experience and generates image data of the user's facial expressions.
[0167] Step 2:
[0168] The device analyzes the user's facial expression data using an emotion engine. Real-time captured image data of facial expressions is used as input. Here, OpenCV is used to detect facial features and process the data to determine emotional states (e.g., happiness, surprise). The output is emotional data representing the user's emotions.
[0169] Step 3:
[0170] The user taps on an object of interest within the video. The input consists of the user's tap location data and the current video frame data. The device uses machine learning models such as YOLO or TensorFlow to recognize the subject within the video frame and extract feature data. The output is attribute data of the tapped object.
[0171] Step 4:
[0172] The terminal sends attribute data of the acquired object and user sentiment data to the server. As input, the attribute data and sentiment data are packaged together. This data is sent to the server using a secure communication protocol. As output, the data is successfully delivered to the server.
[0173] Step 5:
[0174] The server processes data based on the received data. It receives attribute data and sentiment data as input. The server analyzes this data and searches for relevant information from external storage. Specifically, it executes database queries to retrieve product information that matches the specified criteria. Suggestions are generated as output.
[0175] Step 6:
[0176] The server sends the generated suggestions to the terminal. The input is package data containing the suggested product information and related links. The output is the package data sent to the terminal.
[0177] Step 7:
[0178] The terminal displays the proposed content to the user. Using package data received from the server as input, it generates a user interface that presents the proposed content to the user. The user can view product details and transaction links on the screen. The output is a visual display of information to support the user's decision-making process.
[0179] (Application Example 2)
[0180] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0181] In today's world, there is a lack of flexible, interactive systems that can directly provide product information based on viewers' emotions. This makes it difficult for viewers to instantly acquire information about related products based on the inspiration they gain from videos. Furthermore, the inability to suggest products that take viewers' emotions into account makes it difficult to provide personalized purchasing experiences. Technologies that can solve these problems are needed.
[0182] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0183] In this invention, the server includes means for the device to analyze a scene and extract information about an object when the user selects an object by action while watching a video; means for transmitting emotion data along with the object information to the server, and for the server to search for similar product information from an online information database; and means for generating product information related to the object based on the emotion data and providing the user with a purchase path. As a result, the user receives appropriate product suggestions according to their emotions, enabling a personalized purchasing experience.
[0184] A "user" is the entity that watches a video, selects objects of interest through actions, and interacts with the system.
[0185] "Video footage" is a media format that provides a series of images and is a means of conveying information visually and aurally.
[0186] An "object" is something that is visually represented within a video and that the user shows interest in as a potential selection.
[0187] "Selection by action" refers to an interaction a user performs with a specific object while watching a video, such as tapping the screen.
[0188] "Device" refers to a digital device capable of playing and interacting with video footage, and includes hardware and software for processing and transmitting data.
[0189] A "scene" refers to a specific moment or state within a video, and is the unit in which a user interacts with a particular object.
[0190] "Analysis" refers to the process by which a device extracts information from scenes within video footage and performs processing to identify the characteristics of objects.
[0191] "Information" refers to identifiable data about an object, including content that describes its features and attributes.
[0192] "Extraction" is the process of taking out the features necessary for identifying objects from video footage and organizing them as data.
[0193] "Emotional data" refers to data that expresses a user's emotional state using numbers or codes, and accurately reflects the user's emotions during interaction.
[0194] A "server" is a computer system that processes data via a network and performs specific functions using data received from external devices.
[0195] An "online information collection" is a collection of databases and information sources that exist on the internet, and is the target for which a server searches for product information.
[0196] "Similar product information" refers to data obtained by searching for highly relevant products based on the characteristics of the object selected by the user.
[0197] "Generation" refers to the operation of creating a new data package based on the information obtained and providing it to the user in a useful format.
[0198] A "purchase channel" is a means of directly supporting a user's purchase by providing the necessary information and links to buy a product.
[0199] This invention is an interactive system that allows users to directly interact with objects they are interested in while watching video footage, and provides personalized product information accordingly.
[0200] Users view video footage using digital devices such as smartphones and tablets. During viewing, users interact with specific objects, causing the device to analyze the video scene and acquire information about those objects along with the user's emotional data. This analysis utilizes an AI model that analyzes the user's facial expressions captured by a camera. Specific software such as OpenCV and Dlib can be used.
[0201] The device sends analyzed object information and emotion data to a server. The server receives this data and searches its database for similar product information. The server runs on a platform like Node.js, and its product information collection and customization process includes data processing to consider the user's emotions. As a result, product information that best matches the user's emotions is generated and sent back to the device.
[0202] As a concrete example, suppose a user watches a travel-related video and selects a specific suitcase with a joyful expression. In this case, the system recognizes the suitcase and, based on the user's emotional data, provides information about suitcases that will further enhance their enjoyment of the trip.
[0203] An example of a prompt used in this system is: "The user is tapping an accessory displayed during the movie. The user's current emotion has been detected as 'enjoyment.' Please suggest product information for similar accessories related to enjoyment."
[0204] In this way, users can leverage the immediate inspiration they gain from the videos they are watching and enjoy a customized purchasing experience based on their emotions.
[0205] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0206] Step 1:
[0207] A user watches a video and selects a specific object of interest. This is done by tapping the screen of a digital device. The input is the coordinate data of the user's tap position. The output is the data necessary to identify the selected object.
[0208] Step 2:
[0209] The system analyzes the tap locations received by the device and uses an AI model to identify the corresponding object in the video scene. This object detection algorithm is applied using TensorFlow or PyTorch. Input consists of tap location coordinates and video frames, while output is feature data of the identified object.
[0210] Step 3:
[0211] The terminal camera detects the user's emotions, and real-time facial expression analysis is performed using OpenCV and Dlib. The input is camera footage, and the output is user emotion data. This allows the user's emotional state during viewing to be captured as numerical data.
[0212] Step 4:
[0213] The device sends characteristic data of the identified object and user emotion data to the server. This data is sent to the server via an HTTP request in JSON format. The input is data on the identified object and emotion state, and the output is the transmission of data to the server.
[0214] Step 5:
[0215] The server searches for similar product information from an online database based on the received object's feature data and sentiment data. This process uses Node.js and queries the database. The input is feature data and sentiment data, and the output is similar product information.
[0216] Step 6:
[0217] The server constructs an information package based on the user's sentiment, using the acquired product information. The package includes the product name, price, and purchase link. The input is similar product information, and the output is the constructed information package. This process enables product recommendations tailored to the user's sentiment.
[0218] Step 7:
[0219] The device displays an information package received from the server to the user. React Native is used to visually present the information on the smartphone screen. The input is an information package from the server, and the output is the display of product information to the user. This allows the user to instantly view detailed information about products they are interested in.
[0220] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0221] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0222] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0223] [Second Embodiment]
[0224] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0225] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0226] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0227] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0228] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0229] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0230] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0231] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0232] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0233] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0234] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0235] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0236] This invention constructs an interactive system that directly provides product information from video content, thereby stimulating viewers' purchasing intent. Users can view video content using mobile devices, tablets, etc. If a user becomes interested in a particular item on this device, they tap the screen, and the system collects information about that item.
[0237] An AI model running on the device recognizes subjects within the video frame, analyzes tapped locations, and extracts features related to those subjects. This feature data is sent from the device to a server, where it performs a more detailed product information search. The server accesses an online database and applies an algorithm to find similar products. Information on similar products includes product names, prices, and purchase links, and the server formats this information and sends it back to the user's device.
[0238] As a concrete example, consider a scenario where a user is watching a fashion-related video and becomes interested in a particular bag. When the user taps on the bag, the device identifies the bag within the video frame and extracts its features. Based on the feature data sent to the server, the server retrieves information on similar bags from shopping websites. This information, along with a link to the bag's details page, is sent to the user, allowing them to browse and purchase the product.
[0239] Furthermore, this system records user behavior data, which is used to improve AI models and target advertisements. This system allows viewers to gain direct inspiration from videos and instantly access detailed product information, dramatically improving the shopping experience.
[0240] The following describes the processing flow.
[0241] Step 1:
[0242] Users watch videos and tap on products or items that interest them on the screen. The tap locations are then recorded by the device.
[0243] Step 2:
[0244] The device captures a video frame at the tapped location and uses an AI model to recognize the subject within that frame. Feature data of the subject is then extracted.
[0245] Step 3:
[0246] The device sends the extracted feature data to the server. The data includes information such as color, shape, and size.
[0247] Step 4:
[0248] Based on the received data, the server uses an AI algorithm to search online databases and identify information on similar products.
[0249] Step 5:
[0250] The server organizes the retrieved data on similar products and generates an information package that includes product name, price, image, and purchase link.
[0251] Step 6:
[0252] The server sends the generated information package to the user's terminal.
[0253] Step 7:
[0254] The terminal displays the received product information on the user interface and provides the user with product details and a purchase link.
[0255] Step 8:
[0256] Users can review the displayed information and, if interested, tap the provided link to access the product purchase page.
[0257] (Example 1)
[0258] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0259] A system is needed that can quickly and intuitively provide users viewing video information on communication terminals with detailed information about objects they are interested in. Conventional systems require users to manually search for information when they want to know more about items in the video they are watching, lacking convenience and speed. Furthermore, there were challenges in optimizing marketing using user selection history.
[0260] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0261] In this invention, the server includes means for the terminal to analyze an image and extract feature information of an object when the user taps an object while viewing video information; means for transmitting the feature information to a processing unit, which then searches for similar item information from an online storage area; and means for generating information about items related to the object and providing the user with purchase guidance. This allows the user to easily obtain information about related items from the video they are viewing and to proceed with the purchase process interactively.
[0262] "Video information" refers to video data that users can view using their communication devices.
[0263] An "image" is a static visual image extracted from any frame within a video.
[0264] An "object" is a specific object of interest that exists within the video information.
[0265] A "terminal" is a portable or stationary electronic device used by a user to view video information.
[0266] "Feature information" refers to data that represents the characteristics of an object, including its shape, color, and size.
[0267] A "processing device" is a computer system that receives characteristic information and searches for information on similar items.
[0268] "Online storage" refers to a database or data storage that is accessible over a network.
[0269] "Item information" refers to detailed data including the name, price, and purchase link of similar objects.
[0270] A "purchase guide" is an interface that provides users with information and links that enable them to purchase goods.
[0271] This invention constructs an interactive system that directly provides product information from video information, thereby stimulating viewers' purchasing intent. Users can view the video information on a portable or stationary electronic device. For example, suppose a user is watching a fashion-related video and becomes interested in a particular bag. The user can then tap the portion of the image showing that bag on their device.
[0272] The device, upon receiving a tap from the user, uses a generated AI model to recognize objects in the tapped image. The AI model extracts feature information from objects within the image and analyzes attributes such as shape, color, and size. This identifies objects of interest to the user, and that information is sent to the server.
[0273] The server inputs the feature information received from the terminal into the processing unit to access online storage. The processing unit applies an algorithm to search for similar items and finds item information. This item information includes the product name, price, and purchase link. The server returns the formatted information to the user's terminal, allowing the user to view the item details and make a purchase.
[0274] As an example of a prompt, you could instruct the generative AI model to "explain a system that, when a user taps on an item they are interested in within a video, extracts the features of that item and provides information on similar products."
[0275] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0276] Step 1:
[0277] The user views video information using a communication device. They tap on the part of the image that contains an object of interest. The input is the coordinates of the user's tap, and the output is image data of the tapped location. This action initiates the detection of the target object.
[0278] Step 2:
[0279] The device receives image data from the tapped location and uses a generative AI model to recognize the object. The input is the tapped image data, and the output is the object's feature information. The AI model extracts features such as the object's shape, color, and size based on the image data. In this step, the device collects specific attribute information of the detected object.
[0280] Step 3:
[0281] The terminal transmits feature information to the processing unit. This processing unit is connected to a server and performs initial data processing. The input is the feature information of an object, and the output is formatted data for database searching. This process formats the feature information into a format that can be processed by the server.
[0282] Step 4:
[0283] The server compares the formatted feature information with the database in the online storage area for retrieval. The input is the formatted feature data, and the output is the similar item information. The server uses an algorithm to identify the items that match the contents of the database and extracts the necessary information.
[0284] Step 5:
[0285] The server receives the item information and formats it into a form that is easy for the user to understand. The input is the raw data of the similar item, and the output is the formatted information including the product name, price, and purchase link. Through this information formatting, the user can efficiently obtain the necessary information.
[0286] Step 6:
[0287] The server transmits the formatted item information to the terminal. The terminal presents the received information to the user and displays the details of the object tapped in the video information. The input is the formatted item information, and the output is the user's visual interface. Based on this information, the user can access the online store and purchase the items they are interested in.
[0288] (Application Example 1)
[0289] Next, Application Example 1 will be described. In the following description, the data processing device 12 is referred to as the "server", and the smart glasses 214 are referred to as the "terminal".
[0290] During the viewing of video information, there is a problem that the viewer lacks a means to immediately grasp detailed information about a specific object and link it to a purchasing activity, so the viewer's purchasing desire cannot be immediately satisfied and the shopping experience is limited.
[0291] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0292] In this invention, the server includes means for an information processing device to analyze information and extract information about a specific object within video information when a user taps on that object while watching the video information; means for transmitting the information about the object to the information processing device, which then searches for similar product information from online information sources; and means for generating information about products related to the object and providing the user with purchase-inducing information. This makes it possible to instantly obtain product information about an object of interest in the video being watched and to quickly engage in purchasing activities.
[0293] "User is watching video information" refers to a situation where a user is using their device to play a video in real time and review its content.
[0294] "A subject within specific video information" refers to an object or person that the user is interested in within the image displayed at a specific point in time during video playback.
[0295] An "information processing device" is a general term for hardware and software that have the function of receiving data, analyzing it, and providing appropriate information to the user.
[0296] "Extracting information about a target" refers to the process of collecting detailed data about a specific target, such as its shape, color, and characteristics.
[0297] "Online information sources" refer to databases and web services accessible via the internet, and serve as a source of product information.
[0298] "Searching for similar product information" is the act of finding data on products with similar characteristics to the extracted target information from online information sources.
[0299] "Providing purchase-inducing information" means showing users details about products they are interested in and presenting links and pricing information to facilitate the purchase process.
[0300] "Operates as an application installed on a smartphone or tablet device" means that it is a program developed to run on a mobile device.
[0301] This invention is a system that, when a user watches a video using a mobile device such as a smartphone or tablet, instantly obtains product information related to a specific object if the user becomes interested in that object, thereby promoting purchasing activity.
[0302] When a user taps on a specific object in a video they are watching, the device begins processing the information. The device is equipped with an AI model that recognizes the object within the tapped video and extracts its features. AI modeling software such as TensorFlow Lite is used for this process. The extracted feature information is then sent from the device to a server via a network protocol.
[0303] Based on the received feature information, the server accesses a cloud database (e.g., Firebase) to search for similar product information from online sources. This uses REST API communication technology. When similar product information is found, the server composes the information and sends it back to the user's device. The user can then view the product name, price, and purchase link through the device's interface and proceed with the purchase.
[0304] For example, if a user sees a bag in a scene from a drama and becomes interested in it, they can use this feature to check information about the bag on the spot and easily access the purchase page.
[0305] Example of a prompt:
[0306] "How can we easily recognize specific items from videos a user is watching and retrieve information about them? How does an AI app work to extract item features and provide purchase links for similar products?"
[0307] The process of a specific operation in Application Example 1 will be described using Fig. 12.
[0308] Step 1:
[0309] While the user is watching a video, if the user taps on an object of interest on the screen, the input is the tap position of the user. The terminal detects this position and captures the frame of the tapped video information. The output is the target video frame.
[0310] Step 2:
[0311] The terminal inputs the video frame captured by the terminal into the AI model to identify and recognize the object. The input is the video frame, and the output is the feature information of the recognized object. TensorFlow Lite is used as the AI model, and through its operations, the shape, color, size, etc. of specific objects in the video are extracted as features.
[0312] Step 3:
[0313] The terminal transmits the feature information extracted by the terminal to the server via the REST API. The input is the feature information, and the output is the communication success log to the server. Through this communication, the data transfer on the server side is confirmed.
[0314] Step 4:
[0315] Based on the feature information received by the server, the server accesses the cloud database to search for similar product information. The input is the feature information, and the output is the dataset of similar products. The server queries an online information source (e.g., Firebase) to collect information such as the characteristics and prices of similar products.
[0316] Step 5:
[0317] The server searches for product information and sends it back to the user's terminal. The input is product information, and the output is a notification that data transmission to the user's terminal is complete. The server formats the data and returns it for a smooth user experience.
[0318] Step 6:
[0319] The received product information is displayed on the user's device interface. Input is product information sent from the server, and output is a visual display. Users can view product details, pricing, and purchase links, and proceed directly with the purchase.
[0320] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0321] This invention constructs an interactive system that provides product information directly from video content while taking into account the viewer's emotions. Users can view the video content using mobile devices or tablet devices. These devices are equipped with an emotion engine that further enhances the user's viewing experience.
[0322] When a user taps on a product or item they are interested in within a video, an AI model running on the device recognizes the subject within that video frame and extracts its feature data. Additionally, an emotion engine analyzes the user's facial expressions and actions to recognize their emotions while watching. This resulting emotion data, along with the subject's feature data, is then sent to the server.
[0323] Based on the received feature and sentiment data, the server searches an online database to find information on similar products that best match the user's emotions. Specifically, it incorporates sentiment data into the analysis to customize product suggestions based on emotions. The server organizes the acquired product information and generates an information package that includes product names, prices, and purchase links optimized based on emotions.
[0324] For example, if a user is watching a travel-related video and makes a happy face while tapping their favorite suitcase, the emotion engine will sense that happiness. The device will identify the suitcase at the tapped location and extract its features. The server will then perform a search based on this data, retrieving information about suitcases that particularly resonate with the excitement of travel, and providing this information to the user.
[0325] This system not only allows users to get direct inspiration from videos and instantly obtain detailed product information, but also enables emotionally-driven product selection, providing a more personalized shopping experience.
[0326] The following describes the processing flow.
[0327] Step 1:
[0328] The user watches a video and taps on products or items that interest them on the screen. The device records the tap position and obtains the video frame at that point.
[0329] Step 2:
[0330] The device analyzes the acquired video frames using an AI model to identify the subject at the tapped location. This process extracts characteristic data of the subject.
[0331] Step 3:
[0332] The device uses an emotion engine to assess the user's current emotions. Facial recognition technology analyzes the type and intensity of emotions from the user's facial expressions and voice tone.
[0333] Step 4:
[0334] The device transmits subject characteristic data and user emotion data to the server. The transmitted data includes subject identification information and the user's emotional state.
[0335] Step 5:
[0336] The server searches online databases based on the received data to find information on similar products related to the subject. During this process, it optimizes product recommendations by taking into account the user's emotional data.
[0337] Step 6:
[0338] The server organizes product information and generates an information package that reflects sentiment-based customization. This information package includes a detailed product description, images, price, and a purchase link.
[0339] Step 7:
[0340] The server sends the generated information package to the user's terminal.
[0341] Step 8:
[0342] The device displays the received product information to the user. The user can check the product details and easily purchase items they are interested in through the purchase link.
[0343] (Example 2)
[0344] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0345] Traditional video viewing presents challenges such as users being unable to quickly obtain detailed information about specific objects within a video while watching, and being unable to receive personalized product recommendations based on their emotions. Furthermore, there is a lack of effective methods for providing information that takes user emotions into consideration.
[0346] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0347] In this invention, the server includes means for the device to analyze video data and obtain attribute data of an object when the user selects an object in the video data during the viewing experience; means for transmitting the attribute data of the object and the user's emotion data, and for the data processing device to retrieve related information from an external storage device; and means for generating suggested content related to the object and presenting a transaction link to the user. As a result, the user can instantly obtain information about products that have caught their interest while watching a video and receive personalized product suggestions that match their emotions.
[0348] A "user" refers to an individual who uses this system to view digital content and operate specific functions.
[0349] "Viewing experience" refers to the process by which a user views video content through a digital device.
[0350] "Video data" refers to viewable video information that is stored or distributed in digital format.
[0351] "Object" refers to an object in the video data that the user has focused their attention on and which is the subject of identification or analysis.
[0352] "Device" refers to a hardware device used by the user, specifically equipment for analyzing video and processing data.
[0353] "Attribute data" refers to information about the features and characteristics associated with an object, and is obtained through image analysis.
[0354] "Emotional data" refers to information about a user's emotional state and is collected through sentiment analysis.
[0355] A "data processing device" refers to a computer system used to process received data and retrieve or generate information.
[0356] "External storage device" refers to a database or storage system that a data processing device accesses to retrieve information.
[0357] "Proposed content" refers to information about products and services offered to the user, and is customized based on the user's choices and feelings.
[0358] A "transaction link" refers to a web link that a user uses to purchase or use a suggested product or service.
[0359] This invention is built as a system that takes into account the viewer's emotions while they are watching video content and provides product information interactively. Users watch video content using a mobile device or tablet and select items that interest them. This system has the following hardware and software configuration.
[0360] The device is equipped with a high-resolution camera and touchscreen, and runs an emotion engine and machine learning models for object recognition. For example, OpenCV is used for the emotion engine, and YOLO or TensorFlow are used for the machine learning models. This allows the device to analyze the user's facial expressions in real time, recognize objects in videos the moment they are tapped, and acquire attribute data for those objects.
[0361] When a user taps a specific object, the device sends the object's attribute data and the user's sentiment data to a server. The server processes the received data and searches for information stored in external storage devices. The server generates information about related products and services based on the user's sentiment and sends it to the device as a suggestion.
[0362] For example, if a user is enjoying watching a travel video and taps on a suitcase in the video, the emotion engine will determine the user's emotion as "happy." The device retrieves the suitcase's attribute data and sends it to the server. The server uses this data to search its database for travel suitcase information and generates suggestions, including a purchase link. The user can then purchase the suitcase through the information displayed on the screen.
[0363] A concrete example of a prompt message would be: "Explain how to provide related product information when the user taps on a product with a happy expression."
[0364] In this way, this system instantly provides personalized product information based on the objects that the user shows interest in while watching videos, realizing a personalized shopping experience that responds to the user's emotions.
[0365] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0366] Step 1:
[0367] Users view video content on mobile devices or tablets. The input is video data selected by the user. This data is played on the device, and the user's facial expressions are captured by the camera during viewing. The output provides a video viewing experience and generates image data of the user's facial expressions.
[0368] Step 2:
[0369] The device analyzes the user's facial expression data using an emotion engine. Real-time captured image data of facial expressions is used as input. Here, OpenCV is used to detect facial features and process the data to determine emotional states (e.g., happiness, surprise). The output is emotional data representing the user's emotions.
[0370] Step 3:
[0371] The user taps on an object of interest within the video. The input consists of the user's tap location data and the current video frame data. The device uses machine learning models such as YOLO or TensorFlow to recognize the subject within the video frame and extract feature data. The output is attribute data of the tapped object.
[0372] Step 4:
[0373] The terminal sends attribute data of the acquired object and user sentiment data to the server. As input, the attribute data and sentiment data are packaged together. This data is sent to the server using a secure communication protocol. As output, the data is successfully delivered to the server.
[0374] Step 5:
[0375] The server processes data based on the received data. It receives attribute data and sentiment data as input. The server analyzes this data and searches for relevant information from external storage. Specifically, it executes database queries to retrieve product information that matches the specified criteria. Suggestions are generated as output.
[0376] Step 6:
[0377] The server sends the generated suggestions to the terminal. The input is package data containing the suggested product information and related links. The output is the package data sent to the terminal.
[0378] Step 7:
[0379] The terminal displays the proposed content to the user. Using package data received from the server as input, it generates a user interface that presents the proposed content to the user. The user can view product details and transaction links on the screen. The output is a visual display of information to support the user's decision-making process.
[0380] (Application Example 2)
[0381] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0382] In today's world, there is a lack of flexible, interactive systems that can directly provide product information based on viewers' emotions. This makes it difficult for viewers to instantly acquire information about related products based on the inspiration they gain from videos. Furthermore, the inability to suggest products that take viewers' emotions into account makes it difficult to provide personalized purchasing experiences. Technologies that can solve these problems are needed.
[0383] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0384] In this invention, the server includes means for the device to analyze a scene and extract information about an object when the user selects an object by action while watching a video; means for transmitting emotion data along with the object information to the server, and for the server to search for similar product information from an online information database; and means for generating product information related to the object based on the emotion data and providing the user with a purchase path. As a result, the user receives appropriate product suggestions according to their emotions, enabling a personalized purchasing experience.
[0385] A "user" is the entity that watches a video, selects objects of interest through actions, and interacts with the system.
[0386] "Video footage" is a media format that provides a series of images and is a means of conveying information visually and aurally.
[0387] An "object" is something that is visually represented within a video and that the user shows interest in as a potential selection.
[0388] "Selection by action" refers to an interaction a user performs with a specific object while watching a video, such as tapping the screen.
[0389] "Device" refers to a digital device capable of playing and interacting with video footage, and includes hardware and software for processing and transmitting data.
[0390] A "scene" refers to a specific moment or state within a video, and is the unit in which a user interacts with a particular object.
[0391] "Analysis" refers to the process by which a device extracts information from scenes within video footage and performs processing to identify the characteristics of objects.
[0392] "Information" refers to identifiable data about an object, including content that describes its features and attributes.
[0393] "Extraction" is the process of taking out the features necessary for identifying objects from video footage and organizing them as data.
[0394] "Emotional data" refers to data that expresses a user's emotional state using numbers or codes, and accurately reflects the user's emotions during interaction.
[0395] A "server" is a computer system that processes data via a network and performs specific functions using data received from external devices.
[0396] An "online information collection" is a collection of databases and information sources that exist on the internet, and is the target for which a server searches for product information.
[0397] "Similar product information" refers to data obtained by searching for highly relevant products based on the characteristics of the object selected by the user.
[0398] "Generation" refers to the operation of creating a new data package based on the information obtained and providing it to the user in a useful format.
[0399] A "purchase channel" is a means of directly supporting a user's purchase by providing the necessary information and links to buy a product.
[0400] This invention is an interactive system that allows users to directly interact with objects they are interested in while watching video footage, and provides personalized product information accordingly.
[0401] Users view video footage using digital devices such as smartphones and tablets. During viewing, users interact with specific objects, causing the device to analyze the video scene and acquire information about those objects along with the user's emotional data. This analysis utilizes an AI model that analyzes the user's facial expressions captured by a camera. Specific software such as OpenCV and Dlib can be used.
[0402] The device sends analyzed object information and emotion data to a server. The server receives this data and searches its database for similar product information. The server runs on a platform like Node.js, and its product information collection and customization process includes data processing to consider the user's emotions. As a result, product information that best matches the user's emotions is generated and sent back to the device.
[0403] As a concrete example, suppose a user watches a travel-related video and selects a specific suitcase with a joyful expression. In this case, the system recognizes the suitcase and, based on the user's emotional data, provides information about suitcases that will further enhance their enjoyment of the trip.
[0404] An example of a prompt used in this system is: "The user is tapping an accessory displayed during the movie. The user's current emotion has been detected as 'enjoyment.' Please suggest product information for similar accessories related to enjoyment."
[0405] In this way, users can leverage the immediate inspiration they gain from the videos they are watching and enjoy a customized purchasing experience based on their emotions.
[0406] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0407] Step 1:
[0408] A user watches a video and selects a specific object of interest. This is done by tapping the screen of a digital device. The input is the coordinate data of the user's tap position. The output is the data necessary to identify the selected object.
[0409] Step 2:
[0410] The system analyzes the tap locations received by the device and uses an AI model to identify the corresponding object in the video scene. This object detection algorithm is applied using TensorFlow or PyTorch. Input consists of tap location coordinates and video frames, while output is feature data of the identified object.
[0411] Step 3:
[0412] The terminal camera detects the user's emotions, and real-time facial expression analysis is performed using OpenCV and Dlib. The input is camera footage, and the output is user emotion data. This allows the user's emotional state during viewing to be captured as numerical data.
[0413] Step 4:
[0414] The device sends characteristic data of the identified object and user emotion data to the server. This data is sent to the server via an HTTP request in JSON format. The input is data on the identified object and emotion state, and the output is the transmission of data to the server.
[0415] Step 5:
[0416] The server searches for similar product information from an online database based on the received object's feature data and sentiment data. This process uses Node.js and queries the database. The input is feature data and sentiment data, and the output is similar product information.
[0417] Step 6:
[0418] The server constructs an information package based on the user's emotions, using the acquired product information. The package includes the product name, price, and purchase link. The input is similar product information, and the output is the constructed information package. This process enables product recommendations tailored to the user's emotions.
[0419] Step 7:
[0420] The device displays an information package received from the server to the user. React Native is used to visually present the information on the smartphone screen. The input is an information package from the server, and the output is the display of product information to the user. This allows the user to instantly view detailed information about products they are interested in.
[0421] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0422] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet Search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0423] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0424] [Third Embodiment]
[0425] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0426] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0427] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0428] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0429] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0430] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0431] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0432] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0433] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0434] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0435] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0436] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0437] This invention constructs an interactive system that directly provides product information from video content, thereby stimulating viewers' purchasing intent. Users can view video content using mobile devices, tablets, etc. If a user becomes interested in a particular item on this device, they tap the screen, and the system collects information about that item.
[0438] An AI model running on the device recognizes subjects within the video frame, analyzes tapped locations, and extracts features related to those subjects. This feature data is sent from the device to a server, where it performs a more detailed product information search. The server accesses an online database and applies an algorithm to find similar products. Information on similar products includes product names, prices, and purchase links, and the server formats this information and sends it back to the user's device.
[0439] As a concrete example, consider a scenario where a user is watching a fashion-related video and becomes interested in a particular bag. When the user taps on the bag, the device identifies the bag within the video frame and extracts its features. Based on the feature data sent to the server, the server retrieves information on similar bags from shopping websites. This information, along with a link to the bag's details page, is sent to the user, allowing them to browse and purchase the product.
[0440] Furthermore, this system records user behavior data, which is used to improve AI models and target advertisements. This system allows viewers to gain direct inspiration from videos and instantly access detailed product information, dramatically improving the shopping experience.
[0441] The following describes the processing flow.
[0442] Step 1:
[0443] Users watch videos and tap on products or items that interest them on the screen. The tap locations are then recorded by the device.
[0444] Step 2:
[0445] The device captures a video frame at the tapped location and uses an AI model to recognize the subject within that frame. Feature data of the subject is then extracted.
[0446] Step 3:
[0447] The device sends the extracted feature data to the server. The data includes information such as color, shape, and size.
[0448] Step 4:
[0449] Based on the received data, the server uses an AI algorithm to search online databases and identify information on similar products.
[0450] Step 5:
[0451] The server organizes the retrieved data on similar products and generates an information package that includes product name, price, image, and purchase link.
[0452] Step 6:
[0453] The server sends the generated information package to the user's terminal.
[0454] Step 7:
[0455] The terminal displays the received product information on the user interface and provides the user with product details and a purchase link.
[0456] Step 8:
[0457] Users can review the displayed information and, if interested, tap the provided link to access the product purchase page.
[0458] (Example 1)
[0459] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0460] A system is needed that can quickly and intuitively provide users viewing video information on communication terminals with detailed information about objects they are interested in. Conventional systems require users to manually search for information when they want to know more about items in the video they are watching, lacking convenience and speed. Furthermore, there were challenges in optimizing marketing using user selection history.
[0461] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0462] In this invention, the server includes means for the terminal to analyze an image and extract feature information of an object when the user taps an object while viewing video information; means for transmitting the feature information to a processing unit, which then searches for similar item information from an online storage area; and means for generating information about items related to the object and providing the user with purchase guidance. This allows the user to easily obtain information about related items from the video they are viewing and to proceed with the purchase process interactively.
[0463] "Video information" refers to video data that users can view using their communication devices.
[0464] An "image" is a static visual image extracted from any frame within a video.
[0465] An "object" is a specific object of interest that exists within the video information.
[0466] A "terminal" is a portable or stationary electronic device used by a user to view video information.
[0467] "Feature information" refers to data that represents the characteristics of an object, including its shape, color, and size.
[0468] A "processing device" is a computer system that receives characteristic information and searches for information on similar items.
[0469] "Online storage" refers to a database or data storage that is accessible over a network.
[0470] "Item information" refers to detailed data including the name, price, and purchase link of similar objects.
[0471] A "purchase guide" is an interface that provides users with information and links that enable them to purchase goods.
[0472] This invention constructs an interactive system that directly provides product information from video information, thereby stimulating viewers' purchasing intent. Users can view the video information on a portable or stationary electronic device. For example, suppose a user is watching a fashion-related video and becomes interested in a particular bag. The user can then tap the portion of the image showing that bag on their device.
[0473] The device, upon receiving a tap from the user, uses a generated AI model to recognize objects in the tapped image. The AI model extracts feature information from objects within the image and analyzes attributes such as shape, color, and size. This identifies objects of interest to the user, and that information is sent to the server.
[0474] The server inputs the feature information received from the terminal into the processing unit to access online storage. The processing unit applies an algorithm to search for similar items and finds item information. This item information includes the product name, price, and purchase link. The server returns the formatted information to the user's terminal, allowing the user to view the item details and make a purchase.
[0475] As an example of a prompt, you could instruct the generative AI model to "explain a system that, when a user taps on an item they are interested in within a video, extracts the features of that item and provides information on similar products."
[0476] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0477] Step 1:
[0478] The user views video information using a communication device. They tap on the part of the image that contains an object of interest. The input is the coordinates of the user's tap, and the output is image data of the tapped location. This action initiates the detection of the target object.
[0479] Step 2:
[0480] The device receives image data from the tapped location and uses a generative AI model to recognize the object. The input is the tapped image data, and the output is the object's feature information. The AI model extracts features such as the object's shape, color, and size based on the image data. In this step, the device collects specific attribute information of the detected object.
[0481] Step 3:
[0482] The terminal transmits feature information to the processing unit. This processing unit is connected to a server and performs initial data processing. The input is the feature information of an object, and the output is formatted data for database searching. This process formats the feature information into a format that can be processed by the server.
[0483] Step 4:
[0484] The server performs searches by comparing formatted feature information with a database in online storage. The input is formatted feature data, and the output is information about similar items. The server uses an algorithm to identify items that match the database contents and extracts the necessary information.
[0485] Step 5:
[0486] The server receives item information and formats it into a user-friendly format. The input is raw data of similar items, and the output is formatted information including product name, price, and purchase link. This information formatting allows users to efficiently obtain the necessary information.
[0487] Step 6:
[0488] The server sends formatted item information to the terminal. The terminal presents the received information to the user and displays details of the object tapped within the video information. The input is formatted item information, and the output is the user's visual interface. Based on this information, the user can access the online store and purchase items that interest them.
[0489] (Application Example 1)
[0490] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0491] During video viewing, there is a lack of means for viewers to immediately grasp detailed information about specific products or services and translate that information into purchasing decisions. As a result, viewers' purchasing intent cannot be immediately satisfied, and the shopping experience is limited.
[0492] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0493] In this invention, the server includes means for an information processing device to analyze information and extract information about a specific object within video information when a user taps on that object while watching the video information; means for transmitting the information about the object to the information processing device, which then searches for similar product information from online information sources; and means for generating information about products related to the object and providing the user with purchase-inducing information. This makes it possible to instantly obtain product information about an object of interest in the video being watched and to quickly engage in purchasing activities.
[0494] "User is watching video information" refers to a situation where a user is using their device to play a video in real time and review its content.
[0495] "A subject within specific video information" refers to an object or person that the user is interested in within the image displayed at a specific point in time during video playback.
[0496] An "information processing device" is a general term for hardware and software that have the function of receiving data, analyzing it, and providing appropriate information to the user.
[0497] "Extracting information about a target" refers to the process of collecting detailed data about a specific target, such as its shape, color, and characteristics.
[0498] "Online information sources" refer to databases and web services accessible via the internet, and serve as a source of product information.
[0499] "Searching for similar product information" is the act of finding data on products with similar characteristics to the extracted target information from online information sources.
[0500] "Providing purchase-inducing information" means showing users details about products they are interested in and presenting links and pricing information to facilitate the purchase process.
[0501] "Operates as an application installed on a smartphone or tablet device" means that it is a program developed to run on a mobile device.
[0502] This invention is a system that, when a user watches a video using a mobile device such as a smartphone or tablet, instantly obtains product information related to a specific object if the user becomes interested in that object, thereby promoting purchasing activity.
[0503] When a user taps on a specific object in a video they are watching, the device begins processing the information. The device is equipped with an AI model that recognizes the object within the tapped video and extracts its features. AI modeling software such as TensorFlow Lite is used for this process. The extracted feature information is then sent from the device to a server via a network protocol.
[0504] Based on the received feature information, the server accesses a cloud database (e.g., Firebase) to search for similar product information from online sources. This uses REST API communication technology. When similar product information is found, the server composes the information and sends it back to the user's device. The user can then view the product name, price, and purchase link through the device's interface and proceed with the purchase.
[0505] For example, if a user sees a bag in a scene from a drama and becomes interested in it, they can use this feature to check information about the bag on the spot and easily access the purchase page.
[0506] Example of a prompt:
[0507] "How can we easily recognize specific items from videos a user is watching and retrieve information about them? How does an AI app work to extract item features and provide purchase links for similar products?"
[0508] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0509] Step 1:
[0510] While watching a video, the user taps on an object of interest on the screen. The input is the user's tap location, and the device detects this location and captures the frame of the tapped video information. The output is the video frame of the object in question.
[0511] Step 2:
[0512] The system inputs video frames captured by the device into an AI model to identify and recognize objects. The input is video frames, and the output is feature information of the recognized objects. TensorFlow Lite is used as the AI model, and its calculations extract features such as the shape, color, and size of specific objects within the video.
[0513] Step 3:
[0514] The terminal extracts feature information and sends it to the server via a REST API. The input is feature information, and the output is a log of successful communication to the server. This communication allows the server to confirm the transfer of data.
[0515] Step 4:
[0516] Based on the feature information received by the server, it accesses a cloud database to search for similar product information. The input is feature information, and the output is a dataset of similar products. The server queries online information sources (e.g., Firebase) to collect information such as the characteristics and prices of similar products.
[0517] Step 5:
[0518] The server searches for product information and sends it back to the user's terminal. The input is product information, and the output is a notification that data transmission to the user's terminal is complete. The server formats the data and returns it for a smooth user experience.
[0519] Step 6:
[0520] The received product information is displayed on the user's device interface. Input is product information sent from the server, and output is a visual display. Users can view product details, pricing, and purchase links, and proceed directly with the purchase.
[0521] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0522] This invention constructs an interactive system that provides product information directly from video content while taking into account the viewer's emotions. Users can view the video content using mobile devices or tablet devices. These devices are equipped with an emotion engine that further enhances the user's viewing experience.
[0523] When a user taps on a product or item they are interested in within a video, an AI model running on the device recognizes the subject within that video frame and extracts its feature data. Additionally, an emotion engine analyzes the user's facial expressions and actions to recognize their emotions while watching. This resulting emotion data, along with the subject's feature data, is then sent to the server.
[0524] Based on the received feature and sentiment data, the server searches an online database to find information on similar products that best match the user's emotions. Specifically, it incorporates sentiment data into the analysis to customize product suggestions based on emotions. The server organizes the acquired product information and generates an information package that includes product names, prices, and purchase links optimized based on emotions.
[0525] For example, if a user is watching a travel-related video and makes a happy face while tapping their favorite suitcase, the emotion engine will sense that happiness. The device will identify the suitcase at the tapped location and extract its features. The server will then perform a search based on this data, retrieving information about suitcases that particularly resonate with the excitement of travel, and providing this information to the user.
[0526] This system not only allows users to get direct inspiration from videos and instantly obtain detailed product information, but also enables emotionally-driven product selection, providing a more personalized shopping experience.
[0527] The following describes the processing flow.
[0528] Step 1:
[0529] The user watches a video and taps on products or items that interest them on the screen. The device records the tap position and obtains the video frame at that point.
[0530] Step 2:
[0531] The device analyzes the acquired video frames using an AI model to identify the subject at the tapped location. This process extracts characteristic data of the subject.
[0532] Step 3:
[0533] The device uses an emotion engine to assess the user's current emotions. Facial recognition technology analyzes the type and intensity of emotions from the user's facial expressions and voice tone.
[0534] Step 4:
[0535] The device transmits subject characteristic data and user emotion data to the server. The transmitted data includes subject identification information and the user's emotional state.
[0536] Step 5:
[0537] The server searches online databases based on the received data to find information on similar products related to the subject. During this process, it optimizes product recommendations by taking into account the user's emotional data.
[0538] Step 6:
[0539] The server organizes product information and generates an information package that reflects sentiment-based customization. This information package includes a detailed product description, images, price, and a purchase link.
[0540] Step 7:
[0541] The server sends the generated information package to the user's terminal.
[0542] Step 8:
[0543] The device displays the received product information to the user. The user can check the product details and easily purchase items they are interested in through the purchase link.
[0544] (Example 2)
[0545] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0546] Traditional video viewing presents challenges such as users being unable to quickly obtain detailed information about specific objects within a video while watching, and being unable to receive personalized product recommendations based on their emotions. Furthermore, there is a lack of effective methods for providing information that takes user emotions into consideration.
[0547] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0548] In this invention, the server includes means for the device to analyze video data and obtain attribute data of an object when the user selects an object in the video data during the viewing experience; means for transmitting the attribute data of the object and the user's emotion data, and for the data processing device to retrieve related information from an external storage device; and means for generating suggested content related to the object and presenting a transaction link to the user. As a result, the user can instantly obtain information about products that have caught their interest while watching a video and receive personalized product suggestions that match their emotions.
[0549] A "user" refers to an individual who uses this system to view digital content and operate specific functions.
[0550] "Viewing experience" refers to the process by which a user views video content through a digital device.
[0551] "Video data" refers to viewable video information that is stored or distributed in digital format.
[0552] "Object" refers to an object in the video data that the user has focused their attention on and which is the subject of identification or analysis.
[0553] "Device" refers to a hardware device used by the user, specifically equipment for analyzing video and processing data.
[0554] "Attribute data" refers to information about the features and characteristics associated with an object, and is obtained through image analysis.
[0555] "Emotional data" refers to information about a user's emotional state and is collected through sentiment analysis.
[0556] A "data processing device" refers to a computer system used to process received data and retrieve or generate information.
[0557] "External storage device" refers to a database or storage system that a data processing device accesses to retrieve information.
[0558] "Proposed content" refers to information about products and services offered to the user, and is customized based on the user's choices and feelings.
[0559] A "transaction link" refers to a web link that a user uses to purchase or use a suggested product or service.
[0560] This invention is built as a system that takes into account the viewer's emotions while they are watching video content and provides product information interactively. Users watch video content using a mobile device or tablet and select items that interest them. This system has the following hardware and software configuration.
[0561] The device is equipped with a high-resolution camera and touchscreen, and runs an emotion engine and machine learning models for object recognition. For example, OpenCV is used for the emotion engine, and YOLO or TensorFlow are used for the machine learning models. This allows the device to analyze the user's facial expressions in real time, recognize objects in videos the moment they are tapped, and acquire attribute data for those objects.
[0562] When a user taps a specific object, the device sends the object's attribute data and the user's sentiment data to a server. The server processes the received data and searches for information stored in external storage devices. The server generates information about related products and services based on the user's sentiment and sends it to the device as a suggestion.
[0563] For example, if a user is enjoying watching a travel video and taps on a suitcase in the video, the emotion engine will determine the user's emotion as "happy." The device retrieves the suitcase's attribute data and sends it to the server. The server uses this data to search its database for travel suitcase information and generates suggestions, including a purchase link. The user can then purchase the suitcase through the information displayed on the screen.
[0564] A concrete example of a prompt message would be: "Explain how to provide related product information when the user taps on a product with a happy expression."
[0565] In this way, this system instantly provides personalized product information based on the objects that the user shows interest in while watching videos, realizing a personalized shopping experience that responds to the user's emotions.
[0566] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0567] Step 1:
[0568] Users view video content on mobile devices or tablets. The input is video data selected by the user. This data is played on the device, and the user's facial expressions are captured by the camera during viewing. The output provides a video viewing experience and generates image data of the user's facial expressions.
[0569] Step 2:
[0570] The device analyzes the user's facial expression data using an emotion engine. Real-time captured image data of facial expressions is used as input. Here, OpenCV is used to detect facial features and process the data to determine emotional states (e.g., happiness, surprise). The output is emotional data representing the user's emotions.
[0571] Step 3:
[0572] The user taps on an object of interest within the video. The input consists of the user's tap location data and the current video frame data. The device uses machine learning models such as YOLO or TensorFlow to recognize the subject within the video frame and extract feature data. The output is attribute data of the tapped object.
[0573] Step 4:
[0574] The terminal sends attribute data of the acquired object and user sentiment data to the server. As input, the attribute data and sentiment data are packaged together. This data is sent to the server using a secure communication protocol. As output, the data is successfully delivered to the server.
[0575] Step 5:
[0576] The server processes data based on the received data. It receives attribute data and sentiment data as input. The server analyzes this data and searches for relevant information from external storage. Specifically, it executes database queries to retrieve product information that matches the specified criteria. Suggestions are generated as output.
[0577] Step 6:
[0578] The server sends the generated suggestions to the terminal. The input is package data containing the suggested product information and related links. The output is the package data sent to the terminal.
[0579] Step 7:
[0580] The terminal displays the proposed content to the user. Using package data received from the server as input, it generates a user interface that presents the proposed content to the user. The user can view product details and transaction links on the screen. The output is a visual display of information to support the user's decision-making process.
[0581] (Application Example 2)
[0582] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0583] In today's world, there is a lack of flexible, interactive systems that can directly provide product information based on viewers' emotions. This makes it difficult for viewers to instantly acquire information about related products based on the inspiration they gain from videos. Furthermore, the inability to suggest products that take viewers' emotions into account makes it difficult to provide personalized purchasing experiences. Technologies that can solve these problems are needed.
[0584] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0585] In this invention, the server includes means for the device to analyze a scene and extract information about an object when the user selects an object by action while watching a video; means for transmitting emotion data along with the object information to the server, and for the server to search for similar product information from an online information database; and means for generating product information related to the object based on the emotion data and providing the user with a purchase path. As a result, the user receives appropriate product suggestions according to their emotions, enabling a personalized purchasing experience.
[0586] A "user" is the entity that watches a video, selects objects of interest through actions, and interacts with the system.
[0587] "Video footage" is a media format that provides a series of images and is a means of conveying information visually and aurally.
[0588] An "object" is something that is visually represented within a video and that the user shows interest in as a potential selection.
[0589] "Selection by action" refers to an interaction a user performs with a specific object while watching a video, such as tapping the screen.
[0590] "Device" refers to a digital device capable of playing and interacting with video footage, and includes hardware and software for processing and transmitting data.
[0591] A "scene" refers to a specific moment or state within a video, and is the unit in which a user interacts with a particular object.
[0592] "Analysis" refers to the process by which a device extracts information from scenes within video footage and performs processing to identify the characteristics of objects.
[0593] "Information" refers to identifiable data about an object, including content that describes its features and attributes.
[0594] "Extraction" is the process of taking out the features necessary for identifying objects from video footage and organizing them as data.
[0595] "Emotional data" refers to data that expresses a user's emotional state using numbers or codes, and accurately reflects the user's emotions during interaction.
[0596] A "server" is a computer system that processes data via a network and performs specific functions using data received from external devices.
[0597] An "online information collection" is a collection of databases and information sources that exist on the internet, and is the target for which a server searches for product information.
[0598] "Similar product information" refers to data obtained by searching for highly relevant products based on the characteristics of the object selected by the user.
[0599] "Generation" refers to the operation of creating a new data package based on the information obtained and providing it to the user in a useful format.
[0600] A "purchase channel" is a means of directly supporting a user's purchase by providing the necessary information and links to buy a product.
[0601] This invention is an interactive system that allows users to directly interact with objects they are interested in while watching video footage, and provides personalized product information accordingly.
[0602] Users view video footage using digital devices such as smartphones and tablets. During viewing, users interact with specific objects, causing the device to analyze the video scene and acquire information about those objects along with the user's emotional data. This analysis utilizes an AI model that analyzes the user's facial expressions captured by a camera. Specific software such as OpenCV and Dlib can be used.
[0603] The device sends analyzed object information and emotion data to a server. The server receives this data and searches its database for similar product information. The server runs on a platform like Node.js, and its product information collection and customization process includes data processing to consider the user's emotions. As a result, product information that best matches the user's emotions is generated and sent back to the device.
[0604] As a concrete example, suppose a user watches a travel-related video and selects a specific suitcase with a joyful expression. In this case, the system recognizes the suitcase and, based on the user's emotional data, provides information about suitcases that will further enhance their enjoyment of the trip.
[0605] An example of a prompt used in this system is: "The user is tapping an accessory displayed during the movie. The user's current emotion has been detected as 'enjoyment.' Please suggest product information for similar accessories related to enjoyment."
[0606] In this way, users can leverage the immediate inspiration they gain from the videos they are watching and enjoy a customized purchasing experience based on their emotions.
[0607] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0608] Step 1:
[0609] A user watches a video and selects a specific object of interest. This is done by tapping the screen of a digital device. The input is the coordinate data of the user's tap position. The output is the data necessary to identify the selected object.
[0610] Step 2:
[0611] The system analyzes the tap locations received by the device and uses an AI model to identify the corresponding object in the video scene. This object detection algorithm is applied using TensorFlow or PyTorch. Input consists of tap location coordinates and video frames, while output is feature data of the identified object.
[0612] Step 3:
[0613] The terminal camera detects the user's emotions, and real-time facial expression analysis is performed using OpenCV and Dlib. The input is camera footage, and the output is user emotion data. This allows the user's emotional state during viewing to be captured as numerical data.
[0614] Step 4:
[0615] The device sends characteristic data of the identified object and user emotion data to the server. This data is sent to the server via an HTTP request in JSON format. The input is data on the identified object and emotion state, and the output is the transmission of data to the server.
[0616] Step 5:
[0617] The server searches for similar product information from an online database based on the received object's feature data and sentiment data. This process uses Node.js and queries the database. The input is feature data and sentiment data, and the output is similar product information.
[0618] Step 6:
[0619] The server constructs an information package based on the user's emotions, using the acquired product information. The package includes the product name, price, and purchase link. The input is similar product information, and the output is the constructed information package. This process enables product recommendations tailored to the user's emotions.
[0620] Step 7:
[0621] The device displays an information package received from the server to the user. React Native is used to visually present the information on the smartphone screen. The input is an information package from the server, and the output is the display of product information to the user. This allows the user to instantly view detailed information about products they are interested in.
[0622] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0623] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet Search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0624] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[0625] [Fourth Embodiment]
[0626] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0627] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0628] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0629] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0630] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0631] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0632] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0633] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0634] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0635] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0636] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0637] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0638] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0639] This invention constructs an interactive system that directly provides product information from video content, thereby stimulating viewers' purchasing intent. Users can view video content using mobile devices, tablets, etc. If a user becomes interested in a particular item on this device, they tap the screen, and the system collects information about that item.
[0640] An AI model running on the device recognizes subjects within the video frame, analyzes tapped locations, and extracts features related to those subjects. This feature data is sent from the device to a server, where it performs a more detailed product information search. The server accesses an online database and applies an algorithm to find similar products. Information on similar products includes product names, prices, and purchase links, and the server formats this information and sends it back to the user's device.
[0641] As a concrete example, consider a scenario where a user is watching a fashion-related video and becomes interested in a particular bag. When the user taps on the bag, the device identifies the bag within the video frame and extracts its features. Based on the feature data sent to the server, the server retrieves information on similar bags from shopping websites. This information, along with a link to the bag's details page, is sent to the user, allowing them to browse and purchase the product.
[0642] Furthermore, this system records user behavior data, which is used to improve AI models and target advertisements. This system allows viewers to gain direct inspiration from videos and instantly access detailed product information, dramatically improving the shopping experience.
[0643] The following describes the processing flow.
[0644] Step 1:
[0645] Users watch videos and tap on products or items that interest them on the screen. The tap locations are then recorded by the device.
[0646] Step 2:
[0647] The device captures a video frame at the tapped location and uses an AI model to recognize the subject within that frame. Feature data of the subject is then extracted.
[0648] Step 3:
[0649] The device sends the extracted feature data to the server. The data includes information such as color, shape, and size.
[0650] Step 4:
[0651] Based on the received data, the server uses an AI algorithm to search online databases and identify information on similar products.
[0652] Step 5:
[0653] The server organizes the retrieved data on similar products and generates an information package that includes product name, price, image, and purchase link.
[0654] Step 6:
[0655] The server sends the generated information package to the user's terminal.
[0656] Step 7:
[0657] The terminal displays the received product information on the user interface and provides the user with product details and a purchase link.
[0658] Step 8:
[0659] Users can review the displayed information and, if interested, tap the provided link to access the product purchase page.
[0660] (Example 1)
[0661] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0662] A system is needed that can quickly and intuitively provide users viewing video information on communication terminals with detailed information about objects they are interested in. Conventional systems require users to manually search for information when they want to know more about items in the video they are watching, lacking convenience and speed. Furthermore, there were challenges in optimizing marketing using user selection history.
[0663] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0664] In this invention, the server includes means for the terminal to analyze an image and extract feature information of an object when the user taps an object while viewing video information; means for transmitting the feature information to a processing unit, which then searches for similar item information from an online storage area; and means for generating information about items related to the object and providing the user with purchase guidance. This allows the user to easily obtain information about related items from the video they are viewing and to proceed with the purchase process interactively.
[0665] "Video information" refers to video data that users can view using their communication devices.
[0666] An "image" is a static visual image extracted from any frame within a video.
[0667] An "object" is a specific object of interest that exists within the video information.
[0668] A "terminal" is a portable or stationary electronic device used by a user to view video information.
[0669] "Feature information" refers to data that represents the characteristics of an object, including its shape, color, and size.
[0670] A "processing device" is a computer system that receives characteristic information and searches for information on similar items.
[0671] "Online storage" refers to a database or data storage that is accessible over a network.
[0672] "Item information" refers to detailed data including the name, price, and purchase link of similar objects.
[0673] A "purchase guide" is an interface that provides users with information and links that enable them to purchase goods.
[0674] This invention constructs an interactive system that directly provides product information from video information, thereby stimulating viewers' purchasing intent. Users can view the video information on a portable or stationary electronic device. For example, suppose a user is watching a fashion-related video and becomes interested in a particular bag. The user can then tap the portion of the image showing that bag on their device.
[0675] The device, upon receiving a tap from the user, uses a generated AI model to recognize objects in the tapped image. The AI model extracts feature information from objects within the image and analyzes attributes such as shape, color, and size. This identifies objects of interest to the user, and that information is sent to the server.
[0676] The server inputs the feature information received from the terminal into the processing unit to access online storage. The processing unit applies an algorithm to search for similar items and finds item information. This item information includes the product name, price, and purchase link. The server returns the formatted information to the user's terminal, allowing the user to view the item details and make a purchase.
[0677] As an example of a prompt, you could instruct the generative AI model to "explain a system that, when a user taps on an item they are interested in within a video, extracts the features of that item and provides information on similar products."
[0678] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0679] Step 1:
[0680] The user views video information using a communication device. They tap on the part of the image that contains an object of interest. The input is the coordinates of the user's tap, and the output is image data of the tapped location. This action initiates the detection of the target object.
[0681] Step 2:
[0682] The device receives image data from the tapped location and uses a generative AI model to recognize the object. The input is the tapped image data, and the output is the object's feature information. The AI model extracts features such as the object's shape, color, and size based on the image data. In this step, the device collects specific attribute information of the detected object.
[0683] Step 3:
[0684] The terminal transmits feature information to the processing unit. This processing unit is connected to a server and performs initial data processing. The input is the feature information of an object, and the output is formatted data for database searching. This process formats the feature information into a format that can be processed by the server.
[0685] Step 4:
[0686] The server performs searches by comparing formatted feature information with a database in online storage. The input is formatted feature data, and the output is information about similar items. The server uses an algorithm to identify items that match the database contents and extracts the necessary information.
[0687] Step 5:
[0688] The server receives item information and formats it into a user-friendly format. The input is raw data of similar items, and the output is formatted information including product name, price, and purchase link. This information formatting allows users to efficiently obtain the necessary information.
[0689] Step 6:
[0690] The server sends formatted item information to the terminal. The terminal presents the received information to the user and displays details of the object tapped within the video information. The input is formatted item information, and the output is the user's visual interface. Based on this information, the user can access the online store and purchase items that interest them.
[0691] (Application Example 1)
[0692] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0693] During video viewing, there is a lack of means for viewers to immediately grasp detailed information about specific products or services and translate that information into purchasing decisions. As a result, viewers' purchasing intent cannot be immediately satisfied, and the shopping experience is limited.
[0694] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0695] In this invention, the server includes means for an information processing device to analyze information and extract information about a specific object within video information when a user taps on that object while watching the video information; means for transmitting the information about the object to the information processing device, which then searches for similar product information from online information sources; and means for generating information about products related to the object and providing the user with purchase-inducing information. This makes it possible to instantly obtain product information about an object of interest in the video being watched and to quickly engage in purchasing activities.
[0696] "User is watching video information" refers to a situation where a user is using their device to play a video in real time and review its content.
[0697] "A subject within specific video information" refers to an object or person that the user is interested in within the image displayed at a specific point in time during video playback.
[0698] An "information processing device" is a general term for hardware and software that have the function of receiving data, analyzing it, and providing appropriate information to the user.
[0699] "Extracting information about a target" refers to the process of collecting detailed data about a specific target, such as its shape, color, and characteristics.
[0700] "Online information sources" refer to databases and web services accessible via the internet, and serve as a source of product information.
[0701] "Searching for similar product information" is the act of finding data on products with similar characteristics to the extracted target information from online information sources.
[0702] "Providing purchase-inducing information" means showing users details about products they are interested in and presenting links and pricing information to facilitate the purchase process.
[0703] "Operates as an application installed on a smartphone or tablet device" means that it is a program developed to run on a mobile device.
[0704] This invention is a system that, when a user watches a video using a mobile device such as a smartphone or tablet, instantly obtains product information related to a specific object if the user becomes interested in that object, thereby promoting purchasing activity.
[0705] When a user taps on a specific object in a video they are watching, the device begins processing the information. The device is equipped with an AI model that recognizes the object within the tapped video and extracts its features. AI modeling software such as TensorFlow Lite is used for this process. The extracted feature information is then sent from the device to a server via a network protocol.
[0706] Based on the received feature information, the server accesses a cloud database (e.g., Firebase) to search for similar product information from online sources. This uses REST API communication technology. When similar product information is found, the server composes the information and sends it back to the user's device. The user can then view the product name, price, and purchase link through the device's interface and proceed with the purchase.
[0707] For example, if a user sees a bag in a scene from a drama and becomes interested in it, they can use this feature to check information about the bag on the spot and easily access the purchase page.
[0708] Example of a prompt:
[0709] "How can we easily recognize specific items from videos a user is watching and retrieve information about them? How does an AI app work to extract item features and provide purchase links for similar products?"
[0710] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0711] Step 1:
[0712] While watching a video, the user taps on an object of interest on the screen. The input is the user's tap location, and the device detects this location and captures the frame of the tapped video information. The output is the video frame of the object in question.
[0713] Step 2:
[0714] The system inputs video frames captured by the device into an AI model to identify and recognize objects. The input is video frames, and the output is feature information of the recognized objects. TensorFlow Lite is used as the AI model, and its calculations extract features such as the shape, color, and size of specific objects within the video.
[0715] Step 3:
[0716] The terminal extracts feature information and sends it to the server via a REST API. The input is feature information, and the output is a log of successful communication to the server. This communication allows the server to confirm the transfer of data.
[0717] Step 4:
[0718] Based on the feature information received by the server, it accesses a cloud database to search for similar product information. The input is feature information, and the output is a dataset of similar products. The server queries online information sources (e.g., Firebase) to collect information such as the characteristics and prices of similar products.
[0719] Step 5:
[0720] The server searches for product information and sends it back to the user's terminal. The input is product information, and the output is a notification that data transmission to the user's terminal is complete. The server formats the data and returns it for a smooth user experience.
[0721] Step 6:
[0722] The received product information is displayed on the user's device interface. Input is product information sent from the server, and output is a visual display. Users can view product details, pricing, and purchase links, and proceed directly with the purchase.
[0723] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0724] This invention constructs an interactive system that provides product information directly from video content while taking into account the viewer's emotions. Users can view the video content using mobile devices or tablet devices. These devices are equipped with an emotion engine that further enhances the user's viewing experience.
[0725] When a user taps on a product or item they are interested in within a video, an AI model running on the device recognizes the subject within that video frame and extracts its feature data. Additionally, an emotion engine analyzes the user's facial expressions and actions to recognize their emotions while watching. This resulting emotion data, along with the subject's feature data, is then sent to the server.
[0726] Based on the received feature and sentiment data, the server searches an online database to find information on similar products that best match the user's emotions. Specifically, it incorporates sentiment data into the analysis to customize product suggestions based on emotions. The server organizes the acquired product information and generates an information package that includes product names, prices, and purchase links optimized based on emotions.
[0727] For example, if a user is watching a travel-related video and makes a happy face while tapping their favorite suitcase, the emotion engine will sense that happiness. The device will identify the suitcase at the tapped location and extract its features. The server will then perform a search based on this data, retrieving information about suitcases that particularly resonate with the excitement of travel, and providing this information to the user.
[0728] This system not only allows users to get direct inspiration from videos and instantly obtain detailed product information, but also enables emotionally-driven product selection, providing a more personalized shopping experience.
[0729] The following describes the processing flow.
[0730] Step 1:
[0731] The user watches a video and taps on products or items that interest them on the screen. The device records the tap position and obtains the video frame at that point.
[0732] Step 2:
[0733] The device analyzes the acquired video frames using an AI model to identify the subject at the tapped location. This process extracts characteristic data of the subject.
[0734] Step 3:
[0735] The device uses an emotion engine to assess the user's current emotions. Facial recognition technology analyzes the type and intensity of emotions from the user's facial expressions and voice tone.
[0736] Step 4:
[0737] The device transmits subject characteristic data and user emotion data to the server. The transmitted data includes subject identification information and the user's emotional state.
[0738] Step 5:
[0739] The server searches online databases based on the received data to find information on similar products related to the subject. During this process, it optimizes product recommendations by taking into account the user's emotional data.
[0740] Step 6:
[0741] The server organizes product information and generates an information package that reflects sentiment-based customization. This information package includes a detailed product description, images, price, and a purchase link.
[0742] Step 7:
[0743] The server sends the generated information package to the user's terminal.
[0744] Step 8:
[0745] The device displays the received product information to the user. The user can check the product details and easily purchase items they are interested in through the purchase link.
[0746] (Example 2)
[0747] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0748] Traditional video viewing presents challenges such as users being unable to quickly obtain detailed information about specific objects within a video while watching, and being unable to receive personalized product recommendations based on their emotions. Furthermore, there is a lack of effective methods for providing information that takes user emotions into consideration.
[0749] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0750] In this invention, the server includes means for the device to analyze video data and obtain attribute data of an object when the user selects an object in the video data during the viewing experience; means for transmitting the attribute data of the object and the user's emotion data, and for the data processing device to retrieve related information from an external storage device; and means for generating suggested content related to the object and presenting a transaction link to the user. As a result, the user can instantly obtain information about products that have caught their interest while watching a video and receive personalized product suggestions that match their emotions.
[0751] A "user" refers to an individual who uses this system to view digital content and operate specific functions.
[0752] "Viewing experience" refers to the process by which a user views video content through a digital device.
[0753] "Video data" refers to viewable video information that is stored or distributed in digital format.
[0754] "Object" refers to an object in the video data that the user has focused their attention on and which is the subject of identification or analysis.
[0755] "Device" refers to a hardware device used by the user, specifically equipment for analyzing video and processing data.
[0756] "Attribute data" refers to information about the features and characteristics associated with an object, and is obtained through image analysis.
[0757] "Emotional data" refers to information about a user's emotional state and is collected through sentiment analysis.
[0758] A "data processing device" refers to a computer system used to process received data and retrieve or generate information.
[0759] "External storage device" refers to a database or storage system that a data processing device accesses to retrieve information.
[0760] "Proposed content" refers to information about products and services offered to the user, and is customized based on the user's choices and feelings.
[0761] A "transaction link" refers to a web link that a user uses to purchase or use a suggested product or service.
[0762] This invention is built as a system that takes into account the viewer's emotions while they are watching video content and provides product information interactively. Users watch video content using a mobile device or tablet and select items that interest them. This system has the following hardware and software configuration.
[0763] The device is equipped with a high-resolution camera and touchscreen, and runs an emotion engine and machine learning models for object recognition. For example, OpenCV is used for the emotion engine, and YOLO or TensorFlow are used for the machine learning models. This allows the device to analyze the user's facial expressions in real time, recognize objects in videos the moment they are tapped, and acquire attribute data for those objects.
[0764] When a user taps a specific object, the device sends the object's attribute data and the user's sentiment data to a server. The server processes the received data and searches for information stored in external storage devices. The server generates information about related products and services based on the user's sentiment and sends it to the device as a suggestion.
[0765] For example, if a user is enjoying watching a travel video and taps on a suitcase in the video, the emotion engine will determine the user's emotion as "happy." The device retrieves the suitcase's attribute data and sends it to the server. The server uses this data to search its database for travel suitcase information and generates suggestions, including a purchase link. The user can then purchase the suitcase through the information displayed on the screen.
[0766] A concrete example of a prompt message would be: "Explain how to provide related product information when the user taps on a product with a happy expression."
[0767] In this way, this system instantly provides personalized product information based on the objects that the user shows interest in while watching videos, realizing a personalized shopping experience that responds to the user's emotions.
[0768] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0769] Step 1:
[0770] Users view video content on mobile devices or tablets. The input is video data selected by the user. This data is played on the device, and the user's facial expressions are captured by the camera during viewing. The output provides a video viewing experience and generates image data of the user's facial expressions.
[0771] Step 2:
[0772] The device analyzes the user's facial expression data using an emotion engine. Real-time captured image data of facial expressions is used as input. Here, OpenCV is used to detect facial features and process the data to determine emotional states (e.g., happiness, surprise). The output is emotional data representing the user's emotions.
[0773] Step 3:
[0774] The user taps on an object of interest within the video. The input consists of the user's tap location data and the current video frame data. The device uses machine learning models such as YOLO or TensorFlow to recognize the subject within the video frame and extract feature data. The output is attribute data of the tapped object.
[0775] Step 4:
[0776] The terminal sends attribute data of the acquired object and user sentiment data to the server. As input, the attribute data and sentiment data are packaged together. This data is sent to the server using a secure communication protocol. As output, the data is successfully delivered to the server.
[0777] Step 5:
[0778] The server processes data based on the received data. It receives attribute data and sentiment data as input. The server analyzes this data and searches for relevant information from external storage. Specifically, it executes database queries to retrieve product information that matches the specified criteria. Suggestions are generated as output.
[0779] Step 6:
[0780] The server sends the generated suggestions to the terminal. The input is package data containing the suggested product information and related links. The output is the package data sent to the terminal.
[0781] Step 7:
[0782] The terminal displays the proposed content to the user. Using package data received from the server as input, it generates a user interface that presents the proposed content to the user. The user can view product details and transaction links on the screen. The output is a visual display of information to support the user's decision-making process.
[0783] (Application Example 2)
[0784] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0785] In today's world, there is a lack of flexible, interactive systems that can directly provide product information based on viewers' emotions. This makes it difficult for viewers to instantly acquire information about related products based on the inspiration they gain from videos. Furthermore, the inability to suggest products that take viewers' emotions into account makes it difficult to provide personalized purchasing experiences. Technologies that can solve these problems are needed.
[0786] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0787] In this invention, the server includes means for the device to analyze a scene and extract information about an object when the user selects an object by action while watching a video; means for transmitting emotion data along with the object information to the server, and for the server to search for similar product information from an online information database; and means for generating product information related to the object based on the emotion data and providing the user with a purchase path. As a result, the user receives appropriate product suggestions according to their emotions, enabling a personalized purchasing experience.
[0788] A "user" is the entity that watches a video, selects objects of interest through actions, and interacts with the system.
[0789] "Video footage" is a media format that provides a series of images and is a means of conveying information visually and aurally.
[0790] An "object" is something that is visually represented within a video and that the user shows interest in as a potential selection.
[0791] "Selection by action" refers to an interaction a user performs with a specific object while watching a video, such as tapping the screen.
[0792] "Device" refers to a digital device capable of playing and interacting with video footage, and includes hardware and software for processing and transmitting data.
[0793] A "scene" refers to a specific moment or state within a video, and is the unit in which a user interacts with a particular object.
[0794] "Analysis" refers to the process by which a device extracts information from scenes within video footage and performs processing to identify the characteristics of objects.
[0795] "Information" refers to identifiable data about an object, including content that describes its features and attributes.
[0796] "Extraction" is the process of taking out the features necessary for identifying objects from video footage and organizing them as data.
[0797] "Emotional data" refers to data that expresses a user's emotional state using numbers or codes, and accurately reflects the user's emotions during interaction.
[0798] A "server" is a computer system that processes data via a network and performs specific functions using data received from external devices.
[0799] An "online information collection" is a collection of databases and information sources that exist on the internet, and is the target for which a server searches for product information.
[0800] "Similar product information" refers to data obtained by searching for highly relevant products based on the characteristics of the object selected by the user.
[0801] "Generation" refers to the operation of creating a new data package based on the information obtained and providing it to the user in a useful format.
[0802] A "purchase channel" is a means of directly supporting a user's purchase by providing the necessary information and links to buy a product.
[0803] This invention is an interactive system that allows users to directly interact with objects they are interested in while watching video footage, and provides personalized product information accordingly.
[0804] Users view video footage using digital devices such as smartphones and tablets. During viewing, users interact with specific objects, causing the device to analyze the video scene and acquire information about those objects along with the user's emotional data. This analysis utilizes an AI model that analyzes the user's facial expressions captured by a camera. Specific software such as OpenCV and Dlib can be used.
[0805] The device sends analyzed object information and emotion data to a server. The server receives this data and searches its database for similar product information. The server runs on a platform like Node.js, and its product information collection and customization process includes data processing to consider the user's emotions. As a result, product information that best matches the user's emotions is generated and sent back to the device.
[0806] As a concrete example, suppose a user watches a travel-related video and selects a specific suitcase with a joyful expression. In this case, the system recognizes the suitcase and, based on the user's emotional data, provides information about suitcases that will further enhance their enjoyment of the trip.
[0807] An example of a prompt used in this system is: "The user is tapping an accessory displayed during the movie. The user's current emotion has been detected as 'enjoyment.' Please suggest product information for similar accessories related to enjoyment."
[0808] In this way, users can leverage the immediate inspiration they gain from the videos they are watching and enjoy a customized purchasing experience based on their emotions.
[0809] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0810] Step 1:
[0811] A user watches a video and selects a specific object of interest. This is done by tapping the screen of a digital device. The input is the coordinate data of the user's tap position. The output is the data necessary to identify the selected object.
[0812] Step 2:
[0813] The system analyzes the tap locations received by the device and uses an AI model to identify the corresponding object in the video scene. This object detection algorithm is applied using TensorFlow or PyTorch. Input consists of tap location coordinates and video frames, while output is feature data of the identified object.
[0814] Step 3:
[0815] The terminal camera detects the user's emotions, and real-time facial expression analysis is performed using OpenCV and Dlib. The input is camera footage, and the output is user emotion data. This allows the user's emotional state during viewing to be captured as numerical data.
[0816] Step 4:
[0817] The device sends characteristic data of the identified object and user emotion data to the server. This data is sent to the server via an HTTP request in JSON format. The input is data on the identified object and emotion state, and the output is the transmission of data to the server.
[0818] Step 5:
[0819] The server searches for similar product information from an online database based on the received object's feature data and sentiment data. This process uses Node.js and queries the database. The input is feature data and sentiment data, and the output is similar product information.
[0820] Step 6:
[0821] The server constructs an information package based on the user's emotions, using the acquired product information. The package includes the product name, price, and purchase link. The input is similar product information, and the output is the constructed information package. This process enables product recommendations tailored to the user's emotions.
[0822] Step 7:
[0823] The device displays an information package received from the server to the user. React Native is used to visually present the information on the smartphone screen. The input is an information package from the server, and the output is the display of product information to the user. This allows the user to instantly view detailed information about products they are interested in.
[0824] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0825] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet Search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0826] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[0827] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0828] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0829] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0830] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0831] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0832] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0833] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0834] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[0835] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[0836] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0837] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0838] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0839] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0840] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0841] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0842] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0843] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0844] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted as being incorporated by reference.
[0845] The following is further disclosed regarding the embodiments described above.
[0846] (Claim 1)
[0847] A means by which, while a user is watching video content, the device analyzes a specific video frame and extracts information about the subject when the user taps on the subject within that frame.
[0848] A means for transmitting information about the subject to a server, and for the server to search for similar product information from an online database,
[0849] A means for generating information on products related to the subject and providing a purchase link to the user,
[0850] A system that includes this.
[0851] (Claim 2)
[0852] The system according to claim 1, wherein the server collects user behavior data and optimizes advertising targeting based on the generated information.
[0853] (Claim 3)
[0854] The system according to claim 1, wherein the device uses an AI model to recognize an object within a frame and performs a process to extract feature data.
[0855] "Example 1"
[0856] (Claim 1)
[0857] A means by which, while a user is viewing video information, the device analyzes the image and extracts characteristic information of the object by tapping on a specific object within that image.
[0858] A means for transmitting the aforementioned characteristic information to a processing device, and for the processing device to search for similar item information from an online storage area,
[0859] A means for generating information on items related to the aforementioned object and providing purchase information to the user,
[0860] A means of recording the user's selection history and contributing to algorithm improvement,
[0861] A system that includes this.
[0862] (Claim 2)
[0863] The system according to claim 1, wherein the processing device collects user behavior data and optimizes advertising based on the generated information.
[0864] (Claim 3)
[0865] The system according to claim 1, wherein the terminal uses an AI algorithm to recognize objects in an image and performs a process to extract feature data.
[0866] "Application Example 1"
[0867] (Claim 1)
[0868] A means by which, when a user taps on a specific object within video information while watching video information, an information processing device analyzes that information and extracts information about that object.
[0869] A means for transmitting the aforementioned target information to an information processing device, and for the information processing device to search for similar product information from an online information source,
[0870] A means for generating information on products related to the aforementioned target and providing users with information to induce them to purchase,
[0871] A means of operating as an application installed on a smartphone or tablet device,
[0872] A system that includes this.
[0873] (Claim 2)
[0874] The system according to claim 1, wherein the information processing device collects user behavior data and optimizes advertising targeting based on the generated information.
[0875] (Claim 3)
[0876] The system according to claim 1, wherein the information processing device uses an AI program to recognize an object in the information and performs a process to extract feature information.
[0877] "Example 2 of combining an emotion engine"
[0878] (Claim 1)
[0879] A means by which a user selects an object within video data during the viewing experience, and the device analyzes that video data to obtain attribute data of the object,
[0880] The means includes transmitting attribute data of the object and user sentiment data, and a data processing device retrieving related information from an external storage device.
[0881] A means for generating proposals related to the aforementioned object and presenting a transaction link to the user,
[0882] A system that includes this.
[0883] (Claim 2)
[0884] The system according to claim 1, wherein the data processing device collects the user's emotional state and optimizes recommendation information based on the generated suggestion content.
[0885] (Claim 3)
[0886] The system according to claim 1, wherein the device uses a machine learning model to recognize objects in video data and performs a process to acquire attribute data.
[0887] "Application example 2 when combining with an emotional engine"
[0888] (Claim 1)
[0889] A means by which, while a user is watching a video, the device analyzes the scene and extracts information about an object by selecting that object through an action,
[0890] A means for transmitting emotional data along with information about the object to a server, and for the server to search for similar product information from an online information collection,
[0891] A means for generating product information related to the object based on the aforementioned emotional data and providing the user with a purchase route,
[0892] A system that includes this.
[0893] (Claim 2)
[0894] The system according to claim 1, wherein the server collects the user's emotional state and personalizes the targeting of advertisements based on the generated information.
[0895] (Claim 3)
[0896] The system according to claim 1, wherein the device uses an artificial intelligence model to recognize objects in a scene and performs a process to extract characteristic data. [Explanation of Symbols]
[0897] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. A means by which, when a user taps on a subject within a specific video frame while watching video content, the device analyzes that frame and extracts information about the subject. A means for transmitting information about the subject to a server, and for the server to search for similar product information from an online database, A means for generating information on products related to the subject and providing a purchase link to the user, A system that includes this.
2. The system according to claim 1, wherein the server collects user behavior data and optimizes advertising targeting based on the generated information.
3. The system according to claim 1, wherein the device uses an AI model to recognize an object within a frame and performs a process to extract feature data.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A