System

The system addresses the challenge of consecutive video viewing by using generative AI to analyze user requests and automatically generate playlists, enhancing the viewing experience through seamless video playback.

JP2026016225APending Publication Date: 2026-02-03SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024117315
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-22
Publication Date
2026-02-03

AI Technical Summary

Technical Problem

Conventional video viewing services lack mechanisms for automatically providing videos that meet user requests, making it difficult for users to watch multiple videos consecutively based on specific content or context, leading to a cumbersome viewing experience.

Method used

A system that includes means for user request input, analysis, video search, selection, and automatic playback, utilizing generative AI to understand and analyze user requests and generate playlists of related videos for seamless viewing.

Benefits of technology

Enables users to smoothly watch videos according to a specific theme or context in succession, improving viewer convenience and efficiency by eliminating the need for manual video selection and search.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026016225000001_ABST
    Figure 2026016225000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system comprising: means for allowing a user to input a request; means for analyzing the input request; means for searching for a related moving image based on a result of the analysis; means for selecting the searched moving image and generating a playlist; and means for reproducing the generated playlist.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] In conventional video viewing services, it is difficult for users to watch multiple videos consecutively based on specific content or context. Users must manually search for related videos, which often leads to a cumbersome viewing experience. Furthermore, there is a lack of mechanisms for automatically providing videos that meet user requests. Under these circumstances, there is a need for a system that improves the user viewing experience and enables convenient and efficient consecutive viewing of related videos. [Means for solving the problem]

[0005] The present invention provides a system that searches for and selects videos based on user requests and automatically plays those videos. Specifically, the system includes a means for a user to input a request, a means for analyzing the input request, a means for searching for related videos based on the analysis results, a means for selecting the searched videos and generating a playlist, and a means for playing the generated playlist. This allows users to smoothly watch videos according to a specific theme or context in succession, significantly improving viewer convenience.

[0006] "User" refers to a person who operates the system of the present invention and inputs requests.

[0007] "Request input means" refers to a mechanism that provides an interface for users to input the content of the video they wish to watch.

[0008] A "request" refers to a request input by a user indicating the content and conditions of a video they wish to watch.

[0009] "Analysis means" refers to a mechanism that analyzes input requests and extracts relevant keywords and contexts.

[0010] "Generative AI" refers to a program that uses artificial intelligence technology to understand and analyze the content of a request.

[0011] "Search method" refers to the mechanism for finding relevant videos based on the analyzed request.

[0012] "Selection method" refers to the mechanism for selecting videos that match the user's request from the search results.

[0013] A "playlist" refers to a list of selected videos for sequential playback.

[0014] "Playback means" refers to a mechanism that automatically plays videos one after another based on the generated playlist.

[0015] "Video" refers to digital content that includes video and audio.

[0016] The term "system" refers to the entire set of devices and programs that execute a series of processes including request input means, analysis means, search means, selection means, and playback means. [Brief explanation of the drawings]

[0017] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12]FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0018] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0019] First, the terms used in the following description will be explained.

[0020] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0021] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0022] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0023] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0024] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0025] [First embodiment]

[0026] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0027] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0028] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0029] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0030] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0031] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0032] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0033] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0034] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0035] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0036] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0037] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0038] overview

[0039] This invention relates to a system that, when a user requests to watch multiple videos based on specific content, automatically searches for videos that meet the request and plays them sequentially. The following describes in detail an embodiment of the invention.

[0040] User request input

[0041] User

[0042] A user starts a video viewing application on any device (e.g., a smartphone or computer). The application interface includes a Generate AI button for inputting video requests.

[0043] The user clicks the Generate AI button and enters a request in text or voice into the input screen that appears, such as "Show me 50 videos of multiple cats playing."

[0044] Sending and parsing requests

[0045] Terminal

[0046] The terminal sends the request data entered by the user to the server, where the request is converted into an appropriate data format (for example, JSON format).

[0047] The request data sent to the server is analyzed by the AI ​​generator. Specific keywords and contexts are extracted using the analysis method. For example, elements such as "cats," "playing," "video," and "50 items" are identified.

[0048] Video search and selection

[0049] server

[0050] The server searches a video database or an external video providing service based on the analyzed elements.

[0051] From the search results, 50 videos that match the user's request are selected, and the information of the selected videos (title, URL, thumbnail, etc.) is listed.

[0052] Playlist generation and submission

[0053] server

[0054] Based on the selected videos, the server generates a playlist that takes into account the playback order. The playlist includes information such as the URL, title, and thumbnail of each video.

[0055] The playlist is sent to the device in an appropriate data format (e.g., JSON format).

[0056] Receive and autoplay playlists

[0057] Terminal

[0058] The terminal receives the playlist sent from the server.

[0059] Once the playlist has been parsed, the videos will begin playing automatically. Based on the playlist, the user can watch the requested videos in succession.

[0060] Specific examples

[0061] A specific example will be described in which a user requests "Show me 50 videos of multiple cats playing."

[0062] 1. Request input

[0063] The user types a request into the terminal as text and clicks the send button.

[0064] 2. Parsing the Request

[0065] The server receives the request and uses a generation AI to extract the keywords "cat," "playing," "video," and "50."

[0066] 3. Search and select videos

[0067] The server searches a video database based on these keywords and selects 50 videos that meet the criteria.

[0068] 4. Generate a playlist

[0069] The server generates a playlist of 50 videos that match the conditions and sends it to the device.

[0070] 5. Continuous video playback

[0071] The terminal plays the videos based on the received playlist, and the user watches the videos consecutively according to the requested content.

[0072] In this way, users can smoothly watch the content they want, greatly improving convenience. In addition, by using generative AI, it is possible to accurately understand and analyze the content of the request and provide the appropriate video.

[0073] The processing flow will be explained below.

[0074] Step 1: User request input

[0075] The user clicks the Generate AI button displayed on the device, which displays the input screen.

[0076] The user enters a request into the displayed input screen using text or voice, such as "Show me 50 videos of multiple cats playing."

[0077] After completing the input, the user presses the "Send" button to send the input contents.

[0078] Step 2: Submitting the request

[0079] The terminal acquires the request data input by the user.

[0080] The terminal converts the acquired request data into an appropriate data format such as JSON format.

[0081] The terminal transmits the converted request data to the server as an HTTP request.

[0082] Step 3: Receiving and Parsing the Request

[0083] The server receives the request data sent from the terminal.

[0084] The server passes the received request data to the generation AI for analysis.

[0085] The generative AI analyzes the request data and extracts important keywords and context (e.g., "cats," "playing," "video," "50 items").

[0086] The server receives the analysis results from the generation AI.

[0087] Step 4: Search for videos

[0088] Based on the analysis results, the server searches for videos using a video database or an external video search API.

[0089] The server uses the extracted keywords and conditions as a search query.

[0090] The server retrieves a list of related videos from the search results.

[0091] Step 5: Select a video

[0092] The server selects 50 videos from the search results that match the user's request.

[0093] The server lists the information of the selected videos (title, URL, thumbnail, etc.).

[0094] Step 6: Generate a playlist

[0095] The server generates a playlist based on the list of selected videos.

[0096] A playlist contains information such as the playback order, title, URL, and thumbnail of each video.

[0097] Step 7: Submit your playlist

[0098] The server converts the generated playlist into an appropriate data format, such as JSON.

[0099] The server transmits the converted playlist data to the terminal.

[0100] Step 8: Receive and play the playlist

[0101] The terminal receives the playlist data transmitted from the server.

[0102] The device analyzes the received playlist data and obtains the playback order and URL information for each video.

[0103] The device will play the first video in the playlist order, and when that video finishes playing, the next video will automatically play.

[0104] This system allows users to watch videos of their choice continuously without any hassle, greatly improving convenience and the viewing experience.

[0105] Example 1

[0106] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0107] Currently, if a user wants to watch specific content consecutively, they have to manually search for multiple videos and add them to a playlist individually. This makes the process cumbersome for users, making it difficult to quickly watch the content they want. Furthermore, it can be difficult to accurately understand a user's request and provide the appropriate video, which can reduce satisfaction.

[0108] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0109] In this invention, the server includes a means for a user to input a request, a terminal means for transmitting the request input data to the server, a means for analyzing the input request using a generative AI model, a means for searching for related videos from a database or an external video service based on the analysis results, a means for selecting the searched videos and generating a playlist, a means for transmitting the generated playlist to the terminal, and a means for analyzing the generated playlist on the terminal and automatically playing videos. This eliminates the need for complicated user operations and enables users to quickly watch the desired content. Furthermore, the use of the generative AI model makes it possible to accurately understand user requests and provide optimal videos.

[0110] "User" means an individual or corporation that uses the system to input requests and view content.

[0111] "Request input means" refers to an interface or device that allows a user to input a request by text or voice.

[0112] "Terminal means" refers to an electronic device such as a smartphone or computer that sends a user's request to the server.

[0113] A "server" is a computer system that receives requests, analyzes, searches, generates playlists, and transmits playlists.

[0114] A "generative AI model" is an artificial intelligence that analyzes requests entered by users and extracts keywords and context.

[0115] "Analysis means" refers to the process or algorithm that analyzes the input request using a generative AI model and extracts the necessary information.

[0116] "Video Database" means internal or external data storage used to search for relevant videos.

[0117] "Third-party video provider" means an external video hosting platform, such as YouTube.

[0118] "Search method" refers to the process or algorithm for searching for relevant videos from a video database or external video provider service based on the analysis results.

[0119] A "selection method" is a process or algorithm that selects retrieved videos that match the user's request.

[0120] A "playlist generator" is a process or algorithm that compiles selected videos into a playlist, taking into account the playback order.

[0121] "Playlist transmission means" refers to a process or algorithm that transmits the generated playlist to a terminal in an appropriate data format.

[0122] "Autoplay method" refers to the process or algorithm that analyzes the playlist generated by the device and plays videos continuously.

[0123] overview

[0124] This invention relates to a system that automatically searches for and sequentially plays videos that fit a user's request when the user requests to watch multiple videos based on specific content. Specifically, this describes the process flow and device configuration for analyzing the user's request, searching for, selecting, and creating a playlist of related videos, and automatically playing them.

[0125] User request input

[0126] User

[0127] A user launches a video viewing application using a device such as a smartphone or computer.

[0128] The application interface has a Generate AI button, and clicking this will display the input screen.

[0129] The user enters a request into the input screen using text or voice, such as "Show me 50 videos of multiple cats playing," and clicks the send button.

[0130] Sending and parsing requests

[0131] Terminal

[0132] The terminal converts the request data entered by the user into JSON format and sends it to the server, using the HTTP POST method to transfer the data.

[0133] server

[0134] The server uses a generative AI model (e.g., OpenAI's GPT-3) to analyze the received request data.

[0135] The generative AI model extracts necessary keywords and context from the request, such as "cats," "playing," "video," and "50 items."

[0136] Video search and selection

[0137] server

[0138] The server searches its internal video database or external video providers such as YouTube based on the keywords extracted by the generative AI model.

[0139] The server selects 50 videos from the search results that match the user's request. Specifically, it uses the YouTube Data API to search for videos containing tags such as "cat" and "playing" and extracts the appropriate videos.

[0140] List the information of the selected videos (title, URL, thumbnail, etc.).

[0141] Playlist generation and submission

[0142] server

[0143] The server generates a playlist based on the selected video list, taking into consideration the playback order.

[0144] The playlist contains information such as the URL, title, and thumbnail of each video.

[0145] The created playlist is formatted in JSON format and sent to the device as an HTTP response.

[0146] Receive and autoplay playlists

[0147] Terminal

[0148] The terminal receives the playlist sent from the server.

[0149] It analyzes the playlist and automatically starts playing videos based on the information of each video. The videos are played continuously, allowing users to watch the requested videos one after the other.

[0150] Specific examples

[0151] Below is what happens when a user requests "Show me 50 videos of multiple cats playing together."

[0152] 1. Request input

[0153] The user types the text "Show me 50 videos of multiple cats playing" into their device and clicks the send button.

[0154] 2. Parsing the Request

[0155] The server receives the request and uses a generative AI model (such as OpenAI's GPT-3) to extract the keywords "cat," "playing," "video," and "50."

[0156] 3. Search and select videos

[0157] The server searches a video database based on the keywords and selects 50 videos that match the criteria, for example, using the YouTube API.

[0158] 4. Generate a playlist

[0159] The server generates a playlist of 50 videos that match the conditions and sends it to the device.

[0160] 5. Continuous video playback

[0161] The terminal plays the videos based on the received playlist, and the user watches the videos consecutively according to the requested content.

[0162] Prompt Sentence Examples

[0163] "Find 50 videos of multiple cats playing together and create a playlist."

[0164] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0165] Step 1:

[0166] User input of request

[0167] A user launches a video viewing application on a device such as a smartphone or computer.

[0168] Click the Generate AI button on the application interface to display the input screen.

[0169] Users enter a request via text or voice, such as "Show me 50 videos of multiple cats playing," and click the send button.

[0170] Input: The request text or voice data entered by the user.

[0171] Output: The request data is sent to the terminal.

[0172] Step 2:

[0173] Send a request from the device

[0174] The terminal converts the request data entered by the user into JSON format.

[0175] The converted request data is sent to the server using the HTTP POST method.

[0176] Input: Request data from the user (text or voice).

[0177] Output: The request data converted to JSON format is sent to the server.

[0178] Step 3:

[0179] Server parsing of the request

[0180] The server passes the received request data to a generative AI model (e.g., OpenAI's GPT-3).

[0181] The generative AI model extracts keywords and context from the request, such as "cat," "playing," "video," and "50 items."

[0182] Input: Request data in JSON format.

[0183] Output: Keywords and context extracted by the generative AI model.

[0184] Step 4:

[0185] Server-based video search and selection

[0186] The server searches an internal video database or an external video provider (e.g., YouTube) based on the keywords extracted by the generative AI model.

[0187] The server selects 50 videos from the search results that match the user's request. For example, it uses the YouTube Data API to search for videos containing the tags "cat" and "playing."

[0188] Input: Keywords extracted by the generative AI model.

[0189] Output: Information about the selected video (title, URL, thumbnail, etc.).

[0190] Step 5:

[0191] Server-generated and transmitted playlists

[0192] The server generates a playlist based on the selected video list, taking into account the playback order.

[0193] The playlist contents are formatted in JSON format.

[0194] The server sends the generated playlist to the terminal as an HTTP response.

[0195] Input: Information about the selected video (title, URL, thumbnail, etc.).

[0196] Output: Playlist data in JSON format.

[0197] Step 6:

[0198] Device receives and autoplays playlist

[0199] The terminal receives the playlist sent from the server.

[0200] Analyze the playlist and get information about each video.

[0201] The device will automatically start playing videos based on the playlist, allowing users to watch the videos they want in sequence.

[0202] Input: Playlist data in JSON format.

[0203] Output: Automatic continuous playback of videos.

[0204] This allows users to view the content they desire efficiently. The accuracy and speed of server-dependent processing are the keys to improving the overall performance of the system.

[0205] (Application example 1)

[0206] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0207] Conventional video streaming systems have difficulty quickly and accurately selecting relevant videos based on specific user requests and playing them continuously. Furthermore, the technology for accepting voice requests, accurately converting that voice into text, and analyzing it using generative AI models has been immature. This has resulted in inconvenient services that cannot meet user needs.

[0208] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0209] In this invention, the server includes means for allowing a user to input a request, means for analyzing the input request, means for searching for related videos based on the analysis results, means for selecting the searched videos and generating a playlist, means for playing the generated playlist, means for converting the user's voice into text using a voice recognition means, and means for retrieving data from an online video database. This makes it possible to accurately select videos using a generative AI model based on the user's voice or text input request and play them continuously.

[0210] "User" refers to a person who uses the System to enter a request and wish to watch a video.

[0211] "Means for inputting requests" refers to an interface through which a user inputs video requests into the system in the form of voice or text.

[0212] "Means for analyzing requests" refers to the process of analyzing input requests using a generative AI model and extracting relevant keywords and context.

[0213] "Means for searching for related videos" refers to a process for searching for appropriate videos from a video database or external video providing service based on the analyzed keywords and context.

[0214] "Means for selecting and generating a playlist" refers to the process of selecting appropriate videos from the search results and organizing the videos into a playlist based on the playback order.

[0215] The "means for playing the generated playlist" refers to a process for playing videos continuously based on the generated playlist.

[0216] "Speech recognition means" refers to technology for converting user-input speech into text.

[0217] "Means for obtaining data from an online video database" refers to a process for obtaining necessary video data from an online video database via the Internet.

[0218] A "generative AI model" refers to an artificial intelligence algorithm that performs highly accurate analysis based on input information.

[0219] A "playlist" is a list of related videos arranged in a particular order.

[0220] overview

[0221] This invention is a system that automatically searches for and plays back videos that fit a user's request for viewing multiple videos based on specific content, improving user convenience and enabling users to smoothly view the content they desire.

[0222] User request input

[0223] A user uses a device of their choice (e.g., a smartphone or head-mounted display) to launch a video viewing application. The application interface has a voice recognition button for inputting video requests. The user clicks this button and inputs a request by voice, such as "Show me 50 videos of multiple cats playing."

[0224] Sending and parsing requests

[0225] The device sends the request data entered by the user to the server. The request is converted into text by a speech recognition means and then converted into an appropriate data format (e.g., JSON format). Specific keywords and context are extracted by an analysis means. For example, elements such as "cats," "playing," "video," and "50 items" are identified. This is where the generative AI model comes into play.

[0226] Video search and selection

[0227] The server searches online video databases and external video providers based on the analyzed elements. From the search results, it selects 50 videos that match the user's request. The information of the selected videos (title, URL, thumbnail, etc.) is then listed.

[0228] Playlist generation and transmission

[0229] The server generates a playlist based on the selected videos, taking into account the playback order. The playlist includes information such as the URL, title, and thumbnail of each video. The generated playlist is sent to the device in the appropriate data format.

[0230] Receive and autoplay playlists

[0231] The device receives the playlist sent from the server. Once the analysis of the playlist is complete, the video starts playing automatically. Based on the playlist, the user can watch the requested videos in succession.

[0232] Specific examples

[0233] Below is a specific example of a user requesting "Show me 50 videos of multiple cats playing." The user voice-types the request into the device and clicks the send button. The device converts the voice to text and uses a generative AI model to extract the keywords "cat," "playing," "video," and "50." The server searches a video database based on these keywords and selects 50 videos that match the criteria. The server then creates a playlist of the 50 videos that match the criteria and sends it to the device. The device plays the videos based on the received playlist, and the user watches the videos consecutively according to the request.

[0234] Prompt Sentence Examples

[0235] "Show me 50 videos of multiple cats playing together"

[0236] This allows users to easily input requests through voice, and the generative AI model performs accurate analysis and search, allowing them to watch the desired videos continuously.

[0237] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0238] Step 1:

[0239] The user launches a video viewing application and clicks the voice input button. The user then voices their request.

[0240] Input: User's voice request

[0241] Output: Audio data

[0242] Specific actions: Requests such as "Show me 50 videos of multiple cats playing" can be made using voice input.

[0243] Step 2:

[0244] The terminal uses a voice recognition means to convert the voice data into text data.

[0245] Input: Audio data

[0246] Output: Text data

[0247] Specific behavior: Uses a speech recognition library to convert speech into text. Example: "Show me 50 videos of multiple cats playing."

[0248] Step 3:

[0249] The device sends the converted text data to a generative AI model that analyzes the request.

[0250] Input: Text data

[0251] Output: Analysis results (keywords)

[0252] Specific operation: The generative AI model extracts keywords such as "cat," "playing," "video," and "50."

[0253] Step 4:

[0254] The server searches an online video database based on the analyzed keywords.

[0255] Input: Analysis results (keywords)

[0256] Output: Video candidates (title, URL, thumbnail information, etc.)

[0257] Specific operation: Using an API request, search the database for the terms "cat" and "playing" and retrieve 50 videos.

[0258] Step 5:

[0259] The server generates a playlist based on the acquired video candidates, taking into account the playback order.

[0260] Input: Video candidate

[0261] Output: Playlist (video URL, title, thumbnail information, etc.)

[0262] Specific operation: The information of the acquired videos is compiled into a list and compiled into a playlist based on the user's request.

[0263] Step 6:

[0264] The server transmits the generated playlist to the terminal.

[0265] Input: Playlist

[0266] Output: Playlist transmission data

[0267] Specific operation: Send the playlist to the device in an appropriate data format, such as JSON.

[0268] Step 7:

[0269] The device will analyze the received playlist and begin automatically playing the videos.

[0270] Input: Playlist transmission data

[0271] Output: Start of continuous playback

[0272] Specific behavior: Based on the playlist data, videos are played continuously using playback software such as VLC media player.

[0273] Through these steps, users can simply input their requests through voice, and the generative AI model will perform accurate analysis and search, allowing them to watch the desired videos continuously.

[0274] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0275] overview

[0276] This invention relates to a system that, when a user requests to watch multiple videos based on specific content, automatically searches for videos that match the request and the user's emotions and plays them sequentially. The following describes in detail the mode for carrying out the invention.

[0277] User request input and emotion recognition

[0278] User

[0279] The user clicks the Generate AI button on the device, which displays the input screen.

[0280] The user inputs a request into the displayed input screen by text or voice, such as "Show me 50 videos of multiple cats playing." At the same time, the emotion engine analyzes the user's emotions using voice or facial recognition.

[0281] The emotion engine recognizes the user's emotions in real time and uses that information to complete requests.

[0282] Sending and parsing requests

[0283] Terminal

[0284] The terminal acquires request data input by the user and emotion data provided by the emotion engine.

[0285] The device converts the acquired request data and emotion data into an appropriate data format such as JSON.

[0286] The terminal sends the converted data to the server as an HTTP request.

[0287] server

[0288] The server receives the request data and emotion data sent from the terminal.

[0289] The server passes the received request data to the generation AI for analysis. The generation AI analyzes the request data and extracts important keywords and context. At the same time, it analyzes sentiment data and identifies factors that influence the request.

[0290] For example, in addition to elements such as "cats," "playing," "video," and "50 items," it identifies whether the user is expressing the emotion of "joy."

[0291] Video search and selection

[0292] server

[0293] The server searches for videos using a video database or an external video search API based on the analyzed elements and emotion data. Based on the user's emotion, it prioritizes searches for videos that are likely to evoke "joy," for example.

[0294] The server retrieves a list of related videos from the search results.

[0295] Playlist generation and submission

[0296] server

[0297] The server selects 50 videos from the search results that match the user's request and emotions.

[0298] The server lists the information of the selected videos (title, URL, thumbnail, emotion tag, etc.) and generates a playlist.

[0299] The playlist is converted into an appropriate data format, such as JSON, and sent to the device.

[0300] Receive and autoplay playlists

[0301] Terminal

[0302] The terminal receives the playlist data transmitted from the server.

[0303] The device analyzes the received playlist data and obtains the playback order and URL information for each video.

[0304] The device will play the first video in the playlist order, and when that video finishes playing, the next video will automatically play.

[0305] Specific examples

[0306] A specific example will be described in which a user requests "Show me 50 videos of multiple cats playing," and the emotion engine recognizes the user's "joy."

[0307] 1. Request input

[0308] The user enters the text "Show me 50 videos of multiple cats playing," and the device, through its emotion engine, recognizes that the user's emotion is "joy."

[0309] 2. Parsing the Request

[0310] The server receives the request and emotion data and uses a generative AI to extract and analyze the keywords "cat," "playing," "video," "50 items," and the emotion "joy."

[0311] 3. Search and select videos

[0312] The server searches a video database based on the keywords and emotion data and selects 50 videos that match the criteria, with priority given to videos that match the emotion of joy.

[0313] 4. Generate a playlist

[0314] The server generates a playlist of 50 videos that match the conditions and sends it to the device.

[0315] 5. Continuous video playback

[0316] The terminal plays videos based on the received playlist, and the user watches a series of videos that give them pleasure as requested.

[0317] In this way, users can smoothly view content that matches their individual emotions, providing a more satisfying viewing experience.

[0318] The processing flow will be explained below.

[0319] Step 1: User request input and emotion recognition

[0320] The user clicks the Generate AI button on the device, which displays the input screen.

[0321] The user enters a request, such as "Show me 50 videos of multiple cats playing," into the displayed input screen using text or voice.

[0322] At the same time, the emotion engine recognizes the user's emotions (e.g., "joy") in real time through the user's voice input and facial expression analysis using a camera.

[0323] Step 2: Sending the request and emotion data

[0324] The terminal collects request data input by the user and emotion data obtained from the emotion engine.

[0325] The device converts the collected request data and emotion data into JSON format and sends it to the server as an HTTP request.

[0326] Step 3: Receive and parse the request and emotion data

[0327] The server receives the request data and emotion data sent from the terminal.

[0328] The server passes the request data to the generation AI for analysis, which then analyzes the request data and extracts important keywords and context (e.g., "cats," "playing," "video," "50 items").

[0329] The server also analyzes emotional data to recognize the user's emotions (e.g., "happiness") and identify factors that influence the request.

[0330] Step 4: Search and select a video

[0331] The server searches for videos using a video database or an external video search API based on the analyzed elements and emotion data. Based on the user's emotion, it will prioritize videos that evoke "joy," for example.

[0332] The server retrieves a list of relevant videos from the search results and selects 50 videos that meet the user's request and match the user's emotions.

[0333] Step 5: Generate a playlist

[0334] The server lists the information of the selected 50 videos (title, URL, thumbnail, emotion tag, etc.) and generates a playlist taking into account the playback order.

[0335] The playlist is converted to JSON format and sent to the device.

[0336] Step 6: Receiving and parsing the playlist

[0337] The terminal receives the playlist data transmitted from the server.

[0338] The device analyzes the received playlist data and obtains information such as the URL, title, and thumbnail of each video.

[0339] Step 7: Autoplay Video

[0340] The device will play the first video in the playlist, and when it finishes playing, the next video will automatically play.

[0341] Users can watch videos continuously, so they can enjoy the videos they want without stress.

[0342] Examples:

[0343] A specific example will be described in which a user requests "Show me 50 videos of multiple cats playing," and the emotion engine recognizes the user's "joy."

[0344] 1. Request input and emotion recognition

[0345] The user enters the text "Show me 50 videos of multiple cats playing," and the device, through its emotion engine, recognizes that the user's emotion is "joy."

[0346] 2. Sending Requests and Emotion Data

[0347] The device collects request data and emotion data, converts it into JSON format, and sends it to the server.

[0348] 3. Request and Sentiment Data Analysis

[0349] The server analyzes the received data and uses generative AI to extract the keywords "cat," "playing," "video," "50," and the emotion "joy."

[0350] 4. Search and select videos

[0351] The server searches a video database based on the keywords and emotion data, and selects 50 videos that match the criteria. Videos that match the emotion of joy are given priority.

[0352] 5. Generate and send playlists

[0353] The server generates a playlist of the selected 50 videos, converts it to JSON format, and sends it to the device.

[0354] 6. Receiving and parsing the playlist

[0355] The terminal analyzes the received playlist data and obtains information about each video.

[0356] 7. Autoplay Videos

[0357] The device automatically plays videos according to the playlist, allowing users to watch a series of videos that bring them pleasure as requested.

[0358] Example 2

[0359] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0360] While conventional video viewing systems can search for and play videos based on user requests, they are unable to consider the user's emotional state, which can result in a less satisfying viewing experience. Furthermore, they lack a mechanism for properly analyzing the request data and emotional data entered by the user and selecting the most appropriate video based on that data. Furthermore, the quality and content of the video desired by the user often do not completely match the request, making it difficult to provide content that meets the user's viewing objectives.

[0361] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes an emotion recognition means for analyzing the user's emotions in real time, a means for converting input request data and emotion data into an appropriate data format, a means for analyzing the request data and emotion data using a generation AI, and a means for selecting suitable videos from the search results and generating a playlist taking the emotion data into consideration. This enables request analysis that reflects the user's emotional state and optimal video selection.

[0362] "User" refers to a person who uses this system to input requests and watch videos.

[0363] A "request" refers to a text or voice input by a user describing the content and conditions of the video they wish to watch.

[0364] "Emotion recognition means" refers to a means for analyzing the user's voice and facial expressions to obtain emotional data in real time.

[0365] "Request data" refers to data including the content and conditions of the video entered by the user.

[0366] "Emotional data" refers to data that represents the analyzed emotional state of a user.

[0367] "Appropriate data format" refers to converting request data and emotion data into a standard format such as JSON format for sending to the server.

[0368] "Generative AI" refers to an artificial intelligence model that analyzes request data and emotional data, extracts important keywords and elements, and selects videos that are appropriate for the request.

[0369] "Playlist" refers to a list of videos curated based on user requests and emotional data.

[0370] "Video search service" refers to a service that allows users to search for related videos using an internal database or external API.

[0371] "Database" refers to the system that stores and manages videos and other related data.

[0372] overview

[0373] This invention relates to a system that, when a user requests to watch multiple videos based on specific content, automatically searches for videos that match the request and the user's emotions and plays them sequentially.

[0374] User request input and emotion recognition

[0375] User

[0376] The user clicks the Generate AI button on the device to display the input screen.

[0377] The user inputs a request into the displayed input screen, such as "Show me 50 videos of multiple cats playing." This can be done by text input or voice input.

[0378] As soon as the user inputs a request, the device activates an emotion engine, which analyzes the user's emotions in real time through voice and facial recognition systems.

[0379] Sending and parsing requests

[0380] Terminal

[0381] The terminal acquires the request data input by the user and the emotion data acquired from the emotion engine.

[0382] The device converts this data into an appropriate data format, such as JSON.

[0383] The terminal sends the converted data to the server as an HTTP request.

[0384] server

[0385] The server receives the request data and emotion data sent from the terminal.

[0386] The server passes the request data to the AI ​​generator for analysis. The AI ​​generator extracts important keywords and context from the request data, and also analyzes sentiment data to identify factors that influence the request.

[0387] For example, in addition to elements such as "cats," "playing," "video," and "50 items," it identifies that the user is expressing the emotion of "joy."

[0388] Video search and selection

[0389] server

[0390] The server searches for videos using a video database or external video search API based on the analyzed elements and emotional data, with a particular focus on videos that are likely to evoke "joy."

[0391] The server retrieves a list of videos that match the conditions.

[0392] Playlist generation and submission

[0393] server

[0394] The server selects 50 videos that match the request and emotion data.

[0395] The server lists the information of the selected videos (title, URL, thumbnail, emotion tag, etc.) and generates a playlist.

[0396] The playlist is converted into an appropriate data format, such as JSON, and sent to the device.

[0397] Receive and autoplay playlists

[0398] Terminal

[0399] The terminal receives the playlist data transmitted from the server.

[0400] The device analyzes the received playlist data and obtains the playback order and URL information for each video.

[0401] The device will play the first video in the playlist order, and when that video finishes playing, the next video will automatically play.

[0402] Examples of concrete examples and prompts

[0403] Specific examples

[0404] A specific example will be described in which a user requests "Show me 50 videos of multiple cats playing," and the emotion engine recognizes the user's "joy."

[0405] 1. Request input

[0406] The user enters the text "Show me 50 videos of multiple cats playing," and the device, through its emotion engine, recognizes that the user's emotion is "joy."

[0407] 2. Parsing the Request

[0408] The server receives the request and emotion data and uses a generative AI to extract and analyze the keywords "cat," "playing," "video," "50 items," and the emotion "joy."

[0409] 3. Search and select videos

[0410] The server searches a video database based on the keywords and emotion data and selects 50 videos that match the criteria, with priority given to videos that match the emotion of joy.

[0411] 4. Generate a playlist

[0412] The server generates a playlist of 50 videos that match the conditions and sends it to the device.

[0413] 5. Continuous video playback

[0414] The terminal plays videos based on the received playlist, and the user watches a series of videos that give them pleasure as requested.

[0415] Prompt Sentence Examples

[0416] "The user inputs a request via text or voice, and the emotion engine analyzes the user's emotion based on that request. The request including that emotion data is then sent to the server, where the generation AI analyzes it and searches for and selects videos. The system then generates a playlist and provides it to the user. Please explain the process."

[0417] In this way, users can smoothly view content that matches their individual emotions, providing a more satisfying viewing experience.

[0418] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0419] System program processing flow

[0420] Step 1:

[0421] User request input and emotion recognition

[0422] Input: The user clicks the Generate AI button on the device and enters a request in text or voice on the input screen, such as "Show me 50 videos of multiple cats playing."

[0423] Processing: The terminal receives the user's request, and at the same time, the emotion engine analyzes the user's emotion data through the user's voice or face recognition system. The emotion engine recognizes the user's emotion in real time and adds the data to the request.

[0424] Output: The user's request data and emotion data are acquired by the terminal.

[0425] Specific behavior:

[0426] When a user says, "Show me 50 videos of cats playing," the speech recognition module converts the speech into text data, and the emotion engine detects "joy" from the user's tone of voice and facial expression.

[0427] Step 2:

[0428] Sending and parsing requests

[0429] Input: The request data and sentiment data obtained in step 1.

[0430] Processing: The device converts the request data and emotion data into an appropriate data format, such as JSON, and sends the converted data to the server as an HTTP request.

[0431] Output: Request data and emotion data sent to the server.

[0432] Specific behavior:

[0433] The device converts the request and emotion data into JSON format as "Request data: 50 videos of cats playing" and "Emotion data: Joy" and sends it to the server using the HTTP POST method.

[0434] Step 3:

[0435] Receiving and parsing request and emotion data

[0436] Input: Request data and emotion data in JSON format sent from the device.

[0437] Processing: The server receives the request data and emotion data and passes it to the generation AI, which extracts important keywords and context from the request data and analyzes the emotion data to identify factors that influence the request.

[0438] Output: Request and sentiment elements analyzed by the generative AI.

[0439] Specific behavior:

[0440] The server extracts and analyzes emotional data related to keywords such as "cat," "playing," "video," "50 items," and "joy." Based on this data, the AI ​​generator identifies the video characteristics desired by the user.

[0441] Step 4:

[0442] Video search and selection

[0443] Input: Keywords and sentiment elements analyzed by the generative AI.

[0444] Processing: Based on the analyzed elements, the server searches for related videos using a video database or an external video search API. It prioritizes videos that are particularly likely to evoke "joy." It then selects videos that match the search criteria from the search results.

[0445] Output: A list of videos that match the criteria.

[0446] Specific behavior:

[0447] The generation AI calls the search API to search for videos containing the tags "cat," "playing," and "joy." It then retrieves information about videos that match the criteria (title, URL, thumbnail, etc.).

[0448] Step 5:

[0449] Playlist generation and submission

[0450] Input: A list of videos that match the criteria.

[0451] Processing: The server selects 50 videos that match the user's request and emotion, lists their information (title, URL, thumbnail, emotion tag, etc.), and generates a playlist. The playlist is converted into an appropriate data format such as JSON and sent to the device.

[0452] Output: The playlist data sent to the device.

[0453] Specific behavior:

[0454] The server selects 50 videos, generates a JSON file containing the title, URL, thumbnail, and emotion tag, and sends it to the terminal as an HTTP response.

[0455] Step 6:

[0456] Receive and autoplay playlists

[0457] Input: Playlist data sent from the server.

[0458] Processing: The device analyzes the received playlist data and obtains the playback order and URL information for each video. It plays the first video in the playlist order, and when that video finishes playing, the next video automatically starts playing.

[0459] Output: The continuous video viewing experience delivered to the user.

[0460] Specific behavior:

[0461] The device parses the received JSON playlist and passes the URL of the first video to the browser or application's video player, starting autoplay. The next video will automatically play after the previous one finishes.

[0462] The above processing allows users to smoothly view content that matches their individual emotions, providing a more satisfying viewing experience.

[0463] (Application example 2)

[0464] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0465] When a user wants to watch a video, it is necessary to efficiently search for videos that meet the user's request and provide videos that are tailored to the user's emotions. However, conventional systems cannot simultaneously consider the user's request and emotions, making it difficult to improve user satisfaction. To solve this problem, a system is needed that analyzes both the user's request and real-time emotional data and provides the optimal video based on that analysis.

[0466] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0467] In this invention, the server includes a means for a user to input a request, a means for analyzing the input request and the user's emotions, and a means for searching for related videos based on the analysis results, thereby enabling efficient searching of related videos based on the user's request and emotions, and improving user satisfaction.

[0468] The "means for users to input requests" refers to an interface that allows users to input the content they request and the type of video they desire by text or voice.

[0469] The "means for analyzing input requests and user emotions" refers to software and hardware for analyzing request data, facial expressions, and voice from users to estimate emotions.

[0470] The "means for searching for relevant videos based on the analysis results" refers to an algorithm and system that searches for appropriate videos from a database or external API based on the analyzed request data and emotion data.

[0471] "Means for selecting searched videos and generating a playlist" refers to an algorithm and system that selects the most appropriate videos from the search results based on the user's requests and emotions, and creates a list for playing them consecutively.

[0472] The "means for playing back the generated playlist" is a player that plays back videos in sequence based on the playlist, allowing the user to view them continuously.

[0473] "Generative AI" refers to models and algorithms that use artificial intelligence to analyze user requests and provide optimal information and content.

[0474] An "emotion recognition engine" is a system or software that analyzes and estimates emotions from a user's facial expressions and voice.

[0475] A "playlist" is a list of videos that are played consecutively in a specified order.

[0476] MODE FOR CARRYING OUT THE INVENTION

[0477] overview

[0478] This invention relates to a system that, when a user wishes to watch videos based on a specific request and the emotions they are feeling at the time, automatically searches for videos that match the request and emotions and plays them sequentially. The following describes in detail an embodiment of this invention.

[0479] User request input and emotion recognition

[0480] User operations

[0481] Users input requests through the device's request input screen. Requests can be entered by text or voice. At this time, an emotion recognition engine analyzes the user's emotions in real time from their facial expressions and voice tone. For example, if a user inputs "Show me 50 videos of multiple cats playing," the system will recognize that the emotion expressed at that time is "joy."

[0482] Sending requests and emotion data

[0483] Terminal handling

[0484] The device receives requests input by the user and emotion data provided by the emotion recognition engine, converts these data into JSON format, and sends it to the server as an HTTP request.

[0485] Parsing requests and sentiment data

[0486] Server Processing

[0487] The server receives the request data and emotion data sent from the device and analyzes the request content and emotion data using a generative AI model. Based on the analysis results, important keywords and contexts are extracted, and the emotion data that influences the request is identified. For example, the keywords "cat," "playing," "video," "50," and "joy" are extracted.

[0488] Video search and selection

[0489] Server Processing

[0490] The server searches for videos using a video database or an external video search API based on the analyzed keywords and emotion data. Videos that match the emotion of joy are searched for first. The most suitable video is selected from the search results.

[0491] Playlist generation and submission

[0492] Server processing and terminal reception

[0493] The server creates a playlist by listing the information of the selected videos (title, URL, thumbnail, emotion tag, etc.). This playlist is converted to JSON format and sent to the device. The device analyzes the received playlist data and plays the videos sequentially, starting from the first video.

[0494] Hardware and software used

[0495] Hardware:

[0496] Smartphone: Enter requests and play videos.

[0497] Camera and microphone: Used to recognize the user's emotions.

[0498] software:

[0499] Emotion recognition engine (e.g., Affectiva, Microsoft Azure Face API): Analyzes user emotions in real time.

[0500] Generative AI models (e.g., GPT-3, BERT): Analyze user requests and extract relevant keywords.

[0501] Video search API (e.g. YouTube Data API, Vimeo API): Search for related videos based on analyzed keywords.

[0502] Specific examples

[0503] If a user requests "Show me 50 videos of multiple cats playing," and the emotion engine recognizes the user's emotion of "joy," the server will search for keywords such as "cat," "playing," "video," "50 cats," and "joy." From the search results, the server will generate a playlist of the 50 best videos and send it to the device. The user will then be able to watch videos that match their emotion in succession, as requested.

[0504] Prompt Sentence Examples

[0505] Animals, Funny Videos, Joy

[0506] Cat, playing, happy

[0507] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0508] Step 1:

[0509] The terminal receives requests from the user and acquires emotion data.

[0510] The user inputs a request using text or voice. The device's emotion recognition engine also uses a camera or microphone to capture the user's emotion data. The input data and emotion data are temporarily stored on the device. The input includes text data or voice data, and the output includes request data and emotion data.

[0511] Step 2:

[0512] The device converts the request data and emotion data into JSON format and sends it to the server.

[0513] The device receives request data input by the user and emotion data from the emotion recognition engine, formats the data into JSON format, and then sends the data to the server using an HTTP request. The input includes the request data and emotion data, and the output includes a JSON-formatted HTTP request.

[0514] Step 3:

[0515] The server receives the request data and emotion data and analyzes it using a generative AI model.

[0516] The server receives JSON-formatted request data and emotion data sent from the device. It uses a generative AI model to analyze the request content (e.g., "cats," "playing," "video," "50") and emotion data (e.g., "joy") and extract important keywords and context. The input includes the JSON-formatted request data and emotion data, and the output includes the analyzed keywords and emotion information.

[0517] Step 4:

[0518] The server performs a video search based on the analysis results.

[0519] The server uses the analyzed keywords and emotional information to search for related videos using a video database or an external video search API. At this time, the search algorithm is adjusted to prioritize videos that match the emotion (e.g., videos that evoke joy). The input includes the analyzed keywords and emotional information, and the output includes a list of videos as search results.

[0520] Step 5:

[0521] The server generates a playlist from the acquired video list and transmits it to the terminal.

[0522] The server selects 50 videos from the search results and lists them as a playlist. The playlist, including information about the selected videos (title, URL, thumbnail, emotion tag, etc.), is converted to JSON format and sent to the terminal. The input contains a list of videos, and the output contains a JSON-formatted playlist.

[0523] Step 6:

[0524] The device automatically plays videos based on the playlist received from the server.

[0525] The device parses the JSON-formatted playlist data received from the server and plays the videos in the list sequentially. The first video in the playlist plays, and when it finishes, the next video automatically plays. The input is the JSON-formatted playlist data, and the output includes the currently playing and already played videos.

[0526] As a specific example, if you request "Show me 50 videos of multiple cats playing" based on the prompt sentences "animals, funny videos, joy" and "cats, playing, happiness," the system will search for videos that have the corresponding emotion of "joy," and the selected videos will be played as a playlist.

[0527] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0528] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0529] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0530] [Second embodiment]

[0531] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0532] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0533] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0534] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0535] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0536] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0537] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0538] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0539] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0540] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0541] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0542] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0543] overview

[0544] This invention relates to a system that, when a user requests to watch multiple videos based on specific content, automatically searches for videos that meet the request and plays them sequentially. The following describes in detail an embodiment of the invention.

[0545] User request input

[0546] User

[0547] A user starts a video viewing application on any device (e.g., a smartphone or computer). The application interface includes a Generate AI button for inputting video requests.

[0548] The user clicks the Generate AI button and enters a request in text or voice into the input screen that appears, such as "Show me 50 videos of multiple cats playing."

[0549] Sending and parsing requests

[0550] Terminal

[0551] The terminal sends the request data entered by the user to the server, where the request is converted into an appropriate data format (for example, JSON format).

[0552] The request data sent to the server is analyzed by the AI ​​generator. Specific keywords and contexts are extracted using the analysis method. For example, elements such as "cats," "playing," "video," and "50 items" are identified.

[0553] Video search and selection

[0554] server

[0555] The server searches a video database or an external video providing service based on the analyzed elements.

[0556] From the search results, 50 videos that match the user's request are selected, and the information of the selected videos (title, URL, thumbnail, etc.) is listed.

[0557] Playlist generation and submission

[0558] server

[0559] Based on the selected videos, the server generates a playlist that takes into account the playback order. The playlist includes information such as the URL, title, and thumbnail of each video.

[0560] The playlist is sent to the device in an appropriate data format (e.g., JSON format).

[0561] Receive and autoplay playlists

[0562] Terminal

[0563] The terminal receives the playlist sent from the server.

[0564] Once the playlist has been parsed, the videos will begin playing automatically. Based on the playlist, the user can watch the requested videos in succession.

[0565] Specific examples

[0566] A specific example will be described in which a user requests "Show me 50 videos of multiple cats playing."

[0567] 1. Request input

[0568] The user types a request into the terminal as text and clicks the send button.

[0569] 2. Parsing the Request

[0570] The server receives the request and uses a generation AI to extract the keywords "cat," "playing," "video," and "50."

[0571] 3. Search and select videos

[0572] The server searches a video database based on these keywords and selects 50 videos that meet the criteria.

[0573] 4. Generate a playlist

[0574] The server generates a playlist of 50 videos that match the conditions and sends it to the device.

[0575] 5. Continuous video playback

[0576] The terminal plays the videos based on the received playlist, and the user watches the videos consecutively according to the requested content.

[0577] In this way, users can smoothly watch the content they want, greatly improving convenience. In addition, by using generative AI, it is possible to accurately understand and analyze the content of the request and provide the appropriate video.

[0578] The processing flow will be explained below.

[0579] Step 1: User request input

[0580] The user clicks the Generate AI button displayed on the device, which displays the input screen.

[0581] The user enters a request into the displayed input screen using text or voice, such as "Show me 50 videos of multiple cats playing."

[0582] After completing the input, the user presses the "Send" button to send the input contents.

[0583] Step 2: Submitting the request

[0584] The terminal acquires the request data input by the user.

[0585] The terminal converts the acquired request data into an appropriate data format such as JSON format.

[0586] The terminal transmits the converted request data to the server as an HTTP request.

[0587] Step 3: Receiving and Parsing the Request

[0588] The server receives the request data sent from the terminal.

[0589] The server passes the received request data to the generation AI for analysis.

[0590] The generative AI analyzes the request data and extracts important keywords and context (e.g., "cats," "playing," "video," "50 items").

[0591] The server receives the analysis results from the generation AI.

[0592] Step 4: Search for videos

[0593] Based on the analysis results, the server searches for videos using a video database or an external video search API.

[0594] The server uses the extracted keywords and conditions as a search query.

[0595] The server retrieves a list of related videos from the search results.

[0596] Step 5: Select a video

[0597] The server selects 50 videos from the search results that match the user's request.

[0598] The server lists the information of the selected videos (title, URL, thumbnail, etc.).

[0599] Step 6: Generate a playlist

[0600] The server generates a playlist based on the list of selected videos.

[0601] A playlist contains information such as the playback order, title, URL, and thumbnail of each video.

[0602] Step 7: Submit your playlist

[0603] The server converts the generated playlist into an appropriate data format, such as JSON.

[0604] The server transmits the converted playlist data to the terminal.

[0605] Step 8: Receive and play the playlist

[0606] The terminal receives the playlist data transmitted from the server.

[0607] The device analyzes the received playlist data and obtains the playback order and URL information for each video.

[0608] The device will play the first video in the playlist order, and when that video finishes playing, the next video will automatically play.

[0609] This system allows users to watch videos of their choice continuously without any hassle, greatly improving convenience and the viewing experience.

[0610] Example 1

[0611] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0612] Currently, if a user wants to watch specific content consecutively, they have to manually search for multiple videos and add them to a playlist individually. This makes the process cumbersome for users, making it difficult to quickly watch the content they want. Furthermore, it can be difficult to accurately understand a user's request and provide the appropriate video, which can reduce satisfaction.

[0613] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0614] In this invention, the server includes a means for a user to input a request, a terminal means for transmitting the request input data to the server, a means for analyzing the input request using a generative AI model, a means for searching for related videos from a database or an external video service based on the analysis results, a means for selecting the searched videos and generating a playlist, a means for transmitting the generated playlist to the terminal, and a means for analyzing the generated playlist on the terminal and automatically playing videos. This eliminates the need for complicated user operations and enables users to quickly watch the desired content. Furthermore, the use of the generative AI model makes it possible to accurately understand user requests and provide optimal videos.

[0615] "User" means an individual or corporation that uses the system to input requests and view content.

[0616] "Request input means" refers to an interface or device that allows a user to input a request by text or voice.

[0617] "Terminal means" refers to an electronic device such as a smartphone or computer that sends a user's request to the server.

[0618] A "server" is a computer system that receives requests, analyzes, searches, generates playlists, and transmits playlists.

[0619] A "generative AI model" is an artificial intelligence that analyzes requests entered by users and extracts keywords and context.

[0620] "Analysis means" refers to the process or algorithm that analyzes the input request using a generative AI model and extracts the necessary information.

[0621] "Video Database" means internal or external data storage used to search for relevant videos.

[0622] "Third-party video provider" means an external video hosting platform, such as YouTube.

[0623] "Search method" refers to the process or algorithm for searching for relevant videos from a video database or external video provider service based on the analysis results.

[0624] A "selection method" is a process or algorithm that selects retrieved videos that match the user's request.

[0625] A "playlist generator" is a process or algorithm that compiles selected videos into a playlist, taking into account the playback order.

[0626] "Playlist transmission means" refers to a process or algorithm that transmits the generated playlist to a terminal in an appropriate data format.

[0627] "Autoplay method" refers to the process or algorithm that analyzes the playlist generated by the device and plays videos continuously.

[0628] overview

[0629] This invention relates to a system that automatically searches for and sequentially plays videos that fit a user's request when the user requests to watch multiple videos based on specific content. Specifically, this describes the process flow and device configuration for analyzing the user's request, searching for, selecting, and creating a playlist of related videos, and automatically playing them.

[0630] User request input

[0631] User

[0632] A user launches a video viewing application using a device such as a smartphone or computer.

[0633] The application interface has a Generate AI button, and clicking this will display the input screen.

[0634] The user enters a request into the input screen using text or voice, such as "Show me 50 videos of multiple cats playing," and clicks the send button.

[0635] Sending and parsing requests

[0636] Terminal

[0637] The terminal converts the request data entered by the user into JSON format and sends it to the server, using the HTTP POST method to transfer the data.

[0638] server

[0639] The server uses a generative AI model (e.g., OpenAI's GPT-3) to analyze the received request data.

[0640] The generative AI model extracts necessary keywords and context from the request, such as "cats," "playing," "video," and "50 items."

[0641] Video search and selection

[0642] server

[0643] The server searches its internal video database or external video providers such as YouTube based on the keywords extracted by the generative AI model.

[0644] The server selects 50 videos from the search results that match the user's request. Specifically, it uses the YouTube Data API to search for videos containing tags such as "cat" and "playing" and extracts the appropriate videos.

[0645] List the information of the selected videos (title, URL, thumbnail, etc.).

[0646] Playlist generation and submission

[0647] server

[0648] The server generates a playlist based on the selected video list, taking into consideration the playback order.

[0649] The playlist contains information such as the URL, title, and thumbnail of each video.

[0650] The created playlist is formatted in JSON format and sent to the device as an HTTP response.

[0651] Receive and autoplay playlists

[0652] Terminal

[0653] The terminal receives the playlist sent from the server.

[0654] It analyzes the playlist and automatically starts playing videos based on the information of each video. The videos are played continuously, allowing users to watch the requested videos one after the other.

[0655] Specific examples

[0656] Below is what happens when a user requests "Show me 50 videos of multiple cats playing together."

[0657] 1. Request input

[0658] The user types the text "Show me 50 videos of multiple cats playing" into their device and clicks the send button.

[0659] 2. Parsing the Request

[0660] The server receives the request and uses a generative AI model (such as OpenAI's GPT-3) to extract the keywords "cat," "playing," "video," and "50."

[0661] 3. Search and select videos

[0662] The server searches a video database based on the keywords and selects 50 videos that match the criteria, for example, using the YouTube API.

[0663] 4. Generate a playlist

[0664] The server generates a playlist of 50 videos that match the conditions and sends it to the device.

[0665] 5. Continuous video playback

[0666] The terminal plays the videos based on the received playlist, and the user watches the videos consecutively according to the requested content.

[0667] Prompt Sentence Examples

[0668] "Find 50 videos of multiple cats playing together and create a playlist."

[0669] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0670] Step 1:

[0671] User input of request

[0672] A user launches a video viewing application on a device such as a smartphone or computer.

[0673] Click the Generate AI button on the application interface to display the input screen.

[0674] Users enter a request via text or voice, such as "Show me 50 videos of multiple cats playing," and click the send button.

[0675] Input: The request text or voice data entered by the user.

[0676] Output: The request data is sent to the terminal.

[0677] Step 2:

[0678] Send a request from the device

[0679] The terminal converts the request data entered by the user into JSON format.

[0680] The converted request data is sent to the server using the HTTP POST method.

[0681] Input: Request data from the user (text or voice).

[0682] Output: The request data converted to JSON format is sent to the server.

[0683] Step 3:

[0684] Server parsing of the request

[0685] The server passes the received request data to a generative AI model (e.g., OpenAI's GPT-3).

[0686] The generative AI model extracts keywords and context from the request, such as "cat," "playing," "video," and "50 items."

[0687] Input: Request data in JSON format.

[0688] Output: Keywords and context extracted by the generative AI model.

[0689] Step 4:

[0690] Server-based video search and selection

[0691] The server searches an internal video database or an external video provider (e.g., YouTube) based on the keywords extracted by the generative AI model.

[0692] The server selects 50 videos from the search results that match the user's request. For example, it uses the YouTube Data API to search for videos containing the tags "cat" and "playing."

[0693] Input: Keywords extracted by the generative AI model.

[0694] Output: Information about the selected video (title, URL, thumbnail, etc.).

[0695] Step 5:

[0696] Server-generated and transmitted playlists

[0697] The server generates a playlist based on the selected video list, taking into account the playback order.

[0698] The playlist contents are formatted in JSON format.

[0699] The server sends the generated playlist to the terminal as an HTTP response.

[0700] Input: Information about the selected video (title, URL, thumbnail, etc.).

[0701] Output: Playlist data in JSON format.

[0702] Step 6:

[0703] Device receives and autoplays playlist

[0704] The terminal receives the playlist sent from the server.

[0705] Analyze the playlist and get information about each video.

[0706] The device will automatically start playing videos based on the playlist, allowing users to watch the videos they want in sequence.

[0707] Input: Playlist data in JSON format.

[0708] Output: Automatic continuous playback of videos.

[0709] This allows users to view the content they desire efficiently. The accuracy and speed of server-dependent processing are the keys to improving the overall performance of the system.

[0710] (Application example 1)

[0711] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0712] Conventional video streaming systems have difficulty quickly and accurately selecting relevant videos based on specific user requests and playing them continuously. Furthermore, the technology for accepting voice requests, accurately converting that voice into text, and analyzing it using generative AI models has been immature. This has resulted in inconvenient services that cannot meet user needs.

[0713] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0714] In this invention, the server includes means for allowing a user to input a request, means for analyzing the input request, means for searching for related videos based on the analysis results, means for selecting the searched videos and generating a playlist, means for playing the generated playlist, means for converting the user's voice into text using a voice recognition means, and means for retrieving data from an online video database. This makes it possible to accurately select videos using a generative AI model based on the user's voice or text input request and play them continuously.

[0715] "User" refers to a person who uses the System to enter a request and wish to watch a video.

[0716] "Means for inputting requests" refers to an interface through which a user inputs video requests into the system in the form of voice or text.

[0717] "Means for analyzing requests" refers to the process of analyzing input requests using a generative AI model and extracting relevant keywords and context.

[0718] "Means for searching for related videos" refers to a process for searching for appropriate videos from a video database or external video providing service based on the analyzed keywords and context.

[0719] "Means for selecting and generating a playlist" refers to the process of selecting appropriate videos from the search results and organizing the videos into a playlist based on the playback order.

[0720] The "means for playing the generated playlist" refers to a process for playing videos continuously based on the generated playlist.

[0721] "Speech recognition means" refers to technology for converting user-input speech into text.

[0722] "Means for obtaining data from an online video database" refers to a process for obtaining necessary video data from an online video database via the Internet.

[0723] A "generative AI model" refers to an artificial intelligence algorithm that performs highly accurate analysis based on input information.

[0724] A "playlist" is a list of related videos arranged in a particular order.

[0725] overview

[0726] This invention is a system that automatically searches for and plays back videos that fit a user's request for viewing multiple videos based on specific content, improving user convenience and enabling users to smoothly view the content they desire.

[0727] User request input

[0728] A user uses a device of their choice (e.g., a smartphone or head-mounted display) to launch a video viewing application. The application interface has a voice recognition button for inputting video requests. The user clicks this button and inputs a request by voice, such as "Show me 50 videos of multiple cats playing."

[0729] Sending and parsing requests

[0730] The device sends the request data entered by the user to the server. The request is converted into text by a speech recognition means and then converted into an appropriate data format (e.g., JSON format). Specific keywords and context are extracted by an analysis means. For example, elements such as "cats," "playing," "video," and "50 items" are identified. This is where the generative AI model comes into play.

[0731] Video search and selection

[0732] The server searches online video databases and external video providers based on the analyzed elements. From the search results, it selects 50 videos that match the user's request. The information of the selected videos (title, URL, thumbnail, etc.) is then listed.

[0733] Playlist generation and transmission

[0734] The server generates a playlist based on the selected videos, taking into account the playback order. The playlist includes information such as the URL, title, and thumbnail of each video. The generated playlist is sent to the device in the appropriate data format.

[0735] Receive and autoplay playlists

[0736] The device receives the playlist sent from the server. Once the analysis of the playlist is complete, the video starts playing automatically. Based on the playlist, the user can watch the requested videos in succession.

[0737] Specific examples

[0738] Below is a specific example of a user requesting "Show me 50 videos of multiple cats playing." The user voice-types the request into the device and clicks the send button. The device converts the voice to text and uses a generative AI model to extract the keywords "cat," "playing," "video," and "50." The server searches a video database based on these keywords and selects 50 videos that match the criteria. The server then creates a playlist of the 50 videos that match the criteria and sends it to the device. The device plays the videos based on the received playlist, and the user watches the videos consecutively according to the request.

[0739] Prompt Sentence Examples

[0740] "Show me 50 videos of multiple cats playing together"

[0741] This allows users to easily input requests through voice, and the generative AI model performs accurate analysis and search, allowing them to watch the desired videos continuously.

[0742] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0743] Step 1:

[0744] The user launches a video viewing application and clicks the voice input button. The user then voices their request.

[0745] Input: User's voice request

[0746] Output: Audio data

[0747] Specific actions: Requests such as "Show me 50 videos of multiple cats playing" can be made using voice input.

[0748] Step 2:

[0749] The terminal uses a voice recognition means to convert the voice data into text data.

[0750] Input: Audio data

[0751] Output: Text data

[0752] Specific behavior: Uses a speech recognition library to convert speech into text. Example: "Show me 50 videos of multiple cats playing."

[0753] Step 3:

[0754] The device sends the converted text data to a generative AI model that analyzes the request.

[0755] Input: Text data

[0756] Output: Analysis results (keywords)

[0757] Specific operation: The generative AI model extracts keywords such as "cat," "playing," "video," and "50."

[0758] Step 4:

[0759] The server searches an online video database based on the analyzed keywords.

[0760] Input: Analysis results (keywords)

[0761] Output: Video candidates (title, URL, thumbnail information, etc.)

[0762] Specific operation: Using an API request, search the database for the terms "cat" and "playing" and retrieve 50 videos.

[0763] Step 5:

[0764] The server generates a playlist based on the acquired video candidates, taking into account the playback order.

[0765] Input: Video candidate

[0766] Output: Playlist (video URL, title, thumbnail information, etc.)

[0767] Specific operation: The information of the acquired videos is compiled into a list and compiled into a playlist based on the user's request.

[0768] Step 6:

[0769] The server transmits the generated playlist to the terminal.

[0770] Input: Playlist

[0771] Output: Playlist transmission data

[0772] Specific operation: Send the playlist to the device in an appropriate data format, such as JSON.

[0773] Step 7:

[0774] The device will analyze the received playlist and begin automatically playing the videos.

[0775] Input: Playlist transmission data

[0776] Output: Start of continuous playback

[0777] Specific behavior: Based on the playlist data, videos are played continuously using playback software such as VLC media player.

[0778] Through these steps, users can simply input their requests through voice, and the generative AI model will perform accurate analysis and search, allowing them to watch the desired videos continuously.

[0779] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0780] overview

[0781] This invention relates to a system that, when a user requests to watch multiple videos based on specific content, automatically searches for videos that match the request and the user's emotions and plays them sequentially. The following describes in detail the mode for carrying out the invention.

[0782] User request input and emotion recognition

[0783] User

[0784] The user clicks the Generate AI button on the device, which displays the input screen.

[0785] The user inputs a request into the displayed input screen by text or voice, such as "Show me 50 videos of multiple cats playing." At the same time, the emotion engine analyzes the user's emotions using voice or facial recognition.

[0786] The emotion engine recognizes the user's emotions in real time and uses that information to complete requests.

[0787] Sending and parsing requests

[0788] Terminal

[0789] The terminal acquires request data input by the user and emotion data provided by the emotion engine.

[0790] The device converts the acquired request data and emotion data into an appropriate data format such as JSON.

[0791] The terminal sends the converted data to the server as an HTTP request.

[0792] server

[0793] The server receives the request data and emotion data sent from the terminal.

[0794] The server passes the received request data to the generation AI for analysis. The generation AI analyzes the request data and extracts important keywords and context. At the same time, it analyzes sentiment data and identifies factors that influence the request.

[0795] For example, in addition to elements such as "cats," "playing," "video," and "50 items," it identifies whether the user is expressing the emotion of "joy."

[0796] Video search and selection

[0797] server

[0798] The server searches for videos using a video database or an external video search API based on the analyzed elements and emotion data. Based on the user's emotion, it prioritizes searches for videos that are likely to evoke "joy," for example.

[0799] The server retrieves a list of related videos from the search results.

[0800] Playlist generation and submission

[0801] server

[0802] The server selects 50 videos from the search results that match the user's request and emotions.

[0803] The server lists the information of the selected videos (title, URL, thumbnail, emotion tag, etc.) and generates a playlist.

[0804] The playlist is converted into an appropriate data format, such as JSON, and sent to the device.

[0805] Receive and autoplay playlists

[0806] Terminal

[0807] The terminal receives the playlist data transmitted from the server.

[0808] The device analyzes the received playlist data and obtains the playback order and URL information for each video.

[0809] The device will play the first video in the playlist order, and when that video finishes playing, the next video will automatically play.

[0810] Specific examples

[0811] A specific example will be described in which a user requests "Show me 50 videos of multiple cats playing," and the emotion engine recognizes the user's "joy."

[0812] 1. Request input

[0813] The user enters the text "Show me 50 videos of multiple cats playing," and the device, through its emotion engine, recognizes that the user's emotion is "joy."

[0814] 2. Parsing the Request

[0815] The server receives the request and emotion data and uses a generative AI to extract and analyze the keywords "cat," "playing," "video," "50 items," and the emotion "joy."

[0816] 3. Search and select videos

[0817] The server searches a video database based on the keywords and emotion data and selects 50 videos that match the criteria, with priority given to videos that match the emotion of joy.

[0818] 4. Generate a playlist

[0819] The server generates a playlist of 50 videos that match the conditions and sends it to the device.

[0820] 5. Continuous video playback

[0821] The terminal plays videos based on the received playlist, and the user watches a series of videos that give them pleasure as requested.

[0822] In this way, users can smoothly view content that matches their individual emotions, providing a more satisfying viewing experience.

[0823] The processing flow will be explained below.

[0824] Step 1: User request input and emotion recognition

[0825] The user clicks the Generate AI button on the device, which displays the input screen.

[0826] The user enters a request, such as "Show me 50 videos of multiple cats playing," into the displayed input screen using text or voice.

[0827] At the same time, the emotion engine recognizes the user's emotions (e.g., "joy") in real time through the user's voice input and facial expression analysis using a camera.

[0828] Step 2: Sending the request and emotion data

[0829] The terminal collects request data input by the user and emotion data obtained from the emotion engine.

[0830] The device converts the collected request data and emotion data into JSON format and sends it to the server as an HTTP request.

[0831] Step 3: Receive and parse the request and emotion data

[0832] The server receives the request data and emotion data sent from the terminal.

[0833] The server passes the request data to the generation AI for analysis, which then analyzes the request data and extracts important keywords and context (e.g., "cats," "playing," "video," "50 items").

[0834] The server also analyzes emotional data to recognize the user's emotions (e.g., "happiness") and identify factors that influence the request.

[0835] Step 4: Search and select a video

[0836] The server searches for videos using a video database or an external video search API based on the analyzed elements and emotion data. Based on the user's emotion, it will prioritize videos that evoke "joy," for example.

[0837] The server retrieves a list of relevant videos from the search results and selects 50 videos that meet the user's request and match the user's emotions.

[0838] Step 5: Generate a playlist

[0839] The server lists the information of the selected 50 videos (title, URL, thumbnail, emotion tag, etc.) and generates a playlist taking into account the playback order.

[0840] The playlist is converted to JSON format and sent to the device.

[0841] Step 6: Receiving and parsing the playlist

[0842] The terminal receives the playlist data transmitted from the server.

[0843] The device analyzes the received playlist data and obtains information such as the URL, title, and thumbnail of each video.

[0844] Step 7: Autoplay Video

[0845] The device will play the first video in the playlist, and when it finishes playing, the next video will automatically play.

[0846] Users can watch videos continuously, so they can enjoy the videos they want without stress.

[0847] Examples:

[0848] A specific example will be described in which a user requests "Show me 50 videos of multiple cats playing," and the emotion engine recognizes the user's "joy."

[0849] 1. Request input and emotion recognition

[0850] The user enters the text "Show me 50 videos of multiple cats playing," and the device, through its emotion engine, recognizes that the user's emotion is "joy."

[0851] 2. Sending Requests and Emotion Data

[0852] The device collects request data and emotion data, converts it into JSON format, and sends it to the server.

[0853] 3. Request and Sentiment Data Analysis

[0854] The server analyzes the received data and uses generative AI to extract the keywords "cat," "playing," "video," "50," and the emotion "joy."

[0855] 4. Search and select videos

[0856] The server searches a video database based on the keywords and emotion data, and selects 50 videos that match the criteria. Videos that match the emotion of joy are given priority.

[0857] 5. Generate and send playlists

[0858] The server generates a playlist of the selected 50 videos, converts it to JSON format, and sends it to the device.

[0859] 6. Receiving and parsing the playlist

[0860] The terminal analyzes the received playlist data and obtains information about each video.

[0861] 7. Autoplay Videos

[0862] The device automatically plays videos according to the playlist, allowing users to watch a series of videos that bring them pleasure as requested.

[0863] Example 2

[0864] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0865] While conventional video viewing systems can search for and play videos based on user requests, they are unable to consider the user's emotional state, which can result in a less satisfying viewing experience. Furthermore, they lack a mechanism for properly analyzing the request data and emotional data entered by the user and selecting the most appropriate video based on that data. Furthermore, the quality and content of the video desired by the user often do not completely match the request, making it difficult to provide content that meets the user's viewing objectives.

[0866] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes an emotion recognition means for analyzing the user's emotions in real time, a means for converting input request data and emotion data into an appropriate data format, a means for analyzing the request data and emotion data using a generation AI, and a means for selecting suitable videos from the search results and generating a playlist taking the emotion data into consideration. This enables request analysis that reflects the user's emotional state and optimal video selection.

[0867] "User" refers to a person who uses this system to input requests and watch videos.

[0868] A "request" refers to a text or voice input by a user describing the content and conditions of the video they wish to watch.

[0869] "Emotion recognition means" refers to a means for analyzing the user's voice and facial expressions to obtain emotional data in real time.

[0870] "Request data" refers to data including the content and conditions of the video entered by the user.

[0871] "Emotional data" refers to data that represents the analyzed emotional state of a user.

[0872] "Appropriate data format" refers to converting request data and emotion data into a standard format such as JSON format for sending to the server.

[0873] "Generative AI" refers to an artificial intelligence model that analyzes request data and emotional data, extracts important keywords and elements, and selects videos that are appropriate for the request.

[0874] "Playlist" refers to a list of videos curated based on user requests and emotional data.

[0875] "Video search service" refers to a service that allows users to search for related videos using an internal database or external API.

[0876] "Database" refers to the system that stores and manages videos and other related data.

[0877] overview

[0878] This invention relates to a system that, when a user requests to watch multiple videos based on specific content, automatically searches for videos that match the request and the user's emotions and plays them sequentially.

[0879] User request input and emotion recognition

[0880] User

[0881] The user clicks the Generate AI button on the device to display the input screen.

[0882] The user inputs a request into the displayed input screen, such as "Show me 50 videos of multiple cats playing." This can be done by text input or voice input.

[0883] As soon as the user inputs a request, the device activates an emotion engine, which analyzes the user's emotions in real time through voice and facial recognition systems.

[0884] Sending and parsing requests

[0885] Terminal

[0886] The terminal acquires the request data input by the user and the emotion data acquired from the emotion engine.

[0887] The device converts this data into an appropriate data format, such as JSON.

[0888] The terminal sends the converted data to the server as an HTTP request.

[0889] server

[0890] The server receives the request data and emotion data sent from the terminal.

[0891] The server passes the request data to the AI ​​generator for analysis. The AI ​​generator extracts important keywords and context from the request data, and also analyzes sentiment data to identify factors that influence the request.

[0892] For example, in addition to elements such as "cats," "playing," "video," and "50 items," it identifies that the user is expressing the emotion of "joy."

[0893] Video search and selection

[0894] server

[0895] The server searches for videos using a video database or external video search API based on the analyzed elements and emotional data, with a particular focus on videos that are likely to evoke "joy."

[0896] The server retrieves a list of videos that match the conditions.

[0897] Playlist generation and submission

[0898] server

[0899] The server selects 50 videos that match the request and emotion data.

[0900] The server lists the information of the selected videos (title, URL, thumbnail, emotion tag, etc.) and generates a playlist.

[0901] The playlist is converted into an appropriate data format, such as JSON, and sent to the device.

[0902] Receive and autoplay playlists

[0903] Terminal

[0904] The terminal receives the playlist data transmitted from the server.

[0905] The device analyzes the received playlist data and obtains the playback order and URL information for each video.

[0906] The device will play the first video in the playlist order, and when that video finishes playing, the next video will automatically play.

[0907] Examples of concrete examples and prompts

[0908] Specific examples

[0909] A specific example will be described in which a user requests "Show me 50 videos of multiple cats playing," and the emotion engine recognizes the user's "joy."

[0910] 1. Request input

[0911] The user enters the text "Show me 50 videos of multiple cats playing," and the device, through its emotion engine, recognizes that the user's emotion is "joy."

[0912] 2. Parsing the Request

[0913] The server receives the request and emotion data and uses a generative AI to extract and analyze the keywords "cat," "playing," "video," "50 items," and the emotion "joy."

[0914] 3. Search and select videos

[0915] The server searches a video database based on the keywords and emotion data and selects 50 videos that match the criteria, with priority given to videos that match the emotion of joy.

[0916] 4. Generate a playlist

[0917] The server generates a playlist of 50 videos that match the conditions and sends it to the device.

[0918] 5. Continuous video playback

[0919] The terminal plays videos based on the received playlist, and the user watches a series of videos that give them pleasure as requested.

[0920] Prompt Sentence Examples

[0921] "The user inputs a request via text or voice, and the emotion engine analyzes the user's emotion based on that request. The request including that emotion data is then sent to the server, where the generation AI analyzes it and searches for and selects videos. The system then generates a playlist and provides it to the user. Please explain the process."

[0922] In this way, users can smoothly view content that matches their individual emotions, providing a more satisfying viewing experience.

[0923] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0924] System program processing flow

[0925] Step 1:

[0926] User request input and emotion recognition

[0927] Input: The user clicks the Generate AI button on the device and enters a request in text or voice on the input screen, such as "Show me 50 videos of multiple cats playing."

[0928] Processing: The terminal receives the user's request, and at the same time, the emotion engine analyzes the user's emotion data through the user's voice or face recognition system. The emotion engine recognizes the user's emotion in real time and adds the data to the request.

[0929] Output: The user's request data and emotion data are acquired by the terminal.

[0930] Specific behavior:

[0931] When a user says, "Show me 50 videos of cats playing," the speech recognition module converts the speech into text data, and the emotion engine detects "joy" from the user's tone of voice and facial expression.

[0932] Step 2:

[0933] Sending and parsing requests

[0934] Input: The request data and sentiment data obtained in step 1.

[0935] Processing: The device converts the request data and emotion data into an appropriate data format, such as JSON, and sends the converted data to the server as an HTTP request.

[0936] Output: Request data and emotion data sent to the server.

[0937] Specific behavior:

[0938] The device converts the request and emotion data into JSON format as "Request data: 50 videos of cats playing" and "Emotion data: Joy" and sends it to the server using the HTTP POST method.

[0939] Step 3:

[0940] Receiving and parsing request and emotion data

[0941] Input: Request data and emotion data in JSON format sent from the device.

[0942] Processing: The server receives the request data and emotion data and passes it to the generation AI, which extracts important keywords and context from the request data and analyzes the emotion data to identify factors that influence the request.

[0943] Output: Request and sentiment elements analyzed by the generative AI.

[0944] Specific behavior:

[0945] The server extracts and analyzes emotional data related to keywords such as "cat," "playing," "video," "50 items," and "joy." Based on this data, the AI ​​generator identifies the video characteristics desired by the user.

[0946] Step 4:

[0947] Video search and selection

[0948] Input: Keywords and sentiment elements analyzed by the generative AI.

[0949] Processing: Based on the analyzed elements, the server searches for related videos using a video database or an external video search API. It prioritizes videos that are particularly likely to evoke "joy." It then selects videos that match the search criteria from the search results.

[0950] Output: A list of videos that match the criteria.

[0951] Specific behavior:

[0952] The generation AI calls the search API to search for videos containing the tags "cat," "playing," and "joy." It then retrieves information about videos that match the criteria (title, URL, thumbnail, etc.).

[0953] Step 5:

[0954] Playlist generation and submission

[0955] Input: A list of videos that match the criteria.

[0956] Processing: The server selects 50 videos that match the user's request and emotion, lists their information (title, URL, thumbnail, emotion tag, etc.), and generates a playlist. The playlist is converted into an appropriate data format such as JSON and sent to the device.

[0957] Output: The playlist data sent to the device.

[0958] Specific behavior:

[0959] The server selects 50 videos, generates a JSON file containing the title, URL, thumbnail, and emotion tag, and sends it to the terminal as an HTTP response.

[0960] Step 6:

[0961] Receive and autoplay playlists

[0962] Input: Playlist data sent from the server.

[0963] Processing: The device analyzes the received playlist data and obtains the playback order and URL information for each video. It plays the first video in the playlist order, and when that video finishes playing, the next video automatically starts playing.

[0964] Output: The continuous video viewing experience delivered to the user.

[0965] Specific behavior:

[0966] The device parses the received JSON playlist and passes the URL of the first video to the browser or application's video player, starting autoplay. The next video will automatically play after the previous one finishes.

[0967] The above processing allows users to smoothly view content that matches their individual emotions, providing a more satisfying viewing experience.

[0968] (Application example 2)

[0969] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0970] When a user wants to watch a video, it is necessary to efficiently search for videos that meet the user's request and provide videos that are tailored to the user's emotions. However, conventional systems cannot simultaneously consider the user's request and emotions, making it difficult to improve user satisfaction. To solve this problem, a system is needed that analyzes both the user's request and real-time emotional data and provides the optimal video based on that analysis.

[0971] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0972] In this invention, the server includes a means for a user to input a request, a means for analyzing the input request and the user's emotions, and a means for searching for related videos based on the analysis results, thereby enabling efficient searching of related videos based on the user's request and emotions, and improving user satisfaction.

[0973] The "means for users to input requests" refers to an interface that allows users to input the content they request and the type of video they desire by text or voice.

[0974] The "means for analyzing input requests and user emotions" refers to software and hardware for analyzing request data, facial expressions, and voice from users to estimate emotions.

[0975] The "means for searching for relevant videos based on the analysis results" refers to an algorithm and system that searches for appropriate videos from a database or external API based on the analyzed request data and emotion data.

[0976] "Means for selecting searched videos and generating a playlist" refers to an algorithm and system that selects the most appropriate videos from the search results based on the user's requests and emotions, and creates a list for playing them consecutively.

[0977] The "means for playing back the generated playlist" is a player that plays back videos in sequence based on the playlist, allowing the user to view them continuously.

[0978] "Generative AI" refers to models and algorithms that use artificial intelligence to analyze user requests and provide optimal information and content.

[0979] An "emotion recognition engine" is a system or software that analyzes and estimates emotions from a user's facial expressions and voice.

[0980] A "playlist" is a list of videos that are played consecutively in a specified order.

[0981] MODE FOR CARRYING OUT THE INVENTION

[0982] overview

[0983] This invention relates to a system that, when a user wishes to watch videos based on a specific request and the emotions they are feeling at the time, automatically searches for videos that match the request and emotions and plays them sequentially. The following describes in detail an embodiment of this invention.

[0984] User request input and emotion recognition

[0985] User operations

[0986] Users input requests through the device's request input screen. Requests can be entered by text or voice. At this time, an emotion recognition engine analyzes the user's emotions in real time from their facial expressions and voice tone. For example, if a user inputs "Show me 50 videos of multiple cats playing," the system will recognize that the emotion expressed at that time is "joy."

[0987] Sending requests and emotion data

[0988] Terminal handling

[0989] The device receives requests input by the user and emotion data provided by the emotion recognition engine, converts these data into JSON format, and sends it to the server as an HTTP request.

[0990] Parsing requests and sentiment data

[0991] Server Processing

[0992] The server receives the request data and emotion data sent from the device and analyzes the request content and emotion data using a generative AI model. Based on the analysis results, important keywords and contexts are extracted, and the emotion data that influences the request is identified. For example, the keywords "cat," "playing," "video," "50," and "joy" are extracted.

[0993] Video search and selection

[0994] Server Processing

[0995] The server searches for videos using a video database or an external video search API based on the analyzed keywords and emotion data. Videos that match the emotion of joy are searched for first. The most suitable video is selected from the search results.

[0996] Playlist generation and submission

[0997] Server processing and terminal reception

[0998] The server creates a playlist by listing the information of the selected videos (title, URL, thumbnail, emotion tag, etc.). This playlist is converted to JSON format and sent to the device. The device analyzes the received playlist data and plays the videos sequentially, starting from the first video.

[0999] Hardware and software used

[1000] Hardware:

[1001] Smartphone: Enter requests and play videos.

[1002] Camera and microphone: Used to recognize the user's emotions.

[1003] software:

[1004] Emotion recognition engine (e.g., Affectiva, Microsoft Azure Face API): Analyzes user emotions in real time.

[1005] Generative AI models (e.g., GPT-3, BERT): Analyze user requests and extract relevant keywords.

[1006] Video search API (e.g. YouTube Data API, Vimeo API): Search for related videos based on analyzed keywords.

[1007] Specific examples

[1008] If a user requests "Show me 50 videos of multiple cats playing," and the emotion engine recognizes the user's emotion of "joy," the server will search for keywords such as "cat," "playing," "video," "50 cats," and "joy." From the search results, the server will generate a playlist of the 50 best videos and send it to the device. The user will then be able to watch videos that match their emotion in succession, as requested.

[1009] Prompt Sentence Examples

[1010] Animals, Funny Videos, Joy

[1011] Cat, playing, happy

[1012] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1013] Step 1:

[1014] The terminal receives requests from the user and acquires emotion data.

[1015] The user inputs a request using text or voice. The device's emotion recognition engine also uses a camera or microphone to capture the user's emotion data. The input data and emotion data are temporarily stored on the device. The input includes text data or voice data, and the output includes request data and emotion data.

[1016] Step 2:

[1017] The device converts the request data and emotion data into JSON format and sends it to the server.

[1018] The device receives request data input by the user and emotion data from the emotion recognition engine, formats the data into JSON format, and then sends the data to the server using an HTTP request. The input includes the request data and emotion data, and the output includes a JSON-formatted HTTP request.

[1019] Step 3:

[1020] The server receives the request data and emotion data and analyzes it using a generative AI model.

[1021] The server receives JSON-formatted request data and emotion data sent from the device. It uses a generative AI model to analyze the request content (e.g., "cats," "playing," "video," "50") and emotion data (e.g., "joy") and extract important keywords and context. The input includes the JSON-formatted request data and emotion data, and the output includes the analyzed keywords and emotion information.

[1022] Step 4:

[1023] The server performs a video search based on the analysis results.

[1024] The server uses the analyzed keywords and emotional information to search for related videos using a video database or an external video search API. At this time, the search algorithm is adjusted to prioritize videos that match the emotion (e.g., videos that evoke joy). The input includes the analyzed keywords and emotional information, and the output includes a list of videos as search results.

[1025] Step 5:

[1026] The server generates a playlist from the acquired video list and transmits it to the terminal.

[1027] The server selects 50 videos from the search results and lists them as a playlist. The playlist, including information about the selected videos (title, URL, thumbnail, emotion tag, etc.), is converted to JSON format and sent to the terminal. The input contains a list of videos, and the output contains a JSON-formatted playlist.

[1028] Step 6:

[1029] The device automatically plays videos based on the playlist received from the server.

[1030] The device parses the JSON-formatted playlist data received from the server and plays the videos in the list sequentially. The first video in the playlist plays, and when it finishes, the next video automatically plays. The input is the JSON-formatted playlist data, and the output includes the currently playing and already played videos.

[1031] As a specific example, if you request "Show me 50 videos of multiple cats playing" based on the prompt sentences "animals, funny videos, joy" and "cats, playing, happiness," the system will search for videos that have the corresponding emotion of "joy," and the selected videos will be played as a playlist.

[1032] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1033] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1034] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[1035] [Third embodiment]

[1036] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[1037] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[1038] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1039] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[1040] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1041] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1042] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1043] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1044] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1045] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1046] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1047] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[1048] overview

[1049] This invention relates to a system that, when a user requests to watch multiple videos based on specific content, automatically searches for videos that meet the request and plays them sequentially. The following describes in detail an embodiment of the invention.

[1050] User request input

[1051] User

[1052] A user starts a video viewing application on any device (e.g., a smartphone or computer). The application interface includes a Generate AI button for inputting video requests.

[1053] The user clicks the Generate AI button and enters a request in text or voice into the input screen that appears, such as "Show me 50 videos of multiple cats playing."

[1054] Sending and parsing requests

[1055] Terminal

[1056] The terminal sends the request data entered by the user to the server, where the request is converted into an appropriate data format (for example, JSON format).

[1057] The request data sent to the server is analyzed by the AI ​​generator. Specific keywords and contexts are extracted using the analysis method. For example, elements such as "cats," "playing," "video," and "50 items" are identified.

[1058] Video search and selection

[1059] server

[1060] The server searches a video database or an external video providing service based on the analyzed elements.

[1061] From the search results, 50 videos that match the user's request are selected, and the information of the selected videos (title, URL, thumbnail, etc.) is listed.

[1062] Playlist generation and submission

[1063] server

[1064] Based on the selected videos, the server generates a playlist that takes into account the playback order. The playlist includes information such as the URL, title, and thumbnail of each video.

[1065] The playlist is sent to the device in an appropriate data format (e.g., JSON format).

[1066] Receive and autoplay playlists

[1067] Terminal

[1068] The terminal receives the playlist sent from the server.

[1069] Once the playlist has been parsed, the videos will begin playing automatically. Based on the playlist, the user can watch the requested videos in succession.

[1070] Specific examples

[1071] A specific example will be described in which a user requests "Show me 50 videos of multiple cats playing."

[1072] 1. Request input

[1073] The user types a request into the terminal as text and clicks the send button.

[1074] 2. Parsing the Request

[1075] The server receives the request and uses a generation AI to extract the keywords "cat," "playing," "video," and "50."

[1076] 3. Search and select videos

[1077] The server searches a video database based on these keywords and selects 50 videos that meet the criteria.

[1078] 4. Generate a playlist

[1079] The server generates a playlist of 50 videos that match the conditions and sends it to the device.

[1080] 5. Continuous video playback

[1081] The terminal plays the videos based on the received playlist, and the user watches the videos consecutively according to the requested content.

[1082] In this way, users can smoothly watch the content they want, greatly improving convenience. In addition, by using generative AI, it is possible to accurately understand and analyze the content of the request and provide the appropriate video.

[1083] The processing flow will be explained below.

[1084] Step 1: User request input

[1085] The user clicks the Generate AI button displayed on the device, which displays the input screen.

[1086] The user enters a request into the displayed input screen using text or voice, such as "Show me 50 videos of multiple cats playing."

[1087] After completing the input, the user presses the "Send" button to send the input contents.

[1088] Step 2: Submitting the request

[1089] The terminal acquires the request data input by the user.

[1090] The terminal converts the acquired request data into an appropriate data format such as JSON format.

[1091] The terminal transmits the converted request data to the server as an HTTP request.

[1092] Step 3: Receiving and Parsing the Request

[1093] The server receives the request data sent from the terminal.

[1094] The server passes the received request data to the generation AI for analysis.

[1095] The generative AI analyzes the request data and extracts important keywords and context (e.g., "cats," "playing," "video," "50 items").

[1096] The server receives the analysis results from the generation AI.

[1097] Step 4: Search for videos

[1098] Based on the analysis results, the server searches for videos using a video database or an external video search API.

[1099] The server uses the extracted keywords and conditions as a search query.

[1100] The server retrieves a list of related videos from the search results.

[1101] Step 5: Select a video

[1102] The server selects 50 videos from the search results that match the user's request.

[1103] The server lists the information of the selected videos (title, URL, thumbnail, etc.).

[1104] Step 6: Generate a playlist

[1105] The server generates a playlist based on the list of selected videos.

[1106] A playlist contains information such as the playback order, title, URL, and thumbnail of each video.

[1107] Step 7: Submit your playlist

[1108] The server converts the generated playlist into an appropriate data format, such as JSON.

[1109] The server transmits the converted playlist data to the terminal.

[1110] Step 8: Receive and play the playlist

[1111] The terminal receives the playlist data transmitted from the server.

[1112] The device analyzes the received playlist data and obtains the playback order and URL information for each video.

[1113] The device will play the first video in the playlist order, and when that video finishes playing, the next video will automatically play.

[1114] This system allows users to watch videos of their choice continuously without any hassle, greatly improving convenience and the viewing experience.

[1115] Example 1

[1116] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1117] Currently, if a user wants to watch specific content consecutively, they have to manually search for multiple videos and add them to a playlist individually. This makes the process cumbersome for users, making it difficult to quickly watch the content they want. Furthermore, it can be difficult to accurately understand a user's request and provide the appropriate video, which can reduce satisfaction.

[1118] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1119] In this invention, the server includes a means for a user to input a request, a terminal means for transmitting the request input data to the server, a means for analyzing the input request using a generative AI model, a means for searching for related videos from a database or an external video service based on the analysis results, a means for selecting the searched videos and generating a playlist, a means for transmitting the generated playlist to the terminal, and a means for analyzing the generated playlist on the terminal and automatically playing videos. This eliminates the need for complicated user operations and enables users to quickly watch the desired content. Furthermore, the use of the generative AI model makes it possible to accurately understand user requests and provide optimal videos.

[1120] "User" means an individual or corporation that uses the system to input requests and view content.

[1121] "Request input means" refers to an interface or device that allows a user to input a request by text or voice.

[1122] "Terminal means" refers to an electronic device such as a smartphone or computer that sends a user's request to the server.

[1123] A "server" is a computer system that receives requests, analyzes, searches, generates playlists, and transmits playlists.

[1124] A "generative AI model" is an artificial intelligence that analyzes requests entered by users and extracts keywords and context.

[1125] "Analysis means" refers to the process or algorithm that analyzes the input request using a generative AI model and extracts the necessary information.

[1126] "Video Database" means internal or external data storage used to search for relevant videos.

[1127] "Third-party video provider" means an external video hosting platform, such as YouTube.

[1128] "Search method" refers to the process or algorithm for searching for relevant videos from a video database or external video provider service based on the analysis results.

[1129] A "selection method" is a process or algorithm that selects retrieved videos that match the user's request.

[1130] A "playlist generator" is a process or algorithm that compiles selected videos into a playlist, taking into account the playback order.

[1131] "Playlist transmission means" refers to a process or algorithm that transmits the generated playlist to a terminal in an appropriate data format.

[1132] "Autoplay method" refers to the process or algorithm that analyzes the playlist generated by the device and plays videos continuously.

[1133] overview

[1134] This invention relates to a system that automatically searches for and sequentially plays videos that fit a user's request when the user requests to watch multiple videos based on specific content. Specifically, this describes the process flow and device configuration for analyzing the user's request, searching for, selecting, and creating a playlist of related videos, and automatically playing them.

[1135] User request input

[1136] User

[1137] A user launches a video viewing application using a device such as a smartphone or computer.

[1138] The application interface has a Generate AI button, and clicking this will display the input screen.

[1139] The user enters a request into the input screen using text or voice, such as "Show me 50 videos of multiple cats playing," and clicks the send button.

[1140] Sending and parsing requests

[1141] Terminal

[1142] The terminal converts the request data entered by the user into JSON format and sends it to the server, using the HTTP POST method to transfer the data.

[1143] server

[1144] The server uses a generative AI model (e.g., OpenAI's GPT-3) to analyze the received request data.

[1145] The generative AI model extracts necessary keywords and context from the request, such as "cats," "playing," "video," and "50 items."

[1146] Video search and selection

[1147] server

[1148] The server searches its internal video database or external video providers such as YouTube based on the keywords extracted by the generative AI model.

[1149] The server selects 50 videos from the search results that match the user's request. Specifically, it uses the YouTube Data API to search for videos containing tags such as "cat" and "playing" and extracts the appropriate videos.

[1150] List the information of the selected videos (title, URL, thumbnail, etc.).

[1151] Playlist generation and submission

[1152] server

[1153] The server generates a playlist based on the selected video list, taking into consideration the playback order.

[1154] The playlist contains information such as the URL, title, and thumbnail of each video.

[1155] The created playlist is formatted in JSON format and sent to the device as an HTTP response.

[1156] Receive and autoplay playlists

[1157] Terminal

[1158] The terminal receives the playlist sent from the server.

[1159] It analyzes the playlist and automatically starts playing videos based on the information of each video. The videos are played continuously, allowing users to watch the requested videos one after the other.

[1160] Specific examples

[1161] Below is what happens when a user requests "Show me 50 videos of multiple cats playing together."

[1162] 1. Request input

[1163] The user types the text "Show me 50 videos of multiple cats playing" into their device and clicks the send button.

[1164] 2. Parsing the Request

[1165] The server receives the request and uses a generative AI model (such as OpenAI's GPT-3) to extract the keywords "cat," "playing," "video," and "50."

[1166] 3. Search and select videos

[1167] The server searches a video database based on the keywords and selects 50 videos that match the criteria, for example, using the YouTube API.

[1168] 4. Generate a playlist

[1169] The server generates a playlist of 50 videos that match the conditions and sends it to the device.

[1170] 5. Continuous video playback

[1171] The terminal plays the videos based on the received playlist, and the user watches the videos consecutively according to the requested content.

[1172] Prompt Sentence Examples

[1173] "Find 50 videos of multiple cats playing together and create a playlist."

[1174] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1175] Step 1:

[1176] User input of request

[1177] A user launches a video viewing application on a device such as a smartphone or computer.

[1178] Click the Generate AI button on the application interface to display the input screen.

[1179] Users enter a request via text or voice, such as "Show me 50 videos of multiple cats playing," and click the send button.

[1180] Input: The request text or voice data entered by the user.

[1181] Output: The request data is sent to the terminal.

[1182] Step 2:

[1183] Send a request from the device

[1184] The terminal converts the request data entered by the user into JSON format.

[1185] The converted request data is sent to the server using the HTTP POST method.

[1186] Input: Request data from the user (text or voice).

[1187] Output: The request data converted to JSON format is sent to the server.

[1188] Step 3:

[1189] Server parsing of the request

[1190] The server passes the received request data to a generative AI model (e.g., OpenAI's GPT-3).

[1191] The generative AI model extracts keywords and context from the request, such as "cat," "playing," "video," and "50 items."

[1192] Input: Request data in JSON format.

[1193] Output: Keywords and context extracted by the generative AI model.

[1194] Step 4:

[1195] Server-based video search and selection

[1196] The server searches an internal video database or an external video provider (e.g., YouTube) based on the keywords extracted by the generative AI model.

[1197] The server selects 50 videos from the search results that match the user's request. For example, it uses the YouTube Data API to search for videos containing the tags "cat" and "playing."

[1198] Input: Keywords extracted by the generative AI model.

[1199] Output: Information about the selected video (title, URL, thumbnail, etc.).

[1200] Step 5:

[1201] Server-generated and transmitted playlists

[1202] The server generates a playlist based on the selected video list, taking into account the playback order.

[1203] The playlist contents are formatted in JSON format.

[1204] The server sends the generated playlist to the terminal as an HTTP response.

[1205] Input: Information about the selected video (title, URL, thumbnail, etc.).

[1206] Output: Playlist data in JSON format.

[1207] Step 6:

[1208] Device receives and autoplays playlist

[1209] The terminal receives the playlist sent from the server.

[1210] Analyze the playlist and get information about each video.

[1211] The device will automatically start playing videos based on the playlist, allowing users to watch the videos they want in sequence.

[1212] Input: Playlist data in JSON format.

[1213] Output: Automatic continuous playback of videos.

[1214] This allows users to view the content they desire efficiently. The accuracy and speed of server-dependent processing are the keys to improving the overall performance of the system.

[1215] (Application example 1)

[1216] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1217] Conventional video streaming systems have difficulty quickly and accurately selecting relevant videos based on specific user requests and playing them continuously. Furthermore, the technology for accepting voice requests, accurately converting that voice into text, and analyzing it using generative AI models has been immature. This has resulted in inconvenient services that cannot meet user needs.

[1218] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1219] In this invention, the server includes means for allowing a user to input a request, means for analyzing the input request, means for searching for related videos based on the analysis results, means for selecting the searched videos and generating a playlist, means for playing the generated playlist, means for converting the user's voice into text using a voice recognition means, and means for retrieving data from an online video database. This makes it possible to accurately select videos using a generative AI model based on the user's voice or text input request and play them continuously.

[1220] "User" refers to a person who uses the System to enter a request and wish to watch a video.

[1221] "Means for inputting requests" refers to an interface through which a user inputs video requests into the system in the form of voice or text.

[1222] "Means for analyzing requests" refers to the process of analyzing input requests using a generative AI model and extracting relevant keywords and context.

[1223] "Means for searching for related videos" refers to a process for searching for appropriate videos from a video database or external video providing service based on the analyzed keywords and context.

[1224] "Means for selecting and generating a playlist" refers to the process of selecting appropriate videos from the search results and organizing the videos into a playlist based on the playback order.

[1225] The "means for playing the generated playlist" refers to a process for playing videos continuously based on the generated playlist.

[1226] "Speech recognition means" refers to technology for converting user-input speech into text.

[1227] "Means for obtaining data from an online video database" refers to a process for obtaining necessary video data from an online video database via the Internet.

[1228] A "generative AI model" refers to an artificial intelligence algorithm that performs highly accurate analysis based on input information.

[1229] A "playlist" is a list of related videos arranged in a particular order.

[1230] overview

[1231] This invention is a system that automatically searches for and plays back videos that fit a user's request for viewing multiple videos based on specific content, improving user convenience and enabling users to smoothly view the content they desire.

[1232] User request input

[1233] A user uses a device of their choice (e.g., a smartphone or head-mounted display) to launch a video viewing application. The application interface has a voice recognition button for inputting video requests. The user clicks this button and inputs a request by voice, such as "Show me 50 videos of multiple cats playing."

[1234] Sending and parsing requests

[1235] The device sends the request data entered by the user to the server. The request is converted into text by a speech recognition means and then converted into an appropriate data format (e.g., JSON format). Specific keywords and context are extracted by an analysis means. For example, elements such as "cats," "playing," "video," and "50 items" are identified. This is where the generative AI model comes into play.

[1236] Video search and selection

[1237] The server searches online video databases and external video providers based on the analyzed elements. From the search results, it selects 50 videos that match the user's request. The information of the selected videos (title, URL, thumbnail, etc.) is then listed.

[1238] Playlist generation and transmission

[1239] The server generates a playlist based on the selected videos, taking into account the playback order. The playlist includes information such as the URL, title, and thumbnail of each video. The generated playlist is sent to the device in the appropriate data format.

[1240] Receive and autoplay playlists

[1241] The device receives the playlist sent from the server. Once the analysis of the playlist is complete, the video starts playing automatically. Based on the playlist, the user can watch the requested videos in succession.

[1242] Specific examples

[1243] Below is a specific example of a user requesting "Show me 50 videos of multiple cats playing." The user voice-types the request into the device and clicks the send button. The device converts the voice to text and uses a generative AI model to extract the keywords "cat," "playing," "video," and "50." The server searches a video database based on these keywords and selects 50 videos that match the criteria. The server then creates a playlist of the 50 videos that match the criteria and sends it to the device. The device plays the videos based on the received playlist, and the user watches the videos consecutively according to the request.

[1244] Prompt Sentence Examples

[1245] "Show me 50 videos of multiple cats playing together"

[1246] This allows users to easily input requests through voice, and the generative AI model performs accurate analysis and search, allowing them to watch the desired videos continuously.

[1247] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1248] Step 1:

[1249] The user launches a video viewing application and clicks the voice input button. The user then voices their request.

[1250] Input: User's voice request

[1251] Output: Audio data

[1252] Specific actions: Requests such as "Show me 50 videos of multiple cats playing" can be made using voice input.

[1253] Step 2:

[1254] The terminal uses a voice recognition means to convert the voice data into text data.

[1255] Input: Audio data

[1256] Output: Text data

[1257] Specific behavior: Uses a speech recognition library to convert speech into text. Example: "Show me 50 videos of multiple cats playing."

[1258] Step 3:

[1259] The device sends the converted text data to a generative AI model that analyzes the request.

[1260] Input: Text data

[1261] Output: Analysis results (keywords)

[1262] Specific operation: The generative AI model extracts keywords such as "cat," "playing," "video," and "50."

[1263] Step 4:

[1264] The server searches an online video database based on the analyzed keywords.

[1265] Input: Analysis results (keywords)

[1266] Output: Video candidates (title, URL, thumbnail information, etc.)

[1267] Specific operation: Using an API request, search the database for the terms "cat" and "playing" and retrieve 50 videos.

[1268] Step 5:

[1269] The server generates a playlist based on the acquired video candidates, taking into account the playback order.

[1270] Input: Video candidate

[1271] Output: Playlist (video URL, title, thumbnail information, etc.)

[1272] Specific operation: The information of the acquired videos is compiled into a list and compiled into a playlist based on the user's request.

[1273] Step 6:

[1274] The server transmits the generated playlist to the terminal.

[1275] Input: Playlist

[1276] Output: Playlist transmission data

[1277] Specific operation: Send the playlist to the device in an appropriate data format, such as JSON.

[1278] Step 7:

[1279] The device will analyze the received playlist and begin automatically playing the videos.

[1280] Input: Playlist transmission data

[1281] Output: Start of continuous playback

[1282] Specific behavior: Based on the playlist data, videos are played continuously using playback software such as VLC media player.

[1283] Through these steps, users can simply input their requests through voice, and the generative AI model will perform accurate analysis and search, allowing them to watch the desired videos continuously.

[1284] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1285] overview

[1286] This invention relates to a system that, when a user requests to watch multiple videos based on specific content, automatically searches for videos that match the request and the user's emotions and plays them sequentially. The following describes in detail the mode for carrying out the invention.

[1287] User request input and emotion recognition

[1288] User

[1289] The user clicks the Generate AI button on the device, which displays the input screen.

[1290] The user inputs a request into the displayed input screen by text or voice, such as "Show me 50 videos of multiple cats playing." At the same time, the emotion engine analyzes the user's emotions using voice or facial recognition.

[1291] The emotion engine recognizes the user's emotions in real time and uses that information to complete requests.

[1292] Sending and parsing requests

[1293] Terminal

[1294] The terminal acquires request data input by the user and emotion data provided by the emotion engine.

[1295] The device converts the acquired request data and emotion data into an appropriate data format such as JSON.

[1296] The terminal sends the converted data to the server as an HTTP request.

[1297] server

[1298] The server receives the request data and emotion data sent from the terminal.

[1299] The server passes the received request data to the generation AI for analysis. The generation AI analyzes the request data and extracts important keywords and context. At the same time, it analyzes sentiment data and identifies factors that influence the request.

[1300] For example, in addition to elements such as "cats," "playing," "video," and "50 items," it identifies whether the user is expressing the emotion of "joy."

[1301] Video search and selection

[1302] server

[1303] The server searches for videos using a video database or an external video search API based on the analyzed elements and emotion data. Based on the user's emotion, it prioritizes searches for videos that are likely to evoke "joy," for example.

[1304] The server retrieves a list of related videos from the search results.

[1305] Playlist generation and submission

[1306] server

[1307] The server selects 50 videos from the search results that match the user's request and emotions.

[1308] The server lists the information of the selected videos (title, URL, thumbnail, emotion tag, etc.) and generates a playlist.

[1309] The playlist is converted into an appropriate data format, such as JSON, and sent to the device.

[1310] Receive and autoplay playlists

[1311] Terminal

[1312] The terminal receives the playlist data transmitted from the server.

[1313] The device analyzes the received playlist data and obtains the playback order and URL information for each video.

[1314] The device will play the first video in the playlist order, and when that video finishes playing, the next video will automatically play.

[1315] Specific examples

[1316] A specific example will be described in which a user requests "Show me 50 videos of multiple cats playing," and the emotion engine recognizes the user's "joy."

[1317] 1. Request input

[1318] The user enters the text "Show me 50 videos of multiple cats playing," and the device, through its emotion engine, recognizes that the user's emotion is "joy."

[1319] 2. Parsing the Request

[1320] The server receives the request and emotion data and uses a generative AI to extract and analyze the keywords "cat," "playing," "video," "50 items," and the emotion "joy."

[1321] 3. Search and select videos

[1322] The server searches a video database based on the keywords and emotion data and selects 50 videos that match the criteria, with priority given to videos that match the emotion of joy.

[1323] 4. Generate a playlist

[1324] The server generates a playlist of 50 videos that match the conditions and sends it to the device.

[1325] 5. Continuous video playback

[1326] The terminal plays videos based on the received playlist, and the user watches a series of videos that give them pleasure as requested.

[1327] In this way, users can smoothly view content that matches their individual emotions, providing a more satisfying viewing experience.

[1328] The processing flow will be explained below.

[1329] Step 1: User request input and emotion recognition

[1330] The user clicks the Generate AI button on the device, which displays the input screen.

[1331] The user enters a request, such as "Show me 50 videos of multiple cats playing," into the displayed input screen using text or voice.

[1332] At the same time, the emotion engine recognizes the user's emotions (e.g., "joy") in real time through the user's voice input and facial expression analysis using a camera.

[1333] Step 2: Sending the request and emotion data

[1334] The terminal collects request data input by the user and emotion data obtained from the emotion engine.

[1335] The device converts the collected request data and emotion data into JSON format and sends it to the server as an HTTP request.

[1336] Step 3: Receive and parse the request and emotion data

[1337] The server receives the request data and emotion data sent from the terminal.

[1338] The server passes the request data to the generation AI for analysis, which then analyzes the request data and extracts important keywords and context (e.g., "cats," "playing," "video," "50 items").

[1339] The server also analyzes emotional data to recognize the user's emotions (e.g., "happiness") and identify factors that influence the request.

[1340] Step 4: Search and select a video

[1341] The server searches for videos using a video database or an external video search API based on the analyzed elements and emotion data. Based on the user's emotion, it will prioritize videos that evoke "joy," for example.

[1342] The server retrieves a list of relevant videos from the search results and selects 50 videos that meet the user's request and match the user's emotions.

[1343] Step 5: Generate a playlist

[1344] The server lists the information of the selected 50 videos (title, URL, thumbnail, emotion tag, etc.) and generates a playlist taking into account the playback order.

[1345] The playlist is converted to JSON format and sent to the device.

[1346] Step 6: Receiving and parsing the playlist

[1347] The terminal receives the playlist data transmitted from the server.

[1348] The device analyzes the received playlist data and obtains information such as the URL, title, and thumbnail of each video.

[1349] Step 7: Autoplay Video

[1350] The device will play the first video in the playlist, and when it finishes playing, the next video will automatically play.

[1351] Users can watch videos continuously, so they can enjoy the videos they want without stress.

[1352] Examples:

[1353] A specific example will be described in which a user requests "Show me 50 videos of multiple cats playing," and the emotion engine recognizes the user's "joy."

[1354] 1. Request input and emotion recognition

[1355] The user enters the text "Show me 50 videos of multiple cats playing," and the device, through its emotion engine, recognizes that the user's emotion is "joy."

[1356] 2. Sending Requests and Emotion Data

[1357] The device collects request data and emotion data, converts it into JSON format, and sends it to the server.

[1358] 3. Request and Sentiment Data Analysis

[1359] The server analyzes the received data and uses generative AI to extract the keywords "cat," "playing," "video," "50," and the emotion "joy."

[1360] 4. Search and select videos

[1361] The server searches a video database based on the keywords and emotion data, and selects 50 videos that match the criteria. Videos that match the emotion of joy are given priority.

[1362] 5. Generate and send playlists

[1363] The server generates a playlist of the selected 50 videos, converts it to JSON format, and sends it to the device.

[1364] 6. Receiving and parsing the playlist

[1365] The terminal analyzes the received playlist data and obtains information about each video.

[1366] 7. Autoplay Videos

[1367] The device automatically plays videos according to the playlist, allowing users to watch a series of videos that bring them pleasure as requested.

[1368] Example 2

[1369] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1370] While conventional video viewing systems can search for and play videos based on user requests, they are unable to consider the user's emotional state, which can result in a less satisfying viewing experience. Furthermore, they lack a mechanism for properly analyzing the request data and emotional data entered by the user and selecting the most appropriate video based on that data. Furthermore, the quality and content of the video desired by the user often do not completely match the request, making it difficult to provide content that meets the user's viewing objectives.

[1371] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes an emotion recognition means for analyzing the user's emotions in real time, a means for converting input request data and emotion data into an appropriate data format, a means for analyzing the request data and emotion data using a generation AI, and a means for selecting suitable videos from the search results and generating a playlist taking the emotion data into consideration. This enables request analysis that reflects the user's emotional state and optimal video selection.

[1372] "User" refers to a person who uses this system to input requests and watch videos.

[1373] A "request" refers to a text or voice input by a user describing the content and conditions of the video they wish to watch.

[1374] "Emotion recognition means" refers to a means for analyzing the user's voice and facial expressions to obtain emotional data in real time.

[1375] "Request data" refers to data including the content and conditions of the video entered by the user.

[1376] "Emotional data" refers to data that represents the analyzed emotional state of a user.

[1377] "Appropriate data format" refers to converting request data and emotion data into a standard format such as JSON format for sending to the server.

[1378] "Generative AI" refers to an artificial intelligence model that analyzes request data and emotional data, extracts important keywords and elements, and selects videos that are appropriate for the request.

[1379] "Playlist" refers to a list of videos curated based on user requests and emotional data.

[1380] "Video search service" refers to a service that allows users to search for related videos using an internal database or external API.

[1381] "Database" refers to the system that stores and manages videos and other related data.

[1382] overview

[1383] This invention relates to a system that, when a user requests to watch multiple videos based on specific content, automatically searches for videos that match the request and the user's emotions and plays them sequentially.

[1384] User request input and emotion recognition

[1385] User

[1386] The user clicks the Generate AI button on the device to display the input screen.

[1387] The user inputs a request into the displayed input screen, such as "Show me 50 videos of multiple cats playing." This can be done by text input or voice input.

[1388] As soon as the user inputs a request, the device activates an emotion engine, which analyzes the user's emotions in real time through voice and facial recognition systems.

[1389] Sending and parsing requests

[1390] Terminal

[1391] The terminal acquires the request data input by the user and the emotion data acquired from the emotion engine.

[1392] The device converts this data into an appropriate data format, such as JSON.

[1393] The terminal sends the converted data to the server as an HTTP request.

[1394] server

[1395] The server receives the request data and emotion data sent from the terminal.

[1396] The server passes the request data to the AI ​​generator for analysis. The AI ​​generator extracts important keywords and context from the request data, and also analyzes sentiment data to identify factors that influence the request.

[1397] For example, in addition to elements such as "cats," "playing," "video," and "50 items," it identifies that the user is expressing the emotion of "joy."

[1398] Video search and selection

[1399] server

[1400] The server searches for videos using a video database or external video search API based on the analyzed elements and emotional data, with a particular focus on videos that are likely to evoke "joy."

[1401] The server retrieves a list of videos that match the conditions.

[1402] Playlist generation and submission

[1403] server

[1404] The server selects 50 videos that match the request and emotion data.

[1405] The server lists the information of the selected videos (title, URL, thumbnail, emotion tag, etc.) and generates a playlist.

[1406] The playlist is converted into an appropriate data format, such as JSON, and sent to the device.

[1407] Receive and autoplay playlists

[1408] Terminal

[1409] The terminal receives the playlist data transmitted from the server.

[1410] The device analyzes the received playlist data and obtains the playback order and URL information for each video.

[1411] The device will play the first video in the playlist order, and when that video finishes playing, the next video will automatically play.

[1412] Examples of concrete examples and prompts

[1413] Specific examples

[1414] A specific example will be described in which a user requests "Show me 50 videos of multiple cats playing," and the emotion engine recognizes the user's "joy."

[1415] 1. Request input

[1416] The user enters the text "Show me 50 videos of multiple cats playing," and the device, through its emotion engine, recognizes that the user's emotion is "joy."

[1417] 2. Parsing the Request

[1418] The server receives the request and emotion data and uses a generative AI to extract and analyze the keywords "cat," "playing," "video," "50 items," and the emotion "joy."

[1419] 3. Search and select videos

[1420] The server searches a video database based on the keywords and emotion data and selects 50 videos that match the criteria, with priority given to videos that match the emotion of joy.

[1421] 4. Generate a playlist

[1422] The server generates a playlist of 50 videos that match the conditions and sends it to the device.

[1423] 5. Continuous video playback

[1424] The terminal plays videos based on the received playlist, and the user watches a series of videos that give them pleasure as requested.

[1425] Prompt Sentence Examples

[1426] "The user inputs a request via text or voice, and the emotion engine analyzes the user's emotion based on that request. The request including that emotion data is then sent to the server, where the generation AI analyzes it and searches for and selects videos. The system then generates a playlist and provides it to the user. Please explain the process."

[1427] In this way, users can smoothly view content that matches their individual emotions, providing a more satisfying viewing experience.

[1428] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1429] System program processing flow

[1430] Step 1:

[1431] User request input and emotion recognition

[1432] Input: The user clicks the Generate AI button on the device and enters a request in text or voice on the input screen, such as "Show me 50 videos of multiple cats playing."

[1433] Processing: The terminal receives the user's request, and at the same time, the emotion engine analyzes the user's emotion data through the user's voice or face recognition system. The emotion engine recognizes the user's emotion in real time and adds the data to the request.

[1434] Output: The user's request data and emotion data are acquired by the terminal.

[1435] Specific behavior:

[1436] When a user says, "Show me 50 videos of cats playing," the speech recognition module converts the speech into text data, and the emotion engine detects "joy" from the user's tone of voice and facial expression.

[1437] Step 2:

[1438] Sending and parsing requests

[1439] Input: The request data and sentiment data obtained in step 1.

[1440] Processing: The device converts the request data and emotion data into an appropriate data format, such as JSON, and sends the converted data to the server as an HTTP request.

[1441] Output: Request data and emotion data sent to the server.

[1442] Specific behavior:

[1443] The device converts the request and emotion data into JSON format as "Request data: 50 videos of cats playing" and "Emotion data: Joy" and sends it to the server using the HTTP POST method.

[1444] Step 3:

[1445] Receiving and parsing request and emotion data

[1446] Input: Request data and emotion data in JSON format sent from the device.

[1447] Processing: The server receives the request data and emotion data and passes it to the generation AI, which extracts important keywords and context from the request data and analyzes the emotion data to identify factors that influence the request.

[1448] Output: Request and sentiment elements analyzed by the generative AI.

[1449] Specific behavior:

[1450] The server extracts and analyzes emotional data related to keywords such as "cat," "playing," "video," "50 items," and "joy." Based on this data, the AI ​​generator identifies the video characteristics desired by the user.

[1451] Step 4:

[1452] Video search and selection

[1453] Input: Keywords and sentiment elements analyzed by the generative AI.

[1454] Processing: Based on the analyzed elements, the server searches for related videos using a video database or an external video search API. It prioritizes videos that are particularly likely to evoke "joy." It then selects videos that match the search criteria from the search results.

[1455] Output: A list of videos that match the criteria.

[1456] Specific behavior:

[1457] The generation AI calls the search API to search for videos containing the tags "cat," "playing," and "joy." It then retrieves information about videos that match the criteria (title, URL, thumbnail, etc.).

[1458] Step 5:

[1459] Playlist generation and submission

[1460] Input: A list of videos that match the criteria.

[1461] Processing: The server selects 50 videos that match the user's request and emotion, lists their information (title, URL, thumbnail, emotion tag, etc.), and generates a playlist. The playlist is converted into an appropriate data format such as JSON and sent to the device.

[1462] Output: The playlist data sent to the device.

[1463] Specific behavior:

[1464] The server selects 50 videos, generates a JSON file containing the title, URL, thumbnail, and emotion tag, and sends it to the terminal as an HTTP response.

[1465] Step 6:

[1466] Receive and autoplay playlists

[1467] Input: Playlist data sent from the server.

[1468] Processing: The device analyzes the received playlist data and obtains the playback order and URL information for each video. It plays the first video in the playlist order, and when that video finishes playing, the next video automatically starts playing.

[1469] Output: The continuous video viewing experience delivered to the user.

[1470] Specific behavior:

[1471] The device parses the received JSON playlist and passes the URL of the first video to the browser or application's video player, starting autoplay. The next video will automatically play after the previous one finishes.

[1472] The above processing allows users to smoothly view content that matches their individual emotions, providing a more satisfying viewing experience.

[1473] (Application example 2)

[1474] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1475] When a user wants to watch a video, it is necessary to efficiently search for videos that meet the user's request and provide videos that are tailored to the user's emotions. However, conventional systems cannot simultaneously consider the user's request and emotions, making it difficult to improve user satisfaction. To solve this problem, a system is needed that analyzes both the user's request and real-time emotional data and provides the optimal video based on that analysis.

[1476] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1477] In this invention, the server includes a means for a user to input a request, a means for analyzing the input request and the user's emotions, and a means for searching for related videos based on the analysis results, thereby enabling efficient searching of related videos based on the user's request and emotions, and improving user satisfaction.

[1478] The "means for users to input requests" refers to an interface that allows users to input the content they request and the type of video they desire by text or voice.

[1479] The "means for analyzing input requests and user emotions" refers to software and hardware for analyzing request data, facial expressions, and voice from users to estimate emotions.

[1480] The "means for searching for relevant videos based on the analysis results" refers to an algorithm and system that searches for appropriate videos from a database or external API based on the analyzed request data and emotion data.

[1481] "Means for selecting searched videos and generating a playlist" refers to an algorithm and system that selects the most appropriate videos from the search results based on the user's requests and emotions, and creates a list for playing them consecutively.

[1482] The "means for playing back the generated playlist" is a player that plays back videos in sequence based on the playlist, allowing the user to view them continuously.

[1483] "Generative AI" refers to models and algorithms that use artificial intelligence to analyze user requests and provide optimal information and content.

[1484] An "emotion recognition engine" is a system or software that analyzes and estimates emotions from a user's facial expressions and voice.

[1485] A "playlist" is a list of videos that are played consecutively in a specified order.

[1486] MODE FOR CARRYING OUT THE INVENTION

[1487] overview

[1488] This invention relates to a system that, when a user wishes to watch videos based on a specific request and the emotions they are feeling at the time, automatically searches for videos that match the request and emotions and plays them sequentially. The following describes in detail an embodiment of this invention.

[1489] User request input and emotion recognition

[1490] User operations

[1491] Users input requests through the device's request input screen. Requests can be entered by text or voice. At this time, an emotion recognition engine analyzes the user's emotions in real time from their facial expressions and voice tone. For example, if a user inputs "Show me 50 videos of multiple cats playing," the system will recognize that the emotion expressed at that time is "joy."

[1492] Sending requests and emotion data

[1493] Terminal handling

[1494] The device receives requests input by the user and emotion data provided by the emotion recognition engine, converts these data into JSON format, and sends it to the server as an HTTP request.

[1495] Parsing requests and sentiment data

[1496] Server Processing

[1497] The server receives the request data and emotion data sent from the device and analyzes the request content and emotion data using a generative AI model. Based on the analysis results, important keywords and contexts are extracted, and the emotion data that influences the request is identified. For example, the keywords "cat," "playing," "video," "50," and "joy" are extracted.

[1498] Video search and selection

[1499] Server Processing

[1500] The server searches for videos using a video database or an external video search API based on the analyzed keywords and emotion data. Videos that match the emotion of joy are searched for first. The most suitable video is selected from the search results.

[1501] Playlist generation and submission

[1502] Server processing and terminal reception

[1503] The server creates a playlist by listing the information of the selected videos (title, URL, thumbnail, emotion tag, etc.). This playlist is converted to JSON format and sent to the device. The device analyzes the received playlist data and plays the videos sequentially, starting from the first video.

[1504] Hardware and software used

[1505] Hardware:

[1506] Smartphone: Enter requests and play videos.

[1507] Camera and microphone: Used to recognize the user's emotions.

[1508] software:

[1509] Emotion recognition engine (e.g., Affectiva, Microsoft Azure Face API): Analyzes user emotions in real time.

[1510] Generative AI models (e.g., GPT-3, BERT): Analyze user requests and extract relevant keywords.

[1511] Video search API (e.g. YouTube Data API, Vimeo API): Search for related videos based on analyzed keywords.

[1512] Specific examples

[1513] If a user requests "Show me 50 videos of multiple cats playing," and the emotion engine recognizes the user's emotion of "joy," the server will search for keywords such as "cat," "playing," "video," "50 cats," and "joy." From the search results, the server will generate a playlist of the 50 best videos and send it to the device. The user will then be able to watch videos that match their emotion in succession, as requested.

[1514] Prompt Sentence Examples

[1515] Animals, Funny Videos, Joy

[1516] Cat, playing, happy

[1517] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1518] Step 1:

[1519] The terminal receives requests from the user and acquires emotion data.

[1520] The user inputs a request using text or voice. The device's emotion recognition engine also uses a camera or microphone to capture the user's emotion data. The input data and emotion data are temporarily stored on the device. The input includes text data or voice data, and the output includes request data and emotion data.

[1521] Step 2:

[1522] The device converts the request data and emotion data into JSON format and sends it to the server.

[1523] The device receives request data input by the user and emotion data from the emotion recognition engine, formats the data into JSON format, and then sends the data to the server using an HTTP request. The input includes the request data and emotion data, and the output includes a JSON-formatted HTTP request.

[1524] Step 3:

[1525] The server receives the request data and emotion data and analyzes it using a generative AI model.

[1526] The server receives JSON-formatted request data and emotion data sent from the device. It uses a generative AI model to analyze the request content (e.g., "cats," "playing," "video," "50") and emotion data (e.g., "joy") and extract important keywords and context. The input includes the JSON-formatted request data and emotion data, and the output includes the analyzed keywords and emotion information.

[1527] Step 4:

[1528] The server performs a video search based on the analysis results.

[1529] The server uses the analyzed keywords and emotional information to search for related videos using a video database or an external video search API. At this time, the search algorithm is adjusted to prioritize videos that match the emotion (e.g., videos that evoke joy). The input includes the analyzed keywords and emotional information, and the output includes a list of videos as search results.

[1530] Step 5:

[1531] The server generates a playlist from the acquired video list and transmits it to the terminal.

[1532] The server selects 50 videos from the search results and lists them as a playlist. The playlist, including information about the selected videos (title, URL, thumbnail, emotion tag, etc.), is converted to JSON format and sent to the terminal. The input contains a list of videos, and the output contains a JSON-formatted playlist.

[1533] Step 6:

[1534] The device automatically plays videos based on the playlist received from the server.

[1535] The device parses the JSON-formatted playlist data received from the server and plays the videos in the list sequentially. The first video in the playlist plays, and when it finishes, the next video automatically plays. The input is the JSON-formatted playlist data, and the output includes the currently playing and already played videos.

[1536] As a specific example, if you request "Show me 50 videos of multiple cats playing" based on the prompt sentences "animals, funny videos, joy" and "cats, playing, happiness," the system will search for videos that have the corresponding emotion of "joy," and the selected videos will be played as a playlist.

[1537] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1538] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1539] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1540] [Fourth embodiment]

[1541] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1542] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1543] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1544] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1545] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1546] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1547] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1548] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1549] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1550] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1551] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1552] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1553] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1554] overview

[1555] This invention relates to a system that, when a user requests to watch multiple videos based on specific content, automatically searches for videos that meet the request and plays them sequentially. The following describes in detail an embodiment of the invention.

[1556] User request input

[1557] User

[1558] A user starts a video viewing application on any device (e.g., a smartphone or computer). The application interface includes a Generate AI button for inputting video requests.

[1559] The user clicks the Generate AI button and enters a request in text or voice into the input screen that appears, such as "Show me 50 videos of multiple cats playing."

[1560] Sending and parsing requests

[1561] Terminal

[1562] The terminal sends the request data entered by the user to the server, where the request is converted into an appropriate data format (for example, JSON format).

[1563] The request data sent to the server is analyzed by the AI ​​generator. Specific keywords and contexts are extracted using the analysis method. For example, elements such as "cats," "playing," "video," and "50 items" are identified.

[1564] Video search and selection

[1565] server

[1566] The server searches a video database or an external video providing service based on the analyzed elements.

[1567] From the search results, 50 videos that match the user's request are selected, and the information of the selected videos (title, URL, thumbnail, etc.) is listed.

[1568] Playlist generation and submission

[1569] server

[1570] Based on the selected videos, the server generates a playlist that takes into account the playback order. The playlist includes information such as the URL, title, and thumbnail of each video.

[1571] The playlist is sent to the device in an appropriate data format (e.g., JSON format).

[1572] Receive and autoplay playlists

[1573] Terminal

[1574] The terminal receives the playlist sent from the server.

[1575] Once the playlist has been parsed, the videos will begin playing automatically. Based on the playlist, the user can watch the requested videos in succession.

[1576] Specific examples

[1577] A specific example will be described in which a user requests "Show me 50 videos of multiple cats playing."

[1578] 1. Request input

[1579] The user types a request into the terminal as text and clicks the send button.

[1580] 2. Parsing the Request

[1581] The server receives the request and uses a generation AI to extract the keywords "cat," "playing," "video," and "50."

[1582] 3. Search and select videos

[1583] The server searches a video database based on these keywords and selects 50 videos that meet the criteria.

[1584] 4. Generate a playlist

[1585] The server generates a playlist of 50 videos that match the conditions and sends it to the device.

[1586] 5. Continuous video playback

[1587] The terminal plays the videos based on the received playlist, and the user watches the videos consecutively according to the requested content.

[1588] In this way, users can smoothly watch the content they want, greatly improving convenience. In addition, by using generative AI, it is possible to accurately understand and analyze the content of the request and provide the appropriate video.

[1589] The processing flow will be explained below.

[1590] Step 1: User request input

[1591] The user clicks the Generate AI button displayed on the device, which displays the input screen.

[1592] The user enters a request into the displayed input screen using text or voice, such as "Show me 50 videos of multiple cats playing."

[1593] After completing the input, the user presses the "Send" button to send the input contents.

[1594] Step 2: Submitting the request

[1595] The terminal acquires the request data input by the user.

[1596] The terminal converts the acquired request data into an appropriate data format such as JSON format.

[1597] The terminal transmits the converted request data to the server as an HTTP request.

[1598] Step 3: Receiving and Parsing the Request

[1599] The server receives the request data sent from the terminal.

[1600] The server passes the received request data to the generation AI for analysis.

[1601] The generative AI analyzes the request data and extracts important keywords and context (e.g., "cats," "playing," "video," "50 items").

[1602] The server receives the analysis results from the generation AI.

[1603] Step 4: Search for videos

[1604] Based on the analysis results, the server searches for videos using a video database or an external video search API.

[1605] The server uses the extracted keywords and conditions as a search query.

[1606] The server retrieves a list of related videos from the search results.

[1607] Step 5: Select a video

[1608] The server selects 50 videos from the search results that match the user's request.

[1609] The server lists the information of the selected videos (title, URL, thumbnail, etc.).

[1610] Step 6: Generate a playlist

[1611] The server generates a playlist based on the list of selected videos.

[1612] A playlist contains information such as the playback order, title, URL, and thumbnail of each video.

[1613] Step 7: Submit your playlist

[1614] The server converts the generated playlist into an appropriate data format, such as JSON.

[1615] The server transmits the converted playlist data to the terminal.

[1616] Step 8: Receive and play the playlist

[1617] The terminal receives the playlist data transmitted from the server.

[1618] The device analyzes the received playlist data and obtains the playback order and URL information for each video.

[1619] The device will play the first video in the playlist order, and when that video finishes playing, the next video will automatically play.

[1620] This system allows users to watch videos of their choice continuously without any hassle, greatly improving convenience and the viewing experience.

[1621] Example 1

[1622] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1623] Currently, if a user wants to watch specific content consecutively, they have to manually search for multiple videos and add them to a playlist individually. This makes the process cumbersome for users, making it difficult to quickly watch the content they want. Furthermore, it can be difficult to accurately understand a user's request and provide the appropriate video, which can reduce satisfaction.

[1624] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1625] In this invention, the server includes a means for a user to input a request, a terminal means for transmitting the request input data to the server, a means for analyzing the input request using a generative AI model, a means for searching for related videos from a database or an external video service based on the analysis results, a means for selecting the searched videos and generating a playlist, a means for transmitting the generated playlist to the terminal, and a means for analyzing the generated playlist on the terminal and automatically playing videos. This eliminates the need for complicated user operations and enables users to quickly watch the desired content. Furthermore, the use of the generative AI model makes it possible to accurately understand user requests and provide optimal videos.

[1626] "User" means an individual or corporation that uses the system to input requests and view content.

[1627] "Request input means" refers to an interface or device that allows a user to input a request by text or voice.

[1628] "Terminal means" refers to an electronic device such as a smartphone or computer that sends a user's request to the server.

[1629] A "server" is a computer system that receives requests, analyzes, searches, generates playlists, and transmits playlists.

[1630] A "generative AI model" is an artificial intelligence that analyzes requests entered by users and extracts keywords and context.

[1631] "Analysis means" refers to the process or algorithm that analyzes the input request using a generative AI model and extracts the necessary information.

[1632] "Video Database" means internal or external data storage used to search for relevant videos.

[1633] "Third-party video provider" means an external video hosting platform, such as YouTube.

[1634] "Search method" refers to the process or algorithm for searching for relevant videos from a video database or external video provider service based on the analysis results.

[1635] A "selection method" is a process or algorithm that selects retrieved videos that match the user's request.

[1636] A "playlist generator" is a process or algorithm that compiles selected videos into a playlist, taking into account the playback order.

[1637] "Playlist transmission means" refers to a process or algorithm that transmits the generated playlist to a terminal in an appropriate data format.

[1638] "Autoplay method" refers to the process or algorithm that analyzes the playlist generated by the device and plays videos continuously.

[1639] overview

[1640] This invention relates to a system that automatically searches for and sequentially plays videos that fit a user's request when the user requests to watch multiple videos based on specific content. Specifically, this describes the process flow and device configuration for analyzing the user's request, searching for, selecting, and creating a playlist of related videos, and automatically playing them.

[1641] User request input

[1642] User

[1643] A user launches a video viewing application using a device such as a smartphone or computer.

[1644] The application interface has a Generate AI button, and clicking this will display the input screen.

[1645] The user enters a request into the input screen using text or voice, such as "Show me 50 videos of multiple cats playing," and clicks the send button.

[1646] Sending and parsing requests

[1647] Terminal

[1648] The terminal converts the request data entered by the user into JSON format and sends it to the server, using the HTTP POST method to transfer the data.

[1649] server

[1650] The server uses a generative AI model (e.g., OpenAI's GPT-3) to analyze the received request data.

[1651] The generative AI model extracts necessary keywords and context from the request, such as "cats," "playing," "video," and "50 items."

[1652] Video search and selection

[1653] server

[1654] The server searches its internal video database or external video providers such as YouTube based on the keywords extracted by the generative AI model.

[1655] The server selects 50 videos from the search results that match the user's request. Specifically, it uses the YouTube Data API to search for videos containing tags such as "cat" and "playing" and extracts the appropriate videos.

[1656] List the information of the selected videos (title, URL, thumbnail, etc.).

[1657] Playlist generation and submission

[1658] server

[1659] The server generates a playlist based on the selected video list, taking into consideration the playback order.

[1660] The playlist contains information such as the URL, title, and thumbnail of each video.

[1661] The created playlist is formatted in JSON format and sent to the device as an HTTP response.

[1662] Receive and autoplay playlists

[1663] Terminal

[1664] The terminal receives the playlist sent from the server.

[1665] It analyzes the playlist and automatically starts playing videos based on the information of each video. The videos are played continuously, allowing users to watch the requested videos one after the other.

[1666] Specific examples

[1667] Below is what happens when a user requests "Show me 50 videos of multiple cats playing together."

[1668] 1. Request input

[1669] The user types the text "Show me 50 videos of multiple cats playing" into their device and clicks the send button.

[1670] 2. Parsing the Request

[1671] The server receives the request and uses a generative AI model (such as OpenAI's GPT-3) to extract the keywords "cat," "playing," "video," and "50."

[1672] 3. Search and select videos

[1673] The server searches a video database based on the keywords and selects 50 videos that match the criteria, for example, using the YouTube API.

[1674] 4. Generate a playlist

[1675] The server generates a playlist of 50 videos that match the conditions and sends it to the device.

[1676] 5. Continuous video playback

[1677] The terminal plays the videos based on the received playlist, and the user watches the videos consecutively according to the requested content.

[1678] Prompt Sentence Examples

[1679] "Find 50 videos of multiple cats playing together and create a playlist."

[1680] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1681] Step 1:

[1682] User input of request

[1683] A user launches a video viewing application on a device such as a smartphone or computer.

[1684] Click the Generate AI button on the application interface to display the input screen.

[1685] Users enter a request via text or voice, such as "Show me 50 videos of multiple cats playing," and click the send button.

[1686] Input: The request text or voice data entered by the user.

[1687] Output: The request data is sent to the terminal.

[1688] Step 2:

[1689] Send a request from the device

[1690] The terminal converts the request data entered by the user into JSON format.

[1691] The converted request data is sent to the server using the HTTP POST method.

[1692] Input: Request data from the user (text or voice).

[1693] Output: The request data converted to JSON format is sent to the server.

[1694] Step 3:

[1695] Server parsing of the request

[1696] The server passes the received request data to a generative AI model (e.g., OpenAI's GPT-3).

[1697] The generative AI model extracts keywords and context from the request, such as "cat," "playing," "video," and "50 items."

[1698] Input: Request data in JSON format.

[1699] Output: Keywords and context extracted by the generative AI model.

[1700] Step 4:

[1701] Server-based video search and selection

[1702] The server searches an internal video database or an external video provider (e.g., YouTube) based on the keywords extracted by the generative AI model.

[1703] The server selects 50 videos from the search results that match the user's request. For example, it uses the YouTube Data API to search for videos containing the tags "cat" and "playing."

[1704] Input: Keywords extracted by the generative AI model.

[1705] Output: Information about the selected video (title, URL, thumbnail, etc.).

[1706] Step 5:

[1707] Server-generated and transmitted playlists

[1708] The server generates a playlist based on the selected video list, taking into account the playback order.

[1709] The playlist contents are formatted in JSON format.

[1710] The server sends the generated playlist to the terminal as an HTTP response.

[1711] Input: Information about the selected video (title, URL, thumbnail, etc.).

[1712] Output: Playlist data in JSON format.

[1713] Step 6:

[1714] Device receives and autoplays playlist

[1715] The terminal receives the playlist sent from the server.

[1716] Analyze the playlist and get information about each video.

[1717] The device will automatically start playing videos based on the playlist, allowing users to watch the videos they want in sequence.

[1718] Input: Playlist data in JSON format.

[1719] Output: Automatic continuous playback of videos.

[1720] This allows users to view the content they desire efficiently. The accuracy and speed of server-dependent processing are the keys to improving the overall performance of the system.

[1721] (Application example 1)

[1722] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1723] Conventional video streaming systems have difficulty quickly and accurately selecting relevant videos based on specific user requests and playing them continuously. Furthermore, the technology for accepting voice requests, accurately converting that voice into text, and analyzing it using generative AI models has been immature. This has resulted in inconvenient services that cannot meet user needs.

[1724] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1725] In this invention, the server includes means for allowing a user to input a request, means for analyzing the input request, means for searching for related videos based on the analysis results, means for selecting the searched videos and generating a playlist, means for playing the generated playlist, means for converting the user's voice into text using a voice recognition means, and means for retrieving data from an online video database. This makes it possible to accurately select videos using a generative AI model based on the user's voice or text input request and play them continuously.

[1726] "User" refers to a person who uses the System to enter a request and wish to watch a video.

[1727] "Means for inputting requests" refers to an interface through which a user inputs video requests into the system in the form of voice or text.

[1728] "Means for analyzing requests" refers to the process of analyzing input requests using a generative AI model and extracting relevant keywords and context.

[1729] "Means for searching for related videos" refers to a process for searching for appropriate videos from a video database or external video providing service based on the analyzed keywords and context.

[1730] "Means for selecting and generating a playlist" refers to the process of selecting appropriate videos from the search results and organizing the videos into a playlist based on the playback order.

[1731] The "means for playing the generated playlist" refers to a process for playing videos continuously based on the generated playlist.

[1732] "Speech recognition means" refers to technology for converting user-input speech into text.

[1733] "Means for obtaining data from an online video database" refers to a process for obtaining necessary video data from an online video database via the Internet.

[1734] A "generative AI model" refers to an artificial intelligence algorithm that performs highly accurate analysis based on input information.

[1735] A "playlist" is a list of related videos arranged in a particular order.

[1736] overview

[1737] This invention is a system that automatically searches for and plays back videos that fit a user's request for viewing multiple videos based on specific content, improving user convenience and enabling users to smoothly view the content they desire.

[1738] User request input

[1739] A user uses a device of their choice (e.g., a smartphone or head-mounted display) to launch a video viewing application. The application interface has a voice recognition button for inputting video requests. The user clicks this button and inputs a request by voice, such as "Show me 50 videos of multiple cats playing."

[1740] Sending and parsing requests

[1741] The device sends the request data entered by the user to the server. The request is converted into text by a speech recognition means and then converted into an appropriate data format (e.g., JSON format). Specific keywords and context are extracted by an analysis means. For example, elements such as "cats," "playing," "video," and "50 items" are identified. This is where the generative AI model comes into play.

[1742] Video search and selection

[1743] The server searches online video databases and external video providers based on the analyzed elements. From the search results, it selects 50 videos that match the user's request. The information of the selected videos (title, URL, thumbnail, etc.) is then listed.

[1744] Playlist generation and transmission

[1745] The server generates a playlist based on the selected videos, taking into account the playback order. The playlist includes information such as the URL, title, and thumbnail of each video. The generated playlist is sent to the device in the appropriate data format.

[1746] Receive and autoplay playlists

[1747] The device receives the playlist sent from the server. Once the analysis of the playlist is complete, the video starts playing automatically. Based on the playlist, the user can watch the requested videos in succession.

[1748] Specific examples

[1749] Below is a specific example of a user requesting "Show me 50 videos of multiple cats playing." The user voice-types the request into the device and clicks the send button. The device converts the voice to text and uses a generative AI model to extract the keywords "cat," "playing," "video," and "50." The server searches a video database based on these keywords and selects 50 videos that match the criteria. The server then creates a playlist of the 50 videos that match the criteria and sends it to the device. The device plays the videos based on the received playlist, and the user watches the videos consecutively according to the request.

[1750] Prompt Sentence Examples

[1751] "Show me 50 videos of multiple cats playing together"

[1752] This allows users to easily input requests through voice, and the generative AI model performs accurate analysis and search, allowing them to watch the desired videos continuously.

[1753] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1754] Step 1:

[1755] The user launches a video viewing application and clicks the voice input button. The user then voices their request.

[1756] Input: User's voice request

[1757] Output: Audio data

[1758] Specific actions: Requests such as "Show me 50 videos of multiple cats playing" can be made using voice input.

[1759] Step 2:

[1760] The terminal uses a voice recognition means to convert the voice data into text data.

[1761] Input: Audio data

[1762] Output: Text data

[1763] Specific behavior: Uses a speech recognition library to convert speech into text. Example: "Show me 50 videos of multiple cats playing."

[1764] Step 3:

[1765] The device sends the converted text data to a generative AI model that analyzes the request.

[1766] Input: Text data

[1767] Output: Analysis results (keywords)

[1768] Specific operation: The generative AI model extracts keywords such as "cat," "playing," "video," and "50."

[1769] Step 4:

[1770] The server searches an online video database based on the analyzed keywords.

[1771] Input: Analysis results (keywords)

[1772] Output: Video candidates (title, URL, thumbnail information, etc.)

[1773] Specific operation: Using an API request, search the database for the terms "cat" and "playing" and retrieve 50 videos.

[1774] Step 5:

[1775] The server generates a playlist based on the acquired video candidates, taking into account the playback order.

[1776] Input: Video candidate

[1777] Output: Playlist (video URL, title, thumbnail information, etc.)

[1778] Specific operation: The information of the acquired videos is compiled into a list and compiled into a playlist based on the user's request.

[1779] Step 6:

[1780] The server transmits the generated playlist to the terminal.

[1781] Input: Playlist

[1782] Output: Playlist transmission data

[1783] Specific operation: Send the playlist to the device in an appropriate data format, such as JSON.

[1784] Step 7:

[1785] The device will analyze the received playlist and begin automatically playing the videos.

[1786] Input: Playlist transmission data

[1787] Output: Start of continuous playback

[1788] Specific behavior: Based on the playlist data, videos are played continuously using playback software such as VLC media player.

[1789] Through these steps, users can simply input their requests through voice, and the generative AI model will perform accurate analysis and search, allowing them to watch the desired videos continuously.

[1790] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1791] overview

[1792] This invention relates to a system that, when a user requests to watch multiple videos based on specific content, automatically searches for videos that match the request and the user's emotions and plays them sequentially. The following describes in detail the mode for carrying out the invention.

[1793] User request input and emotion recognition

[1794] User

[1795] The user clicks the Generate AI button on the device, which displays the input screen.

[1796] The user inputs a request into the displayed input screen by text or voice, such as "Show me 50 videos of multiple cats playing." At the same time, the emotion engine analyzes the user's emotions using voice or facial recognition.

[1797] The emotion engine recognizes the user's emotions in real time and uses that information to complete requests.

[1798] Sending and parsing requests

[1799] Terminal

[1800] The terminal acquires request data input by the user and emotion data provided by the emotion engine.

[1801] The device converts the acquired request data and emotion data into an appropriate data format such as JSON.

[1802] The terminal sends the converted data to the server as an HTTP request.

[1803] server

[1804] The server receives the request data and emotion data sent from the terminal.

[1805] The server passes the received request data to the generation AI for analysis. The generation AI analyzes the request data and extracts important keywords and context. At the same time, it analyzes sentiment data and identifies factors that influence the request.

[1806] For example, in addition to elements such as "cats," "playing," "video," and "50 items," it identifies whether the user is expressing the emotion of "joy."

[1807] Video search and selection

[1808] server

[1809] The server searches for videos using a video database or an external video search API based on the analyzed elements and emotion data. Based on the user's emotion, it prioritizes searches for videos that are likely to evoke "joy," for example.

[1810] The server retrieves a list of related videos from the search results.

[1811] Playlist generation and submission

[1812] server

[1813] The server selects 50 videos from the search results that match the user's request and emotions.

[1814] The server lists the information of the selected videos (title, URL, thumbnail, emotion tag, etc.) and generates a playlist.

[1815] The playlist is converted into an appropriate data format, such as JSON, and sent to the device.

[1816] Receive and autoplay playlists

[1817] Terminal

[1818] The terminal receives the playlist data transmitted from the server.

[1819] The device analyzes the received playlist data and obtains the playback order and URL information for each video.

[1820] The device will play the first video in the playlist order, and when that video finishes playing, the next video will automatically play.

[1821] Specific examples

[1822] A specific example will be described in which a user requests "Show me 50 videos of multiple cats playing," and the emotion engine recognizes the user's "joy."

[1823] 1. Request input

[1824] The user enters the text "Show me 50 videos of multiple cats playing," and the device, through its emotion engine, recognizes that the user's emotion is "joy."

[1825] 2. Parsing the Request

[1826] The server receives the request and emotion data and uses a generative AI to extract and analyze the keywords "cat," "playing," "video," "50 items," and the emotion "joy."

[1827] 3. Search and select videos

[1828] The server searches a video database based on the keywords and emotion data and selects 50 videos that match the criteria, with priority given to videos that match the emotion of joy.

[1829] 4. Generate a playlist

[1830] The server generates a playlist of 50 videos that match the conditions and sends it to the device.

[1831] 5. Continuous video playback

[1832] The terminal plays videos based on the received playlist, and the user watches a series of videos that give them pleasure as requested.

[1833] In this way, users can smoothly view content that matches their individual emotions, providing a more satisfying viewing experience.

[1834] The processing flow will be explained below.

[1835] Step 1: User request input and emotion recognition

[1836] The user clicks the Generate AI button on the device, which displays the input screen.

[1837] The user enters a request, such as "Show me 50 videos of multiple cats playing," into the displayed input screen using text or voice.

[1838] At the same time, the emotion engine recognizes the user's emotions (e.g., "joy") in real time through the user's voice input and facial expression analysis using a camera.

[1839] Step 2: Sending the request and emotion data

[1840] The terminal collects request data input by the user and emotion data obtained from the emotion engine.

[1841] The device converts the collected request data and emotion data into JSON format and sends it to the server as an HTTP request.

[1842] Step 3: Receive and parse the request and emotion data

[1843] The server receives the request data and emotion data sent from the terminal.

[1844] The server passes the request data to the generation AI for analysis, which then analyzes the request data and extracts important keywords and context (e.g., "cats," "playing," "video," "50 items").

[1845] The server also analyzes emotional data to recognize the user's emotions (e.g., "happiness") and identify factors that influence the request.

[1846] Step 4: Search and select a video

[1847] The server searches for videos using a video database or an external video search API based on the analyzed elements and emotion data. Based on the user's emotion, it will prioritize videos that evoke "joy," for example.

[1848] The server retrieves a list of relevant videos from the search results and selects 50 videos that meet the user's request and match the user's emotions.

[1849] Step 5: Generate a playlist

[1850] The server lists the information of the selected 50 videos (title, URL, thumbnail, emotion tag, etc.) and generates a playlist taking into account the playback order.

[1851] The playlist is converted to JSON format and sent to the device.

[1852] Step 6: Receiving and parsing the playlist

[1853] The terminal receives the playlist data transmitted from the server.

[1854] The device analyzes the received playlist data and obtains information such as the URL, title, and thumbnail of each video.

[1855] Step 7: Autoplay Video

[1856] The device will play the first video in the playlist, and when it finishes playing, the next video will automatically play.

[1857] Users can watch videos continuously, so they can enjoy the videos they want without stress.

[1858] Examples:

[1859] A specific example will be described in which a user requests "Show me 50 videos of multiple cats playing," and the emotion engine recognizes the user's "joy."

[1860] 1. Request input and emotion recognition

[1861] The user enters the text "Show me 50 videos of multiple cats playing," and the device, through its emotion engine, recognizes that the user's emotion is "joy."

[1862] 2. Sending Requests and Emotion Data

[1863] The device collects request data and emotion data, converts it into JSON format, and sends it to the server.

[1864] 3. Request and Sentiment Data Analysis

[1865] The server analyzes the received data and uses generative AI to extract the keywords "cat," "playing," "video," "50," and the emotion "joy."

[1866] 4. Search and select videos

[1867] The server searches a video database based on the keywords and emotion data, and selects 50 videos that match the criteria. Videos that match the emotion of joy are given priority.

[1868] 5. Generate and send playlists

[1869] The server generates a playlist of the selected 50 videos, converts it to JSON format, and sends it to the device.

[1870] 6. Receiving and parsing the playlist

[1871] The terminal analyzes the received playlist data and obtains information about each video.

[1872] 7. Autoplay Videos

[1873] The device automatically plays videos according to the playlist, allowing users to watch a series of videos that bring them pleasure as requested.

[1874] Example 2

[1875] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1876] While conventional video viewing systems can search for and play videos based on user requests, they are unable to consider the user's emotional state, which can result in a less satisfying viewing experience. Furthermore, they lack a mechanism for properly analyzing the request data and emotional data entered by the user and selecting the most appropriate video based on that data. Furthermore, the quality and content of the video desired by the user often do not completely match the request, making it difficult to provide content that meets the user's viewing objectives.

[1877] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes an emotion recognition means for analyzing the user's emotions in real time, a means for converting input request data and emotion data into an appropriate data format, a means for analyzing the request data and emotion data using a generation AI, and a means for selecting suitable videos from the search results and generating a playlist taking the emotion data into consideration. This enables request analysis that reflects the user's emotional state and optimal video selection.

[1878] "User" refers to a person who uses this system to input requests and watch videos.

[1879] A "request" refers to a text or voice input by a user describing the content and conditions of the video they wish to watch.

[1880] "Emotion recognition means" refers to a means for analyzing the user's voice and facial expressions to obtain emotional data in real time.

[1881] "Request data" refers to data including the content and conditions of the video entered by the user.

[1882] "Emotional data" refers to data that represents the analyzed emotional state of a user.

[1883] "Appropriate data format" refers to converting request data and emotion data into a standard format such as JSON format for sending to the server.

[1884] "Generative AI" refers to an artificial intelligence model that analyzes request data and emotional data, extracts important keywords and elements, and selects videos that are appropriate for the request.

[1885] "Playlist" refers to a list of videos curated based on user requests and emotional data.

[1886] "Video search service" refers to a service that allows users to search for related videos using an internal database or external API.

[1887] "Database" refers to the system that stores and manages videos and other related data.

[1888] overview

[1889] This invention relates to a system that, when a user requests to watch multiple videos based on specific content, automatically searches for videos that match the request and the user's emotions and plays them sequentially.

[1890] User request input and emotion recognition

[1891] User

[1892] The user clicks the Generate AI button on the device to display the input screen.

[1893] The user inputs a request into the displayed input screen, such as "Show me 50 videos of multiple cats playing." This can be done by text input or voice input.

[1894] As soon as the user inputs a request, the device activates an emotion engine, which analyzes the user's emotions in real time through voice and facial recognition systems.

[1895] Sending and parsing requests

[1896] Terminal

[1897] The terminal acquires the request data input by the user and the emotion data acquired from the emotion engine.

[1898] The device converts this data into an appropriate data format, such as JSON.

[1899] The terminal sends the converted data to the server as an HTTP request.

[1900] server

[1901] The server receives the request data and emotion data sent from the terminal.

[1902] The server passes the request data to the AI ​​generator for analysis. The AI ​​generator extracts important keywords and context from the request data, and also analyzes sentiment data to identify factors that influence the request.

[1903] For example, in addition to elements such as "cats," "playing," "video," and "50 items," it identifies that the user is expressing the emotion of "joy."

[1904] Video search and selection

[1905] server

[1906] The server searches for videos using a video database or external video search API based on the analyzed elements and emotional data, with a particular focus on videos that are likely to evoke "joy."

[1907] The server retrieves a list of videos that match the conditions.

[1908] Playlist generation and submission

[1909] server

[1910] The server selects 50 videos that match the request and emotion data.

[1911] The server lists the information of the selected videos (title, URL, thumbnail, emotion tag, etc.) and generates a playlist.

[1912] The playlist is converted into an appropriate data format, such as JSON, and sent to the device.

[1913] Receive and autoplay playlists

[1914] Terminal

[1915] The terminal receives the playlist data transmitted from the server.

[1916] The device analyzes the received playlist data and obtains the playback order and URL information for each video.

[1917] The device will play the first video in the playlist order, and when that video finishes playing, the next video will automatically play.

[1918] Examples of concrete examples and prompts

[1919] Specific examples

[1920] A specific example will be described in which a user requests "Show me 50 videos of multiple cats playing," and the emotion engine recognizes the user's "joy."

[1921] 1. Request input

[1922] The user enters the text "Show me 50 videos of multiple cats playing," and the device, through its emotion engine, recognizes that the user's emotion is "joy."

[1923] 2. Parsing the Request

[1924] The server receives the request and emotion data and uses a generative AI to extract and analyze the keywords "cat," "playing," "video," "50 items," and the emotion "joy."

[1925] 3. Search and select videos

[1926] The server searches a video database based on the keywords and emotion data and selects 50 videos that match the criteria, with priority given to videos that match the emotion of joy.

[1927] 4. Generate a playlist

[1928] The server generates a playlist of 50 videos that match the conditions and sends it to the device.

[1929] 5. Continuous video playback

[1930] The terminal plays videos based on the received playlist, and the user watches a series of videos that give them pleasure as requested.

[1931] Prompt Sentence Examples

[1932] "The user inputs a request via text or voice, and the emotion engine analyzes the user's emotion based on that request. The request including that emotion data is then sent to the server, where the generation AI analyzes it and searches for and selects videos. The system then generates a playlist and provides it to the user. Please explain the process."

[1933] In this way, users can smoothly view content that matches their individual emotions, providing a more satisfying viewing experience.

[1934] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1935] System program processing flow

[1936] Step 1:

[1937] User request input and emotion recognition

[1938] Input: The user clicks the Generate AI button on the device and enters a request in text or voice on the input screen, such as "Show me 50 videos of multiple cats playing."

[1939] Processing: The terminal receives the user's request, and at the same time, the emotion engine analyzes the user's emotion data through the user's voice or face recognition system. The emotion engine recognizes the user's emotion in real time and adds the data to the request.

[1940] Output: The user's request data and emotion data are acquired by the terminal.

[1941] Specific behavior:

[1942] When a user says, "Show me 50 videos of cats playing," the speech recognition module converts the speech into text data, and the emotion engine detects "joy" from the user's tone of voice and facial expression.

[1943] Step 2:

[1944] Sending and parsing requests

[1945] Input: The request data and sentiment data obtained in step 1.

[1946] Processing: The device converts the request data and emotion data into an appropriate data format, such as JSON, and sends the converted data to the server as an HTTP request.

[1947] Output: Request data and emotion data sent to the server.

[1948] Specific behavior:

[1949] The device converts the request and emotion data into JSON format as "Request data: 50 videos of cats playing" and "Emotion data: Joy" and sends it to the server using the HTTP POST method.

[1950] Step 3:

[1951] Receiving and parsing request and emotion data

[1952] Input: Request data and emotion data in JSON format sent from the device.

[1953] Processing: The server receives the request data and emotion data and passes it to the generation AI, which extracts important keywords and context from the request data and analyzes the emotion data to identify factors that influence the request.

[1954] Output: Request and sentiment elements analyzed by the generative AI.

[1955] Specific behavior:

[1956] The server extracts and analyzes emotional data related to keywords such as "cat," "playing," "video," "50 items," and "joy." Based on this data, the AI ​​generator identifies the video characteristics desired by the user.

[1957] Step 4:

[1958] Video search and selection

[1959] Input: Keywords and sentiment elements analyzed by the generative AI.

[1960] Processing: Based on the analyzed elements, the server searches for related videos using a video database or an external video search API. It prioritizes videos that are particularly likely to evoke "joy." It then selects videos that match the search criteria from the search results.

[1961] Output: A list of videos that match the criteria.

[1962] Specific behavior:

[1963] The generation AI calls the search API to search for videos containing the tags "cat," "playing," and "joy." It then retrieves information about videos that match the criteria (title, URL, thumbnail, etc.).

[1964] Step 5:

[1965] Playlist generation and submission

[1966] Input: A list of videos that match the criteria.

[1967] Processing: The server selects 50 videos that match the user's request and emotion, lists their information (title, URL, thumbnail, emotion tag, etc.), and generates a playlist. The playlist is converted into an appropriate data format such as JSON and sent to the device.

[1968] Output: The playlist data sent to the device.

[1969] Specific behavior:

[1970] The server selects 50 videos, generates a JSON file containing the title, URL, thumbnail, and emotion tag, and sends it to the terminal as an HTTP response.

[1971] Step 6:

[1972] Receive and autoplay playlists

[1973] Input: Playlist data sent from the server.

[1974] Processing: The device analyzes the received playlist data and obtains the playback order and URL information for each video. It plays the first video in the playlist order, and when that video finishes playing, the next video automatically starts playing.

[1975] Output: The continuous video viewing experience delivered to the user.

[1976] Specific behavior:

[1977] The device parses the received JSON playlist and passes the URL of the first video to the browser or application's video player, starting autoplay. The next video will automatically play after the previous one finishes.

[1978] The above processing allows users to smoothly view content that matches their individual emotions, providing a more satisfying viewing experience.

[1979] (Application example 2)

[1980] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1981] When a user wants to watch a video, it is necessary to efficiently search for videos that meet the user's request and provide videos that are tailored to the user's emotions. However, conventional systems cannot simultaneously consider the user's request and emotions, making it difficult to improve user satisfaction. To solve this problem, a system is needed that analyzes both the user's request and real-time emotional data and provides the optimal video based on that analysis.

[1982] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1983] In this invention, the server includes a means for a user to input a request, a means for analyzing the input request and the user's emotions, and a means for searching for related videos based on the analysis results, thereby enabling efficient searching of related videos based on the user's request and emotions, and improving user satisfaction.

[1984] The "means for users to input requests" refers to an interface that allows users to input the content they request and the type of video they desire by text or voice.

[1985] The "means for analyzing input requests and user emotions" refers to software and hardware for analyzing request data, facial expressions, and voice from users to estimate emotions.

[1986] The "means for searching for relevant videos based on the analysis results" refers to an algorithm and system that searches for appropriate videos from a database or external API based on the analyzed request data and emotion data.

[1987] "Means for selecting searched videos and generating a playlist" refers to an algorithm and system that selects the most appropriate videos from the search results based on the user's requests and emotions, and creates a list for playing them consecutively.

[1988] The "means for playing back the generated playlist" is a player that plays back videos in sequence based on the playlist, allowing the user to view them continuously.

[1989] "Generative AI" refers to models and algorithms that use artificial intelligence to analyze user requests and provide optimal information and content.

[1990] An "emotion recognition engine" is a system or software that analyzes and estimates emotions from a user's facial expressions and voice.

[1991] A "playlist" is a list of videos that are played consecutively in a specified order.

[1992] MODE FOR CARRYING OUT THE INVENTION

[1993] overview

[1994] This invention relates to a system that, when a user wishes to watch videos based on a specific request and the emotions they are feeling at the time, automatically searches for videos that match the request and emotions and plays them sequentially. The following describes in detail an embodiment of this invention.

[1995] User request input and emotion recognition

[1996] User operations

[1997] Users input requests through the device's request input screen. Requests can be entered by text or voice. At this time, an emotion recognition engine analyzes the user's emotions in real time from their facial expressions and voice tone. For example, if a user inputs "Show me 50 videos of multiple cats playing," the system will recognize that the emotion expressed at that time is "joy."

[1998] Sending requests and emotion data

[1999] Terminal handling

[2000] The device receives requests input by the user and emotion data provided by the emotion recognition engine, converts these data into JSON format, and sends it to the server as an HTTP request.

[2001] Parsing requests and sentiment data

[2002] Server Processing

[2003] The server receives the request data and emotion data sent from the device and analyzes the request content and emotion data using a generative AI model. Based on the analysis results, important keywords and contexts are extracted, and the emotion data that influences the request is identified. For example, the keywords "cat," "playing," "video," "50," and "joy" are extracted.

[2004] Video search and selection

[2005] Server Processing

[2006] The server searches for videos using a video database or an external video search API based on the analyzed keywords and emotion data. Videos that match the emotion of joy are searched for first. The most suitable video is selected from the search results.

[2007] Playlist generation and submission

[2008] Server processing and terminal reception

[2009] The server creates a playlist by listing the information of the selected videos (title, URL, thumbnail, emotion tag, etc.). This playlist is converted to JSON format and sent to the device. The device analyzes the received playlist data and plays the videos sequentially, starting from the first video.

[2010] Hardware and software used

[2011] Hardware:

[2012] Smartphone: Enter requests and play videos.

[2013] Camera and microphone: Used to recognize the user's emotions.

[2014] software:

[2015] Emotion recognition engine (e.g., Affectiva, Microsoft Azure Face API): Analyzes user emotions in real time.

[2016] Generative AI models (e.g., GPT-3, BERT): Analyze user requests and extract relevant keywords.

[2017] Video search API (e.g. YouTube Data API, Vimeo API): Search for related videos based on analyzed keywords.

[2018] Specific examples

[2019] If a user requests "Show me 50 videos of multiple cats playing," and the emotion engine recognizes the user's emotion of "joy," the server will search for keywords such as "cat," "playing," "video," "50 cats," and "joy." From the search results, the server will generate a playlist of the 50 best videos and send it to the device. The user will then be able to watch videos that match their emotion in succession, as requested.

[2020] Prompt Sentence Examples

[2021] Animals, Funny Videos, Joy

[2022] Cat, playing, happy

[2023] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[2024] Step 1:

[2025] The terminal receives requests from the user and acquires emotion data.

[2026] The user inputs a request using text or voice. The device's emotion recognition engine also uses a camera or microphone to capture the user's emotion data. The input data and emotion data are temporarily stored on the device. The input includes text data or voice data, and the output includes request data and emotion data.

[2027] Step 2:

[2028] The device converts the request data and emotion data into JSON format and sends it to the server.

[2029] The device receives request data input by the user and emotion data from the emotion recognition engine, formats the data into JSON format, and then sends the data to the server using an HTTP request. The input includes the request data and emotion data, and the output includes a JSON-formatted HTTP request.

[2030] Step 3:

[2031] The server receives the request data and emotion data and analyzes it using a generative AI model.

[2032] The server receives JSON-formatted request data and emotion data sent from the device. It uses a generative AI model to analyze the request content (e.g., "cats," "playing," "video," "50") and emotion data (e.g., "joy") and extract important keywords and context. The input includes the JSON-formatted request data and emotion data, and the output includes the analyzed keywords and emotion information.

[2033] Step 4:

[2034] The server performs a video search based on the analysis results.

[2035] The server uses the analyzed keywords and emotional information to search for related videos using a video database or an external video search API. At this time, the search algorithm is adjusted to prioritize videos that match the emotion (e.g., videos that evoke joy). The input includes the analyzed keywords and emotional information, and the output includes a list of videos as search results.

[2036] Step 5:

[2037] The server generates a playlist from the acquired video list and transmits it to the terminal.

[2038] The server selects 50 videos from the search results and lists them as a playlist. The playlist, including information about the selected videos (title, URL, thumbnail, emotion tag, etc.), is converted to JSON format and sent to the terminal. The input contains a list of videos, and the output contains a JSON-formatted playlist.

[2039] Step 6:

[2040] The device automatically plays videos based on the playlist received from the server.

[2041] The device parses the JSON-formatted playlist data received from the server and plays the videos in the list sequentially. The first video in the playlist plays, and when it finishes, the next video automatically plays. The input is the JSON-formatted playlist data, and the output includes the currently playing and already played videos.

[2042] As a specific example, if you request "Show me 50 videos of multiple cats playing" based on the prompt sentences "animals, funny videos, joy" and "cats, playing, happiness," the system will search for videos that have the corresponding emotion of "joy," and the selected videos will be played as a playlist.

[2043] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[2044] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[2045] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[2046] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[2047] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[2048] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[2049] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[2050] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[2051] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[2052] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[2053] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[2054] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[2055] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[2056] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[2057] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[2058] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[2059] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[2060] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[2061] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[2062] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[2063] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[2064] The following is further disclosed regarding the above embodiment.

[2065] (Claim 1)

[2066] a means for a user to input a request;

[2067] means for analyzing the input request;

[2068] A means for searching for related videos based on the analysis results;

[2069] means for selecting and generating a playlist of the retrieved videos;

[2070] means for playing the generated playlist.

[2071] (Claim 2)

[2072] 2. The system according to claim 1, wherein the request input means accepts input of text or voice.

[2073] (Claim 3)

[2074] 2. The system according to claim 1, wherein the analysis means analyzes the content of the request using a generation AI.

[2075] "Example 1"

[2076] (Claim 1)

[2077] a means for a user to input a request;

[2078] a terminal means for transmitting the request input data to a server;

[2079] means for analyzing the input request using a generative AI model;

[2080] A means for searching for related videos from a database or an external video providing service based on the analysis results;

[2081] means for selecting and generating a playlist of the retrieved videos;

[2082] means for transmitting the generated playlist to a terminal;

[2083] The system includes means for analyzing the playlist generated on the terminal and automatically playing videos.

[2084] (Claim 2)

[2085] 2. The system according to claim 1, wherein the request input means accepts input of text or voice.

[2086] (Claim 3)

[2087] 2. The system according to claim 1, wherein the analysis means analyzes the content of the request using a generative AI model.

[2088] "Application Example 1"

[2089] (Claim 1)

[2090] a means for a user to input a request;

[2091] means for analyzing the input request;

[2092] A means for searching for related videos based on the analysis results;

[2093] means for selecting and generating a playlist of the retrieved videos;

[2094] means for playing the generated playlist;

[2095] means for converting a user's speech into text using speech recognition means;

[2096] A system including means for retrieving data from an online video database.

[2097] (Claim 2)

[2098] 2. The system according to claim 1, wherein the request input means accepts input of text or voice.

[2099] (Claim 3)

[2100] 2. The system according to claim 1, wherein the analysis means analyzes the content of the request using a generative AI model.

[2101] "Example 2: Combining Emotion Engines"

[2102] (Claim 1)

[2103] a means for a user to input a request;

[2104] emotion recognition means for analyzing the emotion of the user in real time when the request is input;

[2105] means for converting the input request data and emotion data into an appropriate data format;

[2106] means for transmitting the request data and emotion data to a server;

[2107] A means for analyzing the input request and emotion data using a generation AI;

[2108] A means for searching for related videos based on the analysis results using a database or an external video search service;

[2109] A means for selecting relevant videos from the search results and generating a playlist that takes into account emotional data;

[2110] means for playing the generated playlist.

[2111] (Claim 2)

[2112] 2. The system according to claim 1, wherein the request input means accepts input of text or voice.

[2113] (Claim 3)

[2114] 2. The system according to claim 1, wherein the emotion recognition means acquires emotion data through voice analysis or face recognition of the user.

[2115] (Claim 4)

[2116] 2. The system according to claim 1, wherein the analyzing means analyzes the request and emotion data using a generating AI.

[2117] "Application example 2 when combining emotion engines"

[2118] (Claim 1)

[2119] a means for a user to input a request;

[2120] A means for analyzing input requests and user sentiment;

[2121] A means for searching for related videos based on the analysis results;

[2122] a means for selecting the searched videos and generating a playlist;

[2123] means for playing the generated playlist;

[2124] A system including:

[2125] (Claim 2)

[2126] 2. The system according to claim 1, wherein the request input means accepts input of text or voice.

[2127] (Claim 3)

[2128] 2. The system according to claim 1, wherein the analysis means analyzes the content of the request and the user's emotions using a generation AI. [Explanation of symbols]

[2129] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. a means for a user to input a request; means for analyzing the input request; A means for searching for related videos based on the analysis results; means for selecting and generating a playlist of the retrieved videos; a system including means for playing the generated playlist;

2. 2. The system of claim 1, wherein the request input means accepts text or voice input.

3. 2. The system according to claim 1, wherein the analyzing means analyzes the content of the request using a generating AI.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A